Attention position detection device, attention position detection method, program, and recording medium

The device and method adjust gaze position coordinates based on camera orientation to accurately estimate attention position, addressing inaccuracies in existing methods by aligning with the user's eye orientation, ensuring precise estimation.

JP7775001B2Active Publication Date: 2025-11-25CANON KK
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2021147898
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-09-10
Publication Date
2025-11-25
Estimated Expiration
2041-09-10

AI Technical Summary

Technical Problem

Existing gaze detection methods inaccurately estimate the user's attention position due to differences in eye movement characteristics between vertical and horizontal orientations, leading to mismatches in calculated and actual vertical/horizontal directions when the camera is held differently.

Method used

A device and method that includes gaze detection, orientation detection, and conversion means to adjust the gaze position coordinates based on the user's camera orientation, ensuring accurate estimation of attention position by aligning the coordinate system with the user's eye orientation.

Benefits of technology

Enables precise estimation of the user's attention position, maintaining accuracy regardless of how the camera is held, by converting gaze position coordinates to match the user's eye orientation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007775001000004
    Figure 0007775001000004
  • Figure 0007775001000005
    Figure 0007775001000005
  • Figure 0007775001000006
    Figure 0007775001000006
Patent Text Reader

Abstract

To accurately estimate an attention position.SOLUTION: An attention position detection device includes: visual line detection means for detecting a visual line position of a user who uses an imaging apparatus; estimation means for estimating an attention position of the user from the visual line position detected by the visual line detection means; direction detection means for detecting a direction of eyes of the user to the imaging apparatus; and conversion means for converting a coordinate system of the visual line position detected by the visual line detection means and inputting the converted visual line position to the estimation means when the direction detected by the direction detection means is not a predetermined direction.SELECTED DRAWING: Figure 9
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a focus position detection device, a focus position detection method, a program, and Recording medium On the body Regarding. [Background technology]

[0002] There is a technology that estimates the position at which a user is looking (attention position) based on information about the user's gaze position. Patent Document 1 describes a method for detecting the gaze position of a user looking through a camera viewfinder. Patent Document 2 describes a technology that changes the gaze detection method depending on whether the user is holding the camera vertically or horizontally. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2004-8323 [Patent Document 2] Japanese Patent Application Publication No. 2-64513 Summary of the Invention [Problem to be solved by the invention]

[0004] However, in the method of Patent Document 2, the method of detecting the gaze changes depending on whether the device is held vertically or horizontally, which results in a change in the accuracy of gaze detection.

[0005] Because the movement characteristics of the human eyeball differ between the up-down and left-right directions, it is expected that the gaze position can be estimated accurately by performing calculations that take these differences in characteristics into account. Here, the orientation of the eye relative to the camera differs by 90 degrees when the user holds the camera vertically and when the user holds it horizontally. Therefore, if the user changes the way they hold the camera, the calculated vertical / horizontal direction may not match the vertical / horizontal direction of the eye. If the calculated vertical / horizontal direction does not match the vertical / horizontal direction of the eye, the accuracy of estimating the gaze position will decrease.

[0006] The present invention has been made in view of the above-mentioned problems, and has an object to estimate a focus position with high accuracy. [Means for solving the problem]

[0007] A first aspect of the present invention is an attention position detection device comprising: gaze detection means for detecting the gaze position of a user using an imaging device; estimation means for estimating the user's attention position from the gaze position detected by the gaze detection means; orientation detection means for detecting the orientation of the user's eyes relative to the imaging device; and conversion means for, if the orientation detected by the orientation detection means is not a predetermined orientation, converting a coordinate system of the gaze position detected by the gaze detection means and inputting the converted gaze position to the estimation means.

[0008] A second aspect of the present invention is a gaze position detection method comprising: a gaze detection step of detecting the gaze position of a user using an imaging device; an estimation step of estimating the gaze position of the user from the gaze position detected by the gaze detection step; a direction detection step of detecting the direction of the user's eyes relative to the imaging device; and a conversion step of, if the direction detected by the direction detection step is not a predetermined direction, converting a coordinate system of the gaze position detected by the gaze detection step so that the converted gaze position is used in the estimation step.

[0009] A third aspect of the present invention is a program for causing a computer to execute each step of the above-described method for detecting a position of interest.

[0010] A fourth aspect of the present invention is a computer-readable recording medium storing a program for causing a computer to execute each step of the above-described method for detecting a position of interest. [Effects of the Invention]

[0013] According to the present invention, it is possible to estimate the position of interest with high accuracy. [Brief explanation of the drawings]

[0014] [Figure 1] 1 is a schematic diagram illustrating the configuration of an imaging device according to a first embodiment. [Figure 2] 1 is a block diagram of an imaging device according to a first embodiment. [Figure 3] FIG. 1 is an explanatory diagram of human involuntary eye movement; [Figure 4] FIG. 10 is a diagram for explaining a calculation method of neural network operation. [Figure 5] FIG. 10 is a diagram illustrating an example of an activation function for neural network calculation. [Figure 6] FIG. 10 is a diagram illustrating an example of input and output using a neural network. [Figure 7] FIG. 2 is an explanatory diagram of the vertical and horizontal relationship between the eye and the finder in the first embodiment. [Figure 8] 10 is a flowchart of a process for acquiring a position of interest in the first embodiment. [Figure 9] 10A to 10C are explanatory diagrams of coordinate conversion processing of a gaze position in the first embodiment. [Figure 10] 10 is a flowchart of a learning process in the second embodiment. [Figure 11] 10A and 10B are diagrams illustrating the movement trajectory of a detected object and the movement of the line of sight in Example 2. DETAILED DESCRIPTION OF THE INVENTION

[0015] Example 1 Hereinafter, a first embodiment of the present invention will be described. In the first embodiment, an example in which the interest position detection device of the present invention is applied to a camera (digital still camera, imaging device) will be described. However, the present invention is not limited to digital still cameras, and can also be applied to, for example, digital video cameras and smartphones.

[0016] FIG. 1 is a schematic diagram of a camera 1 according to the present invention. The photographing lens unit 1A includes two lenses 101 and 102, an aperture 111, an aperture driver 112, a lens drive motor 113, a lens drive member 114, a photocoupler 115, a pulse plate 116, a mount contact 117, and a focus adjustment circuit 118. For simplicity's sake, the first embodiment shows two lenses 101 and 102; however, in reality, the photographing lens unit 1A includes more than two lenses. The lens drive member 114 includes a drive gear and other components. The photocoupler 115 detects the rotation of the pulse plate 116, which is linked to the lens drive member 114, and transmits this information to the focus adjustment circuit 118. The focus adjustment circuit 118 drives the lens drive motor 113 based on information from the photocoupler 115 and information from the camera housing 1B (information on the lens drive amount) to move the lens 101 and change the focus position. Mount contact 117 is an interface between photographic lens unit 1A and camera housing 1B.

[0017] The camera housing 1B contains an image sensor 2, a CPU 3, a memory unit 4, a display device 10, a display device drive circuit 11, etc. The image sensor 2 is located at the intended imaging plane of the photographing lens unit 1A. The CPU 3 is the central processing unit of the microcomputer, and controls the entire camera 1. The memory unit 4 stores images captured by the image sensor 2, etc. The display device 10 is composed of a liquid crystal display or the like, and displays the captured image (subject image), etc. on the display surface of the display device 10. The display device drive circuit 11 drives the display device 10. The user can view the display surface of the display device 10 through an eyepiece 12.

[0018] Camera housing 1B also contains light sources 13a and 13b, a beam splitter 15, a light-receiving lens 16, an eye image sensor 17, and the like. Light sources 13a and 13b are light sources conventionally used in single-lens reflex cameras and the like to detect the gaze direction from the relationship between the pupil and a reflection image of light reflected by the cornea (corneal reflection image; Purkinje image), and are light sources for illuminating user's eye 14. Specifically, light sources 13a and 13b are infrared light-emitting diodes or the like that emit infrared light insensitive to the user, and are arranged around eyepiece 12. An optical image of illuminated eye 14 (eye image; an image formed by light emitted from light sources 13a and 13b and reflected by eye 14) passes through eyepiece 12 and is reflected by beam splitter 15. The eye image is then formed by light-receiving lens 16 on eye image sensor 17, which is a two-dimensional array of photoelectric elements such as a CMOS. The light receiving lens 16 positions the pupil of the eyeball 14 and the eyeball image sensor 17 in a conjugate imaging relationship. Using a predetermined algorithm described later, the line of sight direction of the eyeball 14 (the viewpoint (viewpoint position; line of sight position) on the display surface of the display device 10) is detected from the position of the corneal reflection image in the eyeball image formed on the eyeball image sensor 17.

[0019] Figure 2 is a block diagram showing the electrical configuration of camera 1, with the same numbers used for components that are the same as those in Figure 1. Connected to CPU 3 are a gaze detection circuit 201, a photometry circuit 202, an autofocus detection circuit 203, a signal input circuit 204, a display device drive circuit 11, a light source drive circuit 205, an object detection circuit 207, and the like. CPU 3 also transmits signals via mount contacts 117 to a focus adjustment circuit 118 disposed in photographing lens unit 1A and to an aperture control circuit 206 included in aperture drive unit 112 within photographing lens unit 1A. Memory unit 4 associated with CPU 3 has the functions of recording image signals from image sensor 2 and eye image sensor 17, and recording gaze correction data that corrects for individual differences in gaze.

[0020] The gaze detection circuit 201 A / D converts the output of the eyeball image pickup element 17 (CCD-EYE) when an eyeball image is formed on the eyeball image pickup element 17 (eye image obtained by picking up an image of the eye), and transmits the result to the CPU 3. The CPU 3 extracts feature points required for gaze detection from the eye image according to a predetermined algorithm described later, and calculates the user's gaze (the viewpoint on the display surface of the display device 10) from the positions of the feature points.

[0021] The photometry circuit 202 amplifies, logarithmically compresses, and A / D converts the signal obtained from the image sensor 2, which also functions as a photometry sensor, specifically the luminance signal corresponding to the brightness of the field, and sends the result to the CPU 3 as field luminance information.

[0022] The autofocus detection circuit 203 A / D converts signal voltages from a plurality of detection elements (a plurality of pixels) included in the image sensor 2 and used for phase difference detection, and sends the converted signal to the CPU 3. The CPU 3 calculates the distance to the subject corresponding to each focus detection point from the signals from the plurality of detection elements. This is a well-known technique known as image plane phase difference AF. In the first embodiment, as an example, it is assumed that there are focus detection points at 180 locations on the image plane corresponding to the 180 locations shown in the field of view image in the viewfinder (the display surface of the display device 10).

[0023] A switch SW1 that is turned on by the first stroke of the release button and starts photometry, distance measurement, line of sight detection, etc. of the camera 1, and a switch SW2 that is turned on by the second stroke of the release button and starts an imaging operation are connected to the signal input circuit 204. The ON signals from the switches SW1 and SW2 are input to the signal input circuit 204 and sent to the CPU 3.

[0024] The light source drive circuit 205 drives the light sources 13a and 13b.

[0025] The object detection circuit 207 detects people, animals, and other specific objects from the image pickup signals stored in the memory unit 4 (image pickup signals from the image pickup element 2 and the eye image pickup element 17).

[0026] FIG. 3 is an explanatory diagram of human fixational eye movement. FIG. 3 shows an example of the gaze position when a user attempts to keep looking at the center of the attention area 301. In FIG. 3, the arrow indicates the movement direction of the gaze position. As shown in FIG. 3, even if the user tries to keep looking at the center of the attention area 301, unintended behavior occurs in the gaze position. This behavior is called fixational eye movement, and three types are known: tremor, drift, and flick. Due to the occurrence of fixational eye movement, even if the gaze position is accurately detected physically, a situation may arise in which the user's intended position (attention position) is not accurately identified. Therefore, in Example 1, machine learning calculations using a neural network are used to accurately estimate the attention position. The neural network used receives the coordinates of the physical gaze position detected from the movement of the eyeball and outputs the coordinates of the attention position. The gaze position and attention position are expressed in a two-dimensional coordinate system of X and Y coordinates.

[0027] FIG. 4 is a diagram that shows a typical computational process of a neural network in machine learning. The circles labeled X1 to X4 indicate input nodes. The node labeled 1 represents a bias b for the input value. The values ​​output from the nodes labeled X1 to X4 are multiplied by the weighting coefficient w on the edge and input to the next nodes (Y1 to Y3). Specifically, the input values ​​X1 to X4 are calculated into Y1 to Y3 by the matrix operation shown in Equation 1.

number

[0028] Equation 2 is the derivation formula for Y1 to Y3. Calculation results are input to the nodes labeled Y1 to Y3 from multiple nodes. As shown in Equation 2, Y1 to Y3 are calculated by adding the values ​​input from the nodes X1 to X4 and the bias b. The calculation from X1 to X4 to Y1 to Y3 is called a neuron calculation.

number

[0029] The activation function h is used for one-input, one-output calculations, with the results of neuron operations Y1 to Y3 as inputs. The outputs Z1 to Z3 of the activation function become the outputs of the first layer of the neural network, and are then passed as inputs for similar processing in the second layer. Equation 3 is a conversion formula from Y1 to Y3 to Z1 to Z3. The number of input nodes and the number of output nodes in each layer are arbitrary. In this explanation, we will use an example where the number of input nodes is 4 and the number of output nodes is 3.

number

[0030] 5(a) to 5(d) show examples of activation functions. FIG. 5(a) is a sigmoid function, FIG. 5(b) is a hyperbolic tangent function, FIG. 5(c) is a step function, and FIG. 5(d) is a ReLU function. For example, from these activation functions, one suitable for the calculation target is selected. In Example 1, the sigmoid function in FIG. 5(a) is used.

[0031] FIG. 6 shows an example of input and output using the neural network shown in FIG. 4. The calculation unit 604 is a trained neural network (trained model) that inputs a gaze position 601 and outputs a focus position 605. The gaze position 601 is input as a combination of a gaze detection X coordinate 602 and a gaze detection Y coordinate 603. Nodes inside the calculation unit 604 undergo activation function calculation processing at the end of each calculation. Note that in FIG. 6, the activation function calculation in FIG. 4 is omitted. Also, the calculation unit 604 in FIG. 6 is shown as having five input nodes and output nodes and two layers, but the node calculations are not limited to this. The weighting coefficients of the calculation unit 604 may be optimized by learning. The focus position 605 is output as a combination of a focus position X coordinate 606 and a focus position Y coordinate 607. Although a neural network is used as an example in the first embodiment, the form of the trained model is not particularly limited, and a recurrent neural network (RNN) with a time-series recursion or machine learning calculation by reinforcement learning may be used. The attention position may be estimated from the gaze position by calculation without using a trained model.

[0032] Here, the calculation unit 604 has learned by regarding the gaze detection X coordinate as the horizontal (left-right) coordinate of the eye, and the gaze detection Y coordinate as the vertical (up-down) coordinate of the eye. It is known that the movement of the human eyeball has different movement characteristics between horizontal and vertical movements due to factors such as the fact that the muscles in the eyeball are not symmetrical. Therefore, the calculation unit 604 can estimate the gaze position with greater accuracy by performing calculations taking into account the vertical, left-right, and up-down directions of the eye.

[0033] In the first embodiment, the coordinate system of the detected gaze position has a first coordinate axis (X axis) parallel to the left-right direction of the camera 1 and a second coordinate axis (Y axis) parallel to the up-down direction of the camera 1. When the camera 1 is held horizontally, the CPU 3 acquires the coordinates of the gaze position so that the vertical-horizontal (XY) direction of the calculation unit 604 matches the vertical-horizontal (XY) direction of the gaze position. Here, when the user holds the camera 1 vertically, the orientation of the user's eyes relative to the camera 1 rotates, so that the vertical-horizontal direction in the calculation does not match the vertical-horizontal direction of the gaze position (user's eyes). Therefore, when the camera 1 is held vertically, the CPU 3 converts the coordinate system of the gaze position so that the X axis is parallel to the left-right direction of the eyes and the Y axis is parallel to the up-down direction of the eyes. The CPU 3 inputs the converted gaze position to the calculation unit 604. This allows the camera 1 to accurately estimate the gaze position. Details of the process of estimating the gaze position will be described later.

[0034] 7(a) to 7(i) are explanatory diagrams of the vertical-horizontal relationship between the eye and the viewfinder (the portion where the eyepiece 12 is provided). Using FIGS. 7(a) to 7(i), we will explain why the calculated vertical-horizontal direction and the vertical-horizontal direction of the user's eye do not match. FIGS. 7(a), 7(b), and 7(c) show examples of how to hold the camera 1. FIG. 7(a) shows a state in which the orientation of the user's eye relative to the camera 1 is a predetermined orientation (the orientation assumed in the calculation), while FIGS. 7(b) and 7(c) show states in which the orientation of the user's eye relative to the camera 1 differs from the predetermined orientation. Specifically, FIG. 7(a) shows a state in which the camera 1 is held in the normal position (landscape shooting position). FIG. 7(b) shows a state in which the camera 1 is rotated to the right and held vertically (right-rotation vertical shooting position). FIG. 7(c) shows a state in which the camera 1 is rotated to the left and held vertically (left-rotation vertical shooting position). Figures 7(d), 7(e), and 7(f) show examples of how the in-viewfinder field image (the image seen when looking through the viewfinder; the image displayed on the display device 10) appears, and correspond to Figures 7(a), 7(b), and 7(c), respectively. Figures 7(g), 7(h), and 7(i) show examples of the orientation of the eye 702 relative to the viewfinder frame 701 (a rectangular frame surrounding the eyepiece 12), and correspond to Figures 7(a), 7(b), and 7(c), respectively. Here, as shown in Figures 7(g), 7(h), and 7(i), the axis parallel to the long side of the viewfinder frame 701 is defined as the viewfinder X-axis, and the axis parallel to the short side is defined as the viewfinder Y-axis.

[0035] As shown in Figure 7(g), in the normal position, the finder X axis corresponds to the horizontal coordinate of the eye, and the finder Y axis corresponds to the vertical coordinate of the eye. In the calculation unit 604, the finder X axis corresponds to the horizontal coordinate, and the finder Y axis corresponds to the vertical coordinate. In this way, in the normal position, the vertical and horizontal directions of the user's eye match the vertical and horizontal directions in the calculation.

[0036] As shown in FIGS. 7(h) and 7(i), in the right-rotation vertical shooting position and the left-rotation vertical shooting position, the finder X axis corresponds to the vertical coordinate of the eye, and the finder Y axis corresponds to the horizontal coordinate of the eye. When the camera 1 is held vertically, as in the right-rotation vertical shooting position and the left-rotation vertical shooting position, the vertical and horizontal directions of the user's eye do not match the vertical and horizontal directions used in the calculation. Specifically, the finder X axis does not match the horizontal coordinate of the eye, and the finder Y axis does not match the vertical coordinate of the eye. When the vertical and horizontal directions of the eye do not match the vertical and horizontal directions used in the calculation, the accuracy of estimating the attention position decreases. Therefore, in the first embodiment, the coordinate system of the gaze position input to the calculator is converted depending on the orientation of the user's eye relative to the camera 1. This enables the attention position to be estimated more accurately.

[0037] Fig. 8 is a flowchart showing attention position acquisition processing 800 for acquiring an attention position. The attention position acquisition processing 800 will be described below with reference to Fig. 8 and Fig. 9. The attention position acquisition processing 800 starts when the CPU 3 detects the gaze position based on the eye image captured by the eye image sensor 17.

[0038] In step S801, the CPU 3 detects the coordinates (X, Y) of the gaze position of the user using the camera 1 based on the eye image captured by the eye image sensor 17.

[0039] In step S802, the CPU 3 detects the orientation of the user's eyes relative to the camera 1. In the first embodiment, the CPU 3 detects the orientation of the user's eyes relative to the camera 1 by analyzing the image of the user's eyes acquired by the eye image sensor 17. The detection method is not limited to image analysis. For example, the CPU 3 may detect the orientation of the user's eyes relative to the camera 1 based on the detection result of an acceleration sensor. The acceleration sensor may be, for example, a gyro sensor. Furthermore, the CPU 3 may detect the orientation of the user's eyes relative to the camera 1 using a sensor that detects the vertical direction of the camera 1 in addition to the acceleration sensor.

[0040] In step S803, the CPU 3 determines the eye direction (camera Depending on the position of the camera 1 (the direction of the user's eyes relative to the camera 1), the CPU 3 determines whether the camera 1 is held in the normal position (landscape position), right rotation portrait position, or left rotation portrait position. If it is in the right rotation portrait position, the CPU 3 proceeds to step S804. If it is in the left rotation portrait position, the CPU 3 proceeds to step S805. If it is in the normal position, the CPU 3 proceeds to step S806.

[0041] In step S804, CPU 3 performs "X → Y" and "-Y → X" conversion on the coordinates (X, Y) of the gaze position. That is, when camera 1 is held in a right-rotation vertical shooting position, CPU 3 rotates the coordinate system of the detected gaze position by 90 degrees to the left.

[0042] In step S805, CPU 3 performs "-X → Y" and "Y → X" conversion on the coordinates (X, Y) of the gaze position. That is, when camera 1 is held in a left rotation vertical shooting position, CPU 3 rotates the coordinate system of the detected gaze position by 90 degrees to the right.

[0043] In step S806, the CPU 3 executes the processing for estimating the attention position by the calculation unit 604 to acquire the attention position.

[0044] The processing of steps S803 to S806 will now be described with reference to Fig. 9. Fig. 9 is a schematic diagram of the calculation processing (calculation processing in which the gaze position is input and the focus position is output) when the camera 1 is held in the normal position and the right-rotation vertical shooting position. The upper half shows the processing in the normal position, and the lower half shows the processing in the right-rotation vertical shooting position. The gaze position 601 is input to the calculation unit 604 as a combination of gaze detection X coordinate 602 and gaze detection Y coordinate 603.

[0045] In the eye direction 901 in the normal position, the finder X axis corresponds to the horizontal coordinate of the eye, and the finder Y axis corresponds to the vertical coordinate of the eye. In the direction 902 calculated by the calculation unit 604, the finder X axis corresponds to the gaze detection X coordinate 602 (horizontal coordinate), and the finder Y axis corresponds to the gaze detection Y coordinate 603 (vertical coordinate). In this way, in shooting in the normal position, the vertical and horizontal directions of the user's eyes match the vertical and horizontal directions in the calculation. Therefore, CPU 3 inputs the X and Y coordinates of the gaze position 907 detected in step S801 to the calculation unit 604 as the gaze detection X coordinate 602 and the gaze detection Y coordinate 603 without coordinate conversion.

[0046] On the other hand, for the eye direction 904 in the right rotation vertical shooting position, the finder X axis corresponds to the vertical coordinate of the eye, and the finder Y axis corresponds to the horizontal coordinate of the eye. However, the calculation unit 604 corresponds the finder X axis to the gaze detection X coordinate 602 (horizontal coordinate) and the finder Y axis to the gaze detection Y coordinate 603 (vertical coordinate). Therefore, when shooting in the right rotation vertical shooting position, the vertical and horizontal direction of the eye does not match the calculated vertical and horizontal direction. Therefore, the CPU 3 performs a conversion process on the X and Y coordinates of the gaze position 909 detected in step S801. Specifically, the CPU 3 multiplies the Y coordinate of the gaze position by -1 and inputs the value as the gaze detection X coordinate 602 and the X coordinate as the gaze detection Y coordinate 603 to the calculation unit 604 (step S804). As a result, the eye direction 904 in the right rotation vertical shooting position is converted to the calculated direction 905, and the vertical and horizontal direction of the eye matches the calculated vertical and horizontal direction. By converting the X and Y coordinates of the gaze position 909, the vertical coordinate of the eye 702 corresponds to the Y-axis coordinate, and the horizontal coordinate corresponds to the X-axis coordinate in the processing of the calculation unit 604. The weighting coefficients of the calculation unit 604 are used for input / output processing in accordance with the vertical and horizontal directions of the user's eye. As a result, even when the user is using the camera 1 while holding it vertically, the camera 1 can accurately estimate the gaze position.

[0047] Similarly, in the case of the left rotation vertical shooting position, the CPU 3 converts the coordinates of the gaze position so that the vertical and horizontal directions of the eyes match the vertical and horizontal directions in the calculation, and performs calculation by the calculation unit 604. Note that in the case of the left rotation vertical shooting position, the CPU 3 sets the Y coordinate of the gaze position 909 as the gaze detection X coordinate 602 of the calculation unit 604, and sets the value obtained by multiplying the X coordinate by -1 as the gaze detection Y coordinate 603, and calculates it by the calculation unit 604. (Step S805).

[0048] Returning to the description of Figure 8, in step S807, CPU 3 determines whether camera 1 is held in the normal position (landscape position), right rotation vertical position, or left rotation vertical position, depending on the eye direction detected in step S802 (the direction of the user's eye relative to camera 1). If it is the right rotation vertical position, CPU 3 proceeds to step S808. If it is the left rotation vertical position, CPU 3 proceeds to step S809. If it is the normal position, CPU 3 proceeds to step S810.

[0049] In step S808, the CPU 3 performs "X → -Y" and "Y → X" conversion on the coordinates (X, Y) of the position of interest acquired in step S806. That is, the CPU 3 converts the coordinate system of the position of interest estimated by the calculation unit 604 by the inverse conversion of step S804.

[0050] In step S809, the CPU 3 performs "X → Y" and "Y → -X" conversion on the coordinates (X, Y) of the position of interest acquired in step S806. That is, the CPU 3 converts the coordinate system of the position of interest estimated by the calculation unit 604 by the inverse conversion of step S805.

[0051] In step S810, the CPU 3 stores the target position (X, Y) in the memory unit 4.

[0052] Here, the processing of steps S807 to S810 will be described with reference to Fig. 9 again. In Fig. 9, a target position 605 is output from the calculation unit 604 as a combination of a target position X coordinate 606 and a target position Y coordinate 607.

[0053] When camera 1 is in the normal position, the calculated vertical and horizontal directions match the vertical and horizontal directions of the user's eyes, as shown in direction 902 calculated by calculation unit 604 and eye direction in the normal position 903. Therefore, CPU 3 stores focus position 605 as focus position 908 in memory unit 4 without performing coordinate conversion (step S810).

[0054] On the other hand, in the case of the right rotation vertical shooting position, as shown by the calculated orientation 905 and the eye orientation 906 at the right rotation vertical shooting position, the calculated vertical and horizontal directions do not match the vertical and horizontal directions of the user's eyes. Therefore, the CPU 3 performs a conversion process on the focus position 605. Specifically, the CPU 3 sets the focus position Y coordinate 607 as the X coordinate of the focus position 910, and stores the value obtained by multiplying the focus position X coordinate 606 by -1 as the Y coordinate of the focus position 910 in the memory unit 4 (step S808). This converts the calculated orientation 905 to the eye orientation 906 at the right rotation vertical shooting position. The coordinate system of the focus position 605 is converted back to the coordinate system before input to the calculation unit 604 (the coordinate system of the gaze position 909). The CPU 3 stores the converted coordinates of the focus position in the memory unit 4.

[0055] Similarly, in the case of the left rotation vertical shooting position, the CPU 3 converts the coordinates of the focus position so as to return them to the coordinate system of the gaze position before being input to the calculation unit 604, and stores the converted coordinates of the focus position in the memory unit 4. Note that in the case of the left rotation vertical shooting position, the CPU 3 stores in the memory unit 4 a value obtained by multiplying the focus position Y coordinate 607 by −1 as the X coordinate of the focus position 910, and the focus position X coordinate 606 as the Y coordinate of the focus position 910 (step S809).

[0056] In the first embodiment, as an example, a method for estimating a gaze position when the camera 1 is held in a normal position and in positions rotated by ±90 degrees (right rotation vertical shooting position, left rotation vertical shooting position) has been described. Note that the method for estimating a gaze position is also applicable to cases where the rotation angle of the camera 1 is other than ±90 degrees. For example, the CPU 3 may transform the coordinate system of the gaze position according to the difference between the direction of the user's eyes and the direction of the eyes in the normal position. For the transformation, for example, an affine transformation or the like may be used.

[0057] In the first embodiment, the calculation unit 604 estimates the gaze position based on the user's gaze position by calculation taking into account the vertical and horizontal directions of the eyes when the user holds the camera 1 horizontally. When the user holds the camera 1 vertically, the CPU 3 converts the coordinate system of the gaze position and inputs the converted gaze position to the calculation unit 604. This allows the camera 1 to accurately estimate the gaze position whether it is held horizontally or vertically. Note that the orientation in which the user's eyes are horizontal with respect to the camera 1 is set as the predetermined orientation for determining whether to perform coordinate conversion, but the predetermined orientation is not limited to this. The predetermined orientation may be any orientation as long as it is assumed in the calculation by the calculation unit 604.

[0058] <Example 2> A second embodiment of the present invention will be described below. In the second embodiment, a learning process for obtaining the trained model (calculation unit 604; estimator; inference unit) described in the first embodiment and a process for acquiring training data used in the learning process will be described. The configuration of the camera 1 in the second embodiment is the same as that shown in FIGS. 1 and 2 in the first embodiment, but in the second embodiment, a predetermined video (video data) for creating training data for the calculation unit 604 is stored in the memory unit 4. The calculation unit 604 is generated by performing learning using a combination of the position of a predetermined object in the video and the gaze position of a user watching the video as training data. The following description will be made with reference to FIGS. 10 and 11.

[0059] 10 is a flowchart showing the learning process 1000. The learning process 1000 starts when the power of the camera 1 is turned on. The learning process 1000 includes a process for acquiring training data.

[0060] In step S1001, CPU 3 determines whether the user has instructed the start of teacher data acquisition. If the user has instructed the start of teacher data acquisition, CPU 3 proceeds to step S1002, and if not, CPU 3 waits in step S1001. The start of teacher data acquisition may be instructed by the user pressing the acquisition start button, for example, by assigning the function of starting teacher data acquisition to the shooting button of camera 1.

[0061] In step S1002, CPU 3 plays a video for creating training data on display device 10. As the video is played, CPU 3 may display a message or the like on display device 10 to notify the user to continue to follow a predetermined object displayed in the video with their eyes. In Example 2, a dog is used as the predetermined object, and the user continues to follow the dog displayed in the video with their eyes. Furthermore, information indicating which direction is vertically upward may be added as metadata to the video data, and the video may be rotated and displayed in an orientation that corresponds to the way camera 1 is held based on that information.

[0062] In step S1003, the object detection circuit 207 detects the position of the object that the user is following with his or her eyes in the current frame image (the displayed frame) in the moving image.

[0063] In step S1004, the CPU 3 records the coordinates of the center of gravity of the detected object (the dog in the second embodiment) detected in step S1003 in the memory unit 4. The center of gravity of the object that the user continues to follow with his or her eyes may be previously associated with the video and recorded in the memory unit 4. In addition, the memory unit 4 may record not only the center of gravity of the dog but also the position of the dog's face, etc.

[0064] In step S1005, CPU 3 detects the coordinates (X, Y) of the gaze position of the user watching the video, based on the eye image captured by eye image sensor 17.

[0065] In step S1006, CPU 3 detects the orientation of the user's eyes relative to camera 1. In the second embodiment, CPU 3 detects the orientation of the user's eyes relative to camera 1 by analyzing the image of the user's eyes acquired by eye image sensor 17. Note that, without being limited to image analysis, CPU 3 may detect the orientation of the user's eyes based on the detection result of an acceleration sensor or the like.

[0066] In step S1007, CPU 3 determines whether camera 1 is held in the normal position (landscape position), right rotation vertical position, or left rotation vertical position, depending on the eye direction detected in step S1006 (the direction of the user's eye relative to camera 1). If it is the right rotation vertical position, CPU 3 proceeds to step S1008. If it is the left rotation vertical position, CPU 3 proceeds to step S1009. If it is the normal position, CPU 3 proceeds to step S1010.

[0067] In step S1008, the CPU 3 performs "X→Y" and "-Y→X" conversion on the coordinates (X, Y) of the gaze position.

[0068] In step S1009, the CPU 3 performs "-X→Y" and "Y→X" conversion on the coordinates (X, Y) of the gaze position.

[0069] In step S1010, CPU 3 does not convert the coordinates of the gaze position, and maintains the coordinate system of the gaze position.

[0070] In step S1011, the CPU 3 records the coordinates of the gaze position in the memory unit 4.

[0071] In this way, when the orientation of the eyes of the user viewing the video for creating training data is not in a predetermined orientation (normal position in the second embodiment), the CPU 3 converts the coordinate system of the user's gaze position so that the vertical and horizontal directions of the user's eyes match the vertical and horizontal directions calculated by the calculation unit 604. The CPU 3 acquires the combination of the gaze position after the conversion and the position of the detected object in the video as training data.

[0072] In step S1012, CPU 3 determines whether the video has ended. If CPU 3 determines that the video has ended, it proceeds to step S1006. If CPU 3 determines that the video has not ended, it returns to step S1003 and executes the processes from step S1003 to step S1011 again. The processing cycle (processing cycle) for recording training data in memory unit 4 should be determined based on the fastest cycle possible for gaze detection processing. Note that the processing cycle may also be determined based on the fastest cycle possible for object detection processing, a processing cycle necessary and sufficient for learning processing, etc.

[0073] In step S1013, the CPU 3 determines whether or not the acquisition of teacher data has been completed. Here, the memory unit 4 stores multiple videos with different object types and movement paths as videos for creating teacher data. By using the teacher data acquired from multiple videos for learning, a trained model with improved estimation accuracy of the focus position can be obtained. The CPU 3 determines that the acquisition of teacher data has been completed if it has performed the processes of steps S1001 to S1012 for all videos stored in the memory unit 4. Furthermore, if there is a video for which the processes of steps S1001 to S1012 have not been performed, the CPU 3 determines that the acquisition of teacher data has not been completed, and performs the processes of steps S1001 to S1012 for the other videos for which teacher data has not been acquired. Note that if the user does not issue an instruction to start acquiring teacher data for the second or subsequent videos, the CPU 3 may proceed to step S1014. Furthermore, the CPU 3 may perform the processes of S1002 to S1012 for the second or subsequent videos for which teacher data has not been acquired without receiving an instruction from the user. That is, once CPU 3 receives an instruction from the user to start acquiring teacher data, it may control the playback of multiple videos in succession and the acquisition of teacher data for the multiple videos. Note that memory unit 4 may store one video for creating teacher data. When memory unit 4 stores one video for creating teacher data, CPU 3 may perform the processes of steps S1001 to S1012 for the one video and then proceed to the process of step S1014.

[0074] In step S1014, CPU 3 performs a learning process sequence for calculation unit 604 using the training data stored in memory unit 4. In the learning process sequence, for example, the following process is performed. CPU 3 causes calculation unit 604 to learn a plurality of training data (a plurality of combinations of gaze position and detected object position). The gaze position included in the training data is used as an input, and the detected object position included in the training data is used as a correct output. If the output of calculation unit 604 differs from the detected object position, CPU 3 updates the weighting coefficients of calculation unit 604 so that the output of calculation unit 604 matches the detected object position. CPU 3 uses a neural network with a structure that has a recurrent effect on data in a time series, such as an RNN, and updates the weighting coefficients in calculation unit 604 to more appropriate ones through backpropagation processing. CPU 3 repeatedly performs these processes using a plurality of training data.

[0075] In step S1015, the CPU 3 determines whether the learning process has ended. If the learning process sequence has been performed on all of the teacher data stored in the memory unit 4, the CPU 3 determines that the learning process has ended and ends the learning process 1000. If there is teacher data for which the learning process sequence has not been performed, the CPU 3 returns to step S1004 and performs the learning process sequence on the other teacher data stored in the memory unit 4.

[0076] In the second embodiment, the camera 1 performs the learning process 1000, which includes the acquisition of training data. However, the acquisition of training data and the learning process may be performed separately. For example, the processing of steps S1001 to S1013 in FIG. 10 may be performed as the acquisition of training data, and the processing of steps S1014 and S1015 may be performed as the learning process. In this case, the learning process may be performed on a PC, a workstation, or another device or system capable of performing machine learning processing. The conversion process of the coordinates (coordinate system) of the gaze position may also be performed by a device other than the camera 1. For example, if the coordinates of the user's gaze position are recorded in association with information about the user's eye direction and an eye image, the coordinates (coordinate system) of the gaze position can be converted based on the information about the eye direction and the eye image, even on a device other than the camera 1.

[0077] FIG. 11 shows an example of a moving image 1101 for creating training data. The moving image 1101 is stored in advance in the memory unit 4. The moving image 1101 is played back on the display device 10 and is visually recognized by the user. A detected object (a dog in FIG. 11) moves from the position of the detected object frame 1102 to the position of the detected object frame 1103 along a detected object trajectory 1104. The detected object trajectory 1104 is the change over time (movement trajectory) of the center of gravity position of the detected object detected by the object detection circuit 207 in step S1003. Note that the detected object trajectory 1104 may be recorded in the memory unit 4 in advance. The gaze trajectory 1105 is the change over time (movement trajectory) of the user's gaze position detected by the CPU 3 in step S1004. Note that the detected object is not limited to a dog, and may be a single object that is easily recognized by the user.

[0078] In the second embodiment, the calculation unit 604 used to estimate the attention position is generated by learning using a combination of the gaze position of the user watching the video and the position of the detected object in the video as training data. If the user holds the camera 1 vertically when acquiring the training data, the coordinate system of the gaze position is converted, and the gaze position after conversion and the position of the detected object in the video are used as training data. The combination of the target position and the position of the projecting object is acquired. Camera 1 using the trained model generated in this way can accurately estimate the target position whether it is held horizontally or vertically.

[0079] <Other Examples> Although the present invention has been described in detail above based on preferred embodiments thereof, the present invention is not limited to these specific embodiments, and various modifications within the scope of the gist of the present invention are also included in the present invention. Parts of the above-described embodiments may be combined as appropriate.

[0080] The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a recording medium, and having one or more processors in the computer of the system or device read and execute the program. It can also be realized by a circuit (e.g., ASIC) that realizes one or more functions. [Explanation of symbols]

[0081] 1: Camera 3: CPU 17: Eye image sensor 201: Gaze detection circuit 604: Arithmetic unit

Claims

1. a gaze detection means for detecting a gaze position of a user who uses the imaging device; an estimation means for estimating a position of attention of the user from the gaze position detected by the gaze detection means; an orientation detection means for detecting an orientation of the user's eyes relative to the imaging device; a conversion means for converting a coordinate system of the gaze position detected by the gaze detection means when the direction detected by the direction detection means is not a predetermined direction, and inputting the converted gaze position to the estimation means; A target position detection device comprising:

2. the estimation means includes a trained model that receives coordinates of the gaze position and outputs coordinates of the attention position, 2. The target position detecting device according to claim 1.

3. Further comprising: a learning means for generating the trained model using training data that is a combination of the gaze position of the user watching a predetermined video and the position of a predetermined object in the predetermined video; 3. The target position detecting device according to claim 2.

4. When acquiring the teacher data, if the orientation detected by the orientation detection means is not the predetermined orientation, The conversion means converts a coordinate system of a gaze position of the user watching the predetermined video, A combination of the gaze position after conversion by the conversion means and the position of the predetermined object in the predetermined video is acquired as the training data.

4. The target position detecting device according to claim 3.

5. a coordinate system of the gaze position detected by the gaze detection means has a first coordinate axis parallel to a left-right direction of the imaging device and a second coordinate axis parallel to a top-bottom direction of the imaging device, The conversion means converts the first coordinate axis into a coordinate system parallel to the left-right direction of the eye, and the second coordinate axis into a coordinate system parallel to the left-right direction of the eye. a coordinate system of the gaze position is transformed so that the coordinate axis of the gaze position is parallel to the vertical direction of the eye; The target position detecting device according to claim 1 .

6. the conversion means rotates a coordinate system of the gaze position detected by the gaze detection means by 90 degrees when the imaging device is held vertically. The target position detecting device according to claim 1 .

7. The conversion means When the imaging device is rotated to the right and held vertically, a coordinate system of the gaze position detected by the gaze detection means is rotated 90 degrees to the left; When the imaging device is rotated to the left and held vertically, the coordinate system of the line of sight position detected by the line of sight detection means is rotated 90 degrees to the right. The target position detecting device according to claim 1 .

8. The method further includes acquiring an eye image of the user's eye, the orientation detection means detects the orientation of the user's eyes relative to the imaging device by analyzing the eye images. The target position detecting device according to claim 1 .

9. an acceleration sensor; the orientation detection means detects the orientation of the user's eyes relative to the imaging device based on the detection result of the acceleration sensor. The target position detecting device according to claim 1 .

10. the conversion means further converts the coordinate system of the gaze position estimated by the estimation means by inversely converting the coordinate system of the gaze position when the direction detected by the direction detection means is not the predetermined direction. The target position detecting device according to claim 1 .

11. 11. The attention position detection device according to claim 1, wherein the conversion means converts the coordinate system of the gaze position in accordance with a difference between the orientation of the user's eyes relative to the imaging device detected by the orientation detection means and the predetermined orientation.

12. A method for detecting a position of interest, comprising: a gaze detection step of detecting a gaze position of a user who uses the imaging device; an estimation step of estimating a position of attention of the user from the gaze position detected by the gaze detection step; an orientation detection step of detecting an orientation of the user's eyes relative to the imaging device; a conversion step of converting a coordinate system of the gaze position detected by the gaze detection step when the direction detected by the direction detection step is not a predetermined direction, so that the converted gaze position is used in the estimation step; A method for detecting a position of interest, comprising:

13. A program for causing a computer to execute each step of the method for detecting a position of interest according to claim 12.

14. 13. A computer-readable recording medium storing a program for causing a computer to execute each step of the method for detecting a position of interest according to claim 12.

Citation Information

Patent Citations

  • Camera with viewing line inputting device

    JP1990064513A

  • Optical device with optic axis detector

    JP1995255676A

  • Optical device with visual axis function

    JP2004008323A

  • Electronic device, control methods for electronic device, program, and storage medium

    JP2021108447A

  • Predictive eye tracking systems and methods for foveated rendering for electronic displays

    WO2021112990A1