electronic machinery

The electronic device addresses the issue of mismatched object selection by using gaze position tracking and size comparison to align with user intentions, enhancing selection accuracy and reducing display clutter.

JP7830101B2Active Publication Date: 2026-03-16CANON KK
View PDF 12 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-12-13
Publication Date
2026-03-16

AI Technical Summary

Technical Problem

Conventional technologies fail to distinguish different ways of looking by a user, leading to mismatches in selecting the intended object or region, thus failing to meet user intentions.

Method used

An electronic device that acquires the user's gaze position and changes over time, estimates attention span, detects objects in the line of sight, compares the size of the gaze area with detected objects, and selects one object based on the closest size ratio to the gaze area.

Benefits of technology

Enables selections that align with the user's intentions, reducing clutter by displaying only relevant information and improving information acquisition efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007830101000001
    Figure 0007830101000001
  • Figure 0007830101000002
    Figure 0007830101000002
  • Figure 0007830101000003
    Figure 0007830101000003
Patent Text Reader

Abstract

To provide a technique that makes it possible to make a selection that matches a user's intention.SOLUTION: An electronic device has estimating means for estimating a gazing range of a user, detection means for detecting one or more objects existing in the direction of the gaze of the user, and selecting means for selecting one object from the one or more objects on the basis of the gazing range.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to electronic devices such as head-mounted displays (HMDs) and cameras, and to a technique for selecting an object or region.

Background Art

[0002] Patent Document 1 discloses that a glasses-type wearable device, which is a type of head-mounted display (HMD), displays a three-dimensional drawing so as to be superimposed on a real object. Patent Document 2 discloses that an in-vehicle device selects a fixation target object based on the direction of a user's line of sight and the movement history. In electronic devices such as HMDs and cameras, by selecting an object or region that the user is looking at and performing control according to the selection result, the convenience of the electronic device can be improved.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0004] However, even when the user is looking at the same position, depending on the way of looking, the object or region that the user wants to select (the object or region that the user is looking at) is different. In conventional technologies such as the technology disclosed in Patent Document 2, such different ways of looking cannot be distinguished, and thus a selection that matches the user's intention may not be made.

[0005] An object of the present invention is to provide a technique that enables a selection that matches the user's intention.

Means for Solving the Problems

[0006] The electronic device of the present invention is An acquisition means for acquiring the user's gaze position, and the change in the gaze position acquired by the acquisition means over a predetermined time period. User's attention span Size Estimation means for estimating, and detection means for detecting one or more objects present in the user's line of sight, When multiple objects are detected by the detection means, the size of the gaze area is compared with the size of each of the multiple objects, and the object whose size ratio to the gaze area is closest to 1 is selected from among the multiple objects. It is characterized by having a selection means for selecting one object. [Effects of the Invention]

[0007] According to the present invention, it becomes possible to make choices that match the user's intentions. [Brief explanation of the drawing]

[0008] [Figure 1] This is an external view of a wearable device. [Figure 2] This is a block diagram showing an example configuration of a wearable device. [Figure 3] This flowchart shows an example of how a wearable device works. [Figure 4] This is a diagram illustrating a specific example of how wearable devices work. [Figure 5] This is an external view of a digital camera. [Figure 6] This is a block diagram showing examples of digital camera configurations. [Figure 7] This flowchart shows an example of how a digital camera works. [Figure 8] This is a diagram illustrating a specific example of how a digital camera works. [Figure 9] This is a diagram showing an example of a digital camera display. [Figure 10] This is a diagram showing an example of a digital camera display. [Modes for carrying out the invention]

[0009] <Example 1> The following describes Example 1 of the present invention. Example 1 describes an example in which the present invention is applied to a wearable device such as a head-mounted display (HMD).

[0010] Figure 1 is an external view of a wearable device 110 as an example of an electronic device to which the present invention can be applied. In Figure 1, the right eye imaging unit 150R and the left eye imaging unit 150L each include a lens 103 and an imaging unit 22, which will be described later. The right eye imaging unit 150R and the left eye imaging unit 150L each have a zoom mechanism, and the user can change the zoom magnification of the right eye imaging unit 150R and the left eye imaging unit 150L using the operation unit 70, which will be described later. The wearable device 110 can perform optical zoom, which controls the lens position by the zoom mechanism, electronic zoom, which crops and enlarges a part of the captured image, or a combination thereof as zoom operations. The right eye display unit 160R and the left eye display unit 160L each include an eyepiece 16, an EVF 29 (Electric View Finder), and an eyeball detection unit 161, which will be described later. When the user is wearing the wearable device 110, they view the image displayed on the right eye display unit 160R (right eye image) with their right eye and the image displayed on the left eye display unit 160L (left eye image) with their left eye. For example, the right eye display unit 160R displays an image captured by the right eye imaging unit 150R, and the left eye display unit 160L displays an image captured by the left eye imaging unit 150L.

[0011] Figure 2 is a block diagram showing an example configuration of the wearable device 110. The lens 103 is composed of multiple lenses, but in Figure 2, it is simplified and shown as a single lens. The system control unit 50 communicates with the lens system control circuit 4 and controls the aperture 1 via the aperture drive circuit 2. The system control unit 50 also focuses on the target by displacing the focus lens included in the lens 103 via the AF drive circuit 3.

[0012] The imaging unit 22 is an imaging device (imaging sensor) composed of a CCD, a CMOS element, or the like that converts an optical image into an electrical signal. An A / D converter (not shown) is provided in the imaging unit 22, and the A / D converter is used to convert the analog signal output from the imaging unit 22 into a digital signal. Imaging by the imaging unit 22 is performed in synchronization with the horizontal synchronization signal and the vertical line synchronization signal output from a timing generator (not shown), and the imaging unit 22 outputs the image data of one frame as frame data at the period of the vertical line synchronization signal. While the event sensor 163 described later is an event-based vision sensor (an asynchronous event-based sensor), the imaging unit 22 is a synchronous frame-based sensor.

[0013] The image processing unit 24 performs predetermined processing (such as resizing processing like pixel interpolation and reduction, color conversion processing, etc.) on the data from the imaging unit 22 (A / D converter) or the data from the memory control unit 15 described later. Further, the image processing unit 24 performs predetermined arithmetic processing using the captured image data, and the system control unit 50 performs exposure control and distance measurement control based on the arithmetic result obtained by the image processing unit 24. Thereby, TTL (through-the-lens) type AF (autofocus) processing and AE (automatic exposure) processing are performed. The image processing unit 24 further performs predetermined arithmetic processing using the captured image data and performs TTL type AWB (auto white balance) processing based on the obtained arithmetic result. Also, the image processing unit 24 can perform picture style processing for converting the captured image (image data) into a color image, a monochrome image, or the like.

[0014] The memory control unit 15 controls the transmission and reception of data among the imaging unit 22, the image processing unit 24, and the memory 32 It controls. The output data from the imaging unit 22 is written into the memory 32 via the image processing unit 24 and the memory control unit 15, or via the memory control unit 15 without passing through the image processing unit 24. The memory 32 stores the image data obtained by the imaging unit 22 and the image data for display on the EVF 29. Also, the memory 32 doubles as a memory for image display (video memory). The image data for display written into the memory 32 is displayed by the EVF 29 via the memory control unit 15.

[0015] The EVF 29 performs display according to the signal from the memory control unit 15 on a display such as an LCD or an organic EL. By sequentially transferring and displaying the image data stored in the memory 32 on the EVF 29, a through-display of the captured image can be performed. The eyepiece part 16 is the eyepiece part of an eyepiece finder (a peeping-type finder), and the user can visually recognize the video displayed on the EVF 29 through the eyepiece part 16. The through-display is the same as what is called live view display in a general digital camera, and in the through-display, the captured image is displayed almost without delay. The user can indirectly visually recognize the real space by visually recognizing the video displayed in the through-display.

[0016] The non-volatile memory 56 is an electrically erasable and recordable memory, such as a Flash-ROM. The non-volatile memory 56 stores constants for the operation of the system control unit 50, programs, etc. The program mentioned here is a program for executing the processes of various flowcharts described later.

[0017] The system control unit 50 is a control unit consisting of at least one processor or circuit, and controls the entire wearable device 110. The system control unit 50 implements the processes described later by executing a program recorded in the non-volatile memory 56. The system memory 52 is, for example, RAM, and the system control unit 50 loads constants, variables, and programs read from the non-volatile memory 56 into the system memory 52 for the operation of the system control unit 50. The system control unit 50 also performs display control by controlling the memory 32, EVF 29, etc.

[0018] The system timer 53 is a timekeeping unit that measures the time used for various controls and the time of the built-in clock.

[0019] The operation unit 70 consists of various operating components that act as input units to receive user input (user operation). The operation unit 70 includes, for example, a voice UI and a touchpad. The touchpad is mounted on the side (not shown) of the wearable device 110. The system control unit 50 can detect operations on the touchpad or the state of the touchpad. The position coordinates of the finger touching the touchpad are notified to the system control unit 50 via the internal bus, and the system control unit 50 determines what kind of operation (touch operation) was performed on the touchpad based on the notified information. The touchpad may be of any type from among various types of touch panels, such as resistive, capacitive, surface acoustic wave, infrared, electromagnetic induction, image recognition, and optical sensor types.

[0020] The power switch 72 is an operating component that switches the power of the wearable device 110 ON and OFF.

[0021] The power control unit 31 consists of a battery detection circuit, a DC-DC converter, a switch circuit for switching which blocks are energized, and detects whether a battery is installed, the type of battery, and the remaining battery level. Furthermore, the power control unit 31 controls the DC-DC converter based on the detection results and instructions from the system control unit 50, supplying the necessary voltage to each part, including the recording medium 200, for the required period. The power supply unit 30 supplies primary batteries such as alkaline batteries and lithium batteries, as well as NiCd batteries. It consists of a battery, rechargeable batteries such as NiMH batteries and lithium-ion batteries, and an AC adapter.

[0022] The recording medium I / F17 is an interface to the recording medium 200, such as a memory card or hard disk. The recording medium 200 is a recording medium such as a memory card for recording captured images, and is composed of semiconductor memory or magnetic disks.

[0023] The communication unit 54 transmits and receives video and audio signals to and from external devices connected wirelessly or via wired cables. The communication unit 54 can also transmit and receive data to and from an external database server. For example, the database server may store 3D CAD data (3D drawing data) of the work object. In this case, it is possible to display the 3D drawing superimposed on the work object (actual object) on the EVF 29, for example, by displaying an image of the work object with the 3D drawing superimposed on it on the EVF 29. The communication unit 54 can also connect to a wireless LAN (Local Area Network) or the internet. Furthermore, the communication unit 54 can communicate with external devices using Bluetooth® or Bluetooth Low Energy.

[0024] The attitude detection unit 55 detects the attitude of the wearable device 110 relative to the direction of gravity. An acceleration sensor or a gyroscope can be used as the attitude detection unit 55. The attitude detection unit 55 makes it possible to detect the movement of the wearable device 110 (pan, tilt, roll, whether it is stationary or not, etc.).

[0025] The eyepiece detection unit 57 is a wear detection sensor that detects whether or not the wearable device 110 is being worn by a user. The system control unit 50 can switch the wearable device 110 on (power on) / off (power off) according to the state detected by the eyepiece detection unit 57. The eyepiece detection unit 57 may also be configured to detect the approach of some object to the eyepiece 16, for example, by using an infrared proximity sensor. When an object approaches, infrared light emitted from the light emitter (not shown) of the eyepiece detection unit 57 is reflected and received by the light receiver (not shown) of the infrared proximity sensor. The amount of infrared light received can be used to determine how close the object is to the eyepiece 16. Note that the infrared proximity sensor is just one example, and other sensors such as a capacitive sensor may be used in the eyepiece detection unit 57.

[0026] The feature point extraction unit 33 extracts (detects) feature points (point clouds) from the image data processed by the image processing unit 24. Information on the feature points extracted by the feature point extraction unit 33 for each frame is stored in the memory 32. Methods such as SIFT (Scale-Invariant Feature Transform), FAST (Features from accelerated segment test), and ORB (Oriented FAST and Rotated BRIEF) are used as feature point extraction methods. Information on the extracted feature points includes feature descriptors that compute binary vectors capable of low memory usage and high-speed matching, such as BRIEF (Binary Robust Independent Elementary Features).

[0027] The feature point tracking unit 34 reads the feature points (feature point information) stored in the memory 32 and compares them with the feature points extracted from the image data of the newly captured frame. In this way, feature point matching is performed between multiple images. As matching methods, for example, a brute-force matcher, a FLANN-based matcher, and a mean-shift search can be used.

[0028] The 3D shape reconstruction unit 35 reconstructs the 3D shape of a scene (imaging scene) from multiple captured image data. Methods for reconstructing the 3D shape include, for example, SfM (Structure from Motion) and SLAM (Simultaneous Localization and Mapping). Methods such as 3D SLAM (Simulation and Mapping) are used. In any of these methods, the feature point tracking unit 34 performs feature point matching (matching feature points between multiple images) between multiple images of the wearable device 110 with different poses and positions. The 3D shape reconstruction unit 35 then estimates the 3D position (3D point cloud data) of the matched feature points, the position of the wearable device 110, and the pose of the wearable device 110 by performing optimization, for example, by Bundle Adjustment. The 3D shape reconstruction unit 35 may improve the estimation accuracy of the 3D point cloud data, the position of the wearable device 110, and the pose of the wearable device 110 by using a method such as Visual Inertial SLAM, which uses the output of the pose detection unit 55 for optimization.

[0029] The 3D shape matching unit 36 ​​aligns the shape shown by the 3D point cloud (3D point cloud data) generated by the 3D shape reconstruction unit 35 with the shape of the target object shown by the 3D CAD data stored in an external database server or the like, and compares them. An ICP (Iterative Closet Point) algorithm or the like may be used for alignment. This makes it possible to compare which coordinate position each feature point extracted by the feature point extraction unit 33 corresponds to in the coordinate system of the 3D CAD data. Furthermore, the position and orientation of the wearable device 110 at the time of imaging can be estimated from the arrangement of the 3D point cloud and the target object.

[0030] The subject identification unit 167 analyzes the image data obtained by the imaging unit 22 and identifies the type of subject. The subject identification unit 167 can also determine the size of the subject in the image and its position in the image. The subject identification unit 167 performs the above processing by, for example, using a convolutional neural network, which is widely used in image recognition.

[0031] The eyeball detection unit 161 consists of an eyeball detection lens 162, an event sensor 163, and an event data calculation unit 164, which will be described later. The eyeball detection unit 161 is capable of acquiring eyeball information regarding the state of the user's eyes (right eye 170R and left eye 170L) as they look through the viewfinder.

[0032] The infrared light emitted from the infrared light-emitting diode 58 is reflected by the user's eye, and this reflected infrared light passes through the eyeball detection lens 162 and is imaged onto the imaging surface of the event sensor 163.

[0033] The event sensor 163 is an event-based vision sensor that detects changes in the brightness of light incident on each pixel and outputs information about the pixel where the brightness change occurred, asynchronously from other pixels. The data output from the event sensor 163 includes, for example, the position coordinates of the pixel where the brightness change (event) occurred, the polarity (positive or negative) of the brightness change, and timing information corresponding to the time the event occurred. This data will be referred to as event data from now on. Compared to a synchronous frame-based sensor such as the imaging unit 22, the event sensor 163 eliminates redundancy in the output information and features high-speed operation, high dynamic range, and low power consumption. On the other hand, since the event data (information about the pixel where the brightness change occurred) is output asynchronously from other pixels, special processing is required to determine the relationship between event data. In order to determine the relationship between event data, it is necessary to accumulate the event data output from the event sensor 163 at predetermined intervals and perform various calculations on the results.

[0034] The event data calculation unit 164 is a calculation unit for acquiring (detecting) eyeball information based on event data output continuously and asynchronously from the event sensor 163. For example, the event data calculation unit 164 acquires eyeball information by accumulating event data that occurs over a predetermined period of time and processing them as a set of data. By changing the accumulation time for accumulating event data, it is possible to acquire multiple eyeball information items with different occurrence speeds. Eyeball information includes, for example, gaze position (the position the user is looking at). The data includes positional information, saccade information regarding the direction and speed of saccades, and microsaccade information regarding the frequency and amplitude of microsaccades (changes in gaze position). Eyeball information may also include information on eye movements other than saccades and microsaccades, pupil information regarding pupil size and its changes, and blink information regarding blink speed and frequency. This information is merely an example, and eyeball information is not limited to this information. The event data calculation unit 164 may map the event data for the accumulated time as image data for one frame based on the event occurrence coordinates (position coordinates of the pixel where the brightness change (event) occurred), and perform image processing. With this configuration, eyeball information can be obtained from image data for one frame obtained by mapping the event data for the accumulated time using frame-based image processing.

[0035] The user state determination unit 165 is a determination unit that determines the user's state based on eyeball information obtained by the event data calculation unit 164. For example, it can determine the size (width) of the gaze range or the degree of gaze (overall view) from the frequency and amplitude of microsaccades. Here, gaze range is synonymous with attention range and focus range. The degree of gaze is an index that is higher the narrower the gaze range and lower the wider it is. Overall view is defined as the opposite of degree of gaze. In addition, it can determine the user's level of concentration (state of concentration) or fatigue from the frequency and amplitude of microsaccades, pupil size and change amount, and blinking speed and number. Furthermore, the user's level of excitement is related to the speed of microsaccades and pupil diameter, and can be determined from both parameters. The level of excitement is an index that is high when the user is looking at an object of high preference (such as a favorite face) and low when the user is looking at an object of low preference, so it can also be considered as preference. The user state determination unit 165 can be configured as a neural network, for example, which takes parameters such as eyeball information and the identification results of the subject identification unit 167 as inputs and outputs the user state information described above (hereinafter referred to as user state information). However, the configuration of the user state determination unit 165 is not limited to the above configuration. The eyeball information used by the user state determination unit 165 and the determination results of the user state determination unit 165 are not limited to those described above.

[0036] The eye-tracking input setting unit 166, via the system control unit 50, sets whether the processing of the eyeball detection unit 161 is enabled or disabled. The eye-tracking input setting unit 166 can also set parameters and conditions related to the processing of the event data calculation unit 164 and the user state determination unit 165. For example, the user can arbitrarily set these settings from a menu screen or the like.

[0037] Furthermore, the system control unit 50 can obtain information on which area of ​​the EVF 29 the captured subject (object) is displayed in and at what size. In addition, the eyeball detection unit 161 can obtain information on which area of ​​the EVF 29 the user is looking at. This allows the system control unit 50 to determine which area of ​​the subject the user is looking at.

[0038] Figure 3 is a flowchart illustrating an example of the operation of the wearable device 110. Each process in the flowchart of Figure 3 is realized by the system control unit 50 loading a program stored in the non-volatile memory 56 into the system memory 52 and executing it, thereby controlling each functional block. For example, when the wearable device 110 detects that a user is wearing the wearable device 110, it starts the operation shown in Figure 3.

[0039] In step S301, the system control unit 50 acquires the image (image data) captured by the imaging unit 22.

[0040] In step S302, the system control unit 50 displays the image acquired in step S301 on the EVF 29.

[0041] In step S303, the system control unit 50 acquires event data output from the event sensor 163.

[0042] In step S304, the system control unit 50 controls the event data calculation unit 164 to acquire eyeball information based on the event data acquired in step S303.

[0043] In step S305, the system control unit 50 controls the user state determination unit 165 to determine the degree of gaze based on the eyeball information acquired in step S304, and determines whether the degree of gaze is equal to or greater than a predetermined threshold TH. For example, the user state determination unit 165 determines the size of the gaze range from the frequency and amplitude of microsaccades, and determines that a smaller gaze range results in a higher degree of gaze. If the system control unit 50 determines that the degree of gaze is equal to or greater than the threshold TH, it proceeds to step S306, and if it determines that the degree of gaze is less than the threshold TH, it returns to step S301. The system control unit 50 may also determine whether the fixed gaze position time (time during which there is little or no movement of the gaze position) has exceeded a predetermined time. In that case, if the fixed gaze position time has not exceeded the predetermined time, it returns to step S301, and if the fixed gaze position time has exceeded the predetermined time, it proceeds to step S306.

[0044] In step S306, the system control unit 50 determines the gaze position that was determined to be above the threshold TH for high gaze intensity in step S305 as the gaze position.

[0045] In step S307, the system control unit 50 controls the user state determination unit 165 to estimate the gaze range based on the gaze position determined in step S306. As described above, the size of the gaze range can be estimated based on eyeball information. For example, the user state determination unit 165 estimates the gaze range as an area with the estimated size centered on the gaze position.

[0046] In step S308, the system control unit 50 detects one or more objects in the user's line of sight. For example, the system control unit 50 detects objects that overlap with the gaze range estimated in step S307. The system control unit 50 may also detect objects from the image by comparing the image acquired in step S301 with the gaze range estimated in step S307. The system control unit 50 may also control the 3D shape matching unit 36 ​​to detect objects from the 3D drawing by comparing the 3D drawing with the gaze range.

[0047] In steps S309 to S311, the system control unit 50 selects one object from the one or more objects detected in step S308, based on the gaze range estimated in step S307. If one object is detected in step S308, that object should be selected. If multiple objects are detected in step S308, the system control unit 50 compares the size of the gaze range with the size of each of the multiple objects and selects one object.

[0048] In step S309, the system control unit 50 obtains (calculates) the size of each object detected in step S308. The size of the object is, for example, the size on the image obtained in step S301.

[0049] In step S310, the system control unit 50 obtains (calculates) the ratio (size ratio) between the size of the gaze area and the size of each object acquired in step S309. The size ratio is, for example, the size of the gaze area and the size of each object on the image acquired in step S301. This is the ratio to the object's size. The size ratio may be the size of the gaze area divided by the object's size, or the size of the object divided by the size of the gaze area.

[0050] In step S311, the system control unit 50 determines (detects) the object whose size ratio, acquired in step S310, is closest to 1.

[0051] In step S312, the system control unit 50 selects the object determined in step S311.

[0052] In step S313, the system control unit 50 controls the system to perform processing based on the selection result from step S312. For example, the system control unit 50 displays information about the selected object on the EVF 29, or makes the area of ​​the selected object identifiable by highlighting or bordering it. It may also make the estimated gaze range identifiable.

[0053] Figures 4(a) to 4(d) illustrate specific examples of the operation of the wearable device 110. Figure 4(a) is a graph showing an example of a microsaccade waveform when the gaze range is relatively wide, and Figure 4(b) is a graph showing an example of a microsaccade waveform when the gaze range is relatively narrow. The vertical axis in Figures 4(a) and 4(b) represents the pupillary center position (the rotation angle of the eyeball around the center of the eyeball), and the horizontal axis in Figures 4(a) and 4(b) represents time. A microsaccade waveform is a waveform that shows the change in the pupillary center position when a microsaccade occurs. Parts 401 (Figure 4(a)) and 403 (Figure 4(b)) of the microsaccade waveform correspond to the timing when the microsaccade occurred. As shown in Figure 4(a), the wider the gaze range, the larger the amplitude of the microsaccade tends to be, and the wider the gaze range, the higher the oscillating nature (lower the attenuation rate) of the microsaccade tends to be. Furthermore, as evidenced by the six microsaccades observed during period 402, a wider gaze range tends to result in a higher frequency of microsaccades. Additionally, as shown in Figure 4(b), a narrower gaze range tends to result in smaller microsaccade amplitudes and lower oscillation characteristics (higher damping rates). Moreover, a narrower gaze range tends to result in fewer macrosaccades, as evidenced by the three microsaccades observed during period 404, which is the same length as period 402. Because of these tendencies (larger gaze ranges tend to result in larger microsaccades), the size (width) of the gaze range can be estimated from the amplitude and frequency of microsaccades.

[0054] Figure 4(c) shows an example of the EVF29 display when the gaze range is relatively wide. Image 405 is the image displayed in step S302 of Figure 3, and includes the automobile 410. Range 407 is the gaze range estimated in step S307, and the black circle located in the center of range 407 is the gaze position determined in step S306. In Figure 4(c), it is assumed that three objects were detected in step S309: the automobile 410, the front section 420, and the LED lamp 430. Since the object with the size ratio closest to 1 in relation to the gaze range is the front section 420, the front section 420 is selected in step S312. Therefore, in Figure 4(c), information 406 associated with the front section 420 (for example, the model and repair history of the front section 420) is displayed.

[0055] Figure 4(d) shows an example of the EVF29 display when the gaze range is relatively narrow. Range 409 is the gaze range estimated in step S307, and the black circle located in the center of range 409 is the gaze position determined in step S306. In Figure 4(d), as in Figure 4(c), it is assumed that three objects, the automobile 410, the front part 420, and the LED lamp 430, were detected in step S309. However, the object whose size ratio with respect to the gaze range is closest to 1 is the LED lamp 430, not the front part 420, therefore step S In step 312, the LED lamp 430 is selected. Therefore, in Figure 4(d), information 408 associated with the LED lamp 430 (for example, the model number and size of the LED lamp 430) is displayed.

[0056] Thus, when the user's gaze range is large, larger objects are selected than when it is small. In other words, when the gaze range is small, smaller objects are selected than when it is large. As mentioned above, the larger the gaze range, the larger the microsaccades tend to be. Therefore, when the user's microsaccades are large, larger objects are selected than when they are small. For example, when the amplitude of the microsaccades is large, larger objects are selected than when they are small. When the frequency of microsaccades is high, larger objects are selected than when they are low. When the attenuation rate of microsaccades is low (high oscillation), larger objects are selected than when it is high (low oscillation).

[0057] As described above, according to Example 1, one object is selected from one or more objects present in the user's line of sight based on the user's gaze range. This enables selection that matches the user's intent. For example, it becomes possible to estimate the user's intended work target and present only the information the user needs. Unnecessary information, such as information about objects that are not the work target, can be suppressed, thus reducing clutter on the display. On the other hand, since only information related to the intended work target can be displayed, the amount of information related to the work target can be increased. Ultimately, this improves the efficiency of information acquisition by the user (the efficiency with which the user obtains the information they need).

[0058] In Example 1, the system selected the object whose size ratio with respect to the gaze area was closest to 1. However, the present invention is not limited to this configuration. For example, a parameter other than the size ratio may be used to select the object whose size is closest to the size of the gaze area. Alternatively, the system may select an object whose size ratio with respect to the gaze area is within a predetermined range. If the size of the gaze area is larger than a predetermined threshold, the system can determine that the user is viewing the scene from a distance, and in such cases, the system may choose not to select an object. This further reduces the complexity of the display.

[0059] In Example 1, the user viewed an image displayed on the EVF 29, but the present invention is not limited to this configuration. For example, the present invention is also applicable to configurations in which the user directly views real space, such as a see-through head-mounted display. In that case, for example, the size ratio described above can be calculated as the ratio between the actual size of the object on the display surface (the size of the area projected onto the display surface) and the size of the gaze range on the display surface. Alternatively, a virtual surface may be set in real space, and the size ratio can be calculated as the ratio between the size of the object on the virtual surface (the size of the area projected onto the virtual surface) and the size of the gaze range on the virtual surface. These projection calculations are performed by the 3D shape matching unit 36.

[0060] In Example 1, the image captured by the imaging unit 22 was displayed on the EVF 29, but the present invention is not limited to this configuration. For example, the present invention can also be applied to head-mounted displays that do not have an imaging unit 22. Video (video data) stored on the recording medium 200 may be acquired and played back and displayed on the EVF 29.

[0061] In Example 1, the EVF29 was configured to display information about the selected object, but the present invention is not limited to this configuration. For example, the frame rate, brightness, contrast, and display color of the EVF29 may be changed according to the selected object. By setting display parameters according to the selected object, the selected object The visibility of objects and other elements is improved. Foveated rendering may be performed so that the resolution of the selected object is maintained at a high resolution, while the resolution decreases as the distance from the selected object increases. This reduces the amount of data transferred, the rendering process, and consequently, the power consumption. The imaging unit 22 may also be controlled according to the selected object. For example, by changing the frame rate (imaging rate) according to the movement of the selected object, it is possible to reduce power consumption when shooting still subjects. The parameters of AE processing, AF processing, and the image processing unit 24 may also be changed according to the selected object. By performing appropriate exposure control, focus control, and image processing according to the selected object, the visibility of the selected object and other elements is improved.

[0062] <Example 2> The following describes Embodiment 2 of the present invention. In Embodiment 2, components that play a common role with those in Embodiment 1 are given the same reference numerals as in Embodiment 1, and their descriptions are omitted. Components and configurations that differ from those in Embodiment 1 will be described in detail.

[0063] Figures 5(a) and 5(b) show the external appearance of a digital camera 100, an example of an electronic device to which the present invention can be applied. Figure 5(a) is a front perspective view of the digital camera 100, and Figure 5(b) is a rear perspective view of the digital camera 100.

[0064] The display unit 28 is a display unit located on the back of the digital camera 100 and displays images and various information. The touch panel 70a can detect touch operations on the display surface (operation surface) of the display unit 28. The viewfinder-external display unit 43 is a display unit located on the top surface of the digital camera 100 and displays various settings of the digital camera 100, including shutter speed and aperture. The shutter button 61 is an operating member for giving shooting instructions. The mode selector switch 60 is an operating member for switching between various modes. The terminal cover 40 is a cover that protects the connector (not shown) for connecting the digital camera 100 to an external device.

[0065] The main electronic dial 71 is a rotary control element, and by rotating it, settings such as shutter speed and aperture can be changed. The power switch 72 is an operating element that switches the power of the digital camera 100 ON and OFF. The sub electronic dial 73 is a rotary control element, and by rotating it, the selection frame (cursor) can be moved and images can be advanced. The four-way key 74 is configured so that the up, down, left, and right parts can be pressed, and processing can be performed according to the part of the four-way key 74 that is pressed. The SET button 75 is a push button and is mainly used to confirm selection items.

[0066] The video button 76 is used to instruct the start and stop of video recording. The AE lock button 77 is a push button, and by pressing the AE lock button 77 in shooting standby mode, the exposure state can be fixed. The zoom button 78 is an operation button for switching the zoom mode ON and OFF in the live view display (LV display) in shooting mode. By turning the zoom mode ON and then operating the main electronic dial 71, the live view image (LV image) can be enlarged or reduced. In playback mode, the zoom button 78 functions as an operation button for enlarging the playback image or increasing its magnification. The playback button 79 is an operation button for switching between shooting mode and playback mode. By pressing the playback button 79 in shooting mode, the camera switches to playback mode, and the latest image among the images recorded on the recording medium 200 can be displayed on the display unit 28. The menu button 80 is a push button used to instruct the display of the menu screen, and when the menu button 80 is pressed, a menu screen where various settings can be made is displayed on the display unit 28. Users can intuitively configure various settings using the menu screen displayed on the display unit 28, the four-way key 74, and the SET button 75.

[0067] The communication terminal 10 is a communication terminal for the digital camera 100 to communicate with the lens unit 150 (detachable), which will be described later. The eyepiece section 16 is the eyepiece of the eyepiece viewfinder (a type of viewfinder that you look through), and the user can view the image displayed on the internal EVF 29 through the eyepiece section 16. The eyepiece detection section 57 is an eyepiece detection sensor that detects whether or not the user (photographer) is looking through the eyepiece section 16. The cover 202 is the cover of the slot for storing the recording medium 200. The grip section 90 is a holding section shaped to be easy for the user to grip with their right hand when holding the digital camera 100. With the digital camera 100 held by gripping the grip section 90 with the little finger, ring finger, and middle finger of the right hand, the mode switch 60, shutter button 61, and main electronic dial 71 are positioned to be operated by the index finger of the right hand. Also, in the same position, the sub electronic dial 73 is positioned to be operated by the thumb of the right hand.

[0068] Figure 6 is a block diagram showing an example configuration of the digital camera 100. The lens unit 150 is a lens unit equipped with an interchangeable photographic lens. The lens 103 is usually composed of multiple lenses, but in Figure 6 it is simplified and shown as a single lens. Communication terminal 6 is a communication terminal for the lens unit 150 to communicate with the digital camera 100, and communication terminal 10 is a communication terminal for the digital camera 100 to communicate with the lens unit 150. The lens unit 150 communicates with the system control unit 50 via these communication terminals 6 and 10. The lens unit 150 controls the aperture 1 via the aperture drive circuit 2 using an internal lens system control circuit 4. The lens unit 150 also focuses by displacing the lens 103 via the AF drive circuit 3 using the lens system control circuit 4.

[0069] The shutter 101 is a focal-plane shutter that allows the exposure time of the imaging unit 22 to be freely controlled by the system control unit 50.

[0070] The external viewfinder display unit 43 displays various settings of the digital camera 100, including shutter speed and aperture, via the external viewfinder display unit drive circuit 44.

[0071] The communication unit 54 transmits and receives video and audio signals to and from external devices connected wirelessly or via wired cables. The communication unit 54 can also connect to wireless LAN (Local Area Network) and the internet. Furthermore, the communication unit 54 can communicate with external devices using Bluetooth® and Bluetooth Low Energy. The communication unit 54 can transmit images (including LV images) captured by the imaging unit 22 and images recorded on the recording medium 200, and can receive image data and other various information from external devices.

[0072] The eyepiece detection unit 57 is an eyepiece detection sensor that detects the approach (eye-to-eye contact) and departure (eye-away) of the eye 170 from the eyepiece 16 of the viewfinder (proximity detection). The system control unit 50 switches the display (display state) / hidden (hidden state) of the display unit 28 and EVF 29 according to the state detected by the eyepiece detection unit 57. More specifically, at least in the shooting standby state and when the display destination switching is automatic, when the eye is not focused, the display destination is set to the display unit 28 and the display is turned on, and the EVF 29 is hidden. When the eye is focused, the display destination is set to the EVF 29 and the display is turned on, and the display unit 28 is hidden.

[0073] The operation unit 70 is an input unit that receives user operations (user input) and is used to input various operation instructions to the system control unit 50. As shown in Figure 6, the operation unit 70 includes a mode switching switch 60, a shutter button 61, a touch panel 70a, and other operating members 70b. The other operating members 70b include the main electronic dial 71 This includes a power switch 72, a sub-electronic dial 73, a four-way key 74, a SET button 75, a video button 76, an AE lock button 77, a zoom button 78, a play button 79, a menu button 80, and more.

[0074] The mode switch 60 switches the operating mode of the system control unit 50 to one of the following: still image shooting mode, video shooting mode, playback mode, etc. Modes included in the still image shooting mode include auto shooting mode, auto scene detection mode, manual mode, aperture priority mode (Av mode), shutter speed priority mode (Tv mode), and program AE mode (P mode). There are also various scene modes and custom modes that provide shooting settings for different shooting scenes. The mode switch 60 allows the user to switch directly to any of these modes. Alternatively, the user may switch to a list screen of shooting modes using the mode switch 60, and then selectively switch to one of the displayed modes using another operating element. Similarly, the video shooting mode may also include multiple modes.

[0075] The shutter button 61 is equipped with a first shutter switch 62 and a second shutter switch 63. The first shutter switch 62 turns ON during the operation of the shutter button 61, so-called half-press (shooting preparation instruction), and generates a first shutter switch signal SW1. The system control unit 50 starts shooting preparation operations such as AF (autofocus) processing, AE (automatic exposure) processing, AWB (auto white balance) processing, and EF (flash pre-flash) processing based on the first shutter switch signal SW1. In AE processing, the system control unit 50 calculates and sets appropriate aperture value, shutter speed, and ISO sensitivity based on the difference between the exposure amount calculated based on the set aperture value, shutter speed, and ISO sensitivity and a predetermined appropriate exposure amount. The second shutter switch 63 turns ON when the operation of the shutter button 61 is completed, so-called full-press (shooting instruction), and generates a second shutter switch signal SW2. The system control unit 50 starts a series of shooting processes, from reading the signal from the imaging unit 22 to writing the captured image as an image file to the recording medium 200, based on the second shutter switch signal SW2.

[0076] The touch panel 70a and the display unit 28 can be configured as an integrated unit. For example, the touch panel 70a is configured such that its light transmittance does not interfere with the display of the display unit 28, and is mounted on the upper layer of the display surface of the display unit 28. Then, the input coordinates on the touch panel 70a are associated with the display coordinates on the display surface of the display unit 28. This makes it possible to provide a GUI (Graphical User Interface) that makes it seem as if the user can directly operate the screen displayed on the display unit 28.

[0077] Figure 7 is a flowchart illustrating an example of the operation of the digital camera 100. Each process in the flowchart of Figure 7 is realized by the system control unit 50 loading the program stored in the non-volatile memory 56 into the system memory 52 and executing it, thereby controlling each functional block. For example, when the digital camera 100 is started in still image shooting mode, it begins the operation shown in Figure 7.

[0078] The processing in steps S701 to S707 is the same as the processing in steps S301 to S307 in Figure 3.

[0079] In step S708, the system control unit 50 controls the subject identification unit 167 to analyze the image acquired in step S701, identifying the type of subject (object) and determining the size and position of the subject on the image. In other words, the system control unit 50 detects the subject from the image acquired in step S701.

[0080] In step S709, the system control unit 50 selects one or more subjects from the subjects detected in step S708 as AF candidates (candidate subjects for AF processing) and sets the area of ​​the AF candidates as an AF candidate area (candidate area for AF processing). For example, the system control unit 50 selects a specific type of subject as an AF candidate. The system control unit 50 may also select all subjects as AF candidates.

[0081] In step S710, the system control unit 50 selects an AF area (the area to be AF processed; selected area) from one or more AF candidate areas set in step S709, based on the gaze area estimated in step S707. For example, the system control unit 50 selects an AF candidate area that overlaps with the gaze area as the AF area. This process can also be considered as the process of selecting an AF target (the target to be AF processed) from one or more AF candidates. The AF area and gaze area may be made identifiable on the EVF 29 by highlighting or displaying a frame.

[0082] In step S711, the system control unit 50 determines whether or not it has detected the first shutter switch signal SW1, that is, whether or not the shutter button 61 has been half-pressed (instruction to prepare for shooting). If the system control unit 50 determines that it has detected the first shutter switch signal SW1, that is, that the shutter button 61 has been half-pressed (instruction to prepare for shooting), it proceeds to step S712. If the system control unit 50 determines that it has not detected the first shutter switch signal SW1, that is, that the shutter button 61 has not been half-pressed (instruction to prepare for shooting), it returns to step S701.

[0083] In step S712, the system control unit 50 calculates the control parameters for aperture 1 and the amount of drive for the focus lens included in lens 103 so that the subject (AF target) in the AF area selected in step S710 is within the depth of field of lens 103.

[0084] In step S713, the system control unit 50 sets appropriate AWB parameters and AE parameters such as ISO sensitivity and shutter speed for the subject (AF target) in the AF area selected in step S710. The system may also be configured to change the picture style setting or scene mode depending on the AF target.

[0085] In step S714, the system control unit 50 performs shooting preparation operations such as AF processing, AE processing, and AWB processing based on the various parameters set in steps S712 and S713. In this way, the system control unit 50 controls imaging parameters based on the AF area selected in step S710 in response to a half-press of the shutter button 61 (a predetermined user operation). For example, the system control unit 50 controls the focus position to focus on the AF target, controls the exposure so that the exposure amount in the AF area approaches a predetermined exposure amount, and controls the white balance so that the color tone in the AF area approaches a predetermined color tone. Note that the system control unit 50 may control imaging parameters based on the AF area even without a half-press of the shutter button 61. For example, the system control unit 50 may repeat the control of imaging parameters based on the AF area at a predetermined cycle.

[0086] In step S715, the system control unit 50 determines whether or not it has detected the second shutter switch signal SW2, that is, whether or not the shutter button 61 has been fully pressed (a shooting instruction has been given). If the system control unit 50 determines that it has detected the second shutter switch signal SW2, that is, that the shutter button 61 has been fully pressed (a shooting instruction has been given), it proceeds to step S719. If the system control unit 50 determines that it has not detected the second shutter switch signal SW2, that is, that the shutter button 61 has not been fully pressed (a shooting instruction has been given), it proceeds to step S716.

[0087] The processing in steps S716 and S717 is the same as in steps S701 and S702.

[0088] In step S718, the system control unit 50 controls the subject identification unit 167 to detect and track the subject (AF target) in the AF area selected in step S710. At this time as well, the AF area and the focus area may be made identifiable on the EVF 29 by highlighting or displaying a frame.

[0089] In step S719, the system control unit 50 completes a series of shooting operations (still image shooting) up to the point of writing the image captured by the imaging unit 22 as a still image file to the recording medium 200.

[0090] Figures 8(a) and 8(b) illustrate specific examples of the operation of the digital camera 100.

[0091] Figure 8(a) shows an example of the EVF29 display when the gaze range is relatively wide. Image 805 is the image displayed in step S702 of Figure 7, and includes multiple people at different positions in the depth direction. Range 807 is the gaze range estimated in step S707, and the black circle located in the center of range 807 is the gaze position determined in step S706. In Figure 8(a), it is assumed that in step S709, regions 810-812, which are all areas of people's faces, were set as AF candidate regions.

[0092] In the method of selecting the AF area based on a single point of gaze position, even if the user wants to select all of areas 810 to 812 as the AF area, they cannot do so. Also, even if the user wants to select area 810 as the AF area, area 811 overlaps with area 810 (area 811 is close to area 810), so unintended eye movements (such as fixation tremors) may cause area 811 to be selected as the AF area. In other words, it is difficult to make a selection that matches the user's intentions. In Example 2, by considering the gaze range, it becomes possible to make a selection that matches the user's intentions.

[0093] Since all of regions 810 to 812 overlap with the gaze range 807, in step S710, all of regions 810 to 812 are selected as the AF region. Therefore, in Figure 8(a), a solid line frame indicating regions 810 to 812 is displayed.

[0094] Figure 8(b) shows an example of the EVF29 display when the gaze range is relatively narrow. Range 809 is the gaze range estimated in step S707, and the black circle located in the center of range 809 is the gaze position determined in step S706. In Figure 8(b), as in Figure 8(a), it is assumed that in step S709, regions 810-812, which are both areas of a person's face, were set as AF candidate regions. However, since only region 810 (the face region of the person furthest back) overlaps with the gaze range 807, only region 810 is selected as the AF region in step S710. Therefore, in Figure 8(b), a solid line frame indicating region 810 is displayed. In Figure 8(b), dashed line frames indicating regions 811 and 812 are displayed, but such dashed line frames do not necessarily need to be displayed.

[0095] Thus, when the user's gaze range is large, more AF candidate areas (AF candidates) are selected as AF areas (AF targets) than when it is small. In other words, when the gaze range is small, fewer AF candidate areas are selected as AF areas than when it is large. As described in Example 1, the larger the gaze range, the larger the microsaccade tends to be. Therefore, when the user's microsaccade is large, more AF candidate areas are selected as AF areas than when it is small, and the focus position is controlled to focus on the subject in the selected AF area. A large microsaccade occurs when the amplitude of the microsaccade is large, when the frequency of microsaccade occurrence is high, or when the attenuation rate of the microsaccade is small (high vibration). A small microsaccade occurs when the microsaccade is small This can occur when the amplitude of the cade is small, when the frequency of microsaccade occurrence is low, or when the damping rate of the microsaccade is high (low vibration).

[0096] Furthermore, the AF area selection process may not be explicitly performed, and the focus position may be controlled to focus on more subjects when the user's microsaccades are large than when they are small. When the size of the microsaccades is a first size, the focus position is controlled to focus on multiple subjects. When the size of the microsaccades is a second size, which is smaller than the first size, the focus position is controlled to focus only on the furthest subject among the multiple subjects. For example, when the amplitude of the microsaccades is a first amplitude, the focus position is controlled to focus on multiple subjects. When the amplitude of the microsaccades is a second amplitude, which is smaller than the first amplitude, the focus position is controlled to focus only on the furthest subject among the multiple subjects. When the frequency of microsaccade occurrence is a first frequency, the focus position is controlled to focus on multiple subjects. When the frequency of microsaccade occurrence is a second frequency, which is less than the first frequency, the focus position is controlled to focus only on the furthest subject among the multiple subjects. When the microsaccade attenuation rate is a first value, the focal position is controlled to focus on multiple subjects. When the microsaccade attenuation rate is a second value greater than the first value, the focal position is controlled to focus only on the furthest subject among the multiple subjects.

[0097] As described above, according to Embodiment 2, a selection area (AF area) is selected from one or more candidate areas (AF candidate areas) based on the user's gaze range. This makes it possible to make a selection that matches the user's intent. For example, with selection using a pointer or selection based on gaze position, it was difficult to select multiple candidate areas as the selection area, but with the method described in Embodiment 2, multiple candidate areas can be easily selected as the selection area.

[0098] In Example 2, the user viewed the image displayed on the EVF29, but the present invention is not limited to this configuration. For example, the present invention can also be applied when the user looks through an optical viewfinder to view the subject. In that case, the subject (object) may be detected from the imaging range (predetermined range), and the area of ​​the detected subject may be set as a candidate area.

[0099] In Example 2, one or more subjects were selected as AF candidates from the detected subjects, and the area of ​​the AF candidates was set as the AF candidate area. However, the present invention is not limited to this configuration. The AF candidate area may be set in any way. For example, the AF candidate area may be a predetermined fixed area, or it may be an area specified by the user, such as in zone AF.

[0100] In Example 2, in step S710, the system is configured to select AF candidate areas that overlap the gaze area as AF areas. However, the definition of AF candidate areas that overlap the gaze area can be changed as appropriate. For example, the system may be configured to select AF candidate areas that overlap the gaze area even slightly as AF areas, or to not select AF candidate areas that only slightly overlap the gaze area as AF areas. The system may also be configured to select AF candidate areas where a predetermined percentage or more overlaps the gaze area as AF areas. If the size of the gaze area is larger than a predetermined threshold, the system can be configured not to select an AF area in such cases, as it can be determined that the user is viewing the scene from above.

[0101] In Example 2, after detecting the first shutter switch signal SW1 in step S711, the system is configured to always set various parameters based on the AF area selected in step S710. However, the present invention is not limited to this configuration. Even after detecting the first shutter switch signal SW1, the system can continue to estimate the gaze range and update the AF area. good.

[0102] In Example 2, the parameters for aperture, AF, AWB, and AE were set in steps S712 and S713, but the present invention is not limited to this configuration. Other parameters related to the control of the digital camera 100 may also be changed.

[0103] While Example 2 described the case of taking still images, the present invention is also applicable to video recording. In Example 2, the AF area was selected, but the present invention is not limited to this configuration. Based on the user's gaze range, a selection area can be made from one or more candidate areas, and the use of the selected area is not particularly limited.

[0104] In Example 2, the image captured by the imaging unit 22 is displayed on the EVF 29, but the present invention is not limited to this configuration. For example, the present invention can also be applied to electronic devices that do not have an imaging unit 22. Video (video data) stored in the recording medium 200 may be acquired and played back and displayed on the display unit.

[0105] In Example 2, the area of ​​a person's face was set as the AF candidate area, but the present invention is not limited to this configuration. For example, areas such as a person's eyes, torso, or entire body may be set as the AF candidate area. Areas such as the eyes, head, torso, or entire body of other animals may be set as the AF candidate area. Areas such as vehicles, plants, or food may be set as the AF candidate area.

[0106] Figures 9(a) to 9(c) show examples of the EVF29 display. In Figures 9(a) to 9(c), the region 901 of the person's face and the region 902 of the person's body are set as AF candidate regions. In Figure 9(a), both region 901 and region 902 overlap with the gaze range 903, so both region 901 and region 902 are selected as AF regions. In Figure 9(b), only region 901 overlaps with the gaze range 904, so only region 901 is selected as an AF region. In Figure 9(c), only region 902 overlaps with the gaze range 905, so only region 902 is selected as an AF region.

[0107] Figures 10(a) and 10(b) show examples of the EVF29 display. Here, we assume that AF candidate areas where a certain percentage or more overlaps with the gaze area are selected as AF areas. In Figures 10(a) and 10(b), the area 1001 of the person's pupil and the area 1002 of the person's face are set as AF candidate areas. In Figure 10(a), both area 1001 and area 1002 overlap with the gaze area 1003 by a certain percentage or more, so both area 1001 and area 1002 are selected as AF areas. In Figure 10(b), area 1002 overlaps with the gaze area 1003 only slightly (less than the predetermined percentage), and only area 1001 overlaps with the gaze area 1004 by a certain percentage or more. Therefore, only area 1001 is selected as an AF area.

[0108] Although two embodiments of the present invention have been described, the present invention is not limited to these specific embodiments, and various forms that do not depart from the spirit of the invention are also included. Furthermore, the two embodiments described above are merely examples, and it is possible to combine the two embodiments.

[0109] For example, the control of the system control unit 50 may be performed by a single piece of hardware, or multiple pieces of hardware (e.g., multiple processors or circuits) may share the processing to control the entire device.

[0110] In Examples 1 and 2, the gaze range was estimated using the event sensor 163. While frame-based sensors have the advantage of low power consumption and high latency compared to frame-based sensors, the present invention is not limited to this configuration. For example, a high-frame-rate frame-based sensor may be used if the power consumption and data transfer rate are within acceptable limits. An event-based sensor and a frame-based sensor may be used in combination. A method that detects potential fluctuations associated with eye movement from electrodes attached around the orbit, such as the EOG method, may also be used.

[0111] In Examples 1 and 2, the gaze range was estimated by focusing on microsaccade motion. While focusing on microsaccade motion allows for relatively fast estimation of the gaze range, the present invention is not limited to this configuration. The changes in gaze position may be recorded as historical information, and the gaze range may be estimated from the changes in gaze position over a predetermined time.

[0112] The present invention is applicable to any electronic device having the necessary configuration requirements. For example, the present invention may be applied to digital telescopes and digital microscopes. It may also be applied to display devices such as image viewers with cameras. It may also be applied to personal computers, PDAs, mobile phone terminals, game consoles, etc. By applying the present invention to any of these electronic devices, users can make choices that suit their intentions.

[0113] <Other examples> The present invention can also be realized by supplying a program that implements one or more of the functions of the above-described embodiments to a system or device via a network or storage medium, and by having one or more processors in the computer of that system or device read and execute the program. It can also be realized by a circuit (e.g., an ASIC) that implements one or more functions. [Explanation of symbols]

[0114] 110: Wearable devices 100: Digital Camera 50: System Control Unit

Claims

1. An acquisition means for acquiring the user's line of sight, An estimation means for estimating the size of the user's gaze range from the changes in the gaze position acquired by the acquisition means over a predetermined period of time, A detection means for detecting one or more objects located in the user's line of sight, An electronic device characterized by having, when multiple objects are detected by the detection means, a selection means that compares the size of the gaze range with the size of each of the multiple objects and selects one object from the multiple objects whose size ratio with the gaze range is closest to 1.

2. The selection means selects from the plurality of objects an object whose size ratio with respect to the gaze area is within a predetermined range. The electronic device according to feature 1.

3. The detection means detects objects that overlap the gaze range. The electronic device according to claim 1 or 2.

4. The system further includes a control means for controlling the processing to be performed based on the selection result of the selection means. The electronic device according to any one of claims 1 to 3.

5. The control means controls the display unit to display information about the object selected by the selection means. The electronic device according to feature 4.

6. The acquisition means acquires the line of sight position from the detection results of the event-based vision sensor. The electronic device according to any one of claims 1 to 5.

7. The selection means will not select an object if the size of the gaze area is greater than a predetermined threshold. The electronic device according to any one of claims 1 to 6.

8. A step of obtaining the user's line of sight, An estimation step which estimates the size of the user's gaze range from the changes in the gaze position acquired in the acquisition step over a predetermined time, A detection step of detecting one or more objects located in the user's line of sight, If multiple objects are detected by the detection step, a selection step is performed to compare the size of the gaze area with the size of each of the multiple objects and select one object from the multiple objects whose size ratio with the gaze area is closest to 1. A method for controlling electronic equipment, characterized by having the following features.

9. A program for causing a computer to function as one of the means of an electronic device according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a program for causing a computer to function as one of the means of an electronic device according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Picture signal processor

    JP1996009237A

  • Imaging apparatus and imaging method

    JP2013005018A

  • Selective enhancement of parts of the display based on eye tracking

    JP2015528120A

  • Drawing projection system, drawing projection method and program

    JP2018163466A

  • Gazing object estimation device, gazing object estimation method and program

    JP2019003312A