Information processing device, information processing method, and program

A dual-image acquisition system with parallel processing of wide-angle and ROI images efficiently detects eye positions, addressing the challenges of high-speed and high-precision eye tracking with reduced power consumption and load, and rapid recovery from eye loss.

WO2026083751A1PCT designated stage Publication Date: 2026-04-23SONY GROUP CORP
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
SONY GROUP CORP
Filing Date
2025-09-18
Publication Date
2026-04-23

AI Technical Summary

Technical Problem

Existing technologies face challenges in efficiently detecting the position of a user's eyes, particularly in applications requiring high-speed and high-precision eye tracking without increasing power consumption or calculation load.

Method used

A dual-image acquisition system is employed, capturing a wide-angle image at a lower frame rate and a region-of-interest (ROI) image at a higher frame rate, allowing for efficient eye position estimation by processing these images in parallel, with the ROI being dynamically adjusted based on the eye position.

Benefits of technology

This approach enables efficient, high-speed, and high-precision detection of eye positions, maintaining a wide detection range while reducing power consumption and processing load, and quickly recovering from eye loss without mode switching.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025032857_23042026_PF_FP_ABST
    Figure JP2025032857_23042026_PF_FP_ABST
Patent Text Reader

Abstract

An information processing device according to one aspect of the present invention comprises an acquisition unit and an estimation unit. The acquisition unit acquires a first image including the face of a user at first intervals, and acquires a second image, which is a part of the first image and includes the eyes of the user, at second intervals shorter than the first intervals. The estimation unit estimates real space eye positions, which are the positions of the eyes of the user in the real space, on the basis of the first image and the second image. In the information processing device, the first image including the face of the user is acquired, and the second image including the eyes of the user is acquired at shorter intervals than the first image. Furthermore, the positions of the eyes of the user are estimated on the basis of the first image and the second image. Therefore, it is possible to efficiently detect the positions of the eyes of the user.
Need to check novelty before this filing date? Find Prior Art

Description

Information Processing Apparatus, Information Processing Method, and Program

[0001] The present technology relates to an information processing apparatus, an information processing method, and a program applicable to detection of the position of eyes, etc.

[0002] In Patent Document 1, there is disclosed an imaging apparatus that exclusively divides pixel groups of an imaging element into a specific region including a person's eyes and mouth, a region including a face, and other regions, and sets a high frame rate for the specific region among these. With this imaging apparatus, it becomes possible to suppress power consumption and calculation load by the above-described processing.

[0003] Japanese Patent Application Laid-Open No. 2018-033192

[0004] Thus, there is a need for a technology that enables efficient detection of the position of a user's eyes.

[0005] In view of the above circumstances, an object of the present technology is to provide an information processing apparatus, an information processing method, and a program that enable efficient detection of the position of a user's eyes.

[0006] To achieve the above object, an information processing apparatus according to one embodiment of the present technology includes an acquisition unit and an estimation unit. The acquisition unit acquires a first image including a user's face at a first period, and acquires a second image that is a part of the first image and includes the user's eyes at a second period shorter than the first period. The estimation unit estimates a real-space eye position that is the position of the user's eyes in real space based on the first image and the second image.

[0007] In this information processing apparatus, a first image including a user's face is acquired, and a second image including the user's eyes is acquired at a period shorter than the first image. Further, based on the first image and the second image, the position of the user's eyes in real space is estimated. Thereby, it becomes possible to efficiently detect the position of the user's eyes.

[0008] The information processing apparatus may further include a target region setting unit that sets a region around the user's eyes as a target region. In this case, the acquisition unit may acquire an image of the target region as the second image.

[0009] The estimation unit may estimate the eye position in the second image. In this case, the target area setting unit may set the target area according to the estimation result of the eye position in the second image.

[0010] The estimation unit may estimate the eye position in the first image. In this case, if the user's eyes are not included in the second image, the target area setting unit may set the target area according to the estimation result of the eye position in the first image.

[0011] The estimation unit may estimate the real-space eye position based on the second image, and correct the estimation result of the real-space eye position with the estimation result of the real-space eye position based on the first image.

[0012] The estimation unit may calculate the amount of movement of the real-space eye position during the time required for estimating the real-space eye position based on the first image, and perform the correction including the amount of movement.

[0013] The estimation unit may estimate the user's facial posture based on the second image and estimate the real-space eye position based on the posture estimation result based on the second image. In this case, the user's facial posture may be estimated based on the first image, and the posture estimation result based on the second image may be corrected by the posture estimation result based on the first image.

[0014] The estimation unit may calculate the amount of change in posture during the time required for the posture estimation process based on the first image, and perform the correction including the amount of change.

[0015] The estimation unit may estimate the interocular distance of the user in the first image. In this case, if the second image includes only one of the user's eyes, the real-space eye position of the user's other eye may be estimated based on the posture and the interocular distance.

[0016] The estimation unit estimates the position of the user's face in the first image, which is the position of the user's face within the first image. In estimating the position of the face in the image, the unit may process only the area surrounding the target region within the first image.

[0017] The information processing device may further include a resolution reduction processing unit that performs resolution reduction processing on the image including the user's face. In this case, the acquisition unit may acquire the image after the resolution reduction processing as the first image.

[0018] The estimation unit does not need to perform any processing on the second image as long as the first image does not contain the user's face.

[0019] The acquisition unit may acquire two images captured simultaneously from different directions as the first image.

[0020] The estimation unit may estimate the user's gaze point based on the second image and the real-space eye position.

[0021] The estimation unit may estimate who the user is, the orientation of the user's face, the user's facial expression, or the user's gestures based on the first image.

[0022] An information processing method according to one embodiment of this technology includes acquiring a first image including the user's face in a first period. A second image, which is a part of the first image and includes the user's eyes, is acquired in a second period shorter than the first period. Based on the first image and the second image, the real-space eye position, which is the real-space position of the user's eyes, is estimated.

[0023] A program according to one embodiment of the present invention causes a computer system to perform the following steps: acquiring a first image including the user's face in a first period; acquiring a second image which is a part of the first image and includes the user's eyes in a second period shorter than the first period; and estimating the real-space eye position, which is the real-space position of the user's eyes, based on the first and second images.

[0024] This is a schematic diagram showing an overview of this technology. This is a schematic diagram showing an example configuration of an information processing system related to this technology. This is a flowchart relating to the processing of this embodiment. This is a schematic diagram showing the timing of output from the camera. This is a flowchart relating to the eye position feedback processing. This is a flowchart relating to the face posture feedback processing. This is a processing flowchart relating to delay compensation. This is a processing flowchart relating to inter-eye distance feedback. This is a schematic diagram showing a state in which only one of the user's eyes is included in the ROI. This is a processing flowchart relating to limiting the face detection range. This is a schematic diagram showing an example configuration relating to decimation processing. This is a schematic diagram showing the output timing from FPGA to information processing device. This is a schematic diagram showing an example configuration relating to a stereo camera. This is a processing flowchart relating to a stereo camera. This is a schematic diagram showing an example configuration relating to variations. This is a schematic diagram showing an example configuration relating to variations. This is a flowchart relating to gaze point estimation. This is a block diagram showing an example hardware configuration of a computer capable of realizing an information processing device, etc.

[0025] <First Embodiment> Hereinafter, embodiments relating to the present technology will be described with reference to the drawings.

[0026] [Overview of this Technology] This section provides an overview of this technology. This technology estimates the position of the user's eyes in real space. For example, in a glasses-free stereoscopic display system that provides stereoscopic vision without requiring the user to wear special glasses, it is necessary to change the display on the display unit according to the position of the user's eyes. This technology can be used in such a system. Of course, a normal image that is not stereoscopic may be displayed on the display unit. Furthermore, there are no limitations to the fields in which this technology can be applied.

[0027] Figure 1 is a schematic diagram illustrating the overview of this technology. In this technology, a camera captures two types of video: wide-angle image 1 and ROI (Region of Interest) image 2, each at a predetermined frame rate.

[0028] Wide-angle image 1 is an image of a fixed region whose position and shape do not change over time. Hereafter, this region will be referred to as the fixed region. The fixed region is a rectangular area and is basically set as a region in which the user's face 3 remains. In other words, wide-angle image 1 is basically an image that includes face 3.

[0029] For example, in the stereoscopic viewing example described above, the space opposite the display unit is set as a fixed area. As long as the user continues to view the display unit, face 3 will be facing the display unit, and therefore face 3 will remain in the fixed area. On the other hand, it is possible that, rarely, the user may leave their seat, resulting in face 3 no longer being in the fixed area, and face 3 not being included in the wide-angle image 1.

[0030] ROI image 2 is an image of a variable region whose position changes with time. Hereinafter, this region will be simply referred to as ROI. The ROI is a rectangular region located inside the fixed region. In this embodiment, the vertical width of the ROI is one-eighth of the fixed region, and the horizontal width is the same as the fixed region. The specific size of the ROI is not limited; for example, the horizontal width may be smaller than that of the fixed region.

[0031] The ROI is continuously updated to track the user's eye 4, and essentially the user's eye 4 remains within it. In other words, ROI image 2 is essentially an image that includes both eyes 4.

[0032] On the other hand, if the user leaves their seat as described above, or if eye 4 fails to track, it is possible that eye 4 will not exist in the ROI or ROI image 2 (hereinafter referred to as "lost").

[0033] The left-right direction in Figure 1 roughly corresponds to the time axis. The ROI image 2 is also schematically shown as a small rectangle. In this embodiment, the wide-angle image 1 is captured at a frame rate of 30 fps, and the ROI image 2 is captured at a frame rate of 240 fps. Therefore, while the wide-angle image 1 is captured once, the ROI image 2 is captured eight times.

[0034] In this technology, initially, the position of eye 4 is estimated based on wide-angle image 1, and an ROI is set to include eye 4. Subsequently, the real-space position of eye 4 is estimated based on ROI image 2, and the ROI is updated according to the estimation result.

[0035] However, if eye 4 suddenly moves significantly, it is possible that eye 4 may no longer be included in ROI image 2 (the "lost" portion in Figure 1). In that case, it becomes impossible to estimate the position of eye 4 and update the ROI. In such cases, the ROI is reset to include eye 4 based on wide-angle image 1. After that, it becomes possible to update the ROI and estimate the position again.

[0036] Wide-angle image 1 corresponds to one embodiment of the first image according to this technology. ROI image 2 corresponds to one embodiment of the second image according to this technology. The imaging frame rate of wide-angle image 1, 30 fps, corresponds to one embodiment of the first period according to this technology. The imaging frame rate of ROI image 2, 240 fps, corresponds to one embodiment of the second period according to this technology. The position of eye 4 in real space corresponds to one embodiment of the real-space eye position according to this technology. ROI corresponds to one embodiment of the target region according to this technology.

[0037] [Information Processing System] Figure 2 is a schematic diagram showing an example configuration of an information processing system 10 related to this technology. The information processing system 10 includes a camera 11, an information processing device 12, and a monitor 13. The camera 11 captures a wide-angle image 1 at 30 fps and an ROI image 2 at 240 fps. In addition, any combination of frame rates may be used such that the frame rate of the wide-angle image 1 is shorter than the frame rate of the ROI image 2. For example, the ratio of frame rates does not have to be 8 times, and the difference is not limited. The fact that the frame rate of the wide-angle image 1 is shorter than the frame rate of the ROI image 2 can also be said to mean that the imaging period of the wide-angle image 1 is longer than the imaging period of the ROI image 2.

[0038] The type of camera 11 can be arbitrary, and its specific configuration is not limited. The camera 11 outputs the wide-angle image 1 and the ROI image 2 as sensor outputs to the information processing device 12.

[0039] The information processing device 12 may be any computer, such as a PC (Personal Computer). The information processing device 12 has hardware circuits necessary for a computer, such as a CPU and memory (RAM, ROM). In this embodiment, the CPU executes a program related to this technology (for example, an application program), thereby realizing the image processing unit 14, control unit 15, memory unit 16, and display unit 17 as functional blocks. These functional blocks then execute the information processing method according to this embodiment. Dedicated hardware such as ICs (integrated circuits) may be used as appropriate to realize each functional block.

[0040] The image processing unit 14 includes an image separation unit 18, a face detection unit 19, a face part detection unit 20, a face pose estimation unit 21, an eye detection unit 22, an eye position estimation unit 23, a viewpoint image generation unit 24, and an ROI determination unit 25. The image separation unit 18 acquires a wide-angle image 1 and an ROI image 2, which are sensor outputs from the camera 11. These acquisitions are also performed at 30 fps and 240 fps. Furthermore, these images are separated, and the wide-angle image 1 is output to the face detection unit 19, and the ROI image 2 is output to the eye detection unit 22.

[0041] The face detection unit 19 detects the user's face 3 included in the wide-angle image 1. The face part detection unit 20 detects parts of the face 3 included in the wide-angle image 1. Any part included in the face 3, such as the user's eyes 4, nose, mouth, ears, etc., may be detected. Feature point detection and processing using a DNN (Deep Neural Network) may also be included.

[0042] The face posture estimation unit 21 estimates the posture (orientation) of the face 3. For example, the rotation of the face 3 relative to a predetermined reference posture is estimated using three values: roll, pitch, and yaw. In addition, the posture of the face 3 may be estimated in any other format. The eye detection unit 22 detects the user's eyes 4 included in the ROI image 2. The eye position estimation unit 23 estimates the three-dimensional position of the eyes 4 in real space.

[0043] The viewpoint image generation unit 24 generates a viewpoint image used for stereoscopic vision or the like according to the three-dimensional position of the eyes 4 estimated by the eye position estimation unit 23, and outputs it to the display unit 17. The ROI determination unit 25 sets, as an ROI, the area around the eyes 4 based on the three-dimensional position of the eyes 4 estimated by the eye position estimation unit 23.

[0044] In addition, the specific configuration of each functional block included in the image processing unit 14 is not limited, and may be appropriately changed within the range where the present technology can be realized. The image separation unit 18 corresponds to an embodiment of the acquisition unit according to the present technology. The face detection unit 19, the face part detection unit 20, the face pose estimation unit 21, the eye detection unit 22, and the eye position estimation unit 23 correspond to an embodiment of the estimation unit according to the present technology. The ROI determination unit 25 corresponds to an embodiment of the target area setting unit according to the present technology.

[0045] The control unit 15 controls the overall operation of the information processing device 12. In particular, the control unit 15 acquires the ROI position set by the ROI determination unit 25 and outputs it to the camera 11. Also, it outputs to the camera 11 the image size and frame rate related to the imaging of the wide-angle image 1 and the ROI image 2. In addition, it outputs basic settings related to imaging such as exposure time and gain.

[0046] The memory unit 16 is a storage device such as a non-volatile memory, and for example, an HDD (Hard Disk Drive), an SSD (Solid State Drive), or the like is used. In addition, any non-transitory computer-readable storage medium may be used. A control program for controlling the overall operation of the information processing device 12 is stored in the memory unit 16. Also, information such as various detection results is stored.

[0047] The display unit 17 controls the display of the viewpoint video generated by the viewpoint video generation unit 24 on the monitor 13. In a naked-eye stereoscopic display, an optical element such as a lenticular lens is attached to a normal display in order to display different videos for the left and right eyes, and it is necessary to display the viewpoint video at the pixel positions visible from the left and right eyes through the lens. Since it is necessary to distribute and display the video at the pixel level and high processing power is required, the display may be controlled by a GUI (Graphical User Interface) or the like.

[0048] The information processing apparatus 12 may be provided with an operation unit, a communication unit, a speaker, etc. not shown in the figure, and its specific configuration is not limited.

[0049] The monitor 13 is a display device using, for example, liquid crystal, EL (Electro-Luminescence), etc., or a naked-eye stereoscopic display using them, and displays an image to the user based on the display control by the display unit 17.

[0050] [Processing Flow] Fig. 3 is a flowchart related to the processing of this embodiment. Fig. 4 is a schematic diagram showing the timing of the output by the camera 11. Fig. 4 shows the timing when an image is output from the camera 11 (sensor) to the image separation unit 18. The horizontal direction in the figure corresponds to the time axis. The marks with patterns such as shading represent the wide-angle image 1, and the single-color marks represent the ROI image 2. The horizontal width of each mark represents the time required for the processing, but the processing time is not limited to that of this example.

[0051] As shown in Figure 3, in this embodiment, after the wide-angle image 1 and ROI image 2 are separated (step 001), the processing of sequence 1, indicated by the dashed frame on the left, is executed for the wide-angle image 1, and the processing of sequence 2, indicated by the dashed frame on the right, is executed for the ROI image 2. Processing of sequence 1 starts when the wide-angle image 1 is input to the face detection unit 19. Processing of sequence 2 starts when the ROI image 2 is input to the eye detection unit 22. Therefore, these processes start at the same cycle as the output from the camera 11. That is, processing of sequence 1 starts at a cycle of 30 fps, and processing of sequence 2 starts at a cycle of 240 fps. However, for the generation of viewpoint video (step 203), the frame rate may be less than 240 fps depending on the rendering time of the image processing. Also, display to the monitor (step 204) may be executed asynchronously with the cycle in which sequence 2 starts, such as at 60 fps or 120 fps.

[0052] In Figure 4, the output of wide-angle image A, etc., is performed approximately once every 33 ms, or at a period of 30 fps. Although the numbers are not shown in the diagram, the output of ROI images A1 to A8, etc., is performed approximately once every 4 ms, or at a period of 240 fps.

[0053] Although sequences 1 and 2 are executed in parallel, the time required for sequence 1 is generally longer than the time required for sequence 2. Therefore, it is possible that the processing of sequence 2 for multiple ROI images 2 may be completed before the processing of sequence 1 for one wide-angle image 1 is completed.

[0054] [(1) Initial Processing] The initial processing will be explained below. Initial processing refers to the time when the user's face 3, which had not been captured until immediately before, begins to be captured. For example, the first time the user's face 3 is captured after the information processing system 10 is started is considered initial processing. At this time, the ROI is set to a predetermined initial value. Also, even if the face 3 had been captured, if the face 3 moves out of the field of view of the wide-angle image 1 and then re-enters the field of view and begins to be captured again is also considered initial processing. At this time, the ROI is set to the value it was at when the face 3 moved out of the field of view.

[0055] Hereafter, these ROIs will be referred to as initial ROIs. Initial ROIs do not inherently include the current eye 4. Therefore, initially, a process is executed to set a new ROI that includes eye 4 based on wide-angle image 1.

[0056] Images are acquired and separated (step 001). A wide-angle image 1 is captured by the camera 11, and an ROI image 2 is captured under the initial ROI. The capture is performed under predetermined image size and frame rate (30 fps, 240 fps) acquired from the control unit 15. After the captured images are separated by the image separation unit 18, the wide-angle image 1 is input to the face detection unit 19, and the ROI image 2 is input to the eye detection unit 22.

[0057] The processing of Sequence 2 will now be explained. When the ROI image 2 is input to the eye detection unit 22, the processing of Sequence 2 begins. The eye detection unit 22 attempts to determine the two-dimensional position of the eye 4 within the ROI image 2 (step 201). However, since this ROI image 2 was captured under the initial ROI, the eye 4 is not included in the ROI image 2. Therefore, the eye detection unit 22 is unable to determine the position of the eye 4, and the processing of Sequence 2 ends at that point.

[0058] The processing of Sequence 1 will now be explained. When the wide-angle image 1 is input to the face detection unit 19, the processing of Sequence 1 begins. The face detection unit 19 identifies the position of the user's face 3 (step 101). Specifically, the range in which the face 3 exists within the wide-angle image 1 is identified in two-dimensional coordinates. Such identification can be performed, for example, by using face detection technology. The range in which the face 3 exists within the wide-angle image 1 corresponds to one embodiment of the in-image face position according to this technology.

[0059] The facial part detection unit 20 identifies the positions of the facial parts 3 (step 102). For example, the two-dimensional position of the eye 4 within the wide-angle image 1 is identified. The two-dimensional positions of other facial parts are similarly identified. These identifications are also achieved by using predetermined detection techniques.

[0060] The face posture estimation unit 21 estimates the posture of face 3 (step 103). The face posture estimation unit 21 estimates the posture of face 3 based on the range of face 3 identified by the face detection unit 19 and the positions of the face parts identified by the face part detection unit 20. For example, if the parts are shifted to one side within the range of face 3, the posture of face 3 is estimated based on the idea that face 3 is not facing forward, using predetermined estimation techniques.

[0061] The eye position estimation unit 23 estimates the three-dimensional position of the eye 4 in real space (step 104). The eye position estimation unit 23 estimates the three-dimensional position of the eye 4 based on the two-dimensional position of the eye 4 identified by the face part detection unit 20, the interocular distance between the left and right eyes, and the posture of the face 3 estimated by the face posture estimation unit 21.

[0062] In sequence 2, it is determined whether or not the position of eye 4 has been identified (step 105). Specifically, in step 201, it is determined whether or not the two-dimensional position of eye 4 in ROI image 2 has been identified. As described above, initially, the position of eye 4 is not identified in step 201, so the determination is No.

[0063] As shown in Figure 3, the processing in step 105 may start immediately after step 102. In other words, the processing in step 105 may be executed in parallel with the processing in steps 103 and 104.

[0064] The ROI determination unit 25 determines the position of the ROI (step 106). Specifically, the ROI determination unit 25 sets the ROI according to the two-dimensional position of the eye 4 detected by the face part detection unit 20 in step 102. The ROI is set as the area around the eye 4. For example, the ROI is set so that the area between the eyebrows is located in the center of the ROI, but the specific setting criteria are not limited. The set ROI is transmitted to the camera 11, and the acquisition of the next ROI image 2 is performed under that ROI.

[0065] In the initial stages, the position of eye 4 cannot be determined by the processing in step 201, so the system may be configured in advance to not perform the processing in sequence 2. In other words, as long as the user's face 3 is not included in the wide-angle image 1, processing of the ROI image 2 does not need to be performed. This reduces the processing load in the initial stages and lowers power consumption.

[0066] [(2) Normal Processing] The normal processing, other than the initial and lost state, will be described below. A wide-angle image 1 is captured by the camera 11, and an ROI image 2 is captured under the currently set ROI. The captured images are separated by the image separation unit 18, and the wide-angle image 1 is input to the face detection unit 19, and the ROI image 2 is input to the eye detection unit 22 (step 001).

[0067] The processing steps 101 to 104 of Sequence 1 are the same as described above, so the explanation is omitted. In step 105, it is determined whether or not eye 4 has been identified in Sequence 2. Under normal circumstances, i.e., when there is no loss, eye 4 is included in ROI image 2, so the determination is Yes. Therefore, ROI determination (step 106) is not performed, and the processing related to Sequence 1 ends without the ROI being changed.

[0068] In sequence 2, the eye detection unit 22 first identifies the position of the eye 4 (step 201). Specifically, the two-dimensional position of the eye 4 within the ROI image 2 is identified. Since the eye 4 is normally included in the ROI image 2, it is possible to identify the two-dimensional position of the eye 4.

[0069] The eye position estimation unit 23 estimates the three-dimensional position of the eye 4 in real space (step 202). The eye position estimation unit 23 estimates the three-dimensional position of the eye 4 by calculation or other means based on the two-dimensional position of the eye 4 identified by the eye detection unit and the coordinates of the ROI, etc.

[0070] The viewpoint image generation unit 24 generates a viewpoint image based on the three-dimensional position of the eye 4 (step 203). The display unit 17 controls the display of the viewpoint image and the like on the monitor 13 (step 204).

[0071] It is determined whether the position of eye 4 has moved (step 205). Under normal circumstances, it is assumed that no loss occurs, but the position of eye 4 can move. That is, the two-dimensional position of eye 4 identified in the current step 201 may have changed within the ROI range relative to the two-dimensional position of eye 4 identified in the previous step 201. In such cases, the determination is Yes.

[0072] Note that the processing in step 205 may start immediately after step 201. In other words, the processing in step 205 may be executed in parallel with the processing in steps 202 to 204.

[0073] If the position of eye 4 has moved (Yes in step 205), the ROI determination unit 25 determines a new ROI position (step 206). Specifically, the ROI determination unit 25 sets the ROI based on the two-dimensional position of eye 4 detected by the eye detection unit 22. The ROI is set as the area around the moved eye 4. For example, the ROI is newly set so that the area between the eyebrows is located in the center of the ROI, but the specific setting criteria are not limited. The set ROI is transmitted to the camera 11, and the acquisition of the next ROI image 2 is performed under the new ROI.

[0074] If the position of eye 4 has not moved (No. in step 205), the ROI is not updated, and the next ROI image 2 is acquired under the existing ROI. In this way, under normal circumstances, the ROI changes in accordance with the change in the position of eye 4, and the processing continues.

[0075] [(3) Processing in case of loss] The processing in case of loss will be explained below. When loss occurs, eye 4 is not included in ROI image 2, so the position of eye 4 cannot be identified in step 201, and the processing of sequence 2 is interrupted.

[0076] On the other hand, in sequence 1, the result is determined to be No in step 105, and in step 106, the ROI determination unit 25 sets the ROI according to the two-dimensional position of the eye 4 detected by the face part detection unit 20 in step 102. The ROI is set as the area around the eye 4. For example, (1) Similar to the initial processing, the ROI is reset so that the area between the eyebrows is located in the center of the ROI. The reset ROI is transmitted to the camera 11, and the acquisition of the next ROI image 2 is performed under that ROI.

[0077] During this period, from the start of the loss until the ROI is determined in step 106, it becomes impossible to estimate the three-dimensional position of eye 4 based on ROI image 2 (step 202).

[0078] In such cases, the most recent ROI determined in step 106 is retrieved. If eye 4 is included in the ROI, it is set as the current ROI, and the process in step 202 resumes as usual. On the other hand, if eye 4 is not included in the ROI, the ROI will be set the next time step 106 is performed. Until then, it will not be possible to estimate the three-dimensional position using step 202, but the most recent estimation result from step 104 or step 202 will be repeatedly used as a substitute for the original estimation result.

[0079] In the information processing device 12 according to this embodiment, a wide-angle image 1 including the user's face 3 is acquired, and an ROI image 2 including the user's eyes 4 is acquired at a shorter interval than the wide-angle image 1. Furthermore, the position of the user's eyes 4 in real space is estimated based on the wide-angle image 1 and the ROI image 2. This makes it possible to efficiently detect the position of the user's eyes 4.

[0080] For example, in eye-tracking type autostereoscopic displays and LFDs (Light Field Displays), high-speed and high-precision eye-sensing technology is required to improve the sense of realism. One way to achieve high precision is to acquire high-resolution video, but this increases the amount of data, lowering the frame rate and increasing latency. As a result, high speed is lost.

[0081] In this technology, a wide-angle image 1 is acquired over a wide area to detect the position and state of face 3. Additionally, a region of interest (ROI) image 2, which tracks the area near eye 4, is acquired at a higher frame rate than the wide-angle image 1 to detect the position of eye 4. These acquisitions are performed in parallel, not by switching between them. This ensures a wide detection range while enabling high-speed and high-precision detection of the eye 4's position.

[0082] Furthermore, in this technology, the region surrounding eye 4 is set as an ROI, and an image of the ROI is acquired as ROI image 2. In addition, the three-dimensional position of eye 4 is estimated based on ROI image 2, and the ROI is set according to the estimation result. This makes it possible to estimate the position of eye 4 even more efficiently.

[0083] Furthermore, in this technology, the three-dimensional position of eye 4 is estimated based on the wide-angle image 1, and if eye 4 is no longer included in the ROI image 2, the ROI is reset based on the estimation result. This makes it possible to quickly recover eye 4 when it is lost (lost) without switching camera modes.

[0084] <Second Embodiment> A more detailed embodiment of the information processing system 10 according to this technology will be described as a second embodiment. In the following description, parts that are the same as the configuration and operation of the information processing system 10 described in the above embodiment will be omitted or simplified.

[0085] [Eye Position Feedback] Figure 5 is a flowchart relating to the feedback process for the position of the eye 4. In this embodiment, the eye position estimation unit 23 estimates the three-dimensional position of the eye 4 based on the ROI image 2, and this estimation result is corrected by the estimation result of the three-dimensional position of the eye 4 based on the wide-angle image 1.

[0086] Specifically, as shown by the thick arrows, once the three-dimensional position of eye 4 is estimated based on the wide-angle image 1 (step 104), the three-dimensional position of eye 4 estimated based on the ROI image 2 (step 202) is corrected.

[0087] For example, in Figure 4, if the estimation results are obtained in the order "ROI image A2 → wide-angle image A → ROI image A3", the estimation result based on ROI image A3 will be corrected by the estimation result based on wide-angle image A. The correction is performed, for example, by simply overwriting the estimation result based on ROI image A3 with the estimation result based on wide-angle image A. The content of the correction can be arbitrary. For example, a correction method may be used in which the estimation result based on ROI image A3 remains to some extent in the corrected estimation result.

[0088] As a result, the estimation result based on wide-angle image A is output as the estimation result based on ROI image A3 (corrected estimation result) (step 202). Furthermore, the ROI is set based on this corrected estimation result (step 206). Since the estimation result based on wide-angle image 1 is based on estimation including pose estimation, it is generally more accurate than the estimation result based on ROI image 2. Therefore, the error in the three-dimensional position and ROI of eye 4 is reduced in the corrected estimation result (A3), and the correction effect extends to subsequent estimation results (A4, A5, ...). This makes it possible to reduce the errors accumulated by the estimation based on ROI image 2.

[0089] <Third Embodiment> [Facial Posture Feedback] Figure 6 is a flowchart relating to the feedback process for the posture of the face 3. In sequence 2, the posture of the user's face 3 is estimated by the face posture estimation unit 21 (step 201-1). This estimation is performed based on the ROI image 2 and the two-dimensional position of the eye 4 identified by the eye detection unit 22. Furthermore, the three-dimensional position of the eye 4 is estimated based on the posture estimation result (step 202).

[0090] In sequence 1, the face pose estimation unit 21 estimates the pose of face 3 based on the wide-angle image 1, etc. (step 103). When this estimation is completed, the pose estimation result based on the ROI image 2, etc. is corrected by the pose estimation result based on the wide-angle image 1, etc. The correction may be by any method, for example, simply overwriting the pose result based on the ROI image 2, etc. with the pose estimation result based on the wide-angle image 1, etc.

[0091] In this embodiment, the three-dimensional position of the eye 4 is estimated based on the estimation result corrected based on the posture estimation result based on the wide-angle image 1, etc., making it possible to perform estimation with even greater accuracy. In this example, the three-dimensional position of the eye 4 is also corrected as in Figure 5, but both posture and three-dimensional position correction may be performed, or only posture correction may be performed.

[0092] <Fourth Embodiment> [Delay Compensation] Figure 7 is a flowchart of the processing related to delay compensation. In this embodiment, correction of three-dimensional position and orientation, including delay compensation, is performed.

[0093] In the three-dimensional position correction, the amount of movement of the three-dimensional position of eye 4 within the time required for the three-dimensional position estimation process based on wide-angle image 1 (steps 101 to 104) is calculated, and a correction including this amount of movement is performed. Specifically, for example, the time required for this process is known in advance, and the movement speed of eye 4 can be estimated from the most recent estimation result of the three-dimensional position of eye 4. Therefore, by taking the product of these, it is possible to calculate the amount of movement within the processing time.

[0094] Furthermore, by reflecting this amount of movement in the three-dimensional position estimation result based on wide-angle image 1, it is possible to perform correction including delay compensation. This makes it possible to correct the three-dimensional position with high accuracy.

[0095] In posture correction, the amount of change in face 3's posture during the time required for posture estimation processing based on wide-angle image 1 (steps 101-103) is calculated, and correction including this change is performed. Specifically, the amount of change during processing time can be calculated by taking the product of the processing time and the speed of posture change estimated from the most recent posture estimation result.

[0096] Furthermore, by reflecting this change in the pose estimation result based on ROI image 2, it is possible to perform correction including delay compensation. This makes it possible to correct the pose with high accuracy.

[0097] In this example, corrections including delay compensation are performed for both the three-dimensional position and orientation. However, for each of the three-dimensional position and orientation, any combination of corrections including delay compensation, corrections without delay compensation, and no corrections may be adopted. This also applies to the embodiments described later. For example, one possible combination is to perform corrections including delay compensation for the three-dimensional position and corrections without delay compensation for the orientation.

[0098] <Fifth Embodiment> [Interocular Distance Feedback] Figure 8 is a processing flowchart related to interocular distance feedback. Figure 9 is a schematic diagram showing a state in which only one of the user's eyes is included in the ROI. In this embodiment, if one eye is not identified, a process is executed to estimate the position of the other eye 4 based on the interocular distance.

[0099] In Figure 9A, the user is tilted relative to the fixed area of ​​wide-angle image 1, and the vertical positions of the right eye (left eye 4 in the figure) and left eye (right eye 4 in the figure) are different. As a result, as shown in Figure 9B, only the right eye is included in the ROI (Region of Interest) of ROI image 2, while the left eye is not. Figure 8 shows an example of processing when only one eye is included in the ROI in this way.

[0100] In step 102, the face part detection unit 20 identifies the parts of the face 3 and simultaneously estimates the user's interpupillary distance (IPD). The interpupillary distance is estimated as the distance between the identified right and left eyes.

[0101] In step 201-2, the eye detection unit 22 identifies the two-dimensional position of one of the eyes 4 included in the ROI. In step 201-3, the face pose estimation unit 21 estimates the pose of the face 3 based on the two-dimensional position of the eye 4, similar to step 201-1. At this time, correction may be made based on the pose estimated in step 103.

[0102] Furthermore, the face posture estimation unit 21 estimates the three-dimensional position of the other eye 4 that is not included in the ROI image 2, based on the estimated posture and interocular distance information. Specifically, for example, it calculates the inclination of the line segment connecting both eyes with respect to the reference direction based on the posture, and calculates the difference in the three-dimensional positions of both eyes based on this inclination and interocular distance. Then, by adding this difference to the three-dimensional position of one eye 4, it is possible to estimate the three-dimensional position of the other eye 4. In addition, the three-dimensional position of the eye 4 may be estimated by any other method based on posture and interocular distance.

[0103] This makes it possible to estimate the three-dimensional position of eye 4 even when face 3 is tilted and one eye is no longer included in the ROI. It also makes it possible to narrow the ROI further. If both eyes are included in the ROI, the normal processing shown in Figure 3, etc., is performed.

[0104] <Sixth Embodiment> [Limitation of Face Detection Range] Figure 10 is a flowchart of the process related to limiting the detection range of face 3. In this embodiment, the face detection unit 19 limits the detection range based on the set ROI.

[0105] In step 106 or 206, if the ROI determination unit 25 determines an ROI, the ROI is transmitted to the face detection unit 19. In step 101, the face detection unit 19 estimates the range in the wide-angle image 1 where the face 3 exists, but in this estimation, it does not process the entire wide-angle image 1, but only the area around the transmitted ROI. The area around the ROI is, for example, the range of points whose distance from the ROI is less than a predetermined value, but the specific range of the area around the ROI is not limited.

[0106] Since the ROI includes the eye 4, it is unlikely that the face 3 will be included in areas far from the ROI. Therefore, by limiting the detection range of the face 3 as in this embodiment, it is possible to improve the processing speed.

[0107] <Seventh Embodiment> [Decline Processing] Figure 11 is a schematic diagram showing an example configuration related to decimation processing. In this embodiment, the information processing system 10 further has an FPGA (Field Programmable Gate Array) 28. The FPGA 28 has a signal processing unit 29 as a functional block, and also has a memory unit 30 and a control unit 31 similar to those in the information processing device 12. The signal processing unit 29 includes an image separation unit 32, an image decimation unit 33, and a format conversion unit 34.

[0108] FPGA 28 is included in the information processing device related to this technology. However, other devices or mechanisms separate from the information processing device 12 may be used instead of FPGA 28.

[0109] The image separation unit 32 separates the captured image, which is the sensor output, into a wide-angle image 1 and an ROI image 2, similar to the image separation unit 18 shown in Figure 2. The image decimation unit 33 performs a decimation process on the wide-angle image 1 to reduce its resolution. For example, a process is performed to compress the wide-angle image 1 to one-quarter of its pixels. The extent to which the wide-angle image 1 is decimated is not limited. The image decimation unit 33 corresponds to one embodiment of the decimation processing unit according to this technology.

[0110] The format conversion unit 34 performs processes on each of the wide-angle image 1 and ROI image 2, such as conversion to a monochrome image, or conversion of the camera signal, which is MIPI (Mobile Industry Processor Interface), into a video signal. The specific content of the conversions is not limited to these processes.

[0111] The signal processing unit 29 performs other processing, such as labeling and merging images near the eyes 4. The face detection unit 19 of the information processing device 12 acquires the converted wide-angle image 1, and the eye detection unit 22 acquires the converted ROI image 2. After that, processing similar to that shown in Figure 2 is performed.

[0112] Figure 12 is a schematic diagram showing the output timing from FPGA 28 to information processing device 12. "Sensor output" represents the output timing from camera 11 to FPGA 28. "Raw Scalar output" represents the output timing of the wide-angle image 1 after downsampling by the image downsampling unit 33. Because processing by the image separation unit 32 and the pixel downsampling unit 33 takes time, the timing of the Raw Scalar output is slightly delayed compared to the timing of the sensor output.

[0113] "EOF (End Of Frame)" is a signal that indicates the timing when the Raw Scalar output has ended. "TG (Timing Generator)" is a signal that indicates the timing when the sensor output has started. The TG corresponding to wide-angle image 1 is output at 30 fps, and the TG corresponding to ROI image 2 is output at 240 fps.

[0114] "USB output" refers to the timing of the USB output from FPGA 28 to information processing device 12. Because the format conversion unit 34 takes time to process, the USB output timing for wide-angle image 1 is slightly delayed compared to the Raw Scalar output timing, and the USB output timing for ROI image 2 is slightly delayed compared to the sensor output timing.

[0115] In this embodiment, as shown in the figure, the wide-angle image 1 is divided and output in segments so that the wide-angle image 1 and the ROI image 2 are output as a single combined image. Of course, the specific output method for the wide-angle image 1 is not limited.

[0116] In this example, sensor output is not being provided for parts of ROI image 2 corresponding to C1 to C3. In this embodiment, imaging is performed using a rolling shutter method, so such frame drops can occur if the position of the ROI changes in the opposite direction to the scanning direction. In this case, for example, as shown by the thin arrow in the figure, the latest output based on ROI image 2 is repeatedly used. That is, in this example, the output based on B8 is repeatedly used instead of the output based on ROI images C1 to C3.

[0117] In this embodiment, the output from the camera 11 is first received by the signal processing unit 29 of the FPGA 28, and then pixels of the wide-angle image 1 are downsampled before being transmitted to the image processing unit 14 of the information processing device 12. Therefore, the amount of data transferred to the information processing device 12 is reduced, and the transfer speed can be improved. In addition, since processing is performed on the downsampled wide-angle image 1 in sequence 1, the computational load and processing delay in sequence 1 can be reduced.

[0118] <Eighth Embodiment> [Stereo Camera] Figure 13 is a schematic diagram showing an example configuration of a stereo camera. In this embodiment, wide-angle images 1 and ROI images 2 are captured by two cameras 11a and 11b, respectively. Cameras 11a and 11b simultaneously capture the user's face 3 from different directions. That is, cameras 11a and 11b are synchronously controlled so that the face 3 is captured from slightly different angles of view. The image separation unit 18 acquires two wide-angle images 1 at the same time. Also, two ROI images 2 are acquired at the same time.

[0119] Figure 14 is a processing flowchart related to the stereo camera. Separation is performed by the image separation unit 18 (step 001), and in sequence 1, two wide-angle images 1 are processed, and in sequence 2, two ROI images 2 are processed. The subsequent processing is basically the same as in Figure 3, etc., but the difference is that stereo ranging is used to detect the three-dimensional position of the eyes, so the pose estimation process of the face 3 (step 103) is not required.

[0120] In this embodiment, stereo distance measurement is performed, which further improves the accuracy of detecting the three-dimensional position of the eye 4. Alternatively, only one of the cameras 11a and 11b may have an ROI function, resulting in two wide-angle images 1 being captured and only one ROI image 2 being captured.

[0121] <Ninth Embodiment> [Variations of Configuration Examples] Figures 15 to 17 are schematic diagrams showing configuration examples related to variations. In Figure 15, of the functional blocks of the image processing unit 14 in Figure 2, only the viewpoint image generation unit 24 is configured within the information processing device 12, while the other blocks are configured within the FPGA 28. In this embodiment, the FPGA 28 performs the estimation of the three-dimensional position of the eye 4, and the output unit 37 outputs the three-dimensional position to the information processing device 12. Subsequently, the information processing device 12 generates a viewpoint image, which is output to the monitor 13.

[0122] This means that the data transferred to the information processing device 12 consists only of three-dimensional position, making it possible to reduce the transfer delay. Note that some of the processing other than viewpoint image generation may be handled by the information processing device 12. For example, in addition to the viewpoint image generation unit 24, a face detection unit 19 and an eye detection unit 22 may be configured within the information processing device 12.

[0123] In Figure 16, the camera 11 contains an imaging unit 40 and an image processing unit 14. The imaging unit 40 performs imaging, and the image processing unit 14 contains functional blocks other than the viewpoint image generation unit 24. In this way, the camera 11 may be configured to have most of the functions included in the information processing system 10. For example, such an embodiment can be applied to an imager with AI functionality.

[0124] In Figure 17, the information processing system 10 includes a cloud 43, which has a viewpoint image generation function 24a. The viewpoint image generation function 24a includes an image processing function 14c. The FPGA 28 performs the process of estimating the three-dimensional position of the eye 4, and the output unit 37 outputs the three-dimensional position to the memory unit 16b of the information processing device 12.

[0125] The information processing device 12 transmits the three-dimensional position to the cloud 43 via the communication unit 44. On the cloud 43, the viewpoint image generation function 24a generates a viewpoint image. The information processing device 12 receives the viewpoint image via the communication unit 44 and controls the display of the viewpoint image on the monitor 13 using the display unit 17. In this embodiment, since the information processing device 12 does not perform the generation of the viewpoint image, etc., the processing load on the information processing device 12 can be reduced. In addition to the configurations shown in Figures 15 to 17, any other configuration may be used as appropriate.

[0126] <Tenth Embodiment> [Gaze Point Estimation] Figure 18 is a flowchart of the process related to gaze point estimation. In this embodiment, gaze point estimation is performed. The term gaze point refers to the point that the user is looking at with their eye 4. Gaze point estimation is sometimes referred to as gaze sensing. In step 201, for example, a face image included in the wide-angle image 1 is transmitted to a gaze point estimation unit (not shown). In step 203-1, the gaze point is estimated based on the face image, the eye image included in the ROI image 2, and the three-dimensional position of the eye 4, etc.

[0127] For example, a separate infrared emitter is provided, and infrared light is emitted towards the user's eye 4. The point of fixation is then estimated based on the positional relationship between the reflected image from the cornea and the pupil. This method is sometimes called the corneal reflection method. Alternatively, the point of fixation is estimated using an image-based method. In this case, feature points such as the center of the pupil and the corners of the eye 4 are extracted. In these methods, the three-dimensional position of the eye 4 is also used for estimation.

[0128] This enables high-speed detection of the user's gaze point. Other detection algorithms, such as CNNs (Convolutional Neural Networks), may also be used, and the gaze point may be estimated by any method. Furthermore, the generation and output of gaze imagery may occur simultaneously with the detection of the gaze point.

[0129] <Other Embodiments> This technology is not limited to the embodiments described above, and various other embodiments can be realized.

[0130] Simultaneously with the detection of the three-dimensional position of the eye 4, estimations of who the user is, the orientation of the user's face 3, the user's facial expression, the user's gestures, etc., may be performed based on the facial image. These estimations are performed, for example, by a user information estimation unit (not shown).

[0131] The specific means for estimating this information based on facial images are not limited. Nor are the methods for using the estimated information. If the user's identity is estimated, information about the user, such as their name, gender, affiliation, and address, may also be obtained at the same time. Regarding the user's facial expressions, for example, the type of expression (e.g., smiling, crying) may be estimated. Regarding the user's gestures, for example, hand signs may be estimated. These advancements will enable the use of this technology in an even wider range of applications.

[0132] The three-dimensional position of the eye 4 estimated in this technology may be used to generate 2D or 3D viewpoint images, as shown in Figure 3, or for any other arbitrary purpose. Furthermore, the information processing device 12 may be configured as a sensor module that only performs the estimation and output of the three-dimensional position of the eye 4.

[0133] In this technology, the ROI image 2 may include the entire face 3. That is, obtaining an ROI image 2 that includes face 3 is equivalent to obtaining an ROI image 2 that includes eyes 4.

[0134] The information processing system 10 may be implemented by multiple computers or by a single computer.

[0135] Figure 19 is a block diagram showing an example of the hardware configuration of a computer 500 capable of realizing an information processing device 12, etc. The computer 500 includes a CPU 501, ROM 502, RAM 503, an input / output interface 505, and a bus 504 connecting these to each other. A display unit 506, an input unit 507, a storage unit 508, a communication unit 509, and a drive unit 510, etc., are connected to the input / output interface 505.

[0136] The display unit 506 is, for example, a display device using liquid crystal, EL, etc., or a glasses-free 3D display using such devices. The input unit 507 is, for example, a keyboard, pointing device, touch panel, or other operating device. If the input unit 507 includes a touch panel, the touch panel may be integrated with the display unit 506. The storage unit 508 is a non-volatile storage device, for example, an HDD, flash memory, or other solid memory. The drive unit 510 is a device capable of driving the removable recording medium 511, for example, an optical recording medium or magnetic recording tape. The communication unit 509 is a modem, router, or other communication device for communicating with other devices, which can be connected to a LAN, WAN, etc. The communication unit 509 may communicate using either wired or wireless methods. The communication unit 509 is often used separately from the computer 500.

[0137] Information processing by the computer 500 having the hardware configuration described above is realized through the cooperation of software stored in the memory unit 508 or ROM 502, etc., and the hardware resources of the computer 500. Specifically, the information processing method related to this technology is realized by loading the programs that constitute the software, stored in the ROM 502, etc., into the RAM 503 and executing them.

[0138] The program is installed on the computer 500, for example, via a removable recording medium 511. Alternatively, the program may be installed on the computer 500 via a global network or the like. In addition, any non-transient storage medium that the computer 500 can read may be used.

[0139] In this disclosure, "system" means a collection of multiple components (devices, modules (parts), etc.), regardless of whether all components are located in the same enclosure. Therefore, multiple devices housed in separate enclosures and connected via a network, and a single device containing multiple modules in one enclosure, are both considered systems.

[0140] The execution of the information processing method related to this technology by a computer system includes both cases where the identification of faces and eyes, estimation of facial features and posture, estimation of the three-dimensional position of eyes, determination of ROI, generation and display of viewpoint images are performed by a single computer, and cases where each process is performed by different computers. Furthermore, the execution of each process by a predetermined computer includes having other computers perform part or all of the process and obtaining the results. In other words, the information processing method related to this technology can also be applied to cloud computing configurations in which a single function is shared and processed jointly by multiple devices via a network.

[0141] The information processing systems, information processing devices, FPGAs, clouds, and processing flows described with reference to each drawing are merely embodiments and can be modified as needed without departing from the spirit of this technology. In other words, other arbitrary configurations and algorithms may be adopted to implement this technology.

[0142] In this disclosure, when the word "abbreviated" is used, it is merely for the purpose of facilitating understanding of the explanation, and there is no special meaning in whether or not the word "abbreviated" is used. In other words, in this disclosure, concepts that define shape, size, positional relationships, states, etc., such as "center," "central," "uniform," "equal," "same," "orthogonal," "parallel," "symmetrical," "extending," "axial," and "rectangular," include concepts such as "substantially centered," "substantially central," "substantially uniform," "substantially equal," "substantially the same," "substantially orthogonal," "substantially parallel," "substantially symmetrical," "substantially extending," "substantially axial," and "substantially rectangular." For example, states that fall within a predetermined range (e.g., a range of ±10%) based on "perfectly centered," "perfectly central," "perfectly uniform," "perfectly equal," "perfectly the same," "perfectly orthogonal," "perfectly parallel," "perfectly symmetrical," "perfectly extending," "perfectly axial," and "perfectly rectangular" are also included. Therefore, even if the word "abbreviated" is not added, concepts that would otherwise be expressed with "abbreviated" added may be included. Conversely, the complete state is not excluded when a state is expressed with "abbreviated."

[0143] In this disclosure, expressions using "greater than A" such as "greater than A" and "less than A" are expressions that comprehensively include both concepts that include cases where something is equivalent to A and concepts that do not include cases where something is equivalent to A. For example, "greater than A" is not limited to cases where something is not equivalent to A, but also includes "greater than or equal to A". Similarly, "less than A" is not limited to "less than A", but also includes "less than or equal to A". When implementing this technology, you may appropriately adopt specific settings from the concepts included in "greater than A" and "less than A" so that the effects described above are achieved.

[0144] It is also possible to combine at least two of the feature features of the present technology described above. In other words, the various feature features described in each embodiment may be combined arbitrarily, regardless of the specific embodiment. Furthermore, the various effects described above are merely examples and not limiting, and other effects may also be exhibited.

[0145] Furthermore, this technology can also be configured as follows: (1) An information processing device comprising: an acquisition unit that acquires a first image including the user's face in a first period, and a second image which is a part of the first image and includes the user's eyes in a second period shorter than the first period; and an estimation unit that estimates the real-space eye position, which is the real-space position of the user's eyes, based on the first image and the second image. (2) An information processing device according to (1), further comprising: a target area setting unit that sets the area around the user's eyes as a target area, and the acquisition unit acquires an image of the target area as the second image. (3) An information processing device according to (2), wherein the estimation unit estimates the eye position in the second image, and the target area setting unit sets the target area according to the estimation result of the eye position in the second image. (4) An information processing device according to (3), wherein the estimation unit estimates the eye position in the first image, and the target area setting unit sets the target area according to the estimation result of the eye position in the first image if the user's eyes are not included in the second image. (5) An information processing device according to any one of (1) to (4), wherein the estimation unit estimates the real-space eye position based on the second image, and corrects the estimation result of the real-space eye position based on the estimation result of the real-space eye position based on the first image. (6) An information processing device according to (5), wherein the estimation unit calculates the amount of movement of the real-space eye position within the time required for the real-space eye position estimation process based on the first image, and performs the correction including the amount of movement. (7) An information processing device according to any one of (1) to (6), wherein the estimation unit estimates the posture of the user's face based on the second image, estimates the real-space eye position based on the posture estimation result based on the second image, estimates the posture of the user's face based on the first image, and corrects the posture estimation result based on the second image with the posture estimation result based on the first image.(8) An information processing device according to (7), wherein the estimation unit calculates the amount of change in posture within the time required for the posture estimation process based on the first image, and performs the correction including the amount of change. (9) An information processing device according to (7) or (8), wherein the estimation unit estimates the interocular distance of the user in the first image, and if the second image includes only one eye of the user, estimates the real-space eye position of the user's other eye based on the posture and the interocular distance. (10) An information processing device according to any one of (2) to (4), wherein the estimation unit estimates the in-image face position, which is the position of the user's face in the first image, and in estimating the in-image face position, processes only the area around the target region in the first image. (11) An information processing device according to any one of (1) to (10), further comprising a decimation processing unit that performs resolution decimation processing on an image including the user's face, wherein the acquisition unit acquires the image after the decimation processing as the first image. (12) An information processing device according to any one of (1) to (11), wherein the estimation unit does not perform processing on the second image while the first image does not include the user's face. (13) An information processing device according to any one of (1) to (12), wherein the acquisition unit acquires two images captured simultaneously from different directions as the first image. (14) An information processing device according to any one of (1) to (13), wherein the estimation unit estimates the user's gaze point based on the second image and the real-space eye position. (15) An information processing device according to any one of (1) to (14), wherein the estimation unit estimates who the user is, the direction of the user's face, the user's facial expression, or the user's gestures based on the first image.(16) An information processing method for acquiring a first image including the user's face in a first period, acquiring a second image which is a part of the first image and includes the user's eyes in a second period shorter than the first period, and estimating the real-space eye position, which is the real-space position of the user's eyes, based on the first image and the second image. (17) A program for causing a computer system to perform the steps of acquiring a first image including the user's face in a first period, acquiring a second image which is a part of the first image and includes the user's eyes in a second period shorter than the first period, and estimating the real-space eye position, which is the real-space position of the user's eyes, based on the first image and the second image.

[0146] 1...Wide-angle image 2...ROI image 3...Face 4...Eyes 10...Information processing system 11, 11a, 11b...Camera 12...Information processing device 14...Image processing unit 18, 32...Image separation unit 19...Face detection unit 20...Face part detection unit 21...Face pose estimation unit 22...Eye detection unit 23...Eye position estimation unit 24...Viewpoint image generation unit 25...ROI determination unit

Claims

1. An information processing device comprising: an acquisition unit that acquires a first image including the user's face in a first period, and a second image which is a part of the first image and includes the user's eyes in a second period shorter than the first period; and an estimation unit that estimates the real-space eye position, which is the real-space position of the user's eyes, based on the first image and the second image.

2. An information processing device according to claim 1, further comprising a target area setting unit for setting the area around the user's eyes as a target area, wherein the acquisition unit acquires an image of the target area as the second image.

3. An information processing apparatus according to claim 2, wherein the estimation unit estimates the eye position in the second image, and the target area setting unit sets the target area according to the estimation result of the eye position in the second image.

4. An information processing apparatus according to claim 3, wherein the estimation unit estimates the position of the eye in the first image, and the target area setting unit sets the target area according to the estimation result of the position of the eye in the first image if the user's eye is not included in the second image.

5. An information processing device according to claim 1, wherein the estimation unit estimates the real-space eye position based on the second image, and corrects the estimation result of the real-space eye position with the estimation result of the real-space eye position based on the first image.

6. An information processing device according to claim 5, wherein the estimation unit calculates the amount of movement of the real-space eye position within the time required for the estimation process of the real-space eye position based on the first image, and performs the correction including the amount of movement.

7. An information processing device according to claim 1, wherein the estimation unit estimates the posture of the user's face based on the second image, estimates the real-space eye position based on the posture estimation result based on the second image, estimates the posture of the user's face based on the first image, and corrects the posture estimation result based on the second image with the posture estimation result based on the first image.

8. An information processing device according to claim 7, wherein the estimation unit calculates the amount of change in posture within the time required for the posture estimation process based on the first image, and performs the correction including the amount of change.

9. An information processing device according to claim 7, wherein the estimation unit estimates the interocular distance of the user in the first image, and if the second image includes only one eye of the user, the information processing device estimates the real-space eye position of the user's other eye based on the posture and the interocular distance.

10. An information processing apparatus according to claim 2, wherein the estimation unit estimates the position of the user's face in the first image, and in the estimation of the position of the face in the image, the information processing apparatus processes only the area surrounding the target region in the first image.

11. An information processing apparatus according to claim 1, further comprising a decimation processing unit that performs resolution decimation processing on an image including the user's face, wherein the acquisition unit acquires the image after the decimation processing as the first image.

12. An information processing apparatus according to claim 1, wherein the estimation unit does not perform any processing on the second image while the first image does not include the user's face.

13. An information processing apparatus according to claim 1, wherein the acquisition unit acquires two images captured simultaneously from different directions as the first image.

14. An information processing device according to claim 1, wherein the estimation unit estimates the user's gaze point based on the second image and the real-space eye position.

15. An information processing device according to claim 1, wherein the estimation unit estimates who the user is, the orientation of the user's face, the user's facial expression, or the user's gestures based on the first image.

16. An information processing method that acquires a first image including the user's face at a first cycle, acquires a second image which is a part of the first image and includes the user's eyes at a second cycle shorter than the first cycle, and estimates the real-space eye position, which is the real-space position of the user's eyes, based on the first image and the second image.

17. A program that causes a computer system to perform the following steps: acquiring a first image including the user's face in a first period; acquiring a second image which is a part of the first image and includes the user's eyes in a second period shorter than the first period; and estimating the real-space eye position, which is the real-space position of the user's eyes, based on the first and second images.

Citation Information

Patent Citations

  • Image display device and image blur preventing method

    JP2004317813A

  • Face image processing device, image observation system, and pupil detection system

    JP2020081756A

  • Information processing system, information processing device, and information processing method

    WO2012001755A1

  • Information processing system, eye state measurement system, information processing method, and non-transitory computer readable medium

    WO2022085276A1