Vergence-Based Gaze Matching for Mixed-Mode Immersive Telepresence Applications
Vergence-based gaze matching techniques adjust user gaze in immersive telepresence by detecting eye convergence and head rotation, addressing misalignment issues to enhance immersion and realism in telepresence applications.
Patent Information
- Application Number
- JP2024546207
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2023-06-08
- Filing Date
- 2023-06-15
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2043-06-15
AI Technical Summary
In immersive telepresence applications, the line of sight between users is often misaligned due to the camera's position not being directly in front of the user, leading to a lack of eye contact and reduced immersion.
Implement vergence-based gaze matching techniques to adjust the gaze position of users by detecting eye convergence and head rotation, and modifying the image to align the user's gaze with the object of interest, using sensors and image processing to enhance the immersive experience.
Enhances user immersion by ensuring eye contact and aligning gaze with the intended object of interest, improving the perception and realism of telepresence interactions.
Smart Images

Figure 0007778950000001 
Figure 0007778950000002 
Figure 0007778950000003
Abstract
Description
[Technical Field]
[0001] This application claims the benefit of priority to U.S. Provisional Application No. 63 / 359,746, entitled "Convergence-Based Gaze Matching for Mixed-Mode Immersive Telepresence Applications," filed July 8, 2022, which claims the benefit of priority to U.S. Patent Application No. 18 / 207,594, entitled "Convergence-Based Gaze Matching for Mixed-Mode Immersive Telepresence Applications," filed June 8, 2023. The disclosures of the prior applications are incorporated herein by reference in their entireties.
[0002] This disclosure describes embodiments generally related to media processing, including real-time immersive telepresence applications. [Background technology]
[0003] The "Background" discussion provided herein is intended to generally present the context for the present disclosure. To the extent described in this Background section, the work of the current inventors, as well as aspects of the disclosure that may not otherwise be admitted as prior art at the time of filing, are not expressly or implicitly admitted as prior art to the present disclosure.
[0004] Real-time immersive telepresence applications, such as video chat, training, and education, enable people in remote locations to engage in real-time conversations, training, and various pedagogical teaching models. In some embodiments, during a real-time immersive telepresence application, a display screen is positioned in front of a first user, and the display screen can display an image of a second user to simulate a face-to-face environment. Summary of the Invention [Problem to be solved by the invention]
[0005] Aspects of the present disclosure provide methods and apparatus for gaze matching. [Means for solving the problem]
[0006] In some embodiments, a processing circuit determines a position of an object of interest of a first user, receives a first one or more images of the first user captured by a camera at a camera position different from the position of the object of interest, detects a first vergence or rotation of eyes of the first user, calculates a mismatch of the first vergence or rotation for viewing the object of interest, and performs gaze correction for the first one or more images based on the mismatch of the first vergence or rotation for viewing the object of interest.
[0007] In some embodiments, the processing circuitry receives a second one or more images of a second user, the second user and the first user engaging in immersive telepresence, derives an interpupillary position of the second user from the second one or more images, displays the second one or more images with the interpupillary position set in a screen plane of a display screen of the first user, and determines the interpupillary position in the screen plane of the display screen as the position of the object of interest of the first user.
[0008] In some embodiments, the processing circuitry determines a center point of the first user's display screen as the location of the first user's object of interest.
[0009] In some embodiments, the processing circuitry receives a second one or more images of a second user engaging in immersive telepresence with the first user, displays the second one or more images on a display screen of the first user, and determines a location on the display screen for displaying the eyes of the second user as the location of the object of interest of the first user.
[0010] In one example, the first vergence or rotation is sensed based on an eye tracking sensor separate from the camera, and in another example, the first vergence or rotation is detected based on image analysis of the first one or more images of the first user.
[0011] In some embodiments, the processing circuitry calculates the first convergence or rotation of the eyes of the first user based on a head position of the first user.
[0012] In some embodiments, the processing circuitry determines a point of gaze according to the first convergence or rotation of the eyes of the first user and a head position of the first user, and calculates the mismatch of the point of gaze relative to the position of the object of interest of the first user.
[0013] In one example, the processing circuitry modifies the head position of the first user in the first one or more images. In another example, the processing circuitry modifies the body position of the first user in the first one or more images. In another example, the processing circuitry modifies the head pose of the first user in the first one or more images. In another example, the processing circuitry modifies the body pose of the first user in the first one or more images. In another example, the processing circuitry modifies the head rotation angle of the first user in the first one or more images. In another example, the processing circuitry modifies the body rotation angle of the first user in the first one or more images. In another example, the processing circuitry modifies the eye convergence angle or rotation angle of the first user in one or more images.
[0014] In some embodiments, the processing circuitry determines the location of the object of interest according to at least one of a facial expression of a person on a display screen, a visual emotion analysis of the person on the display screen, and a mood of the person on the display screen.
[0015] In some examples, the processing circuitry detects the first convergence or rotation of the eyes of the first user using at least one of an inertial measurement unit (IMU), a depth sensor, a light detection and ranging (LiDAR) sensor, a near-infrared (NIR) sensor, and a spatial audio detector.
[0016] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions that, when executed by a computer, cause the computer to perform the method of gaze alignment. [Brief explanation of the drawings]
[0017] [Figure 1] 1 shows an illustration of gaze correction according to some embodiments of the present disclosure. [Figure 2] 1 shows a diagram of an immersive telepresence system according to some embodiments. [Figure 3] 1 is a flowchart outlining a process according to some embodiments of the present disclosure. [Figure 4] FIG. 1 is a schematic diagram of a computer system according to one embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0018] Further features, nature and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings.
[0019] In various scenarios, it may not be possible to position a camera directly in front of a person. Thus, a picture of a person captured by a camera may show a line of sight that differs from the actual line of sight of the real person. According to one aspect of the present disclosure, in immersive telepresence applications, line of sight connection between the real person and the image of the person can improve immersion, enhance perception, and, in some embodiments, can be a ubiquitous cue.
[0020] Some aspects of the present disclosure provide techniques for vergence- or rotation-based image gaze adjustment, such as gaze correction and gaze alignment, which can compensate for camera position and match user gaze in immersive telepresence applications.
[0021] According to some aspects of the present disclosure, image gaze adjustment based on convergence can detect the convergence or rotation of a user's eyes and determine whether the convergence, the user's eye position, and the user's head position are consistent with the user viewing an object of interest. For example, gaze can refer to fixed visual attention by a user, and the user's gaze position (also referred to as gaze position) in an image can be determined based on the user's eye convergence or rotation, the user's eye position, and head pose in the image. Furthermore, gaze correction parameters can be determined to match the gaze position with the position of the object of interest. For example, head rotation parameters and eye rotation parameters can be determined and adjusted to match the gaze position with the position of the object of interest. In some embodiments, one or more images of the user can be captured, and the head position and eye position in the one or more images can be adjusted according to the gaze correction parameters. Thus, the user in one or more images can have, for example, appropriate gaze on the object of interest.
[0022] Figure 1 shows a diagram of gaze correction according to some embodiments of the present disclosure. In the example of Figure 1, a user (101) (e.g., a person) is in front of a display screen (120). The display screen (120) may be part of an electronic system (110), such as a computer system, entertainment system, gaming system, communication system, immersive telepresence system, etc. The components of the electronic system (110) may be integrated into a package or may be separate components connected by wired or wireless connections.
[0023] The electronic system 110 includes a display screen 120 and a camera 130, as well as other suitable components, such as one or more processors (not shown, also referred to as processing circuits), communication components (not shown), etc. In the example of Figure 1, the camera 130 is located above the display screen 120. However, the camera 130 can be located in other suitable locations, such as on the side of the display screen 120.
[0024] Typically, the user 101 is in front of and facing the display screen 120, which is at about face height so that the user 101 can comfortably view the content displayed on the display screen 120. The camera 130 is typically positioned around the periphery of the display screen 120 and is not positioned directly in front of the user 101 so as not to obstruct the user 101's view of the content displayed on the display screen 120.
[0025] According to one aspect of the present disclosure, an image of the user 101 captured by the camera 130 while the user 101 is looking at the display screen 120 may indicate that the user 101 is gazing in a direction other than approximately the center of the display screen 120. For example, when the user 101 is looking at the center of the display screen 120, an image of the user 101 captured by the camera 130 may indicate that the user 101 is looking at a position lower than the center of the display screen 120.
[0026] In another example, a user 101 is a first user communicating with a second user via a telepresence application. The user 101 may see a note posted to the left of the display screen 120 as the camera 130 captures an image of the user 101. When the image is transmitted and displayed on a second display screen in front of the second user, the image of the user 101 on the second display screen does not appear to be looking in the direction of the second user, causing the second user to feel no eye contact with the first user.
[0027] In the example of FIG. 1 , gaze correction based on vergence or rotation can be performed, for example, by the electronic system 110, to compensate for camera position and / or to match the gaze of a second user for eye contact. For example, the electronic system 110 can detect the eye convergence of the user 101. Eye convergence is the movement of the eyes in opposite directions (also called inward rotation) to achieve or maintain single binocular vision. Eye convergence is then used to adjust gaze. Note that other suitable rotational eye movements, such as any combination of vertical rotational eye movements, horizontal rotational eye movements, or rotational eye movements, can be used for gaze correction. In some embodiments, the gaze position 102 of the user 101 can be determined based on eye convergence and other suitable information, such as the position of the user's head and / or eyes. The gaze position 102 is compared to the position 111 of the user's object of interest to determine gaze correction parameters. For example, gaze correction parameters can be applied to the user 101 to adjust the corrected gaze position to the position 111. In some embodiments, the gaze correction parameters can be used to process an image of the user 101 so that the user 101 in the processed image appears to be gazing at an object of interest.
[0028] It should be noted that eye convergence can be detected by a variety of techniques. In one example, a physical eye tracking sensor is used to detect eye convergence. In another example, image analysis can be performed on an image of the user 101 to detect eye convergence. In another example, the user's head position alone is used to detect the user's 101 eye convergence.
[0029] In some embodiments, head position and eye convergence are used to perform gaze adjustment, which can compensate for camera position and match the gaze of a second user.
[0030] It should be noted that the location of the object of interest can be determined by various techniques. In one example, the center point of the display screen (120) is determined to be the object of interest. In another example, an image of a second user is displayed on the display screen (120), and the eye position of the second user on the display screen (120) is determined to be the location of the object of interest. In another example, the location of the object of interest is determined according to at least one of a facial expression of the person on the display screen, a visual emotion analysis of the person on the display screen, and a mood of the person on the display screen.
[0031] In some embodiments, the image of the user 101 captured by the camera 130 is modified. For example, the position, pose, or rotation of the user 101's body in the image may be adjusted. In other embodiments, the position, pose, or rotation of the user 101's head in the image may be adjusted. In other embodiments, the position and rotation of the user 101's eyes in the image may be adjusted.
[0032] It should be noted that the components of FIG. 1 are shown for illustrative purposes, but the electronic system (110) can be implemented using a variety of different components. For example, the electronic system (110) can be configured for immersive telepresence. For example, the display screen (120) can be implemented as a 3D display, such as an 8K autostereoscopic display with left and right view cones. The camera (130) can be implemented as a high-speed RGB camera, a stereo camera, a multi-camera, a depth capture camera, or the like. In one example, the camera (130) is implemented as a 4K stereoscopic RGB camera. The electronic system (110) can include a high-speed transceiver, such as a transceiver capable of live 8K 3D transmission. In some embodiments, the electronic system (110) can include an eye-tracking sensor separate from the camera (130).
[0033] According to some aspects of the present disclosure, vergence-based gaze alignment allows for vergence-based gaze adjustment in immersive telepresence applications to enhance the immersive experience.
[0034] 2 illustrates a diagram of an immersive telepresence system (200) in some embodiments. The immersive telepresence system (200) includes a first electronic system (210A) and a second electronic system (210B) connected via a network (205). The immersive telepresence system (200), in some examples, is capable of two-way, real-time gaze matching. In some examples, an immersive telepresence application is hosted by a server device, and the first electronic system (210A) and the second electronic system (210B) may be client devices of the immersive telepresence application.
[0035] According to some aspects of the present disclosure, a virtual background for an immersive telepresence application can be rendered on a client, such as a first electronic system (210A) and a second electronic system (210B), by a virtual camera, and a dynamic human subject in the foreground can be captured by a stereoscopic camera array. In some embodiments, vergence refers to the inward / outward rotation of the eyes to fixate on an object, and accommodation refers to the eye's focusing mechanism to produce a sharp image on the retina. To avoid mismatches between user vergence and accommodation, a horizontal image translation technique can be used to set the depth position of the user's object of interest (e.g., the person's eyes on the display screen) to the screen plane of the display screen, which is also called the zero-parallax position, thereby matching the user's focus, accommodation, and vergence in some embodiments. When the user is away from the camera, the inward vergence rotation of the user's eyes can be calculated, and gaze correction can be applied when the user is actually looking at the person's eyes on the display screen.
[0036] The first electronic system (210A) and the second electronic system (210B) may each be configured similarly to the electronic system (110).
[0037] For example, the first electronic system (210A) includes a display screen (220A) implemented as a 3D display, such as an 8K autostereoscopic display with left and right view cones or any number of view cones. The first electronic system (210A) includes a camera (230A) implemented as a high-speed RGB camera, a stereo camera, a multi-camera, a depth capture camera, or the like. The first electronic system (210A) includes a transceiver (not shown, wired or wireless) configured to transmit signals to and / or receive signals from the network (205). In some embodiments, the first electronic system (210A) can include an eye tracking sensor (not shown) separate from the camera (230A). The first electronic system (210A) also includes processing circuitry, such as one or more processors (not shown), for image processing.
[0038] Similarly, the second electronic system (210B) includes a display screen (220B) implemented as a 3D display, such as an 8K autostereoscopic display with left and right view cones or any number of view cones. The second electronic system (210B) includes a camera (230B) implemented as a high-speed RGB camera, a stereo camera, a multi-camera, a depth capture camera, or the like. The second electronic system (210B) includes a transceiver (not shown, wired or wireless) configured to transmit signals to and / or receive signals from the network (205). In some embodiments, the second electronic system (210B) can include an eye tracking sensor separate from the camera (230B). The second electronic system (210B) also includes processing circuitry, such as one or more processors (not shown), for image processing.
[0039] The immersive telepresence system (200) is capable of two-way real-time gaze matching. In the example of FIG. 2, the first user (201A) is in front of the display screen (220A) and faces the display screen (220A), which is at about face height so that the first user (201A) can comfortably view the content displayed on the display screen (220A). The camera (230A) is positioned around the periphery of the display screen (220A) and is not positioned directly in front of the first user (201A) so as not to obstruct the first user's view of the content displayed on the display screen (220A). For example, the camera (230A) is positioned to the left of the display screen (220A) in FIG. 2.
[0040] Similarly, the second user 201B is in front of the display screen 220B and faces the display screen 220B, which is at about the same height as the second user 201B's face so that the second user 201B can comfortably view the content displayed on the display screen 220B. The camera 230B is positioned around the periphery of the display screen 220B and is not positioned directly in front of the second user 201B so as not to obstruct the second user 201B's view of the content displayed on the display screen 220B. For example, the camera 230B is positioned above the display screen 220B in FIG. 2.
[0041] According to some aspects of the present disclosure, a camera (230A) captures a first stereoscopic image of a first user (201A). The first stereoscopic image can be processed, for example, according to convergence-based gaze alignment. In one example, the first stereoscopic image is processed by a processing circuit of a first electronic system (210A) according to convergence-based gaze alignment. In another example, the first stereoscopic image is processed by a server (e.g., an immersive telepresence server) in a network (205) according to convergence-based gaze alignment. The processed first stereoscopic image is transmitted to a second electronic system (210B). The processed first stereoscopic image is displayed by a display screen (220B) to show a modified image, such as the display image (202B) of the first user (201A) in FIG. 2. In some embodiments, the first stereoscopic image or the processed first stereoscopic image can be further processed by a processing circuit of the second electronic system (210B) for gaze alignment, and then the processed first stereoscopic image is displayed by a display screen (220B) to show a modified image, such as a display image (202B) of the first user (201A).
[0042] Similarly, the camera (230B) captures a second stereoscopic image of the second user (201B). The second stereoscopic image can be processed, for example, according to convergence-based gaze alignment. In one example, the second stereoscopic image is processed by a processing circuit of the second electronic system (210B) according to convergence-based gaze alignment. In another example, the second stereoscopic image is processed by a server (e.g., an immersive telepresence server) in the network (205) according to convergence-based gaze alignment. The processed second stereoscopic image is transmitted to the first electronic system (210A). The processed second stereoscopic image is displayed by the display screen (220A) to show a modified image, such as the display image (202A) of the second user (201B) in FIG. 2. In some embodiments, the second stereoscopic image or the processed second stereoscopic image can be further processed by processing circuitry of the first electronic system (210A), and then the further processed second stereoscopic image is displayed by the display screen (220A) to show a modified image, such as the displayed image (202A) of the second user (201B).
[0043] In some embodiments, when a first user (201A) looks at a displayed image (202A), the first user (201A) makes eye contact (also called gaze alignment) with a second user (201B)'s displayed image (202A). The eye contact enhances the immersive experience of the first user (201A).
[0044] Similarly, when the second user 201B views the display image 202B, the second user 201B makes eye contact (also known as gaze alignment) with the display image 202B of the first user 201A. The eye contact enhances the immersive experience of the second user 201B. Note that in some embodiments, the first user 201A and the second user 201B do not need to simultaneously view the display image and make eye contact.
[0045] In some embodiments, the first electronic system (210A) receives a processed second stereoscopic image of a second user (201B). In some embodiments, the processed second stereoscopic image includes a pair of images. For each image of the pair of images, the interpupillary position (e.g., the center position of the two pupils) of a person (e.g., the second user (201B)) in the image is determined. For example, the two pupils in the image are determined, and the 3D XYZ coordinates of the interpupillary positions of the two pupils are determined. Using the interpupillary positions of the pair of images, the pair of images are shifted and displayed on the display screen (220A) so that the interpupillary positions overlap and are displayed at one position on the display screen (220A), thereby setting the depth plane of the eye to the screen plane of the display screen (220A). The person's eyes displayed according to the shifted images can then be observed at the screen plane of the display screen (220A), which can be referred to as the zero parallax position (ZPS). In some embodiments, when the eyes are observed at the screen plane, the nose can be observed in front of the screen plane and the back of the head is behind the screen plane. In one example, the location of overlapping interpupillary positions on the display screen (220A), as shown at (203A), is defined as the object of interest of the first user (201A) to achieve gaze alignment.
[0046] Further, in the example of FIG. 2, the eye convergence of the first user (201A) is detected. In one example, an eye tracking sensor can determine the eye movement of the first user (201A). The eye movement can be used to determine a 3D vector of the convergence point of the eyes. In another example, the 3D vector of the convergence point of the eyes is determined according to image analysis of the first stereoscopic image of the first user (201A). In some embodiments, head rotation (or head position, or head pose) can be determined from image analysis of the first stereoscopic image of the first user (201A).
[0047] Eye convergence is then used to perform gaze adjustment. In some embodiments, the gaze position of the first user (201A) can be determined based on the eye convergence of the first user (201A) and other suitable information, such as head rotation and eye position. The gaze position of the first user (201A) is compared with the position of the object of interest of the first user (201A) to determine gaze correction parameters for the first stereoscopic image. The first stereoscopic image can be processed according to the gaze correction parameters.
[0048] In some embodiments, the first electronic system (210A) includes a sensing module configured to detect eye convergence of the first user (201A). In one example, the sensing module can be implemented by a camera (230A), such as a high-speed RGB camera, a stereo camera, a multi-camera, or a depth capture camera. In another example, the sensing module can be implemented by an eye-tracking sensor.
[0049] In some embodiments, the first electronic system (210A) includes a rendering module configured to set a depth plane of a person's eyes in the processed second stereoscopic image relative to a screen plane of the display screen (220A). For example, the rendering module may set a depth plane of a person's eyes in the processed second stereoscopic image relative to a screen plane of the display screen (220A). In some embodiments, the rendering module may be implemented by one or more processors executing software instructions.
[0050] In some embodiments, the first electronic system (210A) includes a registration module configured to determine a disparity between a point of gaze and an object of interest. The point of gaze can be determined based on eye convergence detected by eye tracking and other suitable information. In some embodiments, the registration module can include a compact model and can be implemented by one or more processors executing software instructions.
[0051] In some embodiments, the first electronic system (210A) includes a gaze correction module configured to perform gaze correction on the first stereoscopic image. In some embodiments, the gaze correction module can receive live input from the sensing module and the registration module and perform modifications to the first stereoscopic image. For example, the gaze correction module can use the live input from the sensing module and the registration module to perform 2D restoration and / or 3D rotational gaze correction. In one example, the person in the first stereoscopic image is in the form of a 3D mesh. The 3D mesh can be rotated for gaze correction. In one example, the person's eyes can be rotated for gaze correction. In another example, the person's head can be rotated for gaze correction. In another example, the person's body can be rotated for gaze correction. In some embodiments, the gaze correction module is implemented by a neural network model that is trained to fulfill that purpose.
[0052] In one example, due to the location of the camera (230A), when the first user (201A) looks at the center of the display image (202A), for example, the display screen (220A), the first stereoscopic image captured by the camera (230A) shows that the first user (201A) is looking at the right part of the display screen (220A). The first electronic system (210A) performs gaze matching based on convergence to determine gaze correction. For example, the gaze point of the first user (201A) can be determined based on eye convergence, and the gaze point can be a point on the right part of the display screen (220A). The difference between the gaze point and the position of the object of interest (e.g., the interpupillary position in the screen plane of the display screen (220A)) can be determined. Based on this difference, a rotation angle for rotating the person in the first stereoscopic image to the left can be determined to correct the first stereoscopic image. Then, the mesh representing the person in the first stereoscopic image can be rotated to the left according to the rotation angle for gaze correction.
[0053] In the above description, the first stereoscopic image is processed by the first electronic system (210A) for gaze correction, but it should be noted that in some embodiments gaze correction can also be performed by the second electronic system (210B), although in some embodiments gaze correction can be suitably performed by a server in the network (205).
[0054] Similarly, in some embodiments, the second electronic system (210B) receives a processed first stereoscopic image of the first user (201A). In some embodiments, the processed first stereoscopic image includes a pair of images. For each image of the pair of images, the interpupillary position (e.g., the center position of the two pupils) of the person (first user (201A)) in the image is determined. For example, the two pupils in the image are determined, and the 3D XYZ coordinates of the interpupillary positions of the two pupils are determined. Using the interpupillary positions of the pair of images, the pair of images are shifted (also referred to as horizontal image translation in some embodiments) and displayed on the display screen (220B) so that the interpupillary positions overlap and appear at a single point, and the depth plane of the eyes is set to the screen plane of the display screen (220B). The person's eyes, displayed according to the shifted images, can then be observed at the screen plane of the display screen (220B), which can be referred to as the zero parallax position (ZPS). In some embodiments, when the eyes are observed at the screen plane, the nose can be observed in front of the screen plane and the back of the head is behind the screen plane. In one example, a point of overlapping interpupillary position on the display screen (220B), as shown at (203B), is defined as the object of interest of the second user (201B) to achieve gaze alignment.
[0055] Further, in the example of FIG. 2, eye convergence of the second user (201B) is detected. In one example, an eye tracking sensor can determine eye movements of the second user (201B). The eye movements can be used to determine a 3D vector of a convergence point of the eyes. In another example, the 3D vector of the convergence point of the eyes is determined according to image analysis of the second stereoscopic image of the second user (201B). In some embodiments, head rotation (or head position, or head pose) is determined from image analysis of the second stereoscopic image of the second user (201B).
[0056] Eye convergence is then used to perform gaze adjustment. In some embodiments, the gaze position of the second user (201B) can be determined based on the eye convergence of the second user (201B) and other suitable information, such as head rotation and eye position. The gaze position of the second user (201B) is compared with the position of the object of interest of the second user (201B) to determine gaze correction parameters for the second stereoscopic image. The second stereoscopic image can be processed according to the gaze correction parameters.
[0057] In some embodiments, the second electronic system (210B) includes a sensing module configured to detect eye convergence of the second user (201B). In one example, the sensing module can be implemented by a camera (230B), such as a high-speed RGB camera, a stereo camera, a multi-camera, or a depth capture camera. In another example, the sensing module can be implemented by an eye-tracking sensor.
[0058] In some embodiments, the second electronic system (210B) includes a rendering module configured to set a depth plane of a person's eyes in the processed first stereoscopic image relative to a screen plane of the display screen (220B). For example, the rendering module may set a depth plane of a person's eyes in the processed first stereoscopic image relative to a screen plane of the display screen (220B). In some embodiments, the rendering module may be implemented by one or more processors executing software instructions.
[0059] In some embodiments, the second electronic system (210B) includes a registration module configured to determine a disparity between a point of gaze and an object of interest. The point of gaze can be determined based on eye convergence detected by eye tracking and other suitable information. In some embodiments, the registration module can be implemented by one or more processors executing software instructions.
[0060] In some embodiments, the second electronic system (210B) includes a gaze correction module configured to perform gaze correction on the second stereoscopic image. In some embodiments, the gaze correction module can receive live input from the sensing module and the registration module and perform modifications to the second stereoscopic image. For example, the gaze correction module can use the live input from the sensing module and the registration module to perform 2D inpainting and / or 3D rotational gaze correction. In one example, the person in the second stereoscopic image is in the form of a 3D mesh. The 3D mesh can be rotated for gaze correction. In one example, the person's eyes can be rotated for gaze correction. In another example, the person's head can be rotated for gaze correction. In another example, the person's body can be rotated for gaze correction. In some embodiments, the gaze correction module is implemented by a neural network model that is trained to fulfill that purpose.
[0061] In one example, due to the location of the camera (230B), when the second user (201B) looks at the center of the display image (202A), for example, the display screen (220B), the second stereoscopic image captured by the camera (230B) shows that the second user (201B) is looking at the bottom of the display screen (220B). The second electronic system (210B) performs gaze matching based on convergence to determine gaze correction. For example, the gaze point of the second user (201B) can be determined based on eye convergence, and the gaze point can be a point at the bottom of the display screen (220B). The difference between the gaze point and the position of the object of interest (e.g., the interpupillary position in the screen plane of the display screen (220B)) can be determined. Based on this difference, a rotation angle for rotating the person's head or body upward in the second stereoscopic image can be determined to correct the second stereoscopic image. The mesh representing the person in the second stereo image can then be rotated upward according to the rotation angle for gaze correction.
[0062] In the above description, the second stereoscopic image is processed by the second electronic system (210B) for gaze correction, but it should be noted that in some embodiments gaze correction can also be performed by the first electronic system (210A), although in some embodiments gaze correction can be suitably performed by a server in the network (205).
[0063] 3 shows a flowchart outlining a process (300) according to one embodiment of the present disclosure. In various embodiments, the process (300) can be performed by a processing circuit in an electronic system, such as the electronic system (110), the first electronic system (210A), or the second electronic system (210B). In some embodiments, the process (300) is implemented in software instructions, and thus, the processing circuit performs the process (300) when it executes the software instructions. The process starts at (S301) and proceeds to (S310).
[0064] At (S310), a location of an object of interest of a user, such as a first user, is determined.
[0065] At (S320), a first image or images of a first user captured by a camera at a camera position different from the position of the object of interest are received.
[0066] In (S330), a first convergence or rotation of the first user's eye is determined.
[0067] In (S340), a first vergence or rotation mismatch for viewing the object of interest is determined.
[0068] In (S350), gaze correction is performed on the first one or more images based on a first vergence or rotation mismatch for viewing the object of interest.
[0069] In some embodiments, to determine a location of the object of interest, a second one or more images of a second user engaging in immersive telepresence with the first user are received. An interpupillary position of the second user is then derived from the second one or more images. The second one or more images are displayed with the interpupillary position set in a screen plane of a display screen of the first user. The interpupillary position in the screen plane of the display screen is determined as the location of the object of interest of the first user.
[0070] In some embodiments, a second one or more images of a second user engaging in immersive telepresence with the first user are received, the second one or more images are displayed on a display screen of the first user, and a location on the display screen for displaying the second user's eyes is determined as a location of an object of interest of the first user.
[0071] In some embodiments, the center point of the display screen is determined as the location of the first user's object of interest.
[0072] To detect the first convergence of the eyes of the first user, in one example, the first convergence is determined based on data sensed by an eye tracking sensor separate from the camera. In some embodiments, the first convergence is determined based on image analysis of the first one or more images of the first user.
[0073] In one example, a first convergence of the eyes of the first user is calculated based solely on the head position of the first user.
[0074] In some embodiments, to further calculate a first convergence inconsistency for viewing the object of interest, a gaze point is determined according to the first convergence of the eyes of the first user and the head position of the first user, and then a difference between the gaze point and the position of the object of interest of the first user is calculated to determine the inconsistency.
[0075] In some examples, gaze correction can be performed by adjusting various parameters such as the position of the first user's head and / or body, the pose of the first user's head and / or body, and the rotation angle of the first user's head and / or body. In one example, the head position of the first user in the first one or more images is modified. In another example, the body position of the first user in the first one or more images is modified. In another example, the head pose of the first user in the first one or more images is modified. In another example, the body pose of the first user in the first one or more images is modified. In another example, the head rotation angle of the first user in the first one or more images is modified. In another example, the body rotation angle of the first user in the first one or more images is modified.
[0076] In some embodiments, the location of the object of interest is determined according to at least one of a facial expression of the person on the display screen, a visual emotion analysis of the person on the display screen, and a mood of the person on the display screen.
[0077] In some embodiments, at least one of an inertial measurement unit (IMU), a depth sensor, a light detection and ranging (LiDAR) sensor, a near-infrared (NIR) sensor, and a spatial audio detector is used to detect a first convergence of the eyes of the first user.
[0078] Then, the process proceeds to (S399) and ends.
[0079] The process 300 may be adapted as appropriate. Steps of the process 300 may be modified and / or omitted. Additional steps may be added. Any suitable order of implementation may be used.
[0080] The techniques described above may be implemented as computer software using computer-readable instructions and physically stored on one or more computer-readable media, such as non-transitory computer-readable storage. For example, Figure 4 illustrates a computer system (400) suitable for implementing certain embodiments of the disclosed subject matter.
[0081] Computer software may be encoded using any suitable machine code or computer language, and may be produced by mechanisms such as assembly, compilation, linking, etc. to create code containing instructions that may be executed directly by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., or may be executed by interpretation, microcode execution, etc.
[0082] The instructions may be executed on various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things devices, and the like.
[0083] 4 are exemplary in nature and are not intended to suggest any limitations on the scope of use or functionality of the computer software implementing the embodiments of the present disclosure. The arrangement of components should not be interpreted as having any dependency or requirement relating to any one or combination of components illustrated in the exemplary embodiment of computer system 400.
[0084] The computer system (400) may include certain human interface input devices that can respond to input by one or more users, for example, through tactile input (e.g., keystrokes, swipes, data glove movements, etc.), audio input (e.g., voice, clapping, etc.), visual input (e.g., gestures, etc.), or olfactory input (not shown). The human interface input devices can also be used to capture certain media that are not necessarily directly associated with conscious human input, such as audio (e.g., voice, music, ambient sounds, etc.), images (e.g., scanned images, photographic images obtained from a still camera, etc.), and video (e.g., two-dimensional video, three-dimensional video including stereoscopic vision, etc.).
[0085] The input human interface devices may include one or more (only one of which is shown) of a keyboard (401), a mouse (402), a trackpad (403), a touchscreen (410), a data glove (not shown), a joystick (405), a microphone (406), a scanner (407), and a camera (408).
[0086] The computer system (400) may also include several human interface output devices that may stimulate one or more of the user's senses, for example, through tactile output, sound, light, and smell / taste. Such human interface output devices may include haptic output devices (e.g., haptic feedback via a touchscreen (410), data gloves (not shown), or joystick (405), but may also be haptic feedback devices that do not act as input devices), audio output devices (e.g., speakers (409), headphones (not shown)), visual output devices (e.g., screens (410), including CRT screens, LCD screens, plasma screens, and OLED screens, each of which may or may not have touchscreen input capabilities, each of which may or may not have haptic feedback capabilities, some of which may output two-dimensional visual output or three or more dimensional output by means of stereographic output, virtual or augmented reality glasses (not shown), multi-view displays, holographic displays and smoke tanks (not shown), and printers (not shown)).
[0087] The computer system (400) may also include human-accessible storage devices and their associated media, such as optical media (421), including, for example, CD / DVD ROM / RW (420) with CD / DVD or similar media, thumb drives (422), removable hard drives or solid-state drives (423), legacy magnetic media (not shown), such as tape and floppy disks, and dedicated ROM / ASIC / PLD-based devices (not shown), such as security dongles.
[0088] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the presently disclosed subject matter does not encompass transmission media, carrier waves, or other transitory signals.
[0089] The computer system 400 may also include an interface 454 to one or more communication networks 455. The networks may be, for example, wireless, wired, or optical. The networks may further include local networks, wide area networks, metropolitan area networks, vehicular and industrial networks, real-time networks, delay-tolerant networks, and the like. Examples of networks include Ethernet, WLAN, GSM, cellular networks including 3G, 4G, 5G, LTE, and the like; television wired or wireless wide area digital networks including cable TV, satellite TV, and terrestrial broadcast TV; and vehicular and industrial networks including CAN Bus. Some networks generally require an external network interface adapter connected to a specific general-purpose data port or peripheral bus 449 (e.g., a USB port on the computer system 400), while others are typically integrated into the core of the computer system 400 by connecting to a system bus (e.g., an Ethernet interface to a PC computer system, or a cellular network interface to a smartphone computer system), as described below. Using any of these networks, computer system (400) can communicate with other entities. Such communication may be unidirectional receive-only (e.g., broadcast TV), unidirectional transmit-only (e.g., from the CAN bus to a particular CAN bus device), or bidirectional, e.g., to other computer systems using local or wide-area digital networks. As noted above, specific protocols and protocol stacks may be used with each of these networks and network interfaces.
[0090] The aforementioned human interface devices, human-accessible storage devices, and network interfaces may be connected to the core (440) of the computer system (400).
[0091] The core (440) may include one or more central processing units (CPUs) (441), graphics processing units (GPUs) (442), dedicated programmable processing units in the form of field programmable gate arrays (FPGAs) (443), task-specific hardware accelerators (444), graphics adapters (450), etc. These devices may be connected via a system bus (448), along with read-only memory (ROM) (445), random access memory (446), and internal mass storage (447), such as an internal hard disk or SSD, that is not user accessible. In some computer systems, the system bus (448) is accessible in the form of one or more physical plugs to allow expansion with additional CPUs, GPUs, etc. Peripheral devices may be connected directly to the core's system bus (448) or via a peripheral bus (449). In some embodiments, a screen (410) may be connected to the graphics adapter (450). Peripheral bus architectures include PCI, USB, etc.
[0092] The CPU (441), GPU (442), FPGA (443), and accelerator (444) can execute several instructions, which can be combined to form the computer code described above. The computer code can be stored in ROM (445) or RAM (446). Transient data can also be stored in RAM (446), while permanent data can be stored, for example, in internal mass storage (447). The use of cache memory, which can be closely associated with one or more of the CPU (441), GPU (442), mass storage (447), ROM (445), RAM (446), etc., can enable fast storage and retrieval from any memory device.
[0093] The computer-readable medium may have computer code thereon for performing various computer-implemented operations. The medium and computer code may be those specially designed and constructed for the purposes of the present disclosure, or they may be of the kind well known and available to those skilled in the computer software arts.
[0094] By way of example, and not limitation, a computer system having the architecture (400), and in particular the core (440), can provide functionality as a result of the processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media can be media associated with user-accessible mass storage, as described above, as well as specific storage of the core (440) that is non-transitory in nature, such as the core's internal mass storage (447) or ROM (445). Software implementing various embodiments of the present disclosure can be stored on such devices and executed by the core (440). The computer-readable media can include one or more memory devices or chips, depending on particular needs. The software can cause the core (440), and in particular the processor (including a CPU, GPU, FPGA, etc.) therein, to perform the particular processes or particular portions of the particular processes described herein, including defining data structures stored in RAM (446) and modifying such data structures in accordance with the software-defined processes. Additionally, or alternatively, a computer system may provide functionality as a result of logic hardwired or otherwise embodied in circuitry (e.g., accelerator 444), which may operate in place of or in conjunction with software to perform particular processes or portions of particular processes described herein. Where appropriate, references to software may include logic, and vice versa. Where appropriate, references to computer-readable media may include circuitry (e.g., integrated circuits (ICs)) that stores software for execution, circuitry that embodies logic for execution, or both. The present disclosure encompasses any appropriate combination of hardware and software.
[0095] While this disclosure has described several exemplary embodiments, there are alterations, permutations, and various substitute equivalents that fall within the scope of this disclosure. It will thus be appreciated that those skilled in the art will be able to devise numerous systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and are therefore within its spirit and scope.
Claims
1. determining a location of an object of interest of a first user; receiving a first one or more images of the first user captured by a camera at a camera position different from the position of the object of interest; detecting a first convergence of the eyes of the first user; calculating the first convergence discrepancy for viewing the object of interest; performing gaze correction for the first one or more images based on the inconsistency of the first convergence for viewing the object of interest; Including, The step of performing gaze correction includes: modifying a head position of the first user in the first one or more images; modifying a body position of the first user in the first one or more images; correcting a head pose of the first user in the first one or more images; modifying a body pose of the first user in the first one or more images; correcting a head rotation angle of the first user in the first one or more images; and modifying a body rotation angle of the first user in the first one or more images. A method for gaze alignment characterized by:
2. The step of determining the location of the object of interest comprises: receiving a second one or more images of a second user, the second user and the first user engaging in immersive telepresence; deriving an interpupillary position of the second user from the second one or more images; displaying the second image or images with the interpupillary position set to a screen plane of the display screen of the first user; 2. The method of gaze alignment of claim 1, further comprising determining the interpupillary position in the screen plane of the display screen as the position of the object of interest of the first user.
3. The step of determining the location of the object of interest comprises:
2. The method of gaze alignment of claim 1, further comprising the step of determining a center point of the display screen of the first user as the location of the object of interest of the first user, wherein the camera position is at the periphery of the display screen.
4. The step of determining the location of the object of interest comprises: receiving a second one or more images of a second user, the second user and the first user engaging in immersive telepresence; displaying the second image or images on a display screen of the first user, the camera position being at the periphery of the display screen; 2. The method of gaze alignment of claim 1, further comprising determining a location on the display screen for displaying the second user's eyes as the location of the object of interest of the first user.
5. The step of detecting the first convergence of the eyes of the first user includes: sensing the first convergence based on an eye tracking sensor separate from the camera; and detecting the first convergence based on image analysis of the first one or more images of the first user.
6. The step of detecting the first convergence of the eyes of the first user includes:
2. The method of gaze alignment of claim 1, further comprising: calculating the first convergence of the eyes of the first user based on a head position of the first user.
7. The step of calculating the disparity of the first congestion for viewing the object of interest comprises: determining a gaze point according to the first convergence of the eyes of the first user and a head position of the first user; 2. The method of claim 1, further comprising: calculating the mismatch of the point of gaze relative to the location of the object of interest of the first user.
8. The step of detecting the first convergence of the eyes of the first user includes:
2. The method of gaze alignment of claim 1, comprising detecting the first convergence of the eyes of the first user using at least one of an inertial measurement unit (IMU), a depth sensor, a light detection and ranging (LiDar) sensor, a near-infrared (NIR) sensor, and a spatial audio detector.
9. Determining a location of an object of interest of a first user; receiving a first one or more images of the first user captured by a camera at a camera position different from the position of the object of interest; detecting a first convergence of the eyes of the first user; calculating the first convergence discrepancy for viewing the object of interest; performing gaze correction for the first one or more images based on the inconsistency of the first convergence for viewing the object of interest; Including, The step of determining the location of the object of interest of the first user comprises: determining the location of the object of interest according to at least one of a facial expression of a person on a display screen, a visual emotion analysis of the person on the display screen, and a mood of the person on the display screen; Methods of gaze matching.
10. An electronic system comprising a processing circuit configured to perform the method for gaze alignment according to any one of claims 1 to 9.
Citation Information
Patent Citations
Image display method, terminal and bidirectional interactive system
JP2005354531A
Visual line detecting apparatus and its method
JP2008194146A
Communication system
JP2011097447A
Information processor, specification method, and specification program
JP2015046069A
Eye Tracking Calibration Technique
JP2020522795A