Eye gaze adjustment
By adjusting the presenter's eye gaze in real time during video communication and using machine learning models to identify and simulate natural eye movements, the problem of unnatural eye movements in video communication is solved, improving communication quality and natural interaction effects.
Patent Information
- Application Number
- CN202180074704.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-10-30
- Filing Date
- 2021-08-31
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2041-08-31
AI Technical Summary
In video communication, the presenter's eye movements can easily distract the receiver, especially when reading, causing the receiver to perceive unnatural eye movements and positional changes, thus affecting communication efficiency.
By capturing users' facial video streams in real time, using machine learning models to identify facial features and eye movements, adjusting the direction of eye gaze to simulate natural eye movements, and generating gaze-adjusted images, including saccades, microsaccades, and convergence/divergence movements, the system avoids reliance on template images and personal calibration, and provides automatic adjustment and on/off mechanisms.
It improves the quality of human communication in video communications, maintains the presenter's natural and focused appearance, enhances interaction with the receiver, and avoids unnatural eye movement appearances.
Smart Images

Figure CN116368525B_ABST
Abstract
Description
BACKGROUND
[0001] Video communication is rapidly becoming the primary means of human communication in the business and academic world, with video conferencing and recorded presentations often replacing face-to-face meetings. In such video conferences, a presenter often reads a pre-written script or other text from a display device of their personal computing system while also using a camera on their computing system to record video. However, given the geometry of the setup, including the distance of the presenter from the display device and the magnification of the presenter’s image on the display device of the recipient, the recipient often readily perceives that the presenter’s eyes move and shift while reading. Moreover, if the presenter’s camera is positioned directly above the display device, the recipient can perceive the presenter’s eye gaze as being focused on a point below the level of the recipient’s eyes. This can distract the recipient, thereby reducing the efficiency of the overall communication process. SUMMARY
[0002] The following provides a simplified summary in order to provide a basic understanding of certain aspects described herein. This summary is not an extensive overview of the claimed subject matter. It is not intended to identify key or critical elements of the claimed subject matter nor is it intended to be used as an aid in construing the claimed subject matter. The sole purpose of this summary is to present some concepts of the claimed subject matter in a simplified form as a prelude to the more detailed description presented below.
[0003] In one embodiment, a computing system is described. The computing system includes a camera to capture a video stream including an image of a user of the computing system. The computing system also includes a processor to execute computer executable instructions that cause the processor to receive the image of the user from the camera to detect a facial region of the user within the image and detect a facial feature region of the user within the image based on the detected facial region. The computer executable instructions also cause the processor to determine whether the image represents the user being fully disengaged from the computing system based on the detected facial feature region and, if the image does not represent the user being fully disengaged from the computing system, detect an eye region of the user within the image based on the detected facial feature region. The computer executable instructions also cause the processor to compute a required eye gaze direction of the user based on the detected eye region, generate a gaze-adjusted image based on the required eye gaze direction of the user, wherein the gaze-adjusted image includes at least one of a saccadic eye movement, a microsaccadic eye movement, or a vergence eye movement, and replace the image within the video stream with the gaze-adjusted image.
[0004] In another embodiment, a method for adjusting eye gaze of a user within a video stream is described. The method includes capturing, via a camera of a computing system, a video stream comprising an image of a user of the computing system. The method also includes detecting, via a processor of the computing system, a facial region of the user within the image and detecting, based on the detected facial region, a facial feature region of the user within the image. The method also includes determining, based on the detected facial feature region, whether the image represents the user being fully disengaged from the computing system and, if the image does not represent the user being fully disengaged from the computing system, detecting, based on the detected facial feature region, an eye region of the user within the image. The method also includes calculating, based on the detected eye region, a required eye gaze direction of the user, generating a gaze-adjusted image based on the required eye gaze direction of the user, wherein the gaze-adjusted image comprises at least one of a saccadic eye movement, a microsaccadic eye movement, or a vergence eye movement, and replacing the image within the video stream with the gaze-adjusted image.
[0005] In another embodiment, a computer-readable storage medium is described. A computer-readable storage medium comprising computer-executable instructions that, when executed by a processor of a computing system, cause the processor to capture a video stream comprising an image of a user, detect a facial region of the user within the image, and detect, based on the detected facial region, a facial feature region within the image. The computer-executable instructions also cause the processor to determine, based on the detected facial feature region, whether the image represents the user being fully disengaged from the computing system and, if the image does not represent the user being fully disengaged from the computing system, detect, based on the detected facial feature region, an eye region of the user within the image. The computer-executable instructions also cause the processor to calculate, based on the detected eye region, a required eye gaze direction of the user, generate a gaze-adjusted image based on the required eye gaze direction of the user, wherein the gaze-adjusted image comprises at least one of a saccadic eye movement, a microsaccadic eye movement, or a vergence eye movement, and replace the image within the video stream with the gaze-adjusted image.
[0006] The following description and drawings are illustrative of certain illustrative aspects of the claimed subject matter. However, these aspects are merely examples of the ways in which the principles of the application can be employed and the claimed subject matter is intended to include all such aspects and their equivalents. Other advantages and novel features of the claimed subject matter will become apparent from the following detailed description, when considered in conjunction with the drawings. BRIEF DESCRIPTION OF DRAWINGS
[0007] The following detailed description can be better understood with reference to the accompanying drawings, which contain specific examples of numerous features of the disclosed subject matter.
[0008] Figure 1 is a block diagram of an example network environment suitable for implementing the eye gaze adjustment techniques described herein;
[0009] Figure 2 is a block diagram of an example computing system configured to implement the eye gaze adjustment techniques described herein;
[0010] Figure 3 is a schematic diagram depicting Figure 2 the eye gaze adjustment module of Figure 2 may be implemented in the manner of
[0011] Figure 4 is a schematic diagram depicting an example process for computing a desired eye gaze direction for a user;
[0012] Figure 5 is a schematic diagram depicting an example process for generating a gaze-adjusted image based on a desired eye gaze direction for a user;
[0013] Figure 6 is a schematic diagram depicting an example process for training an image generator as described with reference to Figure 5
[0014] Figure 7A is a schematic diagram of an unadjusted image that a user’s camera can capture when the user is reading text from a display device;
[0015] Figure 7B is a schematic diagram of another unadjusted image that a user’s camera can capture when the user is reading text from a display device;
[0016] Figure 7C is a schematic diagram of a gaze-adjusted image that can be generated in accordance with the eye gaze adjustment techniques described herein;
[0017] Figure 8 is a process flow diagram of a method for adjusting a user’s eye gaze in a video stream; and
[0018] Figure 9 is a process flow diagram of another method for adjusting a user’s eye gaze in a video stream. DETAILED DESCRIPTION
[0019] Attention signals play an important role in human communication. Further, one of the most important signals of attention is eye gaze. Specifically, various psychological studies have shown that when humans are able to make eye contact, they are more likely to effectively communicate with one another in interpersonal interactions. However, in various video communication scenarios, such as video calls, video conferences, video storytelling streams, and recorded speeches / presentations based on pre-written scripts (such as teleprompter scripts or scripts displayed on a display device), this primary signal is lost. Typically, when a video communication includes a presenter reading from a display device, the recipient can perceive the presenter's displaced (or "back and forth") eye movements. Further, if the presenter's camera is positioned directly above the display device, the recipient can perceive the presenter's eye gaze as being focused on a point below the recipient's eye level. Moreover, in some cases, the presenter's eye gaze can appear to be too locked onto one portion of the display device, which can give the presenter's eyes an unnatural appearance. These situations can distract the recipient, thereby reducing the likelihood that the presenter effectively conveys the intended message.
[0020] The present technology provides real-time video modification to adjust a presenter's eye gaze during a video communication. More specifically, the present technology adjusts a presenter's eye gaze in real-time such that suboptimal eye movements (such as those associated with reading) are removed, for example, while still allowing natural eye movements (such as natural eye movements that are not associated with reading). Further, in contrast to previous techniques for modifying eye gaze, the technology described herein goes further than merely fixing the direction of a presenter's eye gaze by providing natural and realistic eye movements that preserve the presenter's livelihood and identity. Thus, this technology improves the quality of human communication that can be achieved via digitally live and / or recorded video sessions.
[0021] In various embodiments, the eye gaze adjustment technology described herein includes capturing a video stream of a user's (or presenter's) face and making adjustments to the images within the video stream to adjust the direction of the user's eye gaze. In some embodiments, this includes identifying the user's eyes as moving in a suboptimal manner (such as moving in a back and forth manner that is typically associated with reading lines of text) and then making changes to the eye gaze (and corresponding eye movements) provided in the images such that the eye gaze (and corresponding eye movements) cause the user to appear to be looking in one direction (such as directly at the camera) without substantial eye movements. In such embodiments, it also includes identifying when the user's eyes are not moving in a back and forth manner that is typically associated with reading lines of text and / or when the user's eyes are completely disengaged from the computing system and determining that eye gaze adjustments are not to be performed in such cases.
[0022] In various embodiments, the eye gaze adjustment described herein is provided at least in part by synthesizing certain types of eye movements by modifying images within a video stream. Specifically, there are at least four types of eye movements that are relevant to the present technology. The first type of eye movement is referred to as a "saccade," which is a rapid, simultaneous movement of both eyes between two foci (or fixation points). Saccadic eye movements are relatively large movements of greater than 0.25°, often movements that scan an entire scene or multiple features of a scene. In other words, in the case of saccadic eye movements, the eyes typically jump from one focal point to another, where each focal point can be separated by several degrees. The second type of eye movement is referred to as a "micro-saccade," which is a rapid, simultaneous movement of both eyes between two foci that are close together. Micro-saccadic eye movements are small movements of 0.25° or less (1° or less when magnified in a low-resolution digital environment), which are often movements that scan a particular object in a scene. In other words, in the case of micro-saccadic eye movements, the eyes typically jump from one region to another nearby region, which can form part of the same overall focal point. The third type of eye movement is referred to as "vergence," which is a simultaneous movement of both eyes in opposite directions to obtain or maintain single binocular vision on a particular focal point. Vergence eye movements include convergent eye movements and divergent eye movements, which are often related to the viewing distance of the eyes from a particular focal point. The fourth type of eye movement is referred to as "complete disengagement," which is a movement of both eyes away from one or more focal points of interest (e.g., in this case, the display device and the camera).
[0023] In various embodiments, the present technology adjusts the eye gaze of a presenter by controlling the presenter's saccadic eye movements, micro-saccadic eye movements, and / or vergence eye movements, while recognizing and allowing for the presenter's complete eye disengagement. The overall goal of this process is to produce eye movements that closely mimic the natural eye movements produced by the human vestibulo-ocular reflex (VOR), which is a reflex that stabilizes eye gaze during head movement. Moreover, by mimicking the presenter's natural VOR in this way, the present technology produces a synthetic eye gaze that appears natural, focused, and dynamic.
[0024] Compared to previous techniques for modifying eye gaze, the present techniques provide several improvements. For example, the present techniques use a trained machine learning model to provide eye gaze synthesis and redirection that does not rely on the sequential selection of previously acquired template images or image sequences. In addition to simplifying the overall process, this has the added benefit of avoiding the stiff, eerie appearance that often results from techniques that correct eye gaze using template images. As another example, the present techniques are not limited to redirecting a presenter's eye gaze to a camera, but are able to aim eye gaze at any desired physical or virtual focal point. As another example, in contrast to previous techniques, the present techniques are able to work automatically without any personal user calibration. As another example, in contrast to previous techniques, the present techniques provide a sophisticated auto-on / off mechanism that prevents adjusting a presenter's eye gaze during periods when the presenter's eye motion and eye motion associated with reading are inconsistent and during periods when the presenter's eye motion appears to be completely detached from the camera and display device. As another example, the present techniques do not rely on the identification of specific eye contour points, but rather rely only on the identification of a general eye region. Moreover, as discussed above, the present techniques provide real-time eye gaze adjustment that preserves a presenter's energy and identity, thereby allowing the presenter to interact with a recipient in a more natural manner.
[0025] As a preliminary matter, some of the figures describe concepts in the context of one or more structural components, referred to as functionality, modules, features, elements, etc. The various components shown in the figures can be implemented in any manner, e.g., by software, hardware (e.g., discrete logic components, etc.), firmware, etc., or any combination of these implementations. In one embodiment, the various components can reflect the use of corresponding components in an actual implementation. In other embodiments, any single component illustrated in the figures can be implemented by a number of actual components. The depiction of any two or more separate components in the figures can reflect different functionality performed by a single actual component.
[0026] Other figures describe concepts in the form of flowcharts. In this form, certain operations are described as constituting distinct blocks performed in a particular order. These implementations are exemplary and non-limiting. Certain blocks described herein can be grouped together and performed in a single operation, certain blocks can be broken apart into multiple component blocks, and certain blocks can be performed in an order different from that described herein (including in a parallel manner). Blocks shown in flowcharts can be implemented by software, hardware, firmware, etc., or any combination of these implementations. As used herein, hardware can include computing systems, discrete logic components such as application specific integrated circuits (ASICs), etc., and any combination thereof.
[0027] Regarding the terminology, the phrase "configured to" encompasses any manner in which any kind of structural component can be constructed to perform the identified operation. A structural component can be configured to perform the operation using software, hardware, firmware, etc., or any combination thereof. For example, the phrase "configured to" can refer to a logic circuit structure of a hardware element used to implement associated functionality. The phrase "configured to" can also refer to a logic circuit structure of a hardware element used for a code design to implement associated functionality of firmware or software. The term "module" refers to a structural element that can be implemented using any suitable hardware (e.g., a processor, etc.), software (e.g., an application, etc.), firmware, or any combination of hardware, software, and firmware.
[0028] The term "logic" encompasses any functionality used to perform a task. For example, each operation illustrated in a flowchart corresponds to the logic used to perform that operation. Operations can be performed using software, hardware, firmware, or any combination thereof.
[0029] As used herein, the terms “component,” “system,” “client,” etc., are intended to refer to computer-related entities that can be hardware, software (e.g., software in execution), and / or firmware, or a combination thereof. For example, a component can be a process, object, executable, program, function, library, subroutine, and / or a computer or a combination of software and hardware running on a processor. By extension, both an application running on a server and the server itself can be components. One or more components may reside within a process, and components may reside on a single computer and / or be distributed across two or more computers.
[0030] Furthermore, the claimed subject matter can be implemented as a method, apparatus, or article of art using standard programming and / or engineering techniques to produce software, firmware, hardware, or any combination thereof for controlling a computer to implement the disclosed subject matter. As used herein, the term "article of art" is intended to cover a computer program accessible from any tangible computer-readable storage medium.
[0031] Furthermore, as used herein, the term "computer-readable storage medium" refers to a product of manufacturing. Generally, a computer-readable storage medium is used to host, store, and / or transport computer-executable instructions and data for subsequent retrieval and / or execution by a computing system. When computer-executable instructions hosted or stored on a computer-readable storage medium are executed by a processor of a computing system, they perform a process that causes, configures, and / or adapts the computing system performing the process to perform various steps, processes, routines, methods, and / or functions, including the steps, processes, routines, methods, and / or functions described herein. Examples of computer-readable storage media include, but are not limited to, optical storage media such as Blu-ray discs, digital video discs (DVDs), compact discs (CDs), optical disc cartridges, and the like, magnetic storage media such as hard disk drives, floppy disks, magnetic tape, and the like, memory storage devices such as random access memory (RAM), read-only memory (ROM), memory cards, thumb drives, and the like, and cloud storage such as online storage services. Computer-readable storage media can deliver computer-executable instructions to a computing system for execution via various transmission means and media including carrier waves and / or propagated signals. However, for purposes of the present disclosure, the term "computer-readable storage medium" specifically refers to non-transitory forms of computer-readable storage media and expressly excludes carrier waves and / or propagated signals.
[0032] Network environments and computing systems for implementing the eye gaze adjustment techniques described herein
[0033] Figure 1 is a block diagram of an example network environment 100 suitable for implementing the eye gaze adjustment techniques described herein. The example network environment 100 includes computing systems 102, 104, and 106. Each of the computing systems 102, 104, and 106 corresponds to one or more users, such as users 108, 110, and 112, respectively.
[0034] In various embodiments, each computing system 102, 104, and 106 is connected to a network 114. The network 114 can be a packet-based network, such as the Internet. Further, in various embodiments, each computing system 102, 104, and 106 includes a display device 116, 118, and 120, respectively, and includes a camera 122, 124, and 126, respectively. The cameras can be built-in components of the computing systems, such as the camera 122 corresponding to the computing system 102 being a tablet computer, and the camera 126 corresponding to the computer system 106 being a laptop computer. Alternatively, the cameras can be external components of the computing systems, such as the camera 124 corresponding to the computing system 104 being a desktop computer. Further, it should be appreciated that the computing systems 102, 104, and / or 106 can take various other forms, such as the form of a mobile phone (e.g., a smartphone), a wearable computing system, a television (e.g., a smart television), a set-top box, and / or a game console. Further, the particular embodiments of the display devices and / or cameras can be tailored to each particular type of computing system.
[0035] At any given time, one or more users 108, 110, and / or 112 can be in communication with any number of other users 108, 100, and / or 112 via video streams transmitted across the network 114. Further, in various embodiments, the video communication can include a particular user (sometimes referred to herein as a "presenter") presenting information to one or more remote users (sometimes referred to herein as "receivers"). For example, if the user 108 is acting as the presenter, the presenter can present such information by reading text from the display device 116 of the computing system 102. In such embodiments, the computing system 102 can be configured to implement the eye gaze adjustment techniques described herein. Thus, the remote users 110 and / or 112 acting as receivers can perceive the adjusted eye gaze of the presenter via their display devices 118 and / or 120, respectively. Based on the adjusted eye gaze of the presenter, the receivers can perceive the eyes of the presenter as having a natural, focused appearance, rather than the displaced appearance typically associated with reading. Reference is made to Figure 2 Further details are described regarding exemplary implementations of the computing system of the presenter (and associated eye gaze adjustment capabilities).
[0036] It should be appreciated that Figure 1 the block diagram of FIG. 1 is not intended to indicate that the network environment 100 includes only the components shown in FIG. 1 in every possible implementation. Figure 1All components shown in the network environment 100. For example, the exact number of users and / or computing systems can vary depending on the specifics of the implementation. Moreover, the designation of each user as a presenter or recipient can change continuously as the video communication proceeds, depending on which user is currently acting as the presenter. Thus, any number of computing systems 102, 104, and / or 106 can be configured to implement the eye gaze adjustment techniques described herein.
[0037] In some embodiments, the eye gaze adjustment techniques are provided by a video streaming service that is configured for each computing system on an as-needed basis. For example, the eye gaze adjustment techniques described herein can be provided as a software licensing and delivery model, sometimes referred to as software as a service (SaaS). In such embodiments, a third-party provider can provide eye gaze adjustment capabilities to consumer computing systems, such as the computing system of a presenter, via a software application running on a cloud infrastructure.
[0038] Moreover, in some embodiments, one or more of the computing systems 102, 104, and / or 106 can have multiple users at any given point in time. Thus, the eye gaze adjustment techniques described herein can include a dominant face tracking function configured to determine which user is the dominant speaker, and thus the current presenter, at any given point in time. Additionally or alternatively, the dominant face tracking function can be configured to treat any (or all) users of a particular computing system as presenters at any given point in time.
[0039] Figure 2 is a block diagram of an example computing system 200 configured to implement the eye gaze adjustment techniques described herein. In various embodiments, the example computing system 200 embodies one or more of the computing systems 102, 104, and 106 described with reference to the network environment 100. Figure 1 In particular, the example computing system 200 embodies a computing system of a user (i.e., a presenter) participating in a video communication involving reading of text from a display device.
[0040] The example computing system 200 includes one or more processors (or processing units), such as the processor 202 and the memory 204. The processor 202 and the memory 204, along with other components, are interconnected by a system bus 206. The memory 204 typically (but not always) includes both a volatile memory 208 and a non-volatile memory 210. The volatile memory 208 retains or stores information as long as power is supplied to the memory. In contrast, the non-volatile memory 210 is able to store (or retain) information even when power is not supplied. Generally, RAM and CPU cache memory are examples of volatile memory 208, while ROM, solid state memory devices, memory storage devices, and / or memory cards are examples of non-volatile memory 210.
[0041] The processor 202 executes instructions retrieved from the memory 204 and / or from a computer-readable storage medium, such as the computer-readable storage medium 212, to implement various functionality, such as the functionality of the eye gaze adjustment techniques described herein. Moreover, the processor 202 can include any of a wide variety of available processors, such as single processors, multi-processors, single-core units, and / or multi-core units.
[0042] The example computing system 200 also includes a network communication component 214 for interconnecting the computing system 200 with other devices and / or services through a computer network, including other computing systems, such as any of the computing systems 102, 104, and / or 106 described with reference to Figure 1 The network communication component 214 (which is sometimes referred to as a network interface card (NIC)) communicates over the network (e.g., the network 114 described with reference to Figure 1 The network communication component 214 (which is sometimes referred to as a network interface card (NIC)) communicates over the network (e.g., the network 114 described with reference to
[0043] The computing system 200 also includes an input / output (I / O) subsystem 216. The I / O subsystem 216 includes a group of hardware, software, and / or firmware components that enable or facilitate communication between a user of the computing system 200 and the computing system 200 processor. Indeed, via the I / O subsystem 216, a user can provide input via one or more input channels, such as, by way of example and not limitation, one or more touchscreens / haptic input devices, one or more buttons, one or more pointing devices, one or more audio input devices, and / or one or more video input devices (such as a camera 218). Moreover, a user can provide output via one or more output channels, such as, by way of example and not limitation, one or more audio output devices, one or more haptic feedback devices, and / or one or more display devices (such as a display device 220).
[0044] In some embodiments, the display device 220 is a built-in display screen of the computing system 200. In other embodiments, the display device 220 is an external display screen. Moreover, in some embodiments, the display device is a touchscreen that serves as both an input device and an output device.
[0045] The camera 218 can be any suitable type of video recording device configured to capture a video stream of a user of the computing system 200. The video stream includes a series of video frames, where each video frame includes a sequence of images. In various embodiments, the camera 218 is located in the vicinity of the display device 220. For example, the camera 218 can be located near an edge of the display device 220, such as directly above or below the display device 220. Moreover, in various embodiments, the camera 218 has an outward-facing image capture component and is capable of capturing a front view of a user’s face as the user views the display device 220. The camera 218 can include, for example, a front-facing camera integrated into the computing system 200 or an external camera attached to the display device 220 in any suitable manner.
[0046] According to embodiments described herein, the computer-readable storage medium 212 includes an eye gaze adjustment module 222. The eye gaze adjustment module 222 includes computer-executable instructions that, when executed by the processor 202, cause the processor 202 to perform a method for adjusting the eye gaze of a user of the computing system 200. In various embodiments, the eye gaze adjustment module 222 receives images extracted from a video stream of the user captured by the camera 218. In some cases, the received images represent a video stream of the user reading text displayed on the display device 220, rather than looking directly at the camera 218. In such cases, the eye gaze adjustment module 222 generates a sequence of images (e.g., video frames) in which the eye gaze of the user has been adjusted to appear as if the user is looking directly at the camera 218. In various embodiments, this method for adjusting the eye gaze is performed in real-time, meaning that there is no significant latency between the recording of the video stream and the delivery of the video stream including the adjusted eye gaze to one or more remote computing systems. In other words, the eye gaze adjustment module 222 is configured to run at substantially the same rate as the frame rate of the camera, without any significant lag time.
[0047] In various embodiments, the eye gaze adjustment module 222 itself includes a plurality of sub-modules (not shown) for performing the method of adjusting the eye gaze. Such sub-modules can include, by way of illustration and not limitation, a face localization sub-module for detecting a face region of the user in an image captured by the camera 218; a facial feature localization sub-module for detecting a facial feature region of the user in the image based on the detected face region; a head pose estimation sub-module for estimating a head pose of the user based on the detected facial feature region; a camera orientation estimation sub-module for estimating an orientation of the camera based on the detected facial feature region; a full disengagement determination sub-module for determining whether the image represents the user being fully disengaged from the computing system; an eye localization sub-module for detecting an eye region of the user within the image based on the detected facial feature region; an eye region classification sub-module for determining whether the eye movement of the user is suboptimal; a required eye gaze determination sub-module for calculating a required eye gaze direction of the user based on the detected eye region; and an eye gaze synthesis sub-module for generating a gaze-adjusted image according to the required eye gaze direction of the user. Reference is made to Figures 4-9 Further details are further described in connection with the functionality of the eye gaze adjustment module 222 (and respective sub-modules) in performing the method for adjusting the eye gaze.
[0048] In various embodiments, the eye gaze adjustment module 222 described herein improves the video functionality provided by the camera 218 of the computing system 200 in several ways. For example, the eye gaze adjustment module 222 allows the user’s eye gaze to be redirected to the camera 218 or any other physical or virtual focal point, regardless of the positioning of the camera 218 relative to the computing system 200. This provides a considerable degree of freedom to the manufacturer and / or user of the computing system 200 with respect to the configuration of the camera 218. As another example, because the eye gaze adjustment module 222 uses a trained machine learning model that does not rely on successive selection of previously acquired template images to perform eye gaze synthesis and redirection, the eye gaze adjustment module 222 can significantly improve the speed of the computing system 200 compared to previous techniques for modifying eye gaze. For example, in some embodiments, the eye gaze adjustment module 222 generates gaze-adjusted images at substantially the same rate as the frame rate of the camera 218. As yet another example, the eye gaze adjustment module 222 allows the computing system 200 to automatically provide gaze-adjusted images (i.e., without requiring any individual user calibration), thereby significantly improving the user’s interaction with the computing system 200.
[0049] As noted herein, in some embodiments, the eye gaze adjustment module 222 does not adjust the user’s eye gaze to make it appear to be looking straight at the camera 218, but rather adjusts the user’s eye gaze to make it appear to be looking at another focal point of interest, such as a virtual focal point on the user’s display device 220. For example, if the video communication involves a presentation of a video to a remote user, the user’s eye gaze can be automatically directed to the portion of the display device 220 that is playing the video. As another example, if multiple remotely located users are participating in a video communication, the users’ respective display devices can be set to split-screen mode. In such situations, the users’ eye gaze can be automatically directed to the portion of the display device 220 that includes the particular remote user who is currently speaking to it. This can provide an important visual cue, thereby further enhancing the overall communication process.
[0050] In some embodiments, the eye gaze adjustment module 222 is used to perform a method for adjusting the eye gaze of a pre-recorded video stream (such as a presentation or a lecture of a pre-recorded event or television program) that is not immediately distributed to a remote computing system. In such embodiments, the video stream including the adjusted eye gaze can not be delivered to the remote computing system in real-time, but rather can be stored in memory (either locally (i.e., in the memory 204) or remotely (e.g., in the cloud)) for subsequent distribution.
[0051] In various embodiments, the eye gaze adjustment module 222 includes an automatic on / off mechanism capable of adjusting the user's eye gaze during periods when the user's eye gaze (and associated eye movements) is determined to be suboptimal, and preventing adjustment of the user's eye gaze during other periods. For example, the on / off mechanism might prevent adjustment of the user's eye gaze during periods when the user's eye movements and eye movements associated with reading are inconsistent, and during periods when the user's eye movements are completely detached from the camera 218 and display device 220. Furthermore, although the eye gaze adjustment module 222 is configured to operate independently of the user in most cases, in some embodiments, the eye gaze adjustment module 222 includes a user-selectable on / off mechanism, thereby allowing the user to prevent the eye gaze adjustment module 222 from performing any eye gaze adjustments during specific portions of the video stream. This can provide the user with the ability to maintain a certain appearance when the user deems the appearance of the text line appropriate for reading.
[0052] Figure 2 The block diagram is not intended to indicate what the computing system 200 should include. Figure 2 All components shown. Conversely, depending on the specific implementation details, computing system 200 may include... Figure 2 Fewer or additional components not shown. Furthermore, any functionality of the eye gaze adjustment module 222 may be implemented, in part or in whole, in the hardware and / or processor 202. For example, any functionality of the eye gaze adjustment module 222 may be implemented using an application-specific integrated circuit (ASIC), logic implemented in the processor 202, and / or any other suitable components or devices.
[0053] As described herein, in some embodiments, the eye gaze adjustment module 222 is provided as a software application licensed to a user and delivered to the user's computing system 200. As another example, in some embodiments, the eye gaze adjustment module 222 is provided as a cloud-based online video streaming service.
[0054] Furthermore, in some embodiments, the eye gaze adjustment module 222 can be implemented on the computing system of a remote user (i.e., the receiver). In such embodiments, the remote computing system can receive a video stream from the presenter's computing system via a network, and the eye gaze adjustment module 222 can adjust the presenter's eye gaze within the video stream before the receiver views the video stream on the receiver's display device. This can be performed in real time on the video stream or on a pre-recorded video stream at a later date.
[0055] Figure 3 It is a description Figure 2 The eye gaze adjustment module 222 can Figure 2A schematic diagram illustrating the implementation within the computer-readable storage medium 212. Items with similar numbering are as shown in the reference. Figure 2 As described. Figure 3 As shown, the eye fixation adjustment module 222 includes computer-readable data 300. The computer-readable data 300 constitutes a set of computer-executable instructions 302, which, when executed by the processor 202, cause the processor 202 to perform one or more methods 304 for adjusting eye fixation, such as referring to... Figure 4 , 5 Any of the exemplary processes 400, 500, and 600 described in section 6, and / or respectively refer to Figure 8 and Figure 9 The exemplary methods 800 and / or 900 are described.
[0056] The process and method for implementing the eye fixation technique described in this article
[0057] As a preliminary point, it should be noted that the exemplary processes 400, 500, and 600, and the exemplary methods 800 and 900 described below are implemented by a computing system, such as reference Figure 2 The described computing system 200 can form part of a network environment, such as as shown in the appendix. Figure 1 The network environment 100 is described. More specifically, exemplary processes 400, 500, and 600, as well as exemplary methods 800 and 900, can be implemented by a computing system of a user (i.e., a presenter) engaging in video communication with one or more remote users (i.e., receivers), wherein the video communication includes images in which the user's eye gaze and associated eye movements are suboptimal, such as images including eye movements corresponding to the user reading lines of text from a display device.
[0058] Figure 4 This is a schematic diagram illustrating an exemplary process 400 for calculating the desired eye gaze direction for a user. In various embodiments, the exemplary process 400 is performed by a processor using a trained neural network. Specifically, the neural network may be trained to prepare and analyze input video image data as part of an inference process for calculating the desired eye gaze direction for the user.
[0059] As depicted in box 402, process 400 begins with input image data, or in other words, with receiving a video stream of images from a user of the computing system. In some embodiments, this includes capturing the video stream using the computing system's camera, while in other embodiments it involves receiving a video stream from a remote computing system over a network.
[0060] As depicted by block 404, face localization can be performed to detect a face region of the user within the image. As depicted by block 406, face feature localization can be performed to detect face feature regions of the user within the image. As depicted by blocks 408 and 410, respectively, the neural network can use the detected face feature regions as input in order to implicitly determine a head pose of the user and an orientation of the camera (and, generally, the computing system). This information can then be used to determine whether the image represents the user being fully disengaged from the computing system. In various embodiments, this involves determining whether the head of the user is oriented too far in one direction relative to the camera orientation (e.g., based on an angular coordinate of the head pose of the user relative to the camera orientation).
[0061] As depicted by block 412, if the image does not represent the user being fully disengaged from the computing system, eye localization can be performed to detect eye regions of the user within the image based on the detected face feature regions. Optionally, in some embodiments, the neural network can then analyze the detected eye regions to determine whether the eye movements of the user represent suboptimal eye movements, such as saccadic eye movements associated with reading. In such embodiments, if the eye movements of the user do represent suboptimal eye movements, a required eye gaze direction of the user is computed based on the detected eye regions, as depicted by block 414. In other embodiments, the required eye gaze direction of the user is computed automatically without determining whether the eye movements of the user represent suboptimal eye movements. This can be particularly useful for embodiments in which the user has manually moved the on / off mechanism of the eye gaze adjustment module to the "on" position. Moreover, in various embodiments, the required eye gaze direction can be computed such that the eye gaze of the user is directed toward the camera or toward a physical or virtual focal point of interest, such as a virtual focal point located on a display device of the user.
[0062] Figure 5 is a schematic diagram depicting an example process 500 for generating a gaze-adjusted image based on a required eye gaze direction of the user. Like-numbered items are as described with reference to Figure 4 In various embodiments, the example process 500 is performed by a processor using a trained neural network. Specifically, the neural network can be a trained image generator 502. In various embodiments, the trained image generator 502 is a generative model within a generative adversarial network (GAN) that is trained in conjunction with a discriminator model, as described in further detail with reference to Figure 6
[0063] In various embodiments, the trained image generator 502 is configured to combine an image of a user's eye region (as depicted by block 412) with a desired eye gaze direction (as depicted by block 414) to generate a gaze-adjusted image (as depicted by block 504). More specifically, the image generator 502 can generate a gaze-adjusted image by (1) analyzing an image to determine a user's natural saccadic eye movement, natural microsaccadic eye movement, and / or natural vergence eye movement within the image; (2) comparing the user's eye gaze within the image to the user's desired eye gaze direction; and (3) modifying or adjusting the saccadic eye movement, the microsaccadic eye movement, or the vergence eye movement of the user within the image to produce the gaze-adjusted image. In some embodiments, the modified saccadic eye movement is used, for example, to accommodate changes in the user's context, while the modified microsaccadic eye movement is used, for example, to add subtle noise to the eye gaze, which can make the synthesized eye gaze appear more natural.
[0064] According to various embodiments described herein, a user's eye movements can be adjusted, modified, simulated, and / or synthesized in any suitable manner to produce a desired gaze-adjusted image. For example, in some embodiments, a particular eye movement is adjusted by pairing an input image with a desired output image to, for example, make the eye movement appear less pronounced or extreme. As one specific example, a saccadic eye movement and / or a microsaccadic eye movement is such that the eye only moves half to the left and / or right. Additionally or alternatively, a particular eye movement can be adjusted by using standard Brownian motion techniques to artificially generate new eye movements that still appear natural and dynamic. Additionally or alternatively, a particular eye movement can be rendered entirely by the trained image generator 502, independent of a user's natural eye movements. For example, the trained image generator 502 can synthesize an eye movement that makes a user's eye appear to move in a natural manner, even when the user's eye gaze is overly locked onto one portion of a display device.
[0065] Further, in some embodiments, the image generator 502 is configured to analyze the gaze-adjusted image generated at block 504 to determine whether the original image within the video stream should be replaced by the gaze-adjusted image. For example, the image generator 502 can include an algorithm for assigning a confidence value (i.e., a non-binary metric) to the gaze-adjusted image (and / or to specific pixels or portions within the gaze-adjusted image). If the confidence value is above a specified threshold, the original image within the video stream can be replaced by the gaze-adjusted image. However, if the confidence value is below the specified threshold, the image generator 502 can determine that the entire eye gaze adjustment process has failed, at which point the entire process can be aborted or repeated.
[0066] Figure 6 is a schematic diagram depicting a process for training a reference Figure 5 A schematic diagram of an example process 600 for the image generator 502 described. The example process 600 can be performed by a processor using a trained image discriminator 602. In various embodiments, the trained image discriminator 602 is a discriminator model used to train the image generator 502 within a generative adversarial network (GAN). As a more specific example, the image discriminator 602 can be a standard convolutional neural network.
[0067] In various embodiments, a plurality of gaze-adjusted images are generated by the image generator 502 during a training phase. These generated gaze-adjusted images are then input into the image discriminator 602 along with corresponding target images, as depicted by blocks 604 and 606, respectively. After a comparison process, the image discriminator 602 outputs a truth value of "true" (as shown by block 608) or "false" (as shown by block 610) for each gaze-adjusted image. This can be accomplished, for example, by using the image discriminator as a classifier to distinguish between two sources, i.e., true images and false images.
[0068] In various embodiments, if the image discriminator 602 assigns a truth value of "false" to a gaze-adjusted image, the image discriminator 602 has identified a flaw in the operation of the image generator. As a result, the image generator 502 can analyze the output from the image discriminator 602 and then update itself, e.g., by adjusting its parameters, to produce more realistic gaze-adjusted images. Moreover, this training process 600 can continue until a predetermined number (or percentage) of the gaze-adjusted images generated by the image generator 502 are classified as "true." Once this occurs, the image generator 502 has converged, and the training process is complete. At this point, the image generator 502 has been trained to produce gaze-adjusted images that are indistinguishable from true images, and thus, the image generator 504 is ready to be used in the eye gaze adjustment techniques described herein.
[0069] Figure 7A is a schematic diagram of an unadjusted image 700 that a user's camera can capture when the user is reading text from a display device. As shown, the unadjusted image 700 can be perceived as if the user's eyes are focused on a point that is below the level of the user's eyes. Figure 7A This is especially true in the case where the user's camera is located directly above the display device that the user is reading.
[0070] Figure 7B is a schematic diagram of another unadjusted image 702 that a user's camera can capture when the user is reading text from a display device. As shown, the unadjusted image 702 can be perceived as if the user's eyes are focused on a point that is above the level of the user's eyes. Figure 7BAs shown, the image 702 without gaze adjustment can be perceived as if the user's eyes are overly locked onto one portion of the user's display device, which can give the user's eyes an unnatural appearance.
[0071] Figure 7C is a schematic diagram of a gaze-adjusted image 704 that can be generated in accordance with the eye gaze adjustment techniques described herein. In particular, as shown, the eye gaze adjustment techniques described herein generate an adjusted eye gaze that looks natural and focused. Figure 7C Further, the eye gaze adjustment techniques described herein also produce natural and authentic eye movements that preserve the user's energy and identity.
[0072] Figure 8 is a process flow diagram of a method 800 for adjusting a user's eye gaze in a video stream. In various embodiments, the method 800 is performed at substantially the same rate as the frame rate of a camera used to capture the video stream. Further, in various embodiments, the method 800 is performed using one or more trained neural networks, such as the neural networks described with reference to Figures 4-6 FIG. 6.
[0073] The method 800 begins at block 802. At block 804, a video stream comprising an image of a user is captured by a camera. At block 806, a facial region of the user is detected within the image. At block 808, a facial feature region of the user is detected within the image based on the detected facial region.
[0074] At block 810, it is determined whether the image represents the user being fully disengaged from a computing system based on the detected facial feature region. In some embodiments, this includes estimating a head pose of the user based on the detected facial feature region; estimating an orientation of the camera based on the detected facial feature region; and determining whether the image represents the full disengagement of the user from the computing system based on the detected facial feature region, the estimated head pose of the user, and the estimated orientation of the camera.
[0075] If the image does represent the user being fully disengaged from the computing system, the method 800 ends at block 812. If the image does not represent the user being fully disengaged from the computing system, the method 800 proceeds to block 814 at which an eye region of the user is detected within the image based on the detected facial feature region.
[0076] At block 816, a required eye gaze direction of the user is computed based on the detected eye region. In some embodiments, this includes using the detected eye region, an estimated head pose of the user, and an estimated orientation of the camera to compute the required eye gaze direction of the user. Further, in various embodiments, this includes computing the required eye gaze direction of the user such that the user’s eye gaze is toward the camera, or computing the required eye gaze direction of the user such that the user’s eye gaze is toward a point of interest located on a display device of the computing system.
[0077] At block 818, a gaze-adjusted image is generated based on the required eye gaze direction of the user. In various embodiments, the gaze-adjusted image includes at least one of a saccadic eye movement, a microsaccadic eye movement, or a vergence eye movement. In various embodiments, the gaze-adjusted image is generated by comparing an original image and the required eye gaze direction via a neural network acting as an image generator, as described with reference to Figure 5 Further, in various embodiments, the gaze-adjusted image is generated by: analyzing the image to determine a saccadic eye movement, a microsaccadic eye movement, or a vergence eye movement of the user within the image; comparing an eye gaze of the user within the image to the required eye gaze direction of the user; and adjusting the saccadic eye movement, the microsaccadic eye movement, or the vergence eye movement of the user within the image to produce the gaze-adjusted image.
[0078] In various embodiments, the gaze-adjusted image is generated using an image generator that can be trained using an image discriminator within a generative adversarial network (GAN). Specifically, in some embodiments, the image generator is trained prior to performing the method 800, where training the image generator includes: (1) inputting a plurality of target images and a plurality of gaze-adjusted images generated by the image generator into an image discriminator; (2) comparing the target images and the gaze-adjusted images using the image discriminator; (3) assigning a truth value of “true” or “false” to each of the gaze-adjusted images; and (4) updating the image generator in response to assigning the truth value of “false” to any of the gaze-adjusted images.
[0079] At box 820, the image in the video stream is replaced with the fixation-adjusted image, and then the method ends at box 822. In some embodiments, the generated fixation-adjusted image is analyzed to assign a confidence value to the fixation-adjusted image. In such embodiments, if the confidence value is higher than a specified threshold, the image in the video stream can be replaced by the fixation-adjusted image, and if the confidence value is lower than the specified threshold, the image in the video stream may not be replaced by the fixation-adjusted image. Furthermore, in various embodiments, the processor can automatically monitor whether the user-selectable on / off mechanism has moved to an "on" or "off" position, and if the user-selectable on / off mechanism has moved to an "off" position, it prevents the image in the video stream from being replaced with the fixation-adjusted image.
[0080] In some embodiments, the video stream includes images of multiple users of the computing system. In such embodiments, method 800 can be performed simultaneously for each user within the image. Alternatively, method 800 can be performed for the user currently presenting information. For example, this could include detecting facial regions for each user within the image, detecting facial feature regions for each user within the image based on the detected facial regions, and analyzing the detected facial feature regions to determine which user is the current presenter. Once the current presenter is identified, the remainder of the method can be performed to generate a gaze-adjusted image of the current presenter.
[0081] Figure 9 This is a process flow diagram for another method used to adjust a user's eye gaze in a video stream. Similar numbered items are shown in the reference. Figure 8 Method 800 describes. Figure 9 Method 900 is similar Figure 8 Method 800. However, in Figure 9 In the exemplary embodiment shown, method 900 includes an additional step for determining whether a user's eye movements represent shifting eye movements associated with reading, as shown in box 902. Figure 9 As shown, this can be performed after the user's eye region is detected within the image at box 816, but before the user's eye gaze direction is detected at box 816. Furthermore, in such embodiments, if the user's eye movements do not represent moving eye movements associated with reading, method 900 terminates at box 904. Conversely, if the user's eye movements do represent displaced eye movements associated with reading, the method proceeds to box 816, where the desired eye gaze direction for the user is calculated based on the detected eye region.
[0082] In various embodiments, if the user’s eye movements indicate that the user is not reading the line of text, the inclusion of this additional step in the method 900 allows the eye gaze adjustment process to be automatically terminated. Moreover, it should be noted that block 902 of the method 900 can be altered to make any determination as to whether the user’s eye gaze (and associated eye movements) is optimal or suboptimal. For example, in some embodiments, block 902 of the method 900 additionally or alternatively includes determining whether the user’s eye gaze is too locked onto one portion of the display device. If the user’s eye gaze is too locked onto one portion of the display device, the method 900 can proceed to block 816. Otherwise, the method 900 can end at block 904.
[0083] It should be noted that while the methods and processes described herein are generally expressed in terms of discrete steps, these steps should be viewed as being essentially logical and can or can not correspond to any particular actual and / or discrete steps of a given implementation. Moreover, unless otherwise stated, the order in which the steps are demonstrated in various methods and processes should not be interpreted as the only order in which the steps can be performed. Furthermore, in some cases, some of the steps can be combined and / or omitted. Those skilled in the art will recognize that the logical demonstration of steps is sufficient to guide the performance of aspects of the claimed subject matter, regardless of any particular development or coding language in which the logical instructions / steps are encoded.
[0084] Of course, while the methods and processes described herein include various novel features of the disclosed subject matter, other steps (not listed) can also be performed in carrying out the subject matter when the methods and processes are implemented. Those skilled in the art will understand that the logical steps of these methods and processes can be combined together or split into additional steps. The steps of the above-described methods and processes can be performed in parallel or serially. Generally, but not exclusively, the functionality of a particular method or process is embodied in software (e.g., applications, system services, libraries, etc.) executed on one or more processors of a computing system. Moreover, in various embodiments, all or some of the various methods and processes can also be embodied in executable hardware modules on a computing system, including but not limited to system-on-chips (SoCs), codecs, specially designed processors, and / or logic circuits, etc.
[0085] As suggested above, each of the methods or processes described herein is typically embodied in computer-executable instructions (or code) modules that are executed by a computer. Specifically, as suggested above, each method or process is a sequence of computer-implemented acts. The acts in each method or process typically represent steps implemented by the computer executable instructions, including by individual routines, functions, loops, selectors, and switches (such as if-then and if-then-else statements), assignments, arithmetic computations, and so on, which collectively integrate the components of the computer system into a programmable device that carries out the particular method or process. As indicated above, however, the exact implementation of the executable statements of each method or process is based on a variety of implementation-specific configurations and decisions, including the programming language, compiler, target processor, operating environment, and linking or binding operation. Those skilled in the art will appreciate that the logical steps identified in these methods and processes can be implemented in a number of ways, and that the logical descriptions above suffice to implement similar results.
[0086] While various novel aspects of the disclosed subject matter have been described above, it should be understood that they have been presented by way of example only, and not limitation. Various changes in form and details can be made to the various aspects without departing from the scope of the disclosed subject matter.
[0087] Examples of Current Technology
[0088] Example 1 is a computing system. The computing system includes a camera to capture a video stream including an image of a user of the computing system. The computing system also includes a processor to execute computer-executable instructions that cause the processor to receive the image of the user from the camera to detect a facial region of the user within the image and to detect a facial feature region of the user within the image based on the detected facial region. The computer-executable instructions further cause the processor to determine whether the image represents a complete disengagement of the user from the computing system based on the detected facial feature region and to detect an eye region of the user within the image based on the detected facial feature region if the image does not represent the complete disengagement of the user from the computing system. The computer-executable instructions further cause the processor to compute a required eye gaze direction of the user based on the detected eye region, to generate a gaze-adjusted image based on the required eye gaze direction of the user, wherein the gaze-adjusted image includes at least one of a saccadic eye movement, a microsaccadic eye movement, or a vergence eye movement, and to replace the image within the video stream with the gaze-adjusted image.
[0089] Example 2 includes the computing system of Example 1, wherein the computer- executable instructions further cause the processor to generate the gaze-adjusted image using a trained image generator, wherein the image generator is trained using an image discriminator within a generative adversarial network (GAN).
[0090] Example 3 includes the computing system of any of Examples 1-2, including or excluding optional features. In this example, the computer-executable instructions further cause the processor to determine whether the image represents the complete disengagement of the user from the computing system by estimating a head pose of the user based on the detected facial feature region, estimate an orientation of the camera based on the detected facial feature region, and determine whether the image represents the complete disengagement of the user from the computing system based on the detected facial feature region, the estimated head pose of the user, and the estimated orientation of the camera.
[0091] Example 4 includes the computing system of Example 3, including or excluding optional features. In this example, the computer-executable instructions further cause the processor to calculate an eye gaze direction required by the user based on the detected eye region, the estimated head pose of the user, and the estimated orientation of the camera.
[0092] Example 5 includes the computing system of any of Examples 1-4, including or excluding optional features. In this example, the video stream includes images of a plurality of users of the computing system, and the computer-executable instructions further cause the processor to generate a gaze-adjusted image of a current presenter by detecting a facial region of each user of the plurality of users within the image, detecting a facial feature region of each user of the plurality of users within the image based on the detected facial region, analyzing the detected facial feature region to determine which user of the plurality of users is the current presenter, and generating the gaze-adjusted image of the current presenter.
[0093] Example 6 includes the computing system of any of Examples 1-5, including or excluding optional features. In this example, the computer-executable instructions further cause the processor to automatically monitor whether a user-selectable on / off mechanism is moved to an "on" position or an "off position, and prevent the replacement of the image within the video stream with the gaze-adjusted image when the user-selectable on / off mechanism is moved to the "off position.
[0094] Example 7 includes the computing system of any of Examples 1-6, including or excluding optional features. In this example, the computer-executable instructions further cause the processor to calculate the eye gaze direction required by the user by calculating the eye gaze direction required by the user such that the eye gaze of the user is directed toward the camera, or calculating the eye gaze direction required by the user such that the eye gaze of the user is directed toward a point of interest located on a display device of the computing system.
[0095] Example 8 includes the computing system of any of Examples 1 to 7, including or excluding optional features. In this example, the computer-executable instructions further cause the processor to generate the gaze-adjusted image based on a desired gaze direction of the user by: (1) analyzing the image to determine at least one of a saccadic eye movement, a microsaccadic eye movement, or a vergence eye movement of the user within the image; (2) comparing an eye gaze of the user within the image to the desired gaze direction of the user; or (3) adjusting at least one of the saccadic eye movement, the microsaccadic eye movement, or the vergence eye movement of the user within the image to produce the gaze-adjusted image.
[0096] Example 9 describes a method for adjusting an eye gaze of a user within a video stream. The method includes capturing, via a camera of a computing system, a video stream that includes an image of a user of the computing system. The method further includes detecting, via a processor of the computing system, a facial region of the user within the image and detecting, based on the detected facial region, a facial feature region of the user within the image. The method further includes determining, based on the detected facial feature region, whether the image represents a complete disengagement of the user from the computing system and, if the image does not represent the complete disengagement of the user from the computing system, detecting, based on the detected facial feature region, an eye region of the user within the image. The method further includes calculating, based on the detected eye region, a desired gaze direction of the user, generating a gaze-adjusted image based on the desired gaze direction of the user, wherein the gaze-adjusted image includes at least one of a saccadic eye movement, a microsaccadic eye movement, or a vergence eye movement, and replacing the image within the video stream with the gaze-adjusted image.
[0097] Example 10 includes the method of Example 9, including or excluding optional features. In this example, the method includes analyzing the detected eye region to determine whether an eye movement of the user represents a disjunctive eye movement associated with reading. The method further includes calculating the desired gaze direction of the user if the eye movement of the user represents the disjunctive eye movement associated with reading or terminating the method if the eye movement of the user does not represent the disjunctive eye movement associated with reading.
[0098] Example 11 includes the method of any of Examples 9-10, including or excluding optional features. In this example, determining whether the image represents the complete disengagement of the user from the computing system based on the detected facial feature region includes: (1) estimating a head pose of the user based on the detected facial feature region; (2) estimating an orientation of a camera based on the detected facial feature region; and (3) determining whether the image represents the complete disengagement of the user from the computing system based on the detected facial feature region, the estimated head pose of the user, and the estimated orientation of the camera.
[0099] Example 12 includes the method of Example 11, including or excluding optional features. In this example, calculating an eye gaze direction required by the user based on the detected eye region includes using the detected eye region, the estimated head pose of the user, and the estimated orientation of the camera to calculate the eye gaze direction required by the user.
[0100] Example 13 includes the method of any of Examples 9-12, including or excluding optional features. In this example, the method includes generating the gaze-adjusted image using a trained image generator. The method further includes training the image generator prior to performing the method of claim 9, wherein training the image generator includes: inputting a plurality of target images and a plurality of gaze-adjusted images generated by the image generator into an image discriminator; comparing the target images and the gaze-adjusted images using the image discriminator; assigning a truth value of "true" or "false" to each of the gaze-adjusted images; and updating the image generator in response to assigning a truth value of "false" to any of the gaze-adjusted images.
[0101] Example 14 includes the method of any of Examples 9-13, including or excluding optional features. In this example, the video stream includes images of a plurality of users of the computing system, and the method is performed for a current presenter by: (1) detecting a facial region of each of the plurality of users within the images; (2) detecting a facial feature region of each of the plurality of users within the images based on the detected facial regions; (3) analyzing the detected facial feature regions to determine which of the plurality of users is the current presenter; and (4) performing the remaining portions of the method to generate a gaze-adjusted image of the current presenter.
[0102] Example 15 includes the method of any of Examples 9-14, including or excluding optional features. In this example, the method includes analyzing the generated gaze-adjusted image to assign a confidence value to the gaze-adjusted image. The method also includes replacing the image within the video stream with the gaze-adjusted image if the confidence value is above a specified threshold, and not replacing the image within the video stream with the gaze-adjusted image if the confidence value is below the specified threshold.
[0103] Example 16 includes the method of any of Examples 9-15, including or excluding optional features. In this example, computing a required eye gaze direction of the user includes computing a required eye gaze direction of the user such that the user's eye gaze is toward a camera, or computing a required eye gaze direction of the user such that the user's eye gaze is toward a point of interest located on a display device of the computing system.
[0104] Example 17 includes the method of any of Examples 9-16, including or excluding optional features. In this example, generating a gaze-adjusted image based on the required eye gaze direction of the user includes: (1) analyzing the image to determine at least one of a saccadic eye movement, a microsaccadic eye movement, or a vergence eye movement of the user within the image; (2) comparing the eye gaze of the user within the image to the required eye gaze direction of the user; or (3) adjusting at least one of the saccadic eye movement, the microsaccadic eye movement, or the vergence eye movement of the user within the image to produce the gaze-adjusted image.
[0105] Example 18 is a computer-readable storage medium. A computer-readable storage medium including computer-executable instructions that, when executed by a processor of a computing system, cause the processor to receive a video stream containing an image of a user, to detect a facial region of the user within the image and to detect a facial feature region within the image based on the detected facial region. The computer-executable instructions also cause the processor to determine, based on the detected facial feature region, whether the image represents a complete disengagement of the user from the computing system, and to detect, based on the detected facial feature region, an eye region of the user within the image if the image does not represent the complete disengagement of the user from the computing system. The computer-executable instructions also cause the processor to compute a required eye gaze direction of the user based on the detected eye region, to generate a gaze-adjusted image based on the required eye gaze direction of the user, wherein the gaze-adjusted image includes at least one of a saccadic eye movement, a microsaccadic eye movement, or a vergence eye movement, and to replace the image within the video stream with the gaze-adjusted image.
[0106] Example 19 includes the computer-readable storage medium of Example 18, including or excluding optional functionality. In this example, the computer-executable instructions further cause the processor to generate the gaze-adjusted image using a trained image generator, wherein the image generator is trained using an image discriminator within a GAN.
[0107] Example 20 includes the computer-readable storage medium of any of Examples 18-19, including or excluding optional features. In this example, the computer-executable instructions further cause the processor to determine whether the image represents the complete disengagement of the user from the computing system based on the detected facial feature region by estimating a head pose of the user based on the detected facial feature region, estimating an orientation of a camera based on the detected facial feature region, and determining whether the image represents the complete disengagement of the user from the computing system based on the detected facial feature region, the estimated head pose of the user, and the estimated orientation of the camera. In addition, the computer-executable instructions further cause the processor to compute an eye gaze direction required by the user using the detected eye region, the estimated head pose of the user, and the estimated orientation of the camera.
[0108] In particular, with respect to the various functions performed by the above-described components, devices, circuits, systems, etc., the terminology used to describe such components, including references to "means," is intended in keeping with the designation of the specified function (e.g., functionally equivalent) of the components performing the functions illustrated in the exemplary aspects of the claimed subject matter described herein, unless otherwise indicated. In this regard, it should also be appreciated that the present innovations include systems and computer-readable storage media having computer-executable instructions for performing the acts and events of the various methods of the claimed subject matter.
[0109] There are a variety of ways to implement the claimed subject matter, e.g., an appropriate API, tool kit, driver code, operating system, control, standalone or downloadable software object, etc. which enables applications and services to use the techniques described herein. The claimed subject matter contemplates such software object, used as a performance of a software or hardware object operating according to the techniques described herein. Thus, the claimed subject matter can be embodied in a variety of ways, including completely hardware, partially hardware and partially software, and completely software.
[0110] The above-described systems have been described with reference to interactions between several components. It can be appreciated that these systems and components can include components or designated sub-components, certain designated components or sub-components, and additional components, and in various permutations and combinations of the above-described components. Sub-components can also be implemented as components communicatively coupled to other components rather than being included within parent components (hierarchical).
[0111] Additionally, it is noted that one or more components can also be combined into a single component providing aggregate functionality of the combined components, or the one or more components can also be divided into several separate sub-components, and that any one or more middle layers, such as an management layer, can be provided to communicatively couple to such sub-components in order to provide integrated functionality. Any components described herein can also interact with one or more other components not specifically described herein but generally known by those of skill in the art.
[0112] Additionally, although a particular feature of the claimed subject matter can have been disclosed with respect to only one of several implementations, other implementations can include the particular feature. Thus, it can be appreciated that features of the claimed subject matter can be combined with other features of the claimed subject matter, or other implementations, as can be desired. Furthermore, as used herein, the terms "includes," "containing," "having," "comprising," and the like are meant to be inclusive in a manner similar to the term "comprising" as an open transition word, without precluding any additional or other elements.
Claims
1. A computing system comprising: a camera to capture a video stream comprising an image of a user of the computing system; and a processor to execute computer-executable instructions that cause the processor to: receive the image of the user from the camera; detect a facial region of the user within the image; detect facial feature regions of the user within the image based on the detected facial region; determine whether the image represents the user being fully disengaged from the computing system based on the detected facial feature regions; if the image does not represent the user being fully disengaged from the computing system, detect eye regions of the user within the image based on the detected facial feature regions; analyze the detected eye regions to determine whether eye movements of the user represent saccadic eye movements associated with reading; and if the eye movements of the user represent the saccadic eye movements associated with reading, compute an eye gaze direction required by the user based on the detected eye regions; generate a gaze-adjusted image based on the eye gaze direction required by the user, wherein the gaze-adjusted image comprises at least one of a saccadic eye movement, a microsaccadic eye movement, or a vergence eye movement; and replace the image within the video stream with the gaze-adjusted image.
2. The computing system of claim 1, wherein, the computer-executable instructions further cause the processor to generate the gaze-adjusted image using a trained image generator, wherein the image generator is trained using an image discriminator within a generative adversarial network (GAN).
3. The computing system of claim 1, wherein, the computer-executable instructions further cause the processor to determine whether the image represents the user being fully disengaged from the computing system based on the detected facial feature regions by: estimating a head pose of the user based on the detected facial feature regions; estimating an orientation of the camera based on the detected facial feature regions; and determining whether the image represents the user being fully disengaged from the computing system based on the detected facial feature regions, the estimated head pose of the user, and the estimated orientation of the camera. the computer-executable instructions further cause the processor to compute the eye gaze direction required by the user based on the detected eye regions, the estimated head pose of the user, and the estimated orientation of the camera.
4. The computing system of claim 3, wherein, the video stream comprises images of a plurality of users of the computing system, and wherein the computer-executable instructions further cause the processor to generate a gaze-adjusted image of a current presenter by:
5. The computing system of claim 1, wherein, detecting a facial region of each user of the plurality of users within the images; detecting facial feature regions of each user of the plurality of users within the images based on the detected facial regions; analyzing the detected facial feature regions to determine which user of the plurality of users is the current presenter; and generating the gaze-adjusted image of the current presenter. the computer-executable instructions further cause the processor to:
6. The computing system of claim 1, wherein, automatically monitoring a user-selectable on / off mechanism to move to an "on" position or an "off position; and preventing replacement of the image within the video stream with the gaze-adjusted image when the user-selectable on / off mechanism is moved to the "off position.
7. The computing system of claim 1, wherein, The computer-executable instructions further cause the processor to calculate the gaze direction required of the user by: calculating the gaze direction required of the user such that the user's gaze is directed toward the camera; or calculating the gaze direction required of the user such that the user's gaze is directed toward a point of interest located on a display device of the computing system.
8. The computing system of claim 1, wherein, The computer-executable instructions further cause the processor to generate the gaze-adjusted image based on the gaze direction required of the user by: analyzing the image to determine at least one of a saccadic eye movement, a microsaccadic eye movement, or a vergence eye movement of the user within the image; comparing the gaze of the user within the image to the gaze direction required of the user; and adjusting at least one of the saccadic eye movement, the microsaccadic eye movement, or the vergence eye movement of the user within the image to produce the gaze-adjusted image.
9. A method for adjusting a gaze of a user within a video stream, comprising: capturing, via a camera of a computing system, a video stream comprising an image of a user of the computing system; detecting, via a processor of the computing system, a facial region of the user within the image; detecting, based on the detected facial region, a facial feature region of the user within the image; determining, based on the detected facial feature region, whether the image represents the user being fully disengaged from the computing system; if the image does not represent the user being fully disengaged from the computing system, detecting, based on the detected facial feature region, an eye region of the user within the image; analyzing the detected eye region to determine whether an eye movement of the user represents a disjunctive eye movement associated with reading; and if the eye movement of the user represents the disjunctive eye movement associated with reading, calculating a gaze direction required of the user based on the detected eye region; generating a gaze-adjusted image based on the gaze direction required of the user, wherein the gaze-adjusted image comprises at least one of a saccadic eye movement, a microsaccadic eye movement, or a vergence eye movement; and replacing the image within the video stream with the gaze-adjusted image.
10. The method of claim 9, wherein, determining, based on the detected facial feature region, whether the image represents the user being fully disengaged from the computing system comprises; estimating a head pose of the user based on the detected facial feature region; estimating an orientation of the camera based on the detected facial feature region; and determining, based on the detected facial feature region, the estimated head pose of the user, and the estimated orientation of the camera, whether the image represents the user being fully disengaged from the computing system.
11. The method of claim 10, wherein, computing the eye gaze direction required by the user based on the detected eye region includes using the detected eye region, an estimated head pose of the user, and an estimated orientation of the camera to compute the eye gaze direction required by the user.
12. The method of claim 9, wherein, comprises: generating the gaze-adjusted image using a trained image generator; and training the image generator prior to performing the method of claim 9, wherein training the image generator comprises: inputting a plurality of target images and a plurality of gaze-adjusted images generated by the image generator into an image discriminator; comparing the plurality of target images and the plurality of gaze-adjusted images using the image discriminator; assigning a truth value of "true" or "false" to each of the plurality of gaze-adjusted images; and updating the image generator in response to assigning a truth value of "false" to any one of the plurality of gaze-adjusted images.
13. The method of claim 9, wherein, the video stream comprises images of a plurality of users of the computing system, and wherein the method is performed for a current presenter by: detecting a face region of each of the plurality of users within the images; detecting a facial feature region of each of the plurality of users within the images based on the detected face region; analyzing the detected facial feature region to determine which of the plurality of users is a current presenter; and performing the remaining portion of the method to generate a gaze-adjusted image of the current presenter.
14. The method of claim 9, wherein, further comprising: analyzing the generated gaze-adjusted image to assign a confidence value to the gaze-adjusted image; if the confidence value is above a specified threshold, replacing the image within the video stream with the gaze-adjusted image; and if the confidence value is below the specified threshold, preventing the replacement of the image within the video stream with the gaze-adjusted image.
15. The method of claim 9, wherein, computing the eye gaze direction required by the user includes: computing the eye gaze direction required by the user such that the user's eye gaze is toward the camera; or computing the eye gaze direction required by the user such that the user's eye gaze is toward a point of interest located on a display device of the computing system.
16. The method of claim 9, wherein, generating the gaze-adjusted image based on the eye gaze direction required by the user includes: analyzing the image to determine at least one of a saccadic eye movement, a microsaccadic eye movement, or a vergence eye movement of the user within the image; comparing the eye gaze of the user within the image to the eye gaze direction required by the user; and adjusting at least one of the saccadic eye movement, the microsaccadic eye movement, or the vergence eye movement of the user within the image to produce the gaze-adjusted image.
17. A computer-readable storage medium comprising computer-executable instructions that, when executed by a processor of a computing system, cause the processor to: receive a video stream comprising an image of a user; detect a face region of the user within the image; detecting a facial feature region within the image based on the detected facial region; determining whether the image represents the user being fully disengaged from the computing system based on the detected facial feature region; if the image does not represent the user being fully disengaged from the computing system, detecting an eye region of the user within the image based on the detected facial feature region; analyzing the detected eye region to determine whether eye movement of the user represents saccadic eye movement associated with reading; and if the eye movement of the user represents the saccadic eye movement associated with reading, calculating an eye gaze direction required by the user based on the detected eye region; generating a gaze-adjusted image based on the eye gaze direction required by the user, wherein the gaze-adjusted image includes at least one of a saccadic eye movement, a microsaccadic eye movement, or a vergence eye movement; and replacing the image within the video stream with the gaze-adjusted image.
18. The computer-readable storage medium of claim 17, wherein, the computer-executable instructions further cause the processor to generate the gaze-adjusted image using a trained image generator, wherein the image generator is trained using an image discriminator within a generative adversarial network (GAN).
19. The computer-readable storage medium of claim 17, wherein, the computer-executable instructions further cause the processor to: determine whether the image represents the user being fully disengaged from the computing system based on the detected facial feature region by: estimating a head pose of the user based on the detected facial feature region; estimating an orientation of a camera based on the detected facial feature region; and determine whether the image represents the user being fully disengaged from the computing system based on the detected facial feature region, the estimated head pose of the user, and the estimated orientation of the camera; and calculate the eye gaze direction required by the user using the detected eye region, the estimated head pose of the user, and the estimated orientation of the camera.
Citation Information
Patent Citations
Eye gaze correction
CN107533640A