Real-time eye contact correction using neural network-based machine learning

A neural network-based machine learning solution addresses the inadequacies of existing eye contact correction technologies by providing real-time, high-quality eye contact correction in video conferencing, adapting input images to achieve natural-looking gaze using a trained neural network with encoding and decoding parts.

DE112017002119B4Active Publication Date: 2026-01-29INTEL CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
DE112017002119
Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Priority Date
2016-04-22
Filing Date
2017-03-17
Publication Date
2026-01-29
Estimated Expiration
2037-03-17

AI Technical Summary

Technical Problem

Current technologies for correcting eye contact issues in video telephony and video conferencing are inadequate, often requiring specialized hardware, are not fast enough for real-time implementation, and lack robustness in quality.

Method used

A neural network-based machine learning approach is employed to correct eye contact in real-time by training a neural network on a database of faces with known gaze angles, using a pre-trained classifier to predict motion vectors that adapt input images to achieve corrected eye contact, utilizing a deep neural network with encoding and decoding parts to generate compressed features and deform eye areas.

Benefits of technology

Provides high-quality, real-time eye contact correction that allows users to appear as if they are looking into the camera while maintaining natural eye contact, robust against local lighting changes, and adaptable to different viewing angles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Machine-implemented method for providing eye contact correction, comprising the following: Receiving several pairs of eye area training images (911), wherein first images (912) of the pairs of eye area training images (911) exhibit a viewing angle difference with respect to second images (913) of the pairs of eye area training images (911). Training a pre-trained neural network (700) based on encoding the first images (912) by the pre-trained neural network (700) to generate compressed training step features, decoding the compressed training step features to generate resulting first images (1413) corresponding to the first images (912), and evaluating the first images (912) and the resulting first images (1413); Obtaining a source image capturing a user (212) with a camera (104), Encoding an eye area (311) of the source image (212) via the pre-trained neural network (700) to generate compressed features (312) corresponding to the eye area (311) of the source image (212); Applying a pre-trained classifier to the compressed features (312) to determine a motion vector field (313) for the eye region (311) of the source image (212); and Deformation of the eye area (311) of the source image (212) based on the motion vector field (313) and integration of the deformed eye area (801) into a remaining part of the source image (212) to produce an image with corrected eye contact (215), wherein the image with corrected eye contact (215) gives the appearance that the user is looking into the camera (104).
Need to check novelty before this filing date? Find Prior Art

Description

CLAIMING A PRIORITY

[0001] This application claims priority over U.S. Patent Application No. 15 / 136,618, filed on April 22, 2016, entitled “EYE CONTACT CORRECTION IN REAL TIME USING NEURAL NETWORK BASED MACHINE LEARNING”, which is incorporated herein by reference in its entirety for all purposes. BACKGROUND

[0002] In video telephony or video conferencing applications on laptops or other devices, the camera, which captures video of the user, and the display, which provides the video to the person or people the user is speaking with, may be offset. For example, the camera may be positioned above the display. This may prevent participants in the video call from simultaneously looking at both the screen (to see the other participant) and the camera (which is desirable for establishing good, natural contact with the other participant).

[0003] Current technologies for correcting such eye contact correction problems are inadequate. For example, current technologies may not be fast enough to support real-time implementation, may require additional specialized camera hardware such as depth cameras or stereo cameras, or may not be sufficiently robust in terms of quality. In light of these and other considerations, it is clear that the improvements presented here are needed. Such improvements could become critical as the implementation of video telephony becomes increasingly widespread in a variety of contexts.

[0004] KONONENKO, Daniil; LEMPITSKY, Victor: Learning to look up: Realtime monocular gaze correction using machine learning. In: 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, 2015. pp. 4667-4675. DOI: 10.1109 / CVPR.2015.7299098, describes a method for eye contact correction in which a gaze direction is first identified, and then a gaze direction upwards is corrected using a random forest decision tree.

[0005] ZHOU, Erjin [et al.]: Extensive Facial Landmark Localization with Coarse-to-Fine Convolutional Network Cascade. In: 2013 IEEE International Conference on Computer Vision Workshops. IEEE, 2013. pp. 386-391. DOI: 10.1109 / ICCVW.2013.58, describes a method in which a neural network is used to determine feature vectors for an eye area.

[0006] US 2015 / 0 339 512 A1 describes a method and apparatus for correcting a gaze direction, comprising identifying first outer points of the eyes and determining second outer points and transforming the eye areas based on the second outer points.

[0007] US 2016 / 0 011 659 A1 describes a method in which a gaze direction is corrected, whereby the correction of the gaze direction is made using previously recorded images in which a user is looking in the correct direction. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] The material described herein is illustrated in the accompanying figures by way of example and without limitation. For the sake of simplicity and clarity, elements illustrated in the figures are not necessarily drawn to scale. For instance, the dimensions of some elements may be exaggerated relative to others for clarity. Furthermore, where deemed appropriate, reference symbols are repeated in the various figures to identify corresponding or analogous elements. In the figures: illustrative Fig. 1. An exemplary environment for providing eye contact correction; illustrative Fig. 2 an exemplary system for providing eye contact correction; illustrative Fig. 3 an exemplary eye contact correction module to provide eye contact correction; illustrative Fig.4 an example input image; illustrative Fig. 5 exemplary face detection data and exemplary face orientation points; illustrative Fig. 6 an example eye area; illustrative Fig. 7 an exemplary neural network; illustrative Fig. 8 an example of a corrected eye area; illustrative Fig. 9 an exemplary system for pre-training an eye contact correction classifier; illustrative Fig. 10 an exemplary source image and an exemplary target image; illustrative Fig. 11 exemplary target face orientation points; illustrative Fig. 12 an exemplary source eye area and an exemplary target eye area; illustrative Fig. 13 an exemplary likelihood map for an exemplary source eye area; illustrative Fig. 14 an exemplary training system for a neural network; illustrative Fig. 15 an exemplary training system for a neural network; illustrative Fig. 16 an exemplary training system for a neural network; is Fig. 17 a flowchart illustrating an exemplary process for pretraining a neural eye contact correction network and classifier; is Fig. 18 a flowchart illustrating an exemplary process for providing eye contact correction; is Fig. 19 an illustrative diagram of an exemplary system for providing eye contact correction; is Fig. 20 an illustrative diagram of an exemplary system; and illustrative Fig.21 an example device with a small form factor, arranged completely according to at least some implementations of the present disclosure. DETAILED DESCRIPTION

[0009] One or more embodiments or implementations will now be described with reference to the accompanying figures. Although specific configurations and arrangements are discussed, it is understood that this serves only illustrative purposes. Those skilled in the art will recognize that other configurations and arrangements can be used without altering the nature and scope of protection of the description. It will be obvious to those skilled in the art that the techniques and / or arrangements described herein may also be used in a range of systems and applications other than those described here.

[0010] Although the following description presents various implementations that may be manifested in architectures such as system-on-a-chip (SoC) architectures, implementations of the techniques and / or arrangements described here are not limited to specific architectures and / or data processing systems and may be implemented by any architecture and / or data processing system for similar purposes. For example, various architectures employing multiple integrated circuit (IC) chips and / or packages, and / or various data processing devices and / or consumer electronics (CE) devices, such as multifunction devices, tablets, smartphones, etc., may implement the techniques and / or arrangements described here.Although the following description may present numerous specific details, such as logic implementations, types and interrelationships of system components, logic partitioning / integration choices, etc., the claimed subject matter can also be exercised without such specific details. In other cases, some material, such as control structures and complete software instruction sequences, may not be shown in detail so as not to obscure the material disclosed herein.

[0011] The material disclosed herein may be implemented in hardware, firmware, software, or any combination thereof. The material disclosed herein may also be implemented as instructions stored on a machine-readable medium that can be read and executed by one or more processors. A machine-readable medium may include any medium and / or any mechanism for storing or transmitting information in a form readable by a machine (e.g., a data processing device). For example, a machine-readable medium may include read-only memory (ROM); random-access memory (RAM); magnetic disk storage media; optical storage media; flash memory devices; electrical, optical, acoustic, or other forms of propagating signals (e.g., carrier waves, infrared signals, digital signals, etc.); and other such media.

[0012] References in the patent specification to "exactly one implementation," "an implementation," "an example implementation," or examples or embodiments, etc., indicate that the described implementation may include a particular feature, structure, or property, but not every embodiment necessarily includes that particular feature, structure, or property. Furthermore, such phrases do not necessarily refer to the same implementation. When a particular feature, structure, or property is described in connection with an embodiment, it is also assumed that it is within the knowledge of those skilled in the art to implement such a feature, structure, or property in connection with other implementations, whether explicitly described herein or not.

[0013] Methods, devices, facilities, data processing platforms and objects relating to eye contact correction are described herein.

[0014] As described above, in video telephony or video conferencing applications on laptops or other devices such as mobile devices, the camera capturing video of the user and the display providing the video of the person or people the user is speaking with may be misaligned. This can result in an undesirable and unnatural portrayal of the user, such as the user not looking at the people they are speaking with. For example, the misalignment between the camera and the display can create a conflict where the user has to choose between looking at the display (which makes them appear to the person they are speaking with as looking away) and looking at the camera (and not at the person they are speaking with).The techniques discussed herein can provide real-time eye contact correction (or gaze correction) using neural network-based machine learning in a camera image signal processing unit for applications such as video telephony, video conferencing, or similar.

[0015] As discussed herein, embodiments can include a training component (e.g., performed offline) and a real-time component (e.g., performed during implementation or execution time) that provides eye contact correction. In the offline training component, a database of faces with a known gaze angle can be used to train a neural network and a model (e.g., a pre-trained classifier such as a random forest) that can be used to predict motion vectors (e.g., a motion vector field) capable of adapting the input image (e.g., with an uncorrected eye area) to a desired output image (e.g., with corrected eye contact) during real-time execution.For example, the pre-trained neural network can be a deep neural network, a convolutional neural network, or similar, comprising an encoding part that generates compressed features based on a given eye region and a decoding part that generates a resulting image based on the compressed features from the encoding part. The neural network can be trained based on error reduction between the given eye regions and resulting images from a training database, based on error reduction between vertically filtered images of the given eye regions and resulting images from the training database, or similar methods, as discussed below.

[0016] During the real-time phase, only the coding portion of the trained neural network can be implemented to generate compressed features based on an eye area of ​​a source image. For example, an eye area of ​​a source image can be encoded using a pre-trained neural network to generate compressed features corresponding to the eye area. The pre-trained classifier (e.g., a pre-trained classifier such as a random forest) can determine motion vectors corresponding to the features based on an optimal random forest tree search or similar method. The motion vector field can be used to deform the eye area to produce the desired output image with corrected eye contact. Such techniques can provide improved images with corrected eyes, robustness (e.g., robust against local lighting changes), real-time capability (e.g., faster operation), and flexibility (e.g.,(to provide training for different viewing angle corrections depending on the relative positioning of the camera to the display in the implementation device).

[0017] Embodiments disclosed herein can provide an image with corrected eye contact from a source image, allowing the user to look at the display. The image with corrected eye contact corrects the user's eye contact in such a way that it appears as if the user is looking into the camera. Such techniques can offer the advantage of allowing the user to see the user with whom they are speaking while maintaining the appearance of natural eye contact. For example, the image with corrected eye contact can be a frame from a video sequence of images or still images that can be encoded and transmitted from a local device (e.g., the user's device) to a remote device (e.g., the device with which the user of the local device is speaking).

[0018] In some embodiments, a source image can be obtained via a camera at a local device. Face detection and face orientation point detection can be provided on the source image to detect faces and landmarks of such faces, such as landmarks corresponding to eyes or eye areas. The source image can be cropped to create eye areas based on the eye orientation points, and the eye areas can be distorted and reinserted into the source image to provide the image with corrected eye contact. As previously discussed, such techniques can be provided on a sequence of images or single images, and the sequence can be encoded and transmitted. In some embodiments, an eye area orEye regions of the source image can be encoded by a pre-trained neural network to generate compressed features corresponding to each eye region of the source image. For example, the pre-trained neural network can be a deep neural network, a convolutional neural network, a fully connected layered neural network with convolutional layers, or similar. In one embodiment, the pre-trained neural network has four layers: a convolutional neural network layer, a second convolutional neural network layer, a fully connected layer, and a second fully connected layer, in that order. A pre-trained classifier can be applied to the feature sets to determine a motion vector field for each eye region of the source image, and the eye region(s) of the source image can be deformed and integrated into the source image based on the motion vector field(s).(into the remaining part of the source image that does not belong to the eye area) to generate the image with corrected eye contact. For example, the pre-trained classifier could be a pre-trained random forest classifier.

[0019] In some embodiments, during a training phase of the neural network and / or the pre-trained classifier, training can be performed based on a training set of images with a known viewing angle difference between them. For example, first images of pairs in the training set and second images of the pairs can have a known viewing angle difference between them. During training, the training set images can be cropped to eye regions, and the eye regions can be aligned to provide pairs of eye region training images. Based on these eye region training image pairs, a likelihood map can be generated for each pixel of each of the first images of the eye region training image pairs, such that the likelihood map includes a sum of absolute differences (SAD) for each of several candidate motion vectors corresponding to the pixel.For example, candidate motion vectors can be defined for a specific pixel of an eye area of ​​a first image, and a sum of the absolute differences can be determined for each of the candidate motion vectors by comparing pixels around the specific pixel (e.g., a window) with a window in the eye area of ​​the second image that is offset by the corresponding motion vector.

[0020] Furthermore, compressed training step features can be determined for the first images of pairs of eye-area training images (e.g., by encoding them with a neural network). The pre-trained neural network can be trained based on the encoding (to generate the compressed training step features), a decoding of the compressed training step features to generate resulting first images corresponding to the first images, and an evaluation of the first images and the resulting first images. For example, the pre-trained neural network can be trained to generate compressed features that, during decoding, can provide resulting images that attempt to match the images provided to the neural network.Evaluating the initial images and the resulting initial images can include assessing any error between the initial images and the resulting initial images, vertically filtering the initial images and the resulting initial images to generate vertically filtered initial images or resulting vertically filtered initial images, and determining any error between the vertically filtered initial images and the vertically filtered resulting initial images, or both. For example, training the neural network can reduce or minimize such errors.

[0021] Based on the likelihood maps and compressed features (e.g., as determined by the coding part of the neural network) for the eye regions in the training set, the pre-trained classifier can be trained to determine an optimal motion vector field based on the feature set during the implementation phase. As discussed, the pre-trained classifier can be a pre-trained random forest classifier. In such embodiments, each leaf of the random forest classifier's tree can represent a likelihood map (e.g., a SAD map) for each pixel in the eye region, and at each branch of the tree, the training process can minimize the entropy between the likelihood maps that led to that branch. The random forest classifier can include any suitable features.In some embodiments, the random forest classifier can have approximately 4 to 8 trees, each with approximately 6 to 7 levels.

[0022] As discussed, during the implementation phase, compressed features for an eye area of ​​a source image can be determined, and the pre-trained classifier can be applied to the eye area to generate a motion vector field corresponding to the eye area. This motion vector field can then be used to deform the eye area, and the deformed eye area can be integrated into the source image to produce an image with corrected eye contact.

[0023] Fig.Figure 1 illustrates an exemplary environment 100 for providing eye contact correction, arranged according to at least some implementations of the present disclosure. As shown, the environment 100 may include a device 101 with a display 102 and a camera 104, operated by a user (not shown) to conduct a video telephony session with a remote user 103. In the example of the Fig.Figure 1 shows the device 101 as a laptop computer. However, the device 101 can comprise any device with a suitable form factor, including a display 102 and a camera 104. For example, the device 101 can be a camera, a smartphone, an ultrabook, a tablet, a portable device, a monitor, a desktop computer, or the like. Furthermore, the eye contact correction techniques discussed can be provided in any suitable context, even when discussed in relation to video telephony or video conferencing.

[0024] During a video call or similar activity, images and / or video of the user of the device 101 can be captured by the camera 104, and, as shown, a remote user 103 can be displayed by the screen 102. During the video call, the user of the device 101 may wish to look at a position 105 of the remote user 103 (e.g., the position 105 corresponding to the eyes of the remote user 103) while images are being captured by the camera 104. As shown, the position 105 and the camera 104 may be offset 106 from each other. As discussed, it may be desirable to modify the images and / or video of the user (not shown) captured by the camera 104 by distorting the eye areas of the captured images and / or video frames to create the appearance of looking or viewing the camera 104.In the example shown, the offset 106 is a vertical offset between the camera 104, which is mounted above the display 102. However, the camera 104 and the display 102 can have any relative position to each other. For example, the relative position between the camera 104 and the display 102 can be a vertical offset with the camera above the display (as shown), a vertical offset with the camera below the display, a horizontal offset with the camera to the left of the display, a horizontal offset with the camera to the right of the display, or any diagonal offset between the camera and the display (e.g., the camera above and to the right of the display) with the camera completely outside the boundaries of the display or within a boundary in one direction (e.g.,The camera 104 can be moved from a central position of the display 102 to a position off-center, but still be within an edge of the display 102.

[0025] Fig. Figure 2 illustrates an exemplary system 200 for providing eye contact correction, arranged according to at least some implementations of the present disclosure. As in Fig. As shown in Figure 2, the system 200 can comprise an image signal processing module 201, a face detection module 202, a face orientation point detection module 203, an eye contact correction module 204, and a video compression module 205. As shown, in some embodiments, the face detection module 202 and the face orientation point detection module 203 can be part of a face detection and face orientation point detection module 301 (as shown in Figure 2). Fig.3), which can provide face detection and face orientation point detection. The System 200 can be implemented via any suitable device, such as Device 101 and / or, for example, a personal computer, laptop computer, tablet, phablet, smartphone, digital camera, game console, portable device, display device, all-in-one device, two-in-one device, or the like. For example, the System 200 can provide an image signal processing pipeline, which may be implemented in hardware, software, or a combination thereof. As discussed, the System 200 can provide real-time eye contact correction (or gaze correction) for video telephony, video conferencing, or the like, to obtain natural-looking corrected images and / or video images. The System 200 can process known scene geometry (e.g.,a relative positioning of camera 104 to display 102 as with reference to . Fig. 1 discussed) and use neural network-based machine learning techniques to provide high-quality real-time eye contact correction.

[0026] As shown, an image signal processing module 201 can receive image source data (ISD) 211. The image source data 211 can comprise any suitable image data, such as image data from an image sensor (not shown) or the like. The image signal processing module 201 can process the image source data 211 to generate an input image (II) 212. The image signal processing module 201 can process the image source data 211 using any suitable technique or techniques, such as demosaicing, gamma correction, color correction, image enhancement, or the like, to generate the input image 212. The input image 212 can be in any suitable color space and can comprise any suitable image data. As used herein, the term "image" can encompass any suitable image data in any suitable context.For example, an image can be a standalone image, an image within a video sequence, a single frame from a video, or something similar. The input image 212 can be characterized as a source image, an image, a single frame, a source single frame, or something similar.

[0027] The input image 212 can be provided to the face detection module 202, which can determine whether faces are present in the input image 212 and, if so, determine the position of the faces. The face detection module 202 can perform such face detection using any suitable technique or techniques to generate face detection data (FD, Face Detection Data) 213, which may comprise any suitable data or data structure representing one or more faces in the input image 212. For example, the face detection data 213 may provide the position and size of a bounding box corresponding to a face detected in the input image 212.The face detection data 213 and / or the input image 212, or parts thereof, can be provided to the face landmark detection module 203, which can determine face landmarks corresponding to the face or faces detected by the face detection module 202. The face landmark detection module 203 can determine such face landmarks (e.g., landmarks corresponding to eyes, a nose, a mouth, etc.) using any suitable technique or techniques to generate face landmarks (FL, Facial Landmarks) 214, which may include any suitable data or data structure representing face landmarks in the input image 212. For example, the face landmarks 214 may include positions of face landmarks and a corresponding descriptor (e.g.,the facial part to which the landmark corresponds) for the facial orientation points.

[0028] As shown, the facial reference points 214 and / or the input image 212, or parts thereof, can be provided to the eye contact correction module 204, which can generate an eye contact corrected image (ECCI) 215 that corresponds to the input image 212. The eye contact correction module 204 can generate an eye contact corrected image 215 using the information provided herein with reference to Fig.3 techniques discussed. As shown, the eye-corrected image 215 or a sequence of eye-corrected images can be provided to the video compression module 205, which can perform image and / or video compression on the eye-corrected image 215 or a sequence of eye-corrected images to provide a compressed bitstream (CB) 216. The video compression module 205 can generate the compressed bitstream 216 using any suitable technique or techniques, such as video coding techniques. In some examples, the compressed bitstream 216 can be a standards-compliant bitstream. For example, the compressed bitstream 216 can conform to the H.264 / MPEG-4 Advanced Video Coding (AVC) standard, the High Efficiency Video Coding (HEVC) standard, or similar.The compressed bitstream 216 can be sent to a remote device (e.g., a device that is like the one in relation to . Fig. 1 discussed, from the device 101) for a presentation to a user of the remote device. For example, the compressed bitstream 216 can be packetized and transmitted to the remote device, which can reassemble the compressed bitstream 216 and decode the compressed bitstream 216 to produce image(s) for a presentation to the user of the remote device.

[0029] Fig. Figure 3 illustrates an exemplary eye contact correction module 204 for providing eye contact correction, arranged according to at least some implementations of the present disclosure. As in Fig. As shown in Figure 2, the eye contact correction module 204 can be provided as part of the system 200 to provide eye contact correction. As shown in Figure 2. Fig. As shown in Figure 3, the input image 212 can be received by the face detection and face orientation point detection module 301, which, as shown in Figure 3, Fig. 2 discussed, which may include the face detection module 202 and the face landmark detection module 203. As shown, the face detection and face landmark detection module 301 can receive an input image 212, and the face detection and face landmark detection module 301 can, as discussed herein, provide face landmarks 214. As in Fig.As shown in Figure 3, the eye contact correction module 204 can include a cropping and resizing module 302, a feature generation module 303 which can include a neural network coding module 304, a random forest classifier 306, and a deformation and integration module 307. For example, the eye contact correction module 204 can receive facial orientation points 214 and the input image 212, and the eye contact correction module 204 can generate an image with corrected eye contact 215.

[0030] Fig. Figure 4 illustrates an exemplary input image 212 arranged according to at least some implementations of the present disclosure. As in Fig.As shown in Figure 4, the input image 212 can include a user 401 and a background 402. For example, image source data 211 can be captured by the camera 104, and the image source data 211 can be processed by the image signal processing module 201 to generate an input image 212. Although in Fig. 4 not shown, the user 401 can have eyes with an unnatural look, so that the user 401, instead of looking into the camera 104, looks at the display 102 (see Fig. 1) looks, which may give a remote user the impression that the user is looking down and away instead of looking at the remote user (see Fig. 6).

[0031] Referring again to Fig.3, the input image 212 can be received by the face detection and face orientation point detection module 301, which, as discussed, can provide face detection and face orientation point detection using any suitable technique or techniques.

[0032] Fig. Figure 5 illustrates exemplary face detection data 213 and exemplary face orientation points 214 arranged according to at least some implementations of the present disclosure. As in Fig.As shown in Figure 5, the face detection and face orientation point detection module 301 can provide face detection to generate face detection data 213, which can provide a bounding box or similar around a detected face area of ​​the user 401. Furthermore, the face detection and face orientation point detection module 301 can provide face orientation points 214, which can provide positions and, if applicable, corresponding descriptors for the detected face orientation points of the user 401. For example, the face orientation points 214 can include eye orientation points 501, which are characterized by positions in the input image 212 and a descriptor or similar indicating that they correspond to a detected eye of the user 401.

[0033] Referring again to Fig.3. As shown, the cropping and resizing module 302 can receive the facial orientation points 214 and the input image 212, and the cropping and resizing module 302 can generate eye regions (ERs) 311 using any suitable technique or techniques. For example, the input image 212 can be cropped to generate one or more eye regions 311 such that the eye regions 311 include all or most of the eye orientation points corresponding to a detected eye and, optionally, a buffer region around the outermost eye orientation points in the horizontal and vertical directions. In one embodiment, the cropping and resizing module 302 can crop the input image 212 such that the eye regions 311 each have a fixed (e.g., predetermined) size.

[0034] Fig.Figure 6 illustrates an exemplary eye region 311 arranged according to at least some implementations of the present disclosure. As in Fig. As shown in Figure 6, the eye area 311 can be a cropped portion of the input image 212 that includes one eye of user 401. Also, as shown, the eye area 311 can include a user eye that, if received by a remote user, would not make eye contact with the remote user (as expected).

[0035] Referring again to Fig. 3. The eye areas 311 can comprise any number of eye areas for any number of users. In an expected implementation, the eye areas 311 can comprise two eye areas for a single user 401. However, the eye areas 311 can comprise a single eye area, multiple eye areas corresponding to multiple users, or something similar.

[0036] With further reference to Fig.3. The feature generation module 303 can receive eye regions 311, and the feature generation module 303 can generate compressed features (CF) 312 such that each set of compressed features 312 corresponds to one eye region of the eye regions 311 (e.g., each eye region can have corresponding compressed features). In one embodiment, the feature generation module 303 can include a neural network coding module 304 that can generate or determine compressed features 312 for each eye region of the eye regions 311. For example, the compressed features 312 can be an output of an output layer of a neural network implemented by the neural network coding module 304. The compressed features 312 can be characterized as a compressed feature set, a feature set, an output of a middle layer, a compressed feature vector, or the like.The neural network coding module 304 can be or implement a deep neural network, a convolutional neural network, or the like. In one embodiment, the neural network coding module 304 can be or implement a coding part of a neural network. In another embodiment, the neural network coding module 304 can be or implement a neural autocoding network that encodes eye areas 311 into sets of compressed features.

[0037] Fig. Figure 7 illustrates an exemplary neural network 700 arranged according to at least some implementations of the present disclosure. As in Fig.As shown in Figure 7, the neural network 700 can comprise an encoding part 710 with layers 701-704 and a decoding part 720 with layers 721-724. As shown, the neural network 700 can receive an eye region 711 via the encoding part 710, and the encoding part 710 can generate compressed features 712 via layers 701-704. Furthermore, the decoding part 720 can receive compressed features 712, and the decoding part 720 can generate a resultant eye region (RER) 713 via layers 721-724. For example, as further discussed herein, the encoding part 710 and the decoding part 720 can be trained in a training phase, and the encoding part 710 can be implemented in an implementation phase.For example, the compressed intermediate features 712 can be used as compressed features for input into a classifier which can generate a motion vector field based on the compressed features, as further described herein with reference to . Fig. 3 is discussed.

[0038] The neural network 700 can comprise any suitable neural network, such as an artificial neural network, a deep neural network, a convolutional neural network, or similar. As in Fig.As shown in Figure 7, the neural network 700 can comprise an encoding part 710 with four layers 701-704 and a decoding part 720 with four layers 721-724. However, the encoding part 710 and the decoding part 720 can have any suitable number of layers. Furthermore, the encoding part 710 and the decoding part 720 can comprise fully connected layers, convolutional layers with max pooling, or a combination thereof. In one embodiment, the encoding part 710 can have two fully connected layers, and the decoding part 720 can have two fully connected layers. In another embodiment, the encoding part 710 can have four fully connected layers, and the decoding part 720 can have four fully connected layers. In yet another embodiment, the encoding part 710 can have six fully connected layers, and the decoding part 720 can have six fully connected layers.In one embodiment, the encoding part 710 can have four layers with a convolutional layer with max-pooling followed by three fully connected layers, and the decoding part 720 can have four fully connected layers. In another embodiment, the encoding part 710 can have two convolutional layers with max-pooling followed by two fully connected layers, and the decoding part 720 can have four fully connected layers. As already discussed, however, any suitable combination can be provided. Furthermore, such properties of the neural network 700 are discussed further herein with regard to training the neural network 700.

[0039] Referring again to Fig.3. The neural network coding module 304 can implement the coding part 710 of the neural network 700. The neural network coding module 304 can implement the coding part 710 of the neural network 700 using any suitable technique or techniques such that the neural network coding module 304 can generate compressed features 312 based on the eye area 311. The compressed features 312 can be generated using any suitable technique or techniques. For example, the eye areas 311 with 50×60 pixels can provide the neural network coding module 304 with 3,000 input nodes, which can implement a neural network encoder to generate compressed features 312 that can comprise any number of features, such as about 30 to 150 features or the like.In one embodiment, the compressed features 312 comprise approximately 100 features or parameters. For example, in the context of the eye areas 311 with 50×60 pixels and the compressed features 312 with approximately 100 features, a compression ratio of 30:1 can be provided. Furthermore, such 100 features, even if not provided via the system 300, can be used in the context of the coding part 720 (see . Fig. 7) provide a resulting eye area 713 that maintains image integrity and quality with respect to eye area 711.

[0040] As shown, the compressed features 312 and the input image 212 can be received by the random forest classifier 306. As discussed further herein, the random forest classifier 306 can be a pre-trained classifier that applies the pre-trained classifier to a compressed set of features 312 to generate a corresponding (and optimal) motion vector field (MVF) 313. Although illustrated with respect to the random forest classifier 306, any suitable pre-trained classifier can be applied, such as a decision tree learning model, a random forest kernel model, or the like.The random forest classifier 306 can receive the compressed features 312 and, for each set of features of the compressed features 312, apply the random forest classifier 306 such that the set of compressed features traverses branches of the forest based on decisions at each branch until a leaf is reached, where the leaf corresponds to the compressed features or provides a movement vector field for them. For example, the movement vector fields 313 can include a movement vector field for each set of features of the compressed features 312 based on an application of the random forest classifier 306 to each set of features.

[0041] The motion vector fields 313 and the input image 212 can be received by the deformation and integration module 307, which can deform the eye areas 311 based on the motion vector fields 313 and integrate the deformed eye area(s) into the input image 212. For example, the motion vector fields 313 can be applied to the eye areas 311 by determining a deformed pixel value for each pixel position of the eye areas 311 as a pixel value corresponding to the pixel value specified by the motion vector for the pixel. The deformed eye areas can then replace the eye areas 311 in the input image 212 (e.g., the deformed eye areas can be integrated into the remaining part of the input image 212) to produce an image with corrected eye contact 215.

[0042] Fig.Figure 8 illustrates an exemplary corrected eye area 801 arranged according to at least some implementations of the present disclosure. As in Fig. As shown in Figure 8, the corrected eye area 801 can provide a corrected view in relation to the eye area 311. Fig. 6, so that the gaze gives the appearance that the receiver of the corrected eye area 801 is being looked at. As discussed, the corrected eye area 801 (e.g., a deformed eye area) can be integrated into a final corrected eye image, which can be encoded and transmitted to a remote user for display.

[0043] Referring again to Fig.3. Any number of corrected eye contact images 215 can be generated, such as a video sequence of corrected eye contact images 215. The techniques discussed herein can provide high-quality, real-time eye contact correction. For example, the discussed processing performed by System 200 can be carried out at such a speed that real-time video telephony, video conferencing, or similar activities can be conducted between users.

[0044] With reference to Fig.1. A user of device 101 (not shown) and a user 103 at a remote device (not shown) can experience a video telephony application, a video conferencing application, or the like, which compensates for the offset 106 between position 105 and camera 104 by adjusting the gaze of one or both users. Furthermore, with respect to device 101, the feature generation module 303 and / or the random forest classifier 306 can utilize pre-training based on the offset 106. For example, as discussed below, training images used to train the neural network coding module 304 and / or the random forest classifier 306 can be selected to match or closely approximate the offset 106. For example, if the offset 106 is approximately 10° (e.g.,If a 10° angle between a line from a user's eyes to position 105 and a second line from the user's eyes to camera 104) is provided, training images with an approximate offset of 10° can be used for the neural network coding module 304 and / or the random forest classifier 306.

[0045] Furthermore, with further reference to System 200, System 200 comprises a single eye contact correction module 204 trained for a single orientation or relative position between the camera 104 and the display 102. In other embodiments, System 200 may have multiple contact correction modules, or the contact correction module 204 may be capable of implementing eye contact correction for multiple orientations or relative positions between the camera 104 and the display 102. For example, the device 101 may include a second camera (not shown) or the ability to reposition the camera 104 relative to the display 102. In one embodiment, System 200 may include a first eye contact correction module for a first relative position between the camera 104 and the display 102 and a second eye contact correction module for a second relative position between the camera 104 and the display 102.In one embodiment, the system 200 can include multiple eye contact correction modules for multiple relative positions between the camera 104 and the display 102. For example, each eye contact correction module can have a different pre-trained neural network coding module 304 and / or a different pre-trained random forest classifier 306. In other examples, the individual eye contact correction module can implement different predetermined variables or data structures (e.g., by loading them from memory) to provide multiple eye contact corrections, each corresponding to the different relative positions between the camera 104 and the display 102. For example, the neural network coding module 304 can implement different pre-trained neural network encoders that respond to or are selectively based on a particular relative position between the camera 104 and the display 102.Additionally or alternatively, the random forest classifier can apply 306 different random forest models that respond to or are selectively based on a specific relative position between the camera 104 and the display 102.

[0046] For example, encoding an eye area of ​​a source image by a pre-trained neural network and applying the pre-trained classifier to the compressed features 312 to determine a motion vector field 313 for the eye area 311 of the input image 212 as discussed above, based on an initial relative position between the camera 104 and the display 102 (e.g., with the camera 104 above the display 102 to provide an offset 106), can be selectively provided. With a second relative position between camera 104 and display 102 (e.g., with camera 104 below, to the left of, to the right of, further above, diagonally to display 102, or similar), the compressed features (e.g., by neural network coding module 304 or another neural network coding module) can be generated into second compressed features (not shown), and a second pre-trained classifier can (e.g.,The second compressed features are applied to the random forest classifier 306 (or another pre-trained classifier, such as another random forest classifier) ​​to determine a second motion vector field (not shown). The second motion vector field can be used to deform the eye area of ​​the input image 212, and the deformed eye area can be integrated into the input image to produce an image with corrected eye contact 215. For example, such techniques can provide the application of different pre-trained models when an orientation or relative position between the camera 104 and the display 102 changes.

[0047] As discussed, various components or modules of System 200 can be pre-trained in a training phase prior to the provision of an implementation phase.

[0048] Fig.Figure 9 illustrates an exemplary system 900 for pre-training an eye contact correction classifier, arranged according to at least some implementations of the present disclosure. As in Fig.As shown in Figure 9, the system 900 can comprise a face detection and face orientation point detection module 901, a face detection and face orientation point detection module 902, a cropping and resizing module 903, a cropping and resizing module 904, a feature determination module 905, a motion vector candidate module 906, and a random forest generation module 907. As shown, in some embodiments, the face detection and face orientation point detection modules 901, 902 and the cropping and resizing modules 903, 904 can be provided separately. In other examples, the face detection and face orientation point detection modules 901, 902 and / or the cropping and resizing modules 903, 904 may be provided together as a single face detection and face orientation point detection module and / or a single cropping and resizing module.System 900 can be implemented via any suitable device, such as Device 101 and / or, for example, a personal computer, laptop computer, tablet, phablet, smartphone, digital camera, game console, portable device, display device, all-in-one device, two-in-one device, or similar. Systems 900 and 200 can be provided separately or together.

[0049] As shown, the system 900 can receive training images 911, where the training images comprise pairs of source images (SI) 912 and target images (TI) 913. For example, each pair of training images 911 can comprise a source image and a corresponding target image, where the source and target images have a known viewing angle difference between them. The system 900 can be used to generate a random forest classifier (RFC) 920 for any number of sets of training images 911 with any suitable viewing angle difference between them.

[0050] Fig. Figure 10 illustrates an exemplary source image 912 and an exemplary target image 913, arranged according to at least some implementations of the present disclosure. As in Fig.As shown in Figure 9, the source image 912 can include an image of a user looking downwards in relation to a camera capturing the source image 912, and the target image 913 can include an image of the user looking into the camera. In the example of the Fig. 10. The source image 912 and the target image 913 have a viewing angle difference of approximately 10° (e.g. a vertical viewing angle difference or an offset of 10°).

[0051] Referring again to Fig. 9, the source images 912 and the target images 913 can be any number of pairs of training images analogous to those in Fig.The 10 depicted include the source images 912, which can be received by the face detection and face landmark detection module 901, which can provide source facial landmarks (SFLs) 914. Likewise, the target images 913 can be received by the face detection and face landmark detection module 902, which can provide target facial landmarks (TFLs) 915. The face detection and face landmark detection modules 901 and 902 can be used as discussed in relation to the face detection and face landmark detection module 301 (see Figure 301). Fig. 3) work. This will not be discussed again for the sake of brevity.

[0052] Fig. Figure 11 illustrates exemplary target face orientation points 915 arranged according to at least some implementations of the present disclosure. As in Fig.As shown in Figure 11, the target face orientation points can comprise 915 eye orientation points 1101. The eye orientation points 1100 can be generated using any suitable technique or techniques and can be characterized by a position within an image and / or descriptors corresponding to the eye orientation points 1100 (e.g., a descriptor of an eye, left eye, or similar). In the example of the Fig. Figure 11 shows target face orientation points 915, which correspond to the target image 913. As already discussed, source face orientation points 914 can also be determined, which correspond to the source image 912.

[0053] Referring again to Fig.9. Source images 912 and source face orientation points 914 can be received by the crop and resize module 903, which can provide source eye regions (SERs) 916. Likewise, target images 913 and target face orientation points 915 can be received by the crop and resize module 903, which can provide target eye regions (TERs) 917. Crop and resize modules 903 and 904 can be used as discussed in relation to crop and resize module 302 (see Fig. 3) work. This will not be discussed again for the sake of brevity.

[0054] Fig. Figure 12 illustrates an exemplary source eye region 916 and an exemplary target eye region 917 arranged according to at least some implementations of the present disclosure. As in Fig.As shown in Figure 12, eye areas 916 and 917 can be cropped portions of the source image 912 and the target image 913, respectively. Furthermore, eye area 916 can encompass a user eye looking downwards, and eye area 917 can encompass the user eye looking forwards or straight ahead.

[0055] Referring again to Fig. In addition to cropping and resizing, system 900 can align the source eye areas 916 and the target eye areas 917. This alignment can, for example, include an affine transformation to compensate for head movement between the corresponding source image 912 and the target image 913.

[0056] As shown, the source eye regions 916 and the target eye regions 917 can be provided to the motion vector candidate module 906. The motion vector candidate module 906 can determine likelihood maps (LMs) 919 based on the source eye regions 916 and the target eye regions 917. For example, for each pixel position of the source eye regions 916, multiple candidate motion vectors can be determined by calculating the sum of the absolute differences between a window around the pixel in the source eye region 916 and a window shifted by one motion vector in the target eye region 917. For example, the window or evaluation block may have a size of 6×6 pixels in a search area of ​​10×10 pixels with fixed spacing, such as every pixel.Although discussed in relation to windows measuring 6×6 pixels, which are searched at each pixel within a search area of ​​10×10 pixels, any window pixel size, any search area, and any spacing can be used.

[0057] For example, the motion vector candidate module 906 can generate a very large dataset of likelihood maps providing candidate motion vectors or inverse probabilities. For instance, a lower sum of absolute differences may correlate with a higher probability that the optimal motion vector of the candidate motion vectors has been found. For example, the lower the sum of absolute differences for a given shift (e.g., as provided by the corresponding motion vector), the more likely the effectively best motion vector is to correspond to that shift. The sum of absolute differences maps (e.g., the likelihood maps 919) can be aggregated across all blocks in the aligned pairs of source eye regions 916 and target eye regions 917, and all pairs of source images 912 and target images 913 in the training images 911.

[0058] Fig. Figure 13 illustrates an exemplary likelihood map 919 for an exemplary target eye area 916, arranged according to at least some implementations of the present disclosure. As in Fig.As shown in Figure 13, a likelihood map 919 for a pixel 1301 of the source eye region 916 can be generated such that, for an entry 1302 of the likelihood map 919, a sum of the absolute differences or a similar similarity measure between a window 1303 around the pixel 1301 and several windows (not shown) in the corresponding target eye region 917 (not shown) can be determined. For example, entry 1302 can correspond to a comparison of window 1303 with a window in the target eye region 917 at a maximum negative x-offset 1311 and a maximum positive y-offset 1312 (e.g., corresponding to a motion vector with a maximum negative x-offset 1311 and a maximum positive y-offset 1312) with respect to pixel 1301. For each combination of x-offsets 1311 and y-offsets 1312, a corresponding entry in the likelihood map 919 can be determined as a sum of the absolute differences or a similar measure of similarity.Furthermore, as discussed, a likelihood map can be generated for each pixel position of the source eye area 916 and even further for each source eye area of ​​the source eye areas 916.

[0059] Referring again to Fig. 9 The feature determination module 905 can receive source eye regions 916, and the feature determination module 905 can determine compressed source feature sets (SCFS, Source Compressed Feature Sets) 918 based on the source eye regions 916. For example, the feature determination module 905 can generate compressed features based on an application of a coding part of a neural network. For example, with reference to Fig.Seven compressed source feature sets 918 are generated by the coding part 710 of the neural network 700, which is implemented by the feature identification module 905, and the resulting compressed source feature sets 918 (e.g., each set corresponding to an eye region of the source eye regions 916) can be provided to the random forest generation module 907. For example, the same coding part 710 of the neural network 700 (e.g., with the same architecture, same features, same parameters, etc.) can be implemented in the training phase and in the implementation phase. Furthermore, as discussed, the neural network 700, which is implemented by the feature identification module 905, can have any suitable architecture and any suitable features.

[0060] As previously discussed, in some examples, source eye areas 916 of 50×60 pixels can be provided to the feature identification module 905, and the compressed source feature sets 918 can each contain 100 features, although any size of eye areas 916 and any number of features can be used. Exemplary neural network architectures and training techniques are further elaborated herein. It is understood that any neural network architecture discussed with respect to training can be implemented in an implementation phase (e.g., as with respect to the neural network coding module 304 of the Fig. 3 discussed).

[0061] Fig. Figure 14 illustrates an exemplary neural network training system 1400, arranged according to at least some implementations of the present disclosure. As in Fig.As shown in Figure 14, the neural network training system 1400 can comprise a neural network (NN) coding module 1401, a neural network decoding module 1402, and an error module 1403. For example, the neural network coding module 1401 can implement a neural network coding part, and the neural network decoding module 1402 can implement a neural network decoding part, such that the neural network coding module 1401 and the neural network decoding module 1402 implement neural network coding-decoding that receives an input image and provides a resulting image that can be compared to the input image. For example, the neural network can be trained to reduce the error between the input image provided to the neural network and the resulting image from the neural network. As discussed, a middle-layer parameter set (e.g.,compressed features) can be used to train a classifier, and the pre-trained coding part of the neural network and the pre-trained classifier can be implemented in real time during an implementation phase to map an input eye area to a motion vector field that can be used to deform the eye area to provide eye contact or gaze correction.

[0062] As shown, the neural network coding module 1401 can receive input data 1411, which can comprise any suitable image data, such as eye regions, source images 912, target images 913, source eye regions 916, target eye regions 917, or the like. For example, the image data 1411 can comprise any suitable corpus of input image data for training the neural network implemented by the neural network coding module 1401 and the neural network decoding module 1402. In one embodiment, the image data 1411 can comprise images with 54×54 pixels. In another embodiment, the image data 1411 can comprise images with 54×66 pixels. The neural network coding module 1401 can receive image data 1411, and the neural network coding module 1401 can generate compressed features 1412 for an image of the image data 1411.The compressed features 1412 can be generated by any suitable neural network coding part, and the compressed features 1412 can comprise any number of features. For example, with reference to . Fig. 7. The neural network coding module 1401 implements the coding part 710 to generate compressed features 1412 analogous to the compressed features 712.

[0063] The neural network decoding module 1402 can receive compressed features 1412, and the neural network decoding module 1402 can generate resultant image data (RID, Resultant Image Data) 1413 based on the compressed features 1412. The resultant image data 1413 can be generated by any suitable neural network decoding part, and the resultant image data can correspond to an image of the image data 1411. For example, the discussed processing can be performed on a large number of images of the image data 1411 to generate a resultant image corresponding to each image. For example, with reference to Fig. 7 the neural network decoding module 1402 implement the decoding part 720 to generate resulting image data 1413 analogous to the resulting eye area 713.

[0064] As shown, the image data 1411 and the resulting image data 1413 can be provided to the error module 1403. The error module 1403 can compare an image or images of the image data 1411 with a resulting image or images of the resulting image data 1413 to generate an evaluation 1421. The error module 1403 can generate the evaluation 1421 using any suitable technique or techniques, and the evaluation 1421 can include any suitable evaluation, any suitable error measurement, or the like. The evaluation 1421 can be used to train the neural network implemented by the neural network encoding module 1401 and the neural network decoding module 1402 such that, for example, an error represented by the evaluation 1421 can be minimized.In one embodiment, the evaluation 1421 can be a Euclidean loss between images of the image data 1411 and resulting images of the resulting image data 1413. In another embodiment, the evaluation 1421 can be an L2 error between images of the image data 1411 and resulting images of the resulting image data 1413.

[0065] Fig. Figure 15 illustrates an exemplary neural network training system 1500, arranged according to at least some implementations of the present disclosure. As in Fig.As shown in Figure 15, a neural network training system 1500 can comprise a neural network coding module 1401, a neural network decoding module 1402, vertical filtering modules 1503, 1504, and error modules 1403, 1506. As discussed, the neural network coding module 1401 can implement a neural network coding part, and the neural network decoding module 1402 can implement a neural network decoding part, such that the neural network coding module 1401 and the neural network decoding module 1402 implement neural network coding-decoding that receives an input image and provides a resulting image that can be compared to the input image. For example, the neural network can be trained to reduce the error between the input image provided to the neural network and the resulting image from the neural network. As discussed, a middle-layer parameter set (e.g.,compressed features) can be used to train a classifier, and the pre-trained coding part of the neural network and the pre-trained classifier can be implemented in real time during an implementation phase to map an input eye area to a motion vector field that can be used to deform the eye area to provide eye contact or gaze correction.

[0066] As shown, the neural network encoding module 1401 can receive image data 1411, which can comprise any suitable image data, and the neural network encoding module 1401 can generate compressed features 1412 as discussed. The neural network decoding module 1402 can receive compressed features 1412, and the neural network decoding module 1402 can, also as discussed, generate resulting image data 1413 based on the compressed features 1412.

[0067] As shown, the image data 1411 and the resulting image data 1413 can be provided to the error module 1403, which can compare an image or images of the image data 1411 with a resulting image or images of the resulting image data 1413 to generate an evaluation 1421. As discussed, the error module 1403 can generate the evaluation 1421 using any suitable technique or techniques, and the evaluation 1421 can include any suitable evaluation, any suitable error measurement, or the like, such as a Euclidean loss or an L2 error between images of the image data 1411 and resulting images of the resulting image data 1413.

[0068] As also shown, the image data 1411 and the resulting image data 1413 can be provided to the vertical filter module 1503 and the vertical filter module 1504, respectively. The vertical filter module 1503 can receive image data 1411 and apply a vertical filter to the image data 1411 to produce vertically filtered image data (VFID) 1514. Likewise, the vertical filter module 1504 can receive resulting image data 1413 and apply a vertical filter to the resulting image data 1413 to produce vertically filtered resulting image data 1515. In one embodiment, the vertical filter modules 1503 and 1504 can apply vertical high-pass filters. The vertical filters applied by the vertical filter modules 1503, 1504 can include one or more suitable vertical filters.Furthermore, the vertical filters applied by the vertical filter modules 1503 and 1504 can be the same or different. Although discussed here in relation to vertical filters, any suitable filters, such as horizontal or diagonal filters, can be applied.

[0069] As shown, vertically filtered image data 1514 and vertically filtered resulting image data (VFRID, Vertical Filtered Resultant Image Data) 1515 can be provided to the fault module 1506. The fault module 1506 can compare a vertically filtered image or images of the vertically filtered image data 1514 with a vertically filtered resulting image or images resulting from vertically filtered image data 1515 to generate an evaluation 1522. The fault module 1506 can generate the evaluation 1522 using any suitable technique or techniques, and the evaluation 1522 can include any suitable evaluation, any suitable fault measurement, or the like.Evaluate 1421 and evaluate 1522 can be used to train the neural network implemented by the neural network encoding module 1401 and the neural network decoding module 1402 such that, for example, the errors represented by evaluate 1421 and evaluate 1522 can be minimized, or a sum, average, weighted average, or similar of evaluate 1421 and evaluate 1522 can be minimized. In one embodiment, evaluate 1522 can be a Euclidean loss between vertically filtered images of the vertically filtered image data 1514 and vertically filtered resulting images of the resulting vertically filtered image data 1515. In another embodiment, evaluate 1522 can be an L2 error between images of the image data 1411 and resulting images of the resulting image data 1413.

[0070] Since, for example, the neural network implemented by the neural network encoding module 1401 and the neural network decoding module 1402 cannot fully reconstruct the images of the image data 1411 when generating resulting images 1413, the implemented neural network may tend to remove high-frequency noise and details. In the context of the eye areas, some eye areas may include eyeglasses, which can be removed or blurred by a neural network. By adding a vertical high-pass filter or similar (e.g., via the vertical filter modules 1503, 1504) and training the neural network based on minimizing a corresponding evaluation 1522, the neural network can be forced to learn to retain high-frequency information such as eyeglasses and other details.

[0071] Fig.Figure 16 illustrates an exemplary neural network training system 1600, arranged according to at least some implementations of the present disclosure. As in Fig.As shown in Figure 16, the neural network training system 1600 can include a neural network coding module 1401, a neural network decoding module 1402, vertical filter modules 1503 and 1504, standard deviation modules 1604 and 1606, and error modules 1403, 1506, and 1608. As discussed, the neural network coding module 1401 can implement a neural network coding part, and the neural network decoding module 1402 can implement a neural network decoding part, such that the neural network coding module 1401 and the neural network decoding module 1402 implement neural network coding-decoding that receives an input image and provides a resulting image that can be compared to the input image. For example, the neural network can be trained to reduce the error between the input image provided to the neural network and the resulting image from the neural network.As discussed, a middle-layer parameter set (e.g., compressed features) can be used to train a classifier, and the pre-trained coding portion of the neural network and the pre-trained classifier can be implemented in real time during an implementation phase to map an input eye area to a motion vector field that can be used to deform the eye area to provide eye contact or gaze correction.

[0072] As shown, the neural network encoding module 1401 can receive image data 1411, which may comprise any suitable image data, and the neural network encoding module 1401 can generate compressed features 1412 as discussed. The neural network decoding module 1402 can receive compressed features 1412, and the neural network decoding module 1402 can, also as discussed, generate resulting image data 1413 based on the compressed features 1412. As also shown, the image data 1411 and the resulting image data 1413 can be provided to the error module 1403, which can compare an image or images of the image data 1411 with a resulting image or images of the resulting image data 1413 to generate an evaluation 1421.As discussed, the fault module 1403 can generate the evaluation 1421 using any suitable technique or techniques, and the evaluation 1421 can include any suitable evaluation, any suitable error measurement, or the like, such as a Euclidean loss or an L2 error between images of the image data 1411 and resulting images of the resulting image data 1413. Furthermore, the image data 1411 and the resulting image data 1413 can each be provided to vertical filter modules 1503 and 1504, respectively, which can apply vertical filters to generate vertically filtered image data 1514 and vertically filtered resulting image data 1515. As previously discussed, vertically filtered image data 1514 and vertically filtered resulting image data 1515 can be provided to the fault module 1506, which can generate an evaluation 1522.The evaluation 1522 can include any suitable evaluation, any suitable error measurement, or similar.

[0073] As also shown, the vertically filtered image data 1514 and the resulting vertically filtered image data 1515 can be provided to the standard deviation module 1604 and the standard deviation module 1608, respectively. The standard deviation module 1604 can receive vertically filtered image data 1514, and the standard deviation module 1604 can determine a standard deviation of the vertically filtered image data 1514 in order to generate a standard deviation of the vertically filtered image data (SDVFID, Standard Deviation of Vertical Filtered Image Data) 1613. Likewise, the standard deviation module 1606 can receive vertically filtered resulting image data 1515 and the standard deviation module 1606 can determine a standard deviation of the vertically filtered resulting image data 1515 in order to generate a standard deviation of the vertically filtered resulting image data (SDVFRID, Standard Deviation of Vertical Filtered Resultant Image Data) 1614.The standard deviation modules 1604, 1606 can determine the standard deviations of the vertically filtered image data 1514 and the vertically filtered resulting image data 1515 using any suitable technique or techniques.

[0074] As shown, the standard deviation of the vertically filtered image data 1613 and the standard deviation of the vertically filtered resulting image data 1614 can be provided to the error module 1608. The error module 1608 can compare the standard deviation of the vertically filtered image data 1613 for an image or images with the standard deviation of the vertically filtered resulting image data 1614 for a resulting image or images to generate an evaluation 1616. The error module 1608 can generate the evaluation 1616 using any suitable technique or techniques, and the evaluation 1616 can include any suitable evaluation, any suitable error measurement, or the like.Evaluate 1421, evaluate 1522, and evaluate 1616 can be used to train the neural network implemented by the neural network encoding module 1401 and the neural network decoding module 1402 such that, for example, the errors represented by evaluate 1421 and evaluate 1522 can be minimized, or a sum, average, weighted average, or similar of evaluate 1421, evaluate 1522, and evaluate 1616 can be minimized. In one embodiment, evaluate 1616 can be a Euclidean loss between the standard deviations of the discussed images. In another embodiment, evaluate 1616 can be an L2 error between the standard deviations of the discussed images.Although illustrated in relation to the fact that the standard deviation modules 1604, 1606 are applied to vertically filtered images and resulting image data, the standard deviation modules 1604, 1606 can be applied to image data 1411 and resulting image data 1413, respectively.

[0075] For example, the evaluation function 1616 can accelerate learning during the training of the neural network implemented by the neural network coding module 1401 and the neural network decoding module 1402. For example, in cases where an image from the image data 1411 was not well coded, the difference between the standard deviations may be high, whereas in the case where the image is well coded, the error value corresponding to the evaluation function 1616 may be negligible.

[0076] As discussed, the neural network implemented by the Neural Network Coding Module 1401 and the Neural Network Decoding Module 1402 (and provided during the implementation phase) can have any suitable properties. With reference to Fig.7. The encoding part 710 can have two to eight fully connected layers with any number of nodes, and the decoding part 720 can have two to eight fully connected layers with any number of nodes. In one embodiment, the encoding part 710 has two fully connected layers with 300 and 30 nodes (e.g., with compressed features 712 with 30 features), and the decoding part 720 has two fully connected layers with 300 and 2916 nodes (e.g., corresponding to a resulting image with 54×54 pixels). In one embodiment, the encoding part 710 has four fully connected layers with 1,000, 500, 250, and 100 nodes (e.g., with compressed features 712 with 100 features), and the decoding part 720 has four fully connected layers with 250, 500, 1,000, and 3,564 nodes (e.g., corresponding to a resulting image with 54×66 pixels). In another embodiment, the encoding part 710 has six fully connected layers with 4,600, 2,000, and 3,564 nodes.200, 1,000, 500, 250 and 100 nodes (e.g., with compressed features 712 with 100 features), and the decoding part 720 has six fully connected layers with 250, 500, 1,000, 2,200, 4,600 and 3,564 nodes (e.g., corresponding to a resulting image with 54×66 pixels).

[0077] In other embodiments, the coding part can have 710 convolutional layers followed by max pooling to reduce the layer size during implementation. For example, each of the fully connected coding part layers discussed above can be replaced by a 3 × 3 × 3 convolutional layer followed by max pooling. Such modifications can provide faster real-time implementation due to fewer multiplications and additions compared to the fully connected layers, at the cost of a slight loss of quality (although such loss of quality may be virtually imperceptible). In one embodiment, the coding part 710 has four interconnected layers with two neural convolutional network layers (the first is 3×3×3 (step size 1) / Max-pooling 3×3 (step size 3) and the second is 3×3×3 (step size 1) / Max-pooling 3×3 (step size 2)) and two fully interconnected layers with 250 and 100 nodes (e.g.with compressed features 712 with 100 features), and the decoding part 720 has four fully connected layers with 250, 500, 1,000 and 3,564 nodes (e.g. corresponding to a resulting image with 54×66 pixels).

[0078] Furthermore, such neural networks can only function with errors between input image and resultant image (such as with regard to...). Fig. 14 discussed), with errors between input image / resulting image and errors between vertically filtered input image / resulting image (such as with reference to Fig. 15 discussed), with error between input image / resulting image, error between vertically filtered input image / resulting image and error of the standard deviation of vertically filtered input image / resulting image (such as with reference to Fig.16 discussed) or other combinations thereof. Table 1 illustrates exemplary neural network architectures or structures and exemplary training techniques according to at least some implementations of the present disclosure. Neural Network / Training Neural network structure and training techniques 1 4 fully connected layers (2 encoding, 2 decoding) 30 compressed features Only Euclidean loss errors Input image / resulting image 2 8 fully connected layers (4 encoding, 4 decoding) 100 compressed features Only Euclidean loss errors Input image / resulting image 3 8 fully connected layers (4 encoding, 4 decoding) 100 compressed features Euclidean loss and vertical filter loss Input image / resulting image 4 12 fully connected layers (6 encoding, 6 decoding) 100 compressed features Euclidean loss and vertical filter loss Input image / resulting image 5 12 fully connected layers (6 encoding, 6 decoding) 100 compressed features Euclidean loss, vertical filter loss and standard deviation of the vertical filter Input image / resulting image 6 8 fully connected layers (6 encoding, 6 decoding) 100 compressed features Euclidean loss, vertical filter loss and standard deviation of the vertical filter Input image / resulting image 7 8 layers (encoding: 1 convolution layer, 3 fully connected layers; decoding: 4 fully connected layers) 100 compressed features Euclidean loss, vertical filter loss and standard deviation of the vertical filter Input image / resulting image 8 8 layers (encoding: 2 convolutional layers, 2 fully connected layers; decoding: 4 fully connected layers) 100 compressed features Euclidean loss, vertical filter loss and standard deviation of the vertical filter input image / resulting image

[0079] Referring again to Fig.The compressed source feature sets 918 can comprise compressed features for each source eye region of the source eye regions 916, and the likelihood maps 919 can comprise a likelihood map for each pixel (or block of pixels) of each source eye region of the source eye regions 916. As shown, the compressed source feature sets 918 and the likelihood maps 919 can be provided to the random forest generation module 907, which can train and determine a random forest classifier 920 based on the compressed source feature sets 918 and the likelihood maps 919. For example, the random forest classifier 920 can be trained to determine optimal motion vector fields from the compressed source feature sets 918 (e.g., compressed features from a coding part of a neural network). The training can be carried out in such a way that each leaf of a tree in the random forest classifier 920 is assigned a likelihood map (e.g.a SAD map) for each pixel in a source eye region. At each branch of a tree, the training process can minimize the entropy between likelihood maps of the training observations that led to that branch of the tree. The Random Forest Classifier 920, as discussed herein, can have any size and data structure, such as 4-8 trees with depths of 6-7 levels. Furthermore, compression can be applied to generate the Random Forest Classifier 920. For example, a classifier pre-trained in a training step can be compressed by parameterized area fitting or similar techniques to generate the Random Forest Classifier 920.

[0080] Fig.Figure 17 is a flowchart illustrating an exemplary process 1700 for pretraining a neural eye contact correction network and classifier, arranged according to at least some implementations of the present disclosure. The process 1700 can, as shown in Fig. As illustrated in Figure 17, one or more operations 1701-1706 may be involved. Process 1700 may constitute at least part of a pre-training technique for a neural eye contact correction network and classifier. According to a non-restrictive example, Process 1700 may constitute at least part of a pre-training technique for a neural eye contact correction network and classifier performed by System 900 discussed herein. Furthermore, Process 1700 may be performed by System 1900, which is described below.

[0081] Process 1700 can begin at operation 1701, where multiple pairs of eye-area training images can be received. For example, the first images of the eye-area training image pairs can have a viewing angle difference with respect to the second images of the eye-area training image pairs. In one embodiment, the multiple pairs of eye-area training images can be received by system 1900. For example, training images 911 or similar can be received.

[0082] Processing can continue at Operation 1702, which generates a likelihood map for each pixel of each of the first images of the pairs of eye-area training images. For example, each likelihood map may include a sum of the absolute differences or another similarity measure for each of multiple candidate motion vectors corresponding to each pixel. The likelihood maps may be generated using any suitable technique or techniques. For example, the likelihood maps may be generated by a System 1900 central processing unit (CPU) 1901.

[0083] Processing can continue at Operation 1703, where a neural network can be trained based on the first images of the pairs of eye-area training images, the second images of the pairs of eye-area training images, and / or other training images. The neural network can be trained using any suitable technique or techniques. For example, the neural network can be trained based on encoding the training images to generate compressed features of a training step, decoding the compressed training step features to generate resulting training images, and evaluating the training images and the resulting training images.For example, the encoding of the training images can be performed by a coding part of a neural network, the decoding of the compressed training step features can be performed by a decoding part of a neural network, and the evaluation can represent an error between the training images and the resulting training images. In one embodiment, the evaluation can additionally or alternatively include vertical filtering of the training images and the resulting training images to generate vertically filtered training images or vertically filtered first training images, and determining an error between the vertically filtered training images and the vertically filtered resulting training images.Furthermore, the evaluation can additionally or alternatively include determining a standard deviation of the vertically filtered training images and a standard deviation of the vertically filtered resulting training images, as well as determining an error between the standard deviations. Such an error or errors can be used as feedback for training the neural network so that the error(s) can be minimized during training. As discussed, a coding portion of the trained neural network can be implemented during an implementation phase or during execution to generate compressed features. In one embodiment, the neural network can be trained by the central processing unit 1901 of the system 1900.

[0084] Processing can continue at Operation 1704, where compressed training step features can be determined for the first images of the pairs of eye-area training images. The compressed features can be determined using any suitable technique or techniques. For example, determining the compressed features may involve applying a coding portion of the trained neural network to the first images of the pairs of eye-area training images to generate a set of compressed features for each of the eye-area training images. In one embodiment, the compressed training step features can be determined by the central processor 1901 of System 1900.

[0085] Processing can continue at Operation 1705, where a pre-trained classifier can be trained based on the likelihood maps determined at Operation 1702 and the compressed training step feature sets determined at Operation 1704. The pre-trained classifier can be any suitable pre-trained classifier, such as a random forest classifier, and the pre-trained classifier can be trained using any suitable technique or techniques. In one embodiment, the pre-trained classifier can be trained by the central processor 1901 of System 1900. In another embodiment, the pre-trained classifier generated at Operation 1905 can be characterized as a pre-trained classifier from the training step or similar.

[0086] Processing can continue at Operation 1706, where the pre-trained classifier generated at Operation 1705 can be compressed. The pre-trained classifier can be compressed using any suitable technique or techniques. In one embodiment, the pre-trained classifier can be compressed based on parameterized surface fitting. In another embodiment, the pre-trained classifier can be compressed by the central processor 1901 of the system 1900. In yet another embodiment, a pre-trained classifier determined in the training step at Operation 1705 can be compressed at Operation 1706 by parameterized surface fitting to generate a pre-trained classifier, such as a random forest classifier, for implementation in an implementation phase to provide eye contact correction.

[0087] Fig. Figure 18 is a flowchart illustrating an exemplary process 1800 for providing eye contact correction, arranged according to at least some implementations of the present disclosure. The process 1800 can, as shown in Fig. As illustrated in Figure 18, one or more operations 1801-1803 may be involved. Process 1800 may constitute at least part of an eye contact correction technique. According to a non-restrictive example, Process 1800 may constitute at least part of an eye contact correction technique performed by System 200 discussed herein. Furthermore, Process 1800 is described herein with reference to System 1900. Fig. 19 described.

[0088] Fig. Figure 19 is an illustrative diagram of an exemplary system 1900 for providing eye contact correction, arranged according to at least some implementations of the present disclosure. As in Fig. As shown in Figure 19, the system 1900 can comprise a central processing unit 1901, an image processor 1902, a memory 1903, and a camera 1904. For example, there can be an offset between the camera 1904 and a display (not shown). As also shown, the central processing unit 1901 can comprise or implement a face detection module 202, a face orientation point detection module 203, an eye contact correction module 204, and a video compression module 205. Such components or modules can be implemented to perform operations discussed herein. The memory 1903 can contain images, image data, input images, image sensor data, face detection data, face orientation points, eye contact correction images, compressed bitstreams, eye regions, compressed features, feature sets, weighting factors and / or parameters of a neural network, pretrained classifier models, or any other data discussed herein.

[0089] As shown, the face detection module 202, the face orientation point detection module 203, the eye contact correction module 204, and the video compression module 205 can be implemented via the central processing unit 1901. In other examples, one or more parts of the face detection module 202, the face orientation point detection module 203, the eye contact correction module 204, and the video compression module 205 can be implemented via the image processor 1902, a video processor, a graphics processor, or similar. In still other examples, one or more parts of the face detection module 202, the face orientation point detection module 203, the eye contact correction module 204, and the video compression module 205 can be implemented via an image or video processing pipeline or unit.

[0090] The 1902 image processor can include any number and type of graphics, image, or video processing units capable of providing the operations discussed herein. In some examples, the 1902 image processor can be an image signal processor. Such operations can be implemented by software, hardware, or a combination thereof. For example, the 1902 image processor can include dedicated circuit arrangements for manipulating still-frame data, image data, or video data obtained from the 1903 memory. The 1901 central processing unit can include any number and type of processing units or modules capable of providing control and other higher-level functions for the 1900 system and / or any of the operations discussed herein. The 1903 memory can be any type of memory, such as volatile memory (e.g.,Static random access memory (SRAM), dynamic random access memory (DRAM), etc., or non-volatile memory (e.g., flash memory, etc.), and so on. In a non-restrictive example, memory 1903 could be implemented as a cache.

[0091] In one embodiment, one or more parts of the face detection module 202, the face orientation point detection module 203, the eye contact correction module 204, and the video compression module 205 can be implemented via an execution unit (EU) of the image processor 1902. The EU can, for example, contain programmable logic or circuit arrangements, such as a logic kernel or kernels, which can provide a wide range of programmable logic functions. In another embodiment, one or more parts of the face detection module 202, the face orientation point detection module 203, the eye contact correction module 204, and the video compression module 205 can be implemented via dedicated hardware, such as fixed-function circuit arrangements or the like.Fixed-function circuit arrangements may include dedicated logic or circuit arrangements and provide a set of fixed-function access points that may be mapped to the dedicated logic for a fixed purpose or function. In some embodiments, one or more parts of the face detection module 202, the face orientation point detection module 203, the eye contact correction module 204, and the video compression module 205 may be implemented via an application-specific integrated circuit (ASIC). The ASIC may comprise an integrated circuit arrangement adapted to perform the operations discussed herein. The camera 1904 may comprise any camera with any suitable number of lenses or the like for capturing images or video.

[0092] Referring again to Fig.18. Process 1800 can begin at operation 1801, in which an eye region of a source image can be encoded by a pre-trained neural network to generate compressed features corresponding to the eye region of the source image. The eye region can be encoded by any suitable pre-trained neural network using any suitable technique or techniques to generate compressed features. In one embodiment, the eye region can be encoded by a coding portion of a neural network. The coding portion of the neural network can comprise any suitable architecture or structure. In one embodiment, the pre-trained neural network is a fully connected neural network, such as a deep neural network.In one embodiment, the pre-trained neural network comprises multiple layers, including at least one convolutional neural network layer. In one embodiment, the pre-trained neural network comprises four layers in the following order: a first convolutional neural network layer, a second convolutional neural network layer, a first fully connected layer, and a second fully connected layer, the second fully connected layer providing the compressed features. In one embodiment, the compressed features can be determined by a neural network such as that implemented by the 1901 central processing unit of the System 1900.

[0093] Furthermore, in some embodiments, the eye area of ​​the source image can be received. In other embodiments, the eye area can be generated from the source image. The eye area can be generated from the source image using any suitable technique or techniques. In one embodiment, the eye area can be generated by providing face detection and face orientation point detection on the source image and cropping the source image based on the face detection and face orientation point detection to generate the eye area from the source image. In another embodiment, the eye area can be determined from the source image by the central processing unit 1901 of the system 1900.

[0094] Processing can continue at Operation 1802, where a pre-trained classifier can be applied to the compressed features to determine a motion vector field for the eye region of the source image. The pre-trained classifier can be any suitable pre-trained classifier, and the pre-trained classifier can be applied using any suitable technique or techniques. In one embodiment, the pre-trained classifier is a pre-trained random forest classifier with one leaf corresponding to the motion vector field. In one embodiment, the pre-trained classifier can be provided at Operation 1706 of Process 1700. In one embodiment, the pre-trained classifier can be applied by the central processor 1901 of System 1900.

[0095] The processing can continue at Operation 1803, in which the eye region of the source image can be deformed based on the motion vector field, and the deformed eye region can be integrated into a remaining portion of the source image to produce an image with corrected eye contact. For example, the remaining portion can be the part of the source image other than the eye region that is deformed. The eye region can be deformed and integrated into the remaining portion of the source image using any suitable technique or techniques to produce the image with corrected eye contact. In one embodiment, the eye region can be deformed and integrated into the remaining portion of the source image by the central processor 1901 of the system 1900 to produce the image with corrected eye contact.As discussed herein, the corrected eye contact image can provide an apparent view of the user toward a camera rather than toward a display (e.g., the corrected eye contact image can correct an offset between a display and a camera), so that a remote user of the image receives a more satisfactory response and the user can see the remote user on the display. In one embodiment, the corrected eye contact image can be encoded and transmitted to a remote device for display to a user (e.g., a remote user of a remote device). For example, the corrected eye contact image can be a frame (or single frame) of a video sequence, and the video sequence can be encoded and transmitted.

[0096] The 1800 process can be repeated for any number of eye regions of a source image, for any number of source images, or for any number of video sequences of source images. Furthermore, operations 1801 and / or 1802 can respond to or be selectively based on a first relative position between a camera and a display. For a second relative position between the camera and the display, operations 1801 and / or 1802 can be performed based on other pre-trained factors (e.g., another neural network and / or another pre-trained classifier).In one embodiment, at a second relative position between the camera and the display, the process 1800 may further include encoding the eye area of ​​the source image by a second pre-trained neural network to generate second compressed features corresponding to the eye area of ​​the source image, applying a second pre-trained classifier to the second compressed features to determine a second motion vector field for the eye area of ​​the source image, and deforming the eye area of ​​the source image based on the second motion vector field and integrating the deformed eye area into the remaining part of the source image to generate the image with corrected eye contact.

[0097] Various components of the system described herein may be implemented in software, firmware, and / or hardware, and / or any combination thereof. For example, various components of the system discussed herein may be provided, at least in part, by the hardware of a system-on-a-chip (SoC) data processing system, such as that found in a data processing system like a smartphone. Those with average technical knowledge will recognize that systems described herein may contain additional components not shown in the relevant figures. For example, the systems discussed herein may include additional components, such as communication modules and the like, which are omitted for clarity.

[0098] Although an implementation of the exemplary processes discussed herein may involve performing all the operations shown in the illustrated order, the present disclosure is not limited in this respect, and an implementation of the exemplary processes herein may in various examples involve only a subset of the operations shown, operations performed in a different order than illustrated, or additional operations.

[0099] Additionally, any one or more of the operations discussed herein can be performed in response to instructions provided by one or more computer program products. Such program products may include signal-carrying media that provide instructions which, when executed by, for example, a processor, can provide the functionality described herein. The computer program products may be provided in any form by one or more machine-readable media. Thus, for example, a processor containing one or more graphics processing units or one or more processor cores may execute one or more of the blocks of exemplary processes herein in response to program code and / or instructions or sets of instructions transmitted to the processor by one or more machine-readable media.In general, a machine-readable medium can transmit software in the form of program code and / or instructions or sets of instructions that can cause any of the devices and / or systems described herein to implement at least parts of the systems discussed herein or of any other module or component as discussed herein.

[0100] As used in any implementation described herein, the term "module" or "component" refers to any combination of software logic, firmware logic, hardware logic, and / or circuit arrangements configured to provide the functionality described herein. The software may be executed as a software package, code, and / or instruction set or instructions, and as used in any implementation described herein, "hardware" may include, for example, individually or in any combination, hardwired circuit arrangements, programmable circuit arrangements, state machine circuit arrangements, fixed-function circuit arrangements, execution unit circuit arrangements, and / or firmware storing instructions that are executed by programmable circuit arrangements.The modules can be implemented collectively or individually as a circuit arrangement that forms part of a larger system, such as an integrated circuit (IC), a system-on-a-chip (SoC), and so on.

[0101] Fig.Figure 20 is an illustrative diagram of an exemplary System 2000 arranged according to at least some implementations of the present disclosure. In various implementations, System 2000 may be a mobile system, although System 2000 is not limited in this respect. System 2000 may implement and / or perform any modules or techniques discussed herein. For example, System 2000 may be implemented in a personal computer (PC), a server, a laptop computer, an ultra-laptop computer, a tablet, a touchpad, a portable computer, a handheld computer, a palmtop computer, a personal digital assistant (PDA), a mobile phone, a mobile phone / PDA combination, a television, a smart device (e.g., a smartphone), a smart device, a smart TV ...A smartphone, a smart tablet, or a smart television), a mobile internet device (MID), a messaging device, a data communication device, cameras (e.g., compact cameras, superzoom cameras, digital single-lens reflex (DSLR) cameras), and so on may be involved. In some examples, System 2000 may be implemented via a cloud computing environment.

[0102] In various implementations, System 2000 includes a Platform 2002 coupled with a Display 2020. The Platform 2002 can receive content from a Content Device, such as one or more Content Service Device(s) 2030, one or more Content Delivery Device(s) 2040, or other similar content sources. A Navigation Controller 2050, comprising one or more navigation features, can be used to interact with, for example, the Platform 2002 and / or the Display 2020. Each of these components is discussed in more detail below.

[0103] In various implementations, the Platform 2002 can include any combination of a Chipset 2005, a Processor 2010, RAM 2012, an Antenna 2013, Memory 2014, a Graphics Subsystem 2015, Applications 2016, and / or Wireless 2018. The Chipset 2005 can provide intercommunication between the Processor 2010, RAM 2012, Memory 2014, Graphics Subsystem 2015, Applications 2016, and / or Wireless 2018. For example, the Chipset 2005 can include a Memory Adapter (not shown) capable of providing intercommunication with Memory 2014.

[0104] The Processor 2010 can be implemented as a Complex Instruction Set Computer (CISC) or Reduced Instruction Set Computer (RISC) processor, an x86-compatible processor, a multi-core or any other microprocessor, or a central processing unit (CPU). In various implementations, the Processor 2010 can be one or more dual-core processors, one or more dual-core mobile processors, and so on.

[0105] In 2012, main memory can be implemented as a volatile storage device, such as random access memory (RAM), dynamic random access memory (DRAM), or static RAM (SRAM).

[0106] Memory 2014 can be implemented as a non-volatile storage device, such as a magnetic disk drive, an optical disk drive, a tape drive, an internal storage device, an attached storage device, flash memory, battery-backed SDRAM (synchronous DRAM), and / or a network-accessible storage device. In various implementations, Memory 2014 can incorporate technology to enhance storage performance protection for valuable digital media, for example, when multiple hard disk drives are included.

[0107] The 2017 image signal processor can be implemented as a specialized digital signal processor or similar device used for single-frame image or video processing. In some examples, the 2017 image signal processor may be implemented based on a single-instruction, multiple-data (SIMD) or multiple-instruction, multiple-data (MDM) architecture, or similar. In some examples, the 2017 image signal processor may be characterized as a media processor. As discussed here, the 2017 image signal processor can be implemented based on a single-chip system-on-a-chip (SOC) architecture and / or a multi-core architecture.

[0108] The Graphics Subsystem 2015 can process images, such as still or video images, for display. The Graphics Subsystem 2015 can be, for example, a graphics processing unit (GPU) or a visual processing unit (VPU). An analog or digital interface can be used to communicate between the Graphics Subsystem 2015 and the Display 2020. The interface can be, for example, a High-Definition Multimedia Interface (HDMI), a DisplayPort, wireless HDMI, and / or other wireless HD-compliant technologies. The Graphics Subsystem 2015 can be integrated into the Processor 2010 or the Chipset 2005. In some implementations, the Graphics Subsystem 2015 can be a standalone device that communicates with the Chipset 2005.

[0109] The graphics and / or video processing techniques described here can be implemented in various hardware architectures. For example, graphics and / or video functionality can be integrated into a chipset. Alternatively, a discrete graphics and / or video processor can be used. Another implementation allows the graphics and / or video functions to be provided by a general-purpose processor containing a multi-core processor. In further embodiments, the functions can be implemented in a consumer electronics device.

[0110] The Radio 2018 may include one or more radio devices capable of transmitting and receiving signals using various suitable wireless communication technologies. Such technologies may involve communication across one or more wireless networks. Examples of wireless networks include (but are not limited to) wireless local area networks (WLANs), wireless personal area networks (WPANs), wireless metropolitan area networks (WMANs), cellular mobile networks, and satellite networks. When communicating across such networks, the Radio 2018 may operate according to one or more applicable standards in any version.

[0111] In various implementations, the Display 2020 can include any type of television-like monitor or television-like display. For example, the Display 2020 can include a computer display screen, a touchscreen display, a video monitor, a television-like device, and / or a television. The Display 2020 can be digital and / or analog. In various implementations, the Display 2020 can be a holographic display. The Display 2020 can also be a transparent surface capable of receiving a visual projection. Such projections can convey various forms of information, images, and / or objects. For example, such projections can be visual overlays for a mobile augmented reality (MAR) application. Under the control of one or more software applications 2016, the Platform 2002 can display a user interface 2022 on the Display 2020.

[0112] In various implementations, the Content Service Device(s) 2030 can be hosted by any national, international, and / or independent service and are thus accessible, for example, to Platform 2002 via the internet. The Content Service Device(s) 2030 can be coupled with Platform 2002 and / or Display 2020. Platform 2002 and / or the Content Service Device(s) 2030 can be coupled with a Network 2060 for communicating (e.g., sending and / or receiving) media information to and from a Network 2060. The Content Delivery Device(s) 2040 can also be coupled with Platform 2002 and / or Display 2020.

[0113] In various implementations, the Content Service Device(s) 2030 may include a cable television box, a personal computer, a network, a telephone, internet-enabled devices, or any device capable of transmitting digital information and / or digital content, and any other similar device capable of unidirectional or bidirectional communication of content between content providers and the Platform 2002 and the Display 2020 via the Network 2060 or directly. It is understood that content can be communicated unidirectionally and / or bidirectionally to and from any component in the System 2000 and a content provider via the Network 2060. Examples of content may include any media information, such as video, music, medical and gaming information, and so on.

[0114] The content service device(s) 2030 can receive content such as cable television programs, including media information, digital information, and / or other content. Examples of content providers include any provider of television, radio, or internet content via cable or satellite. The examples provided are not intended to limit the implementations according to this disclosure in any way.

[0115] In various implementations, the Platform 2002 can receive control signals from a Navigation Controller 2050 with one or more navigation features. The navigation features of the Navigation Controller 2050 can be used, for example, to interact with the User Interface 2022. In various embodiments, the Navigation Controller 2050 can be a pointing device, which can be a computer hardware component (in particular, a user interface device) that allows a user to input spatial (e.g., continuous or multidimensional) data into a computer. Many systems, such as graphical user interfaces (GUIs) and televisions and monitors, allow the user to control data and provide it to the computer using physical gestures.

[0116] Movements of the navigation features of the navigation controller 2050 can be reproduced on a display (e.g., the display 2020) by movements of a pointer, cursor, focus ring, or other visual indicators displayed on the display. For example, under the control of software applications 2016, the navigation features located on the navigation controller 2050 can be mapped to virtual navigation features displayed, for example, on the user interface 2022. In various embodiments, the navigation controller 2050 may not be a separate component but may be integrated into the platform 2002 and / or the display 2020. However, the present disclosure is not limited to the elements or the context shown or described herein.

[0117] In various implementations, drivers (not shown) may include technology that allows users to instantly turn the Platform 2002 on and off, like a television, with the touch of a button after the initial boot-up, if this feature is enabled. Program logic may allow the Platform 2002 to stream content to Media Adapters or other Content Service Device(s) 2030 or Content Delivery Device(s) 2040, even when the platform is switched off. Additionally, the Chipset 2005 may, for example, include hardware and / or software support for 5.1 surround sound audio and / or high-resolution 7.1 surround sound audio. Drivers may include a graphics driver for integrated graphics platforms. In various embodiments, the graphics driver may include a Peripheral Component Interconnect (PCI) Express graphics card.

[0118] In various implementations, any one or more of the components shown in System 2000 may be integrated. For example, Platform 2002 and Content Service Device(s) 2030 may be integrated, or Platform 2002 and Content Delivery Device(s) 2040 may be integrated, or Platform 2002, Content Service Device(s) 2030, and Content Delivery Device(s) 2040 may be integrated. In various embodiments, Platform 2002 and Display 2020 may form an integrated unit. For example, Display 2020 and Content Service Device(s) 2030 may be integrated, or Display 2020 and Content Delivery Device(s) 2040 may be integrated. These examples are not intended to limit the present disclosure.

[0119] In various embodiments, System 2000 can be implemented as a wireless system, a wired system, or a combination of both. When implemented as a wireless system, System 2000 can include components and interfaces suitable for communication over a shared wireless medium, such as one or more antennas, transmitters, receivers, transceivers, amplifiers, filters, control logic, and so on. An example of a shared wireless medium might include portions of a wireless spectrum, such as the RF spectrum.When implemented as a wired system, the system can contain 2000 components and interfaces suitable for communicating via wired communication media, such as input / output (I / O) adapters, physical connectors for connecting the I / O adapter to a corresponding wired communication medium, a network interface card (NIC), a disk controller, a video controller, an audio controller, and similar components. Examples of wired communication media include a wire, a cable, metal conductors, a printed circuit board (PCB), backplanes, switch fabrics, semiconductor material, a twisted pair cable, a coaxial cable, optical fibers, and so on.

[0120] Platform 2002 can create one or more logical or physical channels for communicating information. This information can include media information and control information. Media information can refer to any data representing content intended for a user. Examples of content include data from a voice conversation, video conference, streaming video, email message, voicemail message, alphanumeric symbols, graphics, an image, video, text, and so on. Data from a voice conversation can include speech information, pauses, background noise, comfort sounds, tones, and so on. Control information can refer to any data representing commands, instructions, or control words intended for an automated system.For example, control information can be used to route media information through a system or to instruct a node to process the media information in a predetermined way. However, the embodiments are not limited to the elements or in the context described in . Fig. 20 are shown or described.

[0121] As described above, System 2000 can be implemented in various physical styles or form factors. Fig.Figure 21 illustrates a small form-factor example device 2100 arranged according to at least some implementations of the present disclosure. In some examples, the system 2000 may be implemented via the device 2100. In other examples, other systems discussed herein, or parts thereof, may be implemented via the device 2100. In various embodiments, the device 2100 may, for example, be implemented as a mobile data processing device with wireless capabilities. A mobile data processing device may refer to any device with a processing system and a mobile power source or supply, such as one or more batteries.

[0122] Examples of mobile data processing devices may include a personal computer (PC), a laptop computer, an ultra-laptop computer, a tablet, a touchpad, a portable computer, a handheld computer, a palmtop computer, a personal digital assistant (PDA), a mobile phone, a mobile phone / PDA combination, a smart device (e.g., a smartphone, a smart tablet, or a mobile smart television), a mobile internet device (MID), a messaging device, a data communication device, cameras (e.g., compact cameras, superzoom cameras, digital single-lens reflex (DSLR) cameras), and so on.

[0123] Examples of mobile data processing devices include computers designed to be worn by a person, such as wrist computers, finger computers, ring computers, glasses computers, belt buckle computers, bracelet computers, shoe computers, clothing computers, and other wearable computers. In various embodiments, a mobile data processing device may, for example, be implemented as a smartphone capable of running computer applications as well as voice and / or data communications. Although some embodiments may be described using a mobile data processing device implemented as a smartphone as an example, it is understood that other embodiments may be implemented using other wireless mobile data processing devices. The embodiments are not limited in this respect.

[0124] As in Fig.As shown in Figure 21, the device 2100 can include a housing with a front 2101 and a rear 2102. The device 2100 includes a display 2104, an input / output (I / O) device 2106, a camera 1904, a camera 2105, and an integrated antenna 2108. The device 2100 can also include navigation features 2112. The I / O device 2106 can include any suitable I / O device for inputting information into a mobile data processing device. Examples of the I / O device 2106 include an alphanumeric keypad, a numeric keypad, a touchpad, input keys, buttons, switches, microphones, speakers, a speech recognition device, software, and so on. Information can also be input into the device 2100 by means of a microphone (not shown) or can be digitized by means of a speech recognition device.As shown, the device 2100 can include the camera 2105 and a flash 2110, which are integrated into the rear 2102 (or elsewhere) of the device 2100, and a camera 1904, which is integrated into the front 2101 of the device 2100. In some embodiments, one or both cameras 1904 and 2105 can be movable relative to the display 2104. The camera 1904 and / or the camera 2105 can be components of an imaging module or pipeline for producing color image data processed into a streaming video, which is output at the display 2104 and / or communicated remotely from the device 2100, for example, via the antenna 2108. For example, the camera 1904 can capture input images, and images with corrected eye contact can be provided to the display 2104 and / or communicated remotely from the device 2100 via the antenna 2108.

[0125] Various embodiments can be implemented using hardware elements, software elements, or a combination of both. Examples of hardware elements include processors, microprocessors, circuits, circuit components (e.g., transistors, resistors, capacitors, inductors, and so on), integrated circuits, application-specific integrated circuits (ASICs), programmable logic devices (PLDs), digital signal processors (DSPs), field-programmable gate arrays (FPGAs), logic gates, registers, semiconductor devices, chips, microchips, chipsets, and so on.Examples of software can include software components, programs, applications, computer programs, application programs, system programs, machine programs, operating system software, middleware, firmware, software modules, routines, subroutines, functions, methods, procedures, software interfaces, application program interfaces (APIs), instruction sets, data processing code, computer code, code segments, computer code segments, words, values, symbols, or any combination thereof. Determining whether an embodiment is implemented using hardware elements and / or software elements can vary according to any number of factors, such as desired computing speed, performance levels, thermal tolerances, processing cycle budget, input data rates, output data rates, memory resources, data bus speeds, and other design or performance requirements.

[0126] One or more aspects of at least one embodiment may be implemented by representative instructions stored on a machine-readable medium representing various logics within the processor. When read by a machine, these instructions cause the machine to create logic to execute the techniques described herein. Such representations, known as "IP kernels," may be stored on a physical, machine-readable medium and transmitted to various customers or production sites for loading into the manufacturing machines that actually produce the logic or the processor.

[0127] Although certain features set forth herein have been described in relation to exemplary implementations, this description should not be interpreted in a restrictive sense. Consequently, various modifications of the implementations described herein, as well as other implementations that are recognizable to those skilled in the art in the field to which this disclosure belongs, fall within the scope of the invention and the protection afforded by this disclosure.

[0128] In one or more first embodiments, a machine-implemented method for providing eye contact correction comprises encoding an eye area of ​​a source image via a pre-trained neural network to generate compressed features corresponding to the eye area of ​​the source image, applying a pre-trained classifier to the compressed features to determine a motion vector field for the eye area of ​​the source image, and deforming the eye area of ​​the source image based on the motion vector field and integrating the deformed eye area into a remaining part of the source image to produce an image with corrected eye contact.

[0129] Referring to the first embodiments, the pre-trained neural network comprises several layers, including at least one neural convolutional network layer.

[0130] Referring to the first embodiments, the pre-trained neural network has four layers, wherein the four layers comprise, in the following order: a first neural convolutional network layer, a second neural convolutional network layer, a first fully connected layer and a second fully connected layer, the second fully connected layer providing the compressed features.

[0131] Referring to the first embodiments, the pre-trained classifier comprises a pre-trained random forest classifier with a leaf corresponding to the motion vector field.

[0132] Referring to the first embodiments, wherein the pre-trained neural network comprises several layers including at least one neural convolutional network layer and / or the pre-trained classifier comprises a pre-trained random forest classifier with a leaf corresponding to the motion vector field.

[0133] Referring to the first embodiments, the method further comprises providing face detection and face orientation point detection on the source image and cropping the source image based on the face detection and face orientation point detection to generate the eye area.

[0134] Referring to the first embodiments, the method further comprises encoding and transmitting the final image to a remote device for display to a user.

[0135] Referring to the first embodiments, the method further comprises providing face detection and face orientation point detection on the source image, cropping the source image based on the face detection and face orientation point detection to create the eye area, and / or encoding and transmitting the final image to a remote device for display to a user.

[0136] Referring to the first embodiments, the encoding of the eye area and the application of the pre-trained classifier to the compressed features to determine the motion vector field for the eye area of ​​the source image are selectively provided based on a first relative position between a camera and a display, and the method further comprises, at a second relative position between the camera and the display, encoding the eye area of ​​the source image via a second pre-trained neural network to generate second compressed features corresponding to the eye area of ​​the source image, and applying a second pre-trained classifier to the second compressed features to determine a second motion vector field for the eye area of ​​the source image.and a deformation of the eye area of ​​the source image based on the second motion vector field and an integration of the deformed eye area into the remaining part of the source image to generate the image with corrected eye contact.

[0137] Referring to the first embodiments, the method further comprises receiving multiple pairs of eye area training images, wherein first images of the pairs of eye area training images have a viewing angle difference with respect to second images of the pairs of eye area training images, and training the pre-trained neural network based on encoding the first images by the pre-trained neural network to generate compressed training step features, decoding the compressed training step features to generate resulting first images corresponding to the first images, and evaluating the first images and the resulting first images.

[0138] Referring to the first embodiments, the method further comprises receiving multiple pairs of eye area training images, wherein first images of the pairs of eye area training images have a viewing angle difference with respect to second images of the pairs of eye area training images, and training the pre-trained neural network based on encoding the first images by the pre-trained neural network to generate compressed training step features, decoding the compressed training step features to generate resulting first images corresponding to the first images, and evaluating the first images and the resulting first images, wherein the evaluation of the first images and the resulting first images includes an error between the first images and the resulting first images.

[0139] Referring to the first embodiments, the method further comprises receiving multiple pairs of eye-area training images, wherein first images of the pairs of eye-area training images exhibit a viewing angle difference with respect to second images of the pairs of eye-area training images, and training the pre-trained neural network based on encoding the first images by the pre-trained neural network to generate compressed training step features, decoding the compressed training step features to generate resulting first images corresponding to the first images, and evaluating the first images and the resulting first images, wherein the evaluation of the first images and the resulting first images includes vertical filtering of the first images and the resulting first images to produce vertically filtered first images, respectively.to generate vertically filtered resulting first images, and includes determining an error between the vertically filtered first images and the vertically filtered resulting first images.

[0140] Referring to the first embodiments, the method further comprises receiving multiple pairs of eye-area training images, wherein first images of the pairs of eye-area training images exhibit a viewing angle difference with respect to second images of the pairs of eye-area training images, and training the pre-trained neural network based on encoding the first images by the pre-trained neural network to generate compressed training step features, decoding the compressed training step features to generate resulting first images corresponding to the first images, and evaluating the first images and the resulting first images.wherein evaluating the first images and the resulting first images includes determining the error between the first images and the resulting first images, and / or wherein evaluating the first images and the resulting first images includes vertically filtering the first images and the resulting first images to produce vertically filtered first images or vertically filtered resulting first images, and determining the error between the vertically filtered first images and the vertically filtered resulting first images.

[0141] Referring to the first embodiments, the method further comprises receiving multiple pairs of eye-area training images, wherein first images of the pairs of eye-area training images exhibit a viewing angle difference with respect to second images of the pairs of eye-area training images, training the pre-trained neural network based on encoding the first images by the pre-trained neural network to generate compressed training step features, decoding the compressed training step features to generate resulting first images corresponding to the first images, and evaluating the first images and the resulting first images, generating a likelihood map for each pixel of each of the first images of the pairs of eye-area training images, wherein each likelihood map comprises a sum of the absolute differences for each of the multiple candidate motion vectors corresponding to the pixel.and training a classifier pre-trained in the training step based on the likelihood maps and the compressed training step characteristics.

[0142] Referring to the first embodiments, the method further comprises receiving multiple pairs of eye-area training images, wherein first images of the pairs of eye-area training images exhibit a viewing angle difference with respect to second images of the pairs of eye-area training images, training the pre-trained neural network based on encoding the first images by the pre-trained neural network to generate compressed training step features, decoding the compressed training step features to generate resulting first images corresponding to the first images, and evaluating the first images and the resulting first images, generating a likelihood map for each pixel of each of the first images of the pairs of eye-area training images, wherein each likelihood map comprises a sum of the absolute differences for each of the multiple candidate motion vectors corresponding to the pixel.a training of a classifier pre-trained in the training step based on the likelihood maps and the compressed training step features, and a compression of the classifier pre-trained in the training step by a parameterized area adjustment to generate the pre-trained classifier.

[0143] In one or more second embodiments, a system for providing eye contact correction comprises a memory configured to store a source image and a processor coupled to the memory, wherein the processor encodes an eye area of ​​a source image via a pre-trained neural network to generate compressed features corresponding to the eye area of ​​the source image, applies a pre-trained classifier to the compressed features to determine a motion vector field for the eye area of ​​the source image, and deforms the eye area of ​​the source image based on the motion vector field and integrates the deformed eye area into a remaining part of the source image to produce an image with corrected eye contact.

[0144] Referring to the second embodiments, the pre-trained neural network comprises several layers, including at least one neural convolutional network layer.

[0145] Referring to the second embodiments, the pre-trained neural network has four layers, wherein the four layers comprise, in the following order: a first neural convolutional network layer, a second neural convolutional network layer, a first fully connected layer and a second fully connected layer, wherein the second fully connected layer provides the compressed features.

[0146] Referring to the second embodiments, the pre-trained classifier comprises a pre-trained random forest classifier with a leaf corresponding to the motion vector field.

[0147] Referring to the second embodiments, the processor further provides face detection and face orientation point detection on the source image and crops the source image based on the face detection and face orientation point detection to create the eye area.

[0148] Referring to the second embodiments, the processor further encodes the final image and transmits it to a remote device for display to a user.

[0149] Referring to the second embodiments, the encoding of the eye area and the application of the pre-trained classifier to the compressed features to determine the motion vector field for the eye area of ​​the source image are selectively provided based on a first relative position between a camera and a display, wherein the processor furthermore, at a second relative position between the camera and the display, encodes the eye area of ​​the source image via a second pre-trained neural network to generate second compressed features corresponding to the eye area of ​​the source image, and applies a second pre-trained classifier to the second compressed features to determine a second motion vector field for the eye area of ​​the source image.and deformed the eye area of ​​the source image based on the second motion vector field and integrated the deformed eye area into the remaining part of the source image to generate the image with corrected eye contact.

[0150] Referring to the second embodiments, the processor further receives several pairs of eye area training images, wherein first images of the pairs of eye area training images have a viewing angle difference with respect to second images of the pairs of eye area training images, and trains the pre-trained neural network based on encoding the first images by the pre-trained neural network to generate compressed training step features, decoding the compressed training step features to generate resulting first images corresponding to the first images, and evaluating the first images and the resulting first images.

[0151] Referring to the second embodiments, the processor further receives several pairs of eye-area training images, wherein first images of the pairs of eye-area training images have a viewing angle difference with respect to second images of the pairs of eye-area training images, and trains the pre-trained neural network based on encoding the first images by the pre-trained neural network to generate compressed training step features, decoding the compressed training step features to generate resulting first images corresponding to the first images, and evaluating the first images and the resulting first images, wherein the evaluation of the first images and the resulting first images includes the processor vertically filtering the first images and the resulting first images to produce vertically filtered first images and the resulting first images, respectively.to generate vertically filtered resulting first images, and determine an error between the vertically filtered first images and the vertically filtered resulting first images.

[0152] Referring to the second embodiments, the processor further receives several pairs of eye-area training images, wherein first images of the pairs of eye-area training images have a viewing angle difference with respect to second images of the pairs of eye-area training images, and trains the pre-trained neural network based on encoding the first images by the pre-trained neural network to generate compressed training step features, decoding the compressed training step features to generate resulting first images corresponding to the first images, and evaluating the first images and the resulting first images, wherein the processor further generates a likelihood map for each pixel of each of the first images of the pairs of eye-area training images, wherein each likelihood map comprises a sum of the absolute differences for each of the multiple candidate motion vectors corresponding to the pixel.and trains a classifier pre-trained in the training step based on the likelihood maps and the compressed training step characteristics.

[0153] Referring to the second embodiments, the processor further receives several pairs of eye-area training images, wherein first images of the pairs of eye-area training images have a viewing angle difference with respect to second images of the pairs of eye-area training images, trains the pre-trained neural network based on encoding the first images by the pre-trained neural network to generate compressed training step features, decoding the compressed training step features to generate resulting first images corresponding to the first images, and evaluating the first images and the resulting first images, generates a likelihood map for each pixel of each of the first images of the pairs of eye-area training images, wherein each likelihood map includes a sum of the absolute differences for each of the multiple candidate motion vectors corresponding to the pixel.and trains a classifier pre-trained in the training step based on the likelihood maps and the compressed training step characteristics.

[0154] Referring to the second embodiments, the processor further receives several pairs of eye-area training images, wherein first images of the pairs of eye-area training images have a viewing angle difference with respect to second images of the pairs of eye-area training images, trains the pre-trained neural network based on encoding the first images by the pre-trained neural network to generate compressed training step features, decoding the compressed training step features to generate resulting first images corresponding to the first images, and evaluating the first images and the resulting first images, generates a likelihood map for each pixel of each of the first images of the pairs of eye-area training images, wherein each likelihood map includes a sum of the absolute differences for each of the multiple candidate motion vectors corresponding to the pixel.trains a pre-trained classifier based on the likelihood maps and the compressed training step features, and compresses the pre-trained classifier through a parameterized area adjustment to generate the pre-trained classifier.

[0155] In one or more third embodiments, a system comprises means for encoding an eye area of ​​a source image via a pre-trained neural network to generate compressed features corresponding to the eye area of ​​the source image, means for applying a pre-trained classifier to the compressed features to determine a motion vector field for the eye area of ​​the source image, and means for deforming the eye area of ​​the source image based on the motion vector field and integrating the deformed eye area into a remaining part of the source image to produce an image with corrected eye contact.

[0156] Referring to the third embodiments, the pre-trained neural network comprises several layers, including at least one neural convolutional network layer.

[0157] Referring to the third embodiments, the pre-trained classifier comprises a pre-trained random forest classifier with a leaf corresponding to the motion vector field.

[0158] Referring to the third embodiments, the means for encoding the eye region and the means for applying the pre-trained classifier to the compressed features for determining the motion vector field for the eye region of the source image address a first relative position between a camera and a display, and the system further comprises, at a second relative position between the camera and the display, means for encoding the eye region of the source image via a second pre-trained neural network to generate second compressed features corresponding to the eye region of the source image, and means for applying a second pre-trained classifier to the second compressed features to determine a second motion vector field for the eye region of the source image.and means for deforming the eye area of ​​the source image based on the second motion vector field and integrating the deformed eye area into the remaining part of the source image to generate the image with corrected eye contact.

[0159] Referring to the third embodiments, the system further comprises means for receiving multiple pairs of eye area training images, wherein first images of the pairs of eye area training images have a viewing angle difference with respect to second images of the pairs of eye area training images, and means for training the pre-trained neural network based on encoding the first images by the pre-trained neural network to generate compressed training step features, decoding the compressed training step features to generate resulting first images corresponding to the first images, and evaluating the first images and the resulting first images.

[0160] Referring to the third embodiments, the system further comprises means for receiving multiple pairs of eye-area training images, wherein first images of the pairs of eye-area training images have a viewing angle difference with respect to second images of the pairs of eye-area training images, and means for training the pre-trained neural network based on encoding the first images by the pre-trained neural network to generate compressed training step features, decoding the compressed training step features to generate resulting first images corresponding to the first images, and evaluating the first images and the resulting first images, wherein the evaluation of the first images and the resulting first images includes vertical filtering of the first images and the resulting first images to produce vertically filtered first images and results, respectively.to generate vertically filtered resulting first images, and includes determining an error between the vertically filtered first images and the vertically filtered resulting first images.

[0161] Referring to the third embodiments, the system further comprises means for receiving multiple pairs of eye-area training images, wherein first images of the pairs of eye-area training images have a viewing angle difference with respect to second images of the pairs of eye-area training images, means for training the pre-trained neural network based on encoding the first images by the pre-trained neural network to generate compressed training step features, decoding the compressed training step features to generate resulting first images corresponding to the first images, and evaluating the first images and the resulting first images, and means for generating a likelihood map for each pixel of each of the first images of the pairs of eye-area training images.wherein each likelihood map comprises a sum of the absolute differences for each of the multiple candidate motion vectors corresponding to the pixel, and means for training a pre-trained classifier in the training step based on the likelihood maps and the compressed training step features.

[0162] In one or more fourth embodiments, at least one machine-readable medium comprises several instructions which, in response to an execution on a device, cause the device to perform eye contact correction by encoding an eye area of ​​a source image via a pre-trained neural network to generate compressed features corresponding to the eye area of ​​the source image, applying a pre-trained classifier to the compressed features to determine a motion vector field for the eye area of ​​the source image, and deforming the eye area of ​​the source image based on the motion vector field and integrating the deformed eye area into a remaining part of the source image to produce an image with corrected eye contact.

[0163] Referring to the fourth embodiments, the pre-trained neural network comprises several layers, including at least one neural convolutional network layer.

[0164] Referring to the fourth embodiments, the encoding of the eye area and the application of the pre-trained classifier to the compressed features to determine the motion vector field for the eye area of ​​the source image are provided selectively based on a first relative position between a camera and a display, and the machine-readable medium further comprises several instructions which, in response to an execution on a device, cause the device, at a second relative position between the camera and the display, to perform eye contact correction by encoding the eye area of ​​the source image via a second pre-trained neural network to generate second compressed features corresponding to the eye area of ​​the source image, and by applying a second pre-trained classifier to the second compressed features to determine a second motion vector field for the eye area of ​​the source image.and to provide a deformation of the eye area of ​​the source image based on the second motion vector field and an integration of the deformed eye area into the remaining part of the source image to generate the image with corrected eye contact.

[0165] Referring to the fourth embodiments, the machine-readable medium further comprises several instructions which, in response to an execution on a device, cause the device to perform eye contact correction by receiving several pairs of eye area training images, wherein first images of the pairs of eye area training images have a viewing angle difference with respect to second images of the pairs of eye area training images, and training the pre-trained neural network based on encoding the first images by the pre-trained neural network to generate compressed training step features, decoding the compressed training step features to generate resulting first images corresponding to the first images, and evaluating the first images and the resulting first images.

[0166] Referring to the fourth embodiments, the machine-readable medium further comprises several instructions which, in response to an execution on a device, cause the device to perform eye contact correction by receiving several pairs of eye area training images, wherein first images of the pairs of eye area training images have a viewing angle difference with respect to second images of the pairs of eye area training images, and by training the pre-trained neural network based on encoding the first images by the pre-trained neural network to generate compressed training step features, decoding the compressed training step features to generate resulting first images corresponding to the first images, and evaluating the first images and the resulting first images.wherein the evaluation of the first images and the resulting first images includes vertical filtering of the first images and the resulting first images to produce vertically filtered first images and vertically filtered resulting first images respectively, and includes determining an error between the vertically filtered first images and the vertically filtered resulting first images.

[0167] Referring to the fourth embodiments, the machine-readable medium further comprises several instructions which, in response to an execution on a device, cause the device to perform eye contact correction by receiving several pairs of eye area training images, wherein first images of the pairs of eye area training images have a viewing angle difference with respect to second images of the pairs of eye area training images, training the pre-trained neural network based on encoding the first images by the pre-trained neural network to generate compressed training step features, decoding the compressed training step features to generate resulting first images corresponding to the first images, evaluating the first images and the resulting first images, and generating a likelihood map for each pixel of each of the first images of the pairs of eye area training images.wherein each likelihood map comprises a sum of the absolute differences for each of the multiple candidate motion vectors corresponding to the pixel, providing a training of a pre-trained classifier based on the likelihood maps and the compressed training step features, and a compression of the pre-trained classifier by parameterized surface fitting to generate the pre-trained classifier.

[0168] In one or more fifth embodiments, at least one machine-readable medium may comprise several instructions which, in response to an execution on a data processing device, cause the data processing device to perform a method according to one of the above embodiments.

[0169] In one or more sixth embodiments, a device may include a means for carrying out a method according to one of the above embodiments.

[0170] It is understood that the embodiments are not limited to the embodiments described so far, but can be carried out with modification and alteration without deviating from the scope of protection of the attached claims.

[0171] For example, the embodiments described above may include specific combinations of features. However, the embodiments described above are not limited in this respect, and various implementations may include the implementation of only a subset of such features, the implementation of such features in a different order, the implementation of a different combination of such features, and / or the implementation of additional features beyond those explicitly listed. The scope of protection of the invention should therefore be determined with reference to the attached claims together with the full scope of equivalents to which these claims entitle.

Claims

A machine-implemented method for providing eye contact correction, comprising: receiving multiple pairs of eye area training images (911), wherein first images (912) of the pairs of eye area training images (911) exhibit a viewing angle difference with respect to second images (913) of the pairs of eye area training images (911); training a pre-trained neural network (700) based on encoding the first images (912) by the pre-trained neural network (700) to generate compressed training step features, decoding the compressed training step features to generate resulting first images (1413) corresponding to the first images (912), and evaluating the first images (912) and the resulting first images (1413);Receiving a source image (212) capturing a user with a camera (104); encoding an eye area (311) of the source image (212) via the pre-trained neural network (700) to generate compressed features (312) corresponding to the eye area (311) of the source image (212); applying a pre-trained classifier to the compressed features (312) to determine a motion vector field (313) for the eye area (311) of the source image (212); and deformation of the eye area (311) of the source image (212) based on the motion vector field (313) and integration of the deformed eye area (801) into a remaining part of the source image (212) to produce an image with corrected eye contact (215), wherein the image with corrected eye contact (215) gives the appearance that the user is looking into the camera (104). Method according to claim 1, wherein the pre-trained neural network (700) comprises multiple layers including at least one neural convolution network layer. The method of claim 1, wherein the pre-trained neural network (700) has four layers, the four layers comprising in the following order: a first neural convolutional network layer, a second neural convolutional network layer, a first fully connected layer and a second fully connected layer, wherein the second fully connected layer provides the compressed features (312). Method according to claim 1, wherein the pre-trained classifier comprises a pre-trained random forest classifier (306) with a leaf corresponding to the motion vector field (313). The method of claim 1, further comprising: providing face detection and face orientation point detection on the source image; and cropping the source image (212) based on the face detection and face orientation point detection to generate the eye area (311). The method of claim 1, further comprising: encoding and transferring the final image to a remote device for display to a user. The method of claim 1, wherein the encoding of the eye area (311) and the application of the pre-trained classifier to the compressed features (312) for determining the motion vector field (313) for the eye area (311) of the source image (212) are provided selectively based on a first relative position between a camera (104) and a display (102), wherein, at a second relative position between the camera (104) and the display (102), the method further comprises: encoding the eye area (311) of the source image (212) via a second pre-trained neural network (700) to generate second compressed features corresponding to the eye area (311) of the source image (212); applying a second pre-trained classifier to the second compressed features to determine a second motion vector field (313) for the eye area (311) of the source image. (212);and deformation of the eye area (311) of the source image (212) based on the second motion vector field (313) and integration of the deformed eye area (801) into the remaining part of the source image (212) to generate the image with corrected eye contact (215).; Method according to claim 1, wherein the evaluation of the first images (912) and the resulting first images (1413) comprises an error between the first images (912) and the resulting first images (1413). The method of claim 1, wherein the evaluation of the first images (912) and the resulting first images (1413) comprises: vertically filtering the first images (912) and the resulting first images (1413) to produce vertically filtered first images and vertically filtered resulting first images, respectively; and determining an error between the vertically filtered first images and the vertically filtered resulting first images. The method of claim 1, further comprising: generating a likelihood map for each pixel of each of the first images (912) of the pairs of eye area training images (911), wherein each likelihood map comprises a sum of the absolute differences for each of the multiple candidate motion vectors corresponding to the pixel; and training a pre-trained classifier based on the likelihood maps and the compressed training step features. The method of claim 10, further comprising: compressing the classifier pre-trained in the training step by means of a parameterized surface adjustment to generate the pre-trained classifier. System (1900) for providing eye contact correction, comprising: a camera (1904) for capturing a source image (212) capturing a user; a memory (1903) configured to store the source image; and a processor (1901) coupled to the memory, wherein the processor (1901) receives multiple pairs of eye area training images (911), wherein first images (912) of the pairs of eye area training images (911) exhibit a viewing angle difference with respect to second images (913) of the pairs of eye area training images (911); a pre-trained neural network (700) based on encoding the first images (912) by the pre-trained neural network (700) to generate compressed training step features; and decoding the compressed training step features to generate resulting first images (1413) corresponding to the first images (912).and trained by evaluating the first images (912) and the resulting first images (1413), and encodes an eye area (311) of the source image (212) via the pre-trained neural network (700) to generate compressed features (312) corresponding to the eye area (311) of the source image (212), applies a pre-trained classifier to the compressed features (312) to determine a motion vector field (313) for the eye area (311) of the source image (212), and deforms the eye area (311) of the source image (212) based on the motion vector field (313), and integrates the deformed eye area (801) into a remaining part of the source image (212) to generate an image with corrected eye contact (215), the image with corrected eye contact (215) giving the impression that the user is in the Camera (1904) looks. System (1900) according to claim 12, wherein the pre-trained neural network (700) comprises multiple layers including at least one neural convolution network layer. System (1900) according to claim 12, wherein the pre-trained classifier comprises a pre-trained random forest classifier (306) with a leaf corresponding to the motion vector field (313). System (1900) according to claim 12, wherein the encoding of the eye area (311) and the application of the pre-trained classifier to the compressed features (312) to determine the motion vector field (313) for the eye area (311) of the source image (212) are selectively provided based on a first relative position between a camera and a display, wherein the processor further, at a second relative position between the camera and the display, encodes the eye area (311) of the source image (212) via a second pre-trained neural network to generate second compressed features corresponding to the eye area (311) of the source image (212), and applies a second pre-trained classifier to the second compressed features to determine a second motion vector field (313) for the eye area (311) of the source image (212).and the eye area (311) of the source image (212) is deformed based on the second motion vector field (313) and the deformed eye area (801) is integrated into the remaining part of the source image (212) to generate the image with corrected eye contact (215). System (1900) according to claim 12, wherein the evaluation of the first images (912) and the resulting first images (1413) comprises the processor vertically filtering the first images (912) and the resulting first images (1413) to produce vertically filtered first images and vertically filtered resulting first images respectively, and determining an error between the vertically filtered first images and the vertically filtered resulting first images. System (1900) according to claim 16, wherein the processor further generates a likelihood map for each pixel of each of the first images of the pairs of eye area training images (911), wherein each likelihood map comprises a sum of the absolute differences for each of the multiple candidate motion vectors corresponding to the pixel, and trains a pre-trained classifier based on the likelihood maps and the compressed training step features. At least one machine-readable medium comprising several instructions which, in response to an execution on a device, cause the device to provide eye contact correction by: receiving multiple pairs of eye area training images (911), wherein first images (912) of the pairs of eye area training images (911) have a viewing angle difference with respect to second images (913) of the pairs of eye area training images (911); training a pre-trained neural network (700) based on encoding the first images (912) by the pre-trained neural network (700) to generate compressed training step features, decoding the compressed training step features to generate resulting first images (1413) corresponding to the first images (912), and evaluating the first images (912) and the resulting first images (1413);Receiving a source image (212) capturing a user with a camera (104); encoding an eye area (311) of the source image (212) via the pre-trained neural network (700) to generate compressed features (312) corresponding to the eye area (311) of the source image (212); applying a pre-trained classifier to the compressed features (312) to determine a motion vector field (313) for the eye area (311) of the source image (212); and deformation of the eye area (311) of the source image (212) based on the motion vector field (313) and integration of the deformed eye area (801) into a remaining part of the source image (212) to produce an image with corrected eye contact (215), wherein the image with corrected eye contact (215) gives the appearance that the user is looking into the camera (104). Machine-readable medium according to claim 18, wherein the pre-trained neural network (700) comprises multiple layers including at least one neural convolutional network layer. Machine-readable medium according to claim 18, wherein the encoding of the eye area (311) and the application of the pre-trained classifier to the compressed features (312) for determining the motion vector field (313) for the eye area (311) of the source image (212) are provided selectively based on a first relative position between a camera and a display, wherein the machine-readable medium further comprises several instructions which, in response to an execution on a device, cause the device to provide eye contact correction at a second relative position between the camera and the display by: encoding the eye area (311) of the source image (212) via a second pre-trained neural network to generate second compressed features corresponding to the eye area (311) of the source image (212);Applying a second pre-trained classifier to the second compressed features to determine a second motion vector field for the eye area (311) of the source image (212); and deforming the eye area (311) of the source image (212) based on the second motion vector field and integrating the deformed eye area (801) into the remaining part of the source image (212) to generate the image with corrected eye contact (215). Machine-readable medium according to claim 18, wherein the evaluation of the first images (912) and the resulting first images (1413) comprises: vertical filtering of the first images (912) and the resulting first images (1413) to generate vertically filtered first images and vertically filtered resulting first images, respectively; and determining an error between the vertically filtered first images and the vertically filtered resulting first images. Machine-readable medium according to claim 18, further comprising several instructions which, in response to an execution on a device, cause the device to provide eye contact correction by: generating a likelihood map for each pixel of each of the first images of the pairs of eye area training images (911), wherein each likelihood map comprises a sum of the absolute differences for each of the multiple candidate motion vectors corresponding to the pixel; training a pre-trained classifier based on the likelihood maps and the compressed training step features; and compressing the pre-trained classifier by parameterized area fitting to generate the pre-trained classifier.

Citation Information

Patent Citations

  • Method for correcting user's gaze direction in image, machine-readable storage medium and communication terminal

    US20150339512A1

  • Method and system for correcting gaze offset

    US20160011659A1