Method, processor, and readable medium for generating synthetic infrared images for line of sight estimation machine learning

By using a simulation engine to generate synthetic images of human eye features under infrared light and train a neural network, the problem of tedious and inefficient data collection for machine learning models in existing technologies is solved, and the accuracy and efficiency of line of sight estimation are improved.

CN111696163BActive Publication Date: 2025-09-09NVIDIA CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202010158005.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-03-15
Filing Date
2020-03-09
Publication Date
2025-09-09
Estimated Expiration
2040-09-22

AI Technical Summary

Technical Problem

In existing technologies, training machine learning models requires a large amount of real-world data, and the collection process is tedious and inefficient, making it difficult to scale.

Method used

A simulation engine is used to generate synthetic images that match the characteristics of the human eye under infrared light, and a training engine is used to train the neural network to simulate real-world environmental conditions to improve the accuracy of line of sight estimation.

Benefits of technology

Improves the accuracy and efficiency of machine learning models for estimating human eye gaze under infrared light, reducing dependence on real-world data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111696163B_ABST
    Figure CN111696163B_ABST
Patent Text Reader

Abstract

The present application provides a method for generating synthetic infrared images for use in machine learning for gaze estimation. One embodiment of the method includes computing one or more activation values ​​for one or more neural networks trained to infer eye gaze information based at least in part on eye positions of one or more images of one or more faces indicated by infrared light reflections from the one or more images.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Machine learning models are often trained to generate inferences or predictions related to real-world conditions. For example, a neural network can learn to estimate and / or track the position, orientation, velocity, speed, and / or other physical properties of an object in an image or collection of images. Consequently, the training data for the neural network can include real-world data, such as images and / or sensor readings related to the objects. However, collecting real-world data for training machine learning models can be tedious, inefficient, and / or difficult to scale. BRIEF DESCRIPTION OF THE DRAWINGS

[0002] In order that the manner in which the above-described features of various embodiments may be understood in detail, a more particular description of the inventive concept, briefly summarized above, may be given by reference to various embodiments, some of which are illustrated in the accompanying drawings. It should be noted, however, that the drawings illustrate only typical embodiments of the inventive concept and are therefore not to be considered limiting of the scope in any way, and that there may be other equally effective embodiments.

[0003] Figure 1 A system configured to implement one or more aspects of various embodiments is shown.

[0004] Figure 2 According to various embodiments Figure 1 A more detailed diagram of the simulation engine and training engine.

[0005] Figure 3 is a flowchart of method steps for generating images for training a machine learning model according to various embodiments.

[0006] Figure 4 is a flow chart of method steps for configuring a neural network for use, according to various embodiments.

[0007] Figure 5 is a block diagram illustrating a computer system configured to implement one or more aspects of the various embodiments.

[0008] Figure 6 According to various embodiments, Figure 5 Block diagram of a parallel processing unit (PPU) in a parallel processing subsystem of FIG.

[0009] Figure 7 According to various embodiments, Figure 6 Block diagram of the general processing cluster (GPC) in the parallel processing unit (PPU) of .

[0010] Figure 8 is a block diagram of an exemplary system-on-chip (SoC) integrated circuit in accordance with various embodiments. Detailed description

[0011] In the following description, numerous specific details are set forth to provide a more thorough understanding of various embodiments. However, it will be apparent to one skilled in the art that the present invention can be practiced without one or more of these specific details.

[0012] System Overview

[0013] Figure 1 A computing device 100 configured to implement one or more aspects of the various embodiments is shown. In one embodiment, the computing device 100 can be a desktop computer, a laptop computer, a smart phone, a personal digital assistant (PDA), a tablet computer, or any other type of computing device configured to receive input, process data, and optionally display images and suitable for practicing one or more embodiments. The computing device 100 is configured to run a simulation engine 120 and a training engine 122 located in the memory 116. It should be noted that the computing device described herein is illustrative and any other technically feasible configuration falls within the scope of the present disclosure. For example, multiple instances of the simulation engine 120 and the training engine 122 can be executed on a group of nodes in a distributed system to implement the functionality of the computing device 100.

[0014] In one embodiment, the computing device 100 includes, but is not limited to, an interconnect (bus) 112 connecting one or more processing units 102, an input / output (I / O) device interface 104 (I / O) coupled to one or more input / output devices 108, a memory 116, a storage 114, and a network interface 106. The processing unit 102 can be any suitable processor implemented as a central processing unit (CPU), a graphics processing unit (GPU) application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), an artificial intelligence (AI) accelerator, any other type of processing unit, or a combination of different processing units, such as a CPU configured to operate in conjunction with a GPU. In general, the processing unit 102 can be any technically feasible hardware unit capable of processing data and / or executing software applications. In addition, in the context of the present disclosure, the computing elements shown in the computing device 100 can correspond to physical computing systems (e.g., systems in a data center), or can be virtual computing instances executed within a computing cloud.

[0015] In one embodiment, I / O devices 108 include devices capable of providing input, such as a keyboard, a mouse, a touch-sensitive screen, and devices capable of providing output, such as a display. Furthermore, I / O devices 108 may include devices capable of both receiving input and providing output, such as a touch screen, a Universal Serial Bus (USB) port, and the like. I / O devices 108 may be configured to receive various types of input from an end user (e.g., a designer) of computing device 100, and may also provide various types of output, such as displayed digital images, digital video, or text, to the end user of computing device 100. In some embodiments, one or more I / O devices 108 are configured to couple computing device 100 to network 110.

[0016] In one embodiment, the network 110 is any technically feasible communication network that allows data to be exchanged between the computing device 100 and an external entity or device, such as a network server or another networked computing device. For example, the network 110 may include a wide area network (WAN), a local area network (LAN), a wireless (WiFi) network, and / or the Internet, among others.

[0017] In one embodiment, the memory 114 includes non-volatile memory for applications and data and may include fixed or removable disk drives, flash memory devices, and CD-ROM, DVD-ROM, Blu-ray, HD-DVD, or other magnetic, optical, or solid-state storage devices. The simulation engine 120 and the training engine 122 may be stored in the memory 114 and loaded into the memory 116 when executed.

[0018] In one embodiment, memory 116 includes random access memory (RAM) modules, flash memory cells, or any other type of memory cells, or a combination thereof. Processing unit 102, I / O device interface 104, and network interface 106 are configured to read data from and write data to memory 116. Memory 116 includes various software programs executable by processor 102, including simulation engine 120 and training engine 122, and application data associated with the software programs.

[0019] In one embodiment, the simulation engine 120 includes functionality for generating synthetic images of an object, while the training engine 122 includes functionality for training one or more machine learning models using the synthetic images. For example, the simulation engine 120 may generate a simulated image of the eye region of a face illuminated by infrared light, along with labels representing the location and / or attributes of objects in the eye region. The training engine 122 may input the simulated image and labels as training data for a neural network, and train the neural network to estimate and / or track the gaze and / or pupil location of the eye in the simulated image.

[0020] In some embodiments, the synthetic images generated by the simulation engine 120 are rendered using a geometric representation of the face that accurately simulates the anatomical features of the human eye and / or a rendering setting that simulates infrared light illuminating the eye region. Subsequently, the machine learning model generated by the training engine 122 using the images and labels generated by the simulation engine 120 can perform gaze estimation and / or other types of inferences related to human eyes illuminated under infrared light more accurately than machine learning models trained using other types of synthetic images of the eye region. Figure 2 The simulation engine 120 and the training engine 122 are described in more detail. Synthetic infrared image generation for line of sight estimation machine learning

[0021] Figure 2 According to various embodiments Figure 1 2 is a more detailed illustration of the simulation engine 120 and the training engine 122. In the illustrated embodiment, the simulation engine 122 renders a plurality of images 212 based on the geometric representation 202 of the face 204, one or more textures 208, one or more refractive indices 210, a camera 226, and / or one or more light sources 228. For example, the simulation engine 122 may use ray tracing techniques to determine a composite image 212 of a scene containing the eyes 206 in the face 204 based on the position of the camera 226 in the scene and / or the number and position of the light sources 228 in the scene.

[0022] In one or more embodiments, the simulation engine 120 utilizes various modifications to the geometric representation 202, textures 208, refractive index 210, and / or other components of the rendering pipeline to generate realistic and / or accurate images of the face 204 and / or eyes 206 that simulate the real-world conditions under which the images of the face and / or eyes were captured. For example, the simulation engine 120 may render an image 212 that simulates the capture of the eye region of a person's face under conditions that match those of a near-eye camera in a virtual reality and / or augmented reality headset. Such conditions may include, but are not limited to: illuminating the eye region with infrared wavelengths that are undetectable to humans; placing the camera in a position and / or configuration that allows for the capture of the eye region within the headset; and / or the presence of blur, noise, intensity variations, contrast variations, exposure variations, camera slip, and / or camera misalignment between the images.

[0023] In one embodiment, the simulation engine 120 adjusts the eye 206 and / or the face 204 in the geometric representation 202 to model the anatomical features of the human eye. Such adjustments include, but are not limited to, changing the pose 214 of the eye 206 based on a selected line of sight 220 for the eye 206, rotating the eye 206 216 to account for an axial difference 222 between the pupil axis of the eye 206 and the line of sight 220, and / or changing the pupil position 218 in the eye 206 to reflect movement of the pupil in the eye 206 during pupil constriction.

[0024] For example, the simulation engine 120 can include a geometric model of the human face 204 in the geometric representation 202. The geometric model can be generated using a three-dimensional (3D) scan of a real human face with manual modification, or the geometric model can be generated by the simulation engine 120 and / or another component based on features or characteristics of the real human face. The simulation engine 120 can rescale the face 204 to accommodate an average-sized human eye 206 (e.g., a diameter of 24 mm, a radius of curvature of 7.8 mm at the corneal vertex, and a radius of 10 mm at the scleral boundary) and insert the average-sized human eye 206 into the face 204 in the geometric representation 202.

[0025] Continuing with the above example, the simulation engine 120 can offset the face 204 by a small random offset relative to the camera 226 to simulate the sliding of a head-mounted camera in a real-world environment. The simulation engine 120 can also determine a line of sight 220 from the center of the eye 206 to a randomly selected point of interest (e.g., on a fixed screen at a distance from the face 204) and set the pose 214 (i.e., position and orientation) of the eye 206 to reflect the line of sight 220. The simulation engine 120 can perform a rotation 216 of the eye 206 by approximately 5 degrees in the current orientation (i.e., toward the side of the head) to simulate an anatomical axis disparity 222 between the line of sight 220 and the pupil axis in the eye 206 (i.e., a line perpendicular to the cornea that intersects the center of the pupil).

[0026] Continuing with the above example, the simulation engine 120 can randomly select eyelid positions in the face 204 ranging from fully open to approximately two-thirds closed. For the selected eyelid positions, the simulation engine 120 can cover approximately four times the surface area of ​​the eye 204 with the top eyelid as the surface area covered by the lower eyelid, and simulate the skin of the eyelid in synchronization with the pose 214 and / or rotation 216 of the eye 206 to achieve a physically correct eye appearance during a simulated blink.

[0027] Continuing with the above example, the simulation engine 120 can select a pupil size 224 from a useful range of 2 mm to 8 mm and adjust the pupil position 218 to simulate a supranasal (i.e., toward the forehead and above the nose) shift of the pupil that contracts due to illumination. For an 8 mm dilated pupil in dim light, the simulation engine 120 can adjust the pupil position 218 in the nasal and superior directions by approximately 0.1 mm. For a 4 mm pupil, the simulation engine 120 can adjust the pupil position 218 in the nasal position by approximately 0.2 mm and in the superior position by approximately 0.1 mm. For a 2 mm pupil that contracts in bright light, the simulation engine 120 can adjust the pupil position 218 in the superior position by approximately 0.1 mm.

[0028] In one or more embodiments, simulation engine 120 renders image 212 using rendering settings that match those encountered during near-eye image capture under infrared light. In one embodiment, the rendering settings used by simulation engine 120 to render image 212 include skin and iris textures 208 containing patterns and intensities that match the characteristics of corresponding surfaces observed under monochromatic infrared imaging. For example, simulation engine 120 may increase the brightness and / or uniformity of one or more texture files for the skin in face 204 and use subsurface scattering to simulate the smooth appearance of skin at one or more infrared light frequencies (e.g., 950 nm, 950 nm, 1000 nm, etc.). In another example, simulation engine 120 may use textures 208 and / or rendering techniques to render the sclera in eye 206 without veins, as veins are not visible under infrared light. In a third example, simulation engine 120 may alter the reflectivity of iris texture 208 to produce a reduced variation in iris color encountered under infrared light. In a fourth example, simulation engine 120 may randomly rotate one or more iris textures 208 about the center of the pupil in eye 206 to generate additional variation in the appearance of eye 206 within image 212 .

[0029] In one embodiment, the rendering settings used by simulation engine 120 to render image 212 include a refractive index 210 that matches the refractive index encountered when infrared light illuminates face 204 and / or eye 206. For example, simulation engine 120 may simulate air with a refractive index of unity and the cornea of ​​eye 206 with a refractive index of 1.38 to produce a highly reflective corneal surface in image 212 on which glint (i.e., reflection of light source 228 on the cornea) appears.

[0030] In one embodiment, the rendering settings used by the simulation engine 120 to render the images 212 include random attributes 230 that simulate real-world conditions under which images of a human face may be captured. For example, the simulation engine 120 may independently apply random amounts of exposure, noise, blur, intensity modulation, and / or contrast modulation to the iris, sclera, and / or skin regions of the face 204 in each image. In another example, the simulation engine 120 may randomize the skin color, iris hue, skin texture, and / or iris texture in the images 212. In a third example, the simulation engine 120 may randomize reflections in front of the eyes 206 in the images 212 to simulate reflection artifacts caused by glasses and / or any semi-transparent layers between the image capture device and the face in a real-world environment.

[0031] In one or more embodiments, the simulation engine 120 generates labels 236 associated with objects in the image 212. For example, the simulation engine 120 can output the labels 236 as metadata and store them with or separately from the corresponding image 212. The training engine 122 and / or another component can use the labels 236 and the image 212 to train and / or evaluate the performance of the machine learning model 258 and / or techniques for gaze tracking and / or gaze estimation techniques.

[0032] In one embodiment, the label 236 includes a gaze vector 230 that defines the gaze line 220 of the eye 206 within each image. For example, the gaze vector 230 can be represented as a two-dimensional (2D) viewpoint on a fixed screen at a distance from the face 204 that intersects the gaze line 220. In another example, the gaze vector 230 can include a 3D point that intersects the gaze line 220. In a third example, the gaze vector 230 can include horizontal and vertical gaze angles from a constant reference eye 206 position.

[0033] In one embodiment, label 236 includes location 232 of object in image 212. For example, label 236 may include the 3D location of face 204 and / or eye 206 relative to a reference coordinate system. In another example, label 236 may include the 2D location of the pupil in eye 206.

[0034] In one embodiment, simulation engine 120 may include region map 234 in labels 236 that associates various pixels in image 212 with semantic labels. For example, region map 234 may include regions of pixels in image 212 that are assigned to labels such as “pupil,” “iris,” “sclera,” “skin,” and “glint.”

[0035] In one embodiment, the simulation engine 120 uses the region map 234 to specify the pixel locations of anatomical features in the image 212, regardless of occlusion by the eyelid. For example, the simulation engine 120 can generate a region map for each image that identifies the skin, pupil, iris, sclera, and glint in the image. The simulation engine 120 can also generate another region map for each image that locates the pupil, iris, sclera, and / or glint in the image with the skin removed.

[0036] In one or more embodiments, the training engine 122 uses the images 212 and labels 236 generated by the simulation engine 120 to train the machine learning model 258 to predict and / or estimate one or more attributes associated with the face 204, eyes 206, and / or other objects in the image 212. For example, the training engine 122 can input one or more images 212 into a convolutional neural network and obtain corresponding outputs 260 from the convolutional neural network, which represent predicted gaze 220, pupil location 218, and / or other features associated with the eyes 206 and / or face 204 in the image. The training engine 122 can use the loss function of the convolutional neural network to calculate an error value 262 from the output 260 and the one or more labels 236 for the corresponding image. The training engine 122 can also use gradient descent and / or forward and backward propagation to update the weights of the convolutional neural network through a series of training iterations to reduce the error value 262 over time.

[0037] In some embodiments, error value 262 includes a loss term representing activation of machine learning model 258 on portions of image 212 that do not include eye 206. Continuing with the above example, training engine 122 can use region map 234 to identify skin regions of face 204 in image 212. Training engine 122 can calculate the loss term as the gradient of output 260 of convolutional neural network with respect to the input associated with the skin regions. Because the weights in the convolutional neural network are updated over time to reduce error value 262, the convolutional neural network can learn to ignore skin regions during estimation of gaze 220, pupil location 218, and / or other attributes of eye 206.

[0038] Figure 3 is a flow chart of method steps for generating images for training a machine learning model according to various embodiments. Figure 1 and Figure 2 Although the method steps are described with reference to a system, those skilled in the art will understand that any system configured to perform the method steps in any order falls within the scope of the present disclosure.

[0039] As shown, the simulation engine 120 generates 302 a geometric representation of at least a portion of a face. For example, the simulation engine 120 can generate the geometric representation using a geometric model of a human face that is scaled to accommodate an average-sized human eye. Within the geometric model, the simulation engine 120 can position the eye on the face based on a line of sight from the center of the eye to a randomly selected viewpoint, rotate the eye in a current orientation to simulate parallax between the eye's line of sight and the pupil axis, and / or position the pupil in the eye to simulate pupil constriction.

[0040] Next, simulation engine 120 configures 304 rendering settings that simulate infrared light illuminating the face, and then renders an image of the face according to the rendering settings. Specifically, simulation engine 120 renders 306 surfaces in the image using textures and refractive indices that match the observed properties of the surfaces under monochromatic infrared imaging. For example, simulation engine 120 may load and / or generate skin and iris textures with increased brightness and / or uniformity to simulate illumination of the skin and iris of the face at one or more infrared wavelengths. In another example, simulation engine 120 may set the refractive indices of air and materials in the face to produce a glint in the cornea of ​​the eye.

[0041] Simulation engine 120 also performs 308 subsurface scattering of infrared light in the image based on the surface and / or refractive index. For example, simulation engine 120 may perform subsurface scattering of infrared light relative to skin to generate a smooth appearance of skin in the image.

[0042] The simulation engine 120 further randomizes 310 one or more attributes of the image. For example, the simulation engine 120 may randomize blur, intensity, contrast, exposure, sensor noise, skin color, iris tint, skin texture, and / or reflections in the image. In another example, the simulation engine 120 may randomize the distance from the face to the camera used to render the image.

[0043] Figure 4 is a flow chart of method steps for configuring the use of a neural network according to various embodiments. Figure 1 and Figure 2 Although the method steps are described with reference to a system, those skilled in the art will understand that any system configured to perform the method steps in any order falls within the scope of the present disclosure.

[0044] As shown, the training engine 122 trains 402 one or more neural networks to infer eye gaze information based on the position of eyes in a face indicated by infrared light reflections from an image and labels associated with the image. For example, the training engine 122 may obtain a plurality of composite images of a face illuminated with infrared light that were generated by the simulation engine 120 using the above referenced image. Figure 3The training engine 122 also uses the labels generated for the images by the simulation engine 120 to reduce errors associated with the predictions of gaze, pupil position, and / or other eye gaze information output by the neural network from the images.

[0045] The labels may include semantic labels that map regions of pixels in the image to pupils, irises, sclera, skin, and / or corneal reflections in a face. The labels may also or alternatively include a gaze vector that defines gaze, eye position, and / or pupil position. During training of the neural network, the training engine 122 may use the labels to update weights in the neural network based on a loss term representing activations of the neural network on skin regions in the image, thereby training the neural network to ignore skin regions when inferring gaze information from the image.

[0046] Next, training engine 122 calculates 404 one or more activation values ​​for the neural network(s). For example, training engine 122 may generate activation values ​​based on additional images input to the neural network. The activation values ​​may include gaze, pupil location, and / or other gaze information associated with the face illuminated by infrared light in other images.

[0047] Exemplary Hardware Architecture

[0048] Figure 5 is a block diagram illustrating a computer system 500 configured to implement one or more aspects of various embodiments. In some embodiments, the computer system 500 is a server machine operating in a data center or cloud computing environment that provides scalable computing resources as a service over a network. In some embodiments, the computer system 500 implements Figure 1 The functionality of the computing device 100.

[0049] In various embodiments, computer system 500 includes, but is not limited to, a central processing unit (CPU) 502 and a system memory 504 coupled to a parallel processing subsystem 512 via a memory bridge 505 and a communication path 513. Memory bridge 505 is also coupled to an I / O (input / output) bridge 507 via a communication path 506, and I / O bridge 507 is in turn coupled to a switch 516.

[0050] In one embodiment, I / O bridge 507 is configured to receive user input information from an optional input device 508, such as a keyboard or mouse, and forward the input information to CPU 502 for processing via communication path 506 and memory bridge 505. In some embodiments, computer system 500 may be a server machine in a cloud computing environment. In such an embodiment, computer system 500 may not have input device 508. Instead, computer system 500 may receive equivalent input information by receiving commands in the form of messages sent over a network and received via network adapter 518. In one embodiment, switch 516 is configured to provide connections between I / O bridge 507 and other components of computer system 500, such as network adapter 518 and various plug-in cards 520 and 521.

[0051] In one embodiment, the I / O bridge 507 is coupled to a system disk 514, which can be configured to store content and applications and data for use by the CPU 502 and the parallel processing subsystem 512. In one embodiment, the system disk 514 provides non-volatile storage for applications and data and can include a fixed or removable hard drive, a flash memory device, and a CD-ROM (Compact Disc Read Only Memory), DVD-ROM (Digital Versatile Disc-ROM), Blu-ray, HD-DVD (High Definition DVD), or other magnetic, optical, or solid-state storage device. In various embodiments, other components such as a universal serial bus or other port connection, an optical disk drive, a digital versatile disk drive, a film recording device, etc. can also be connected to the I / O bridge 507.

[0052] In various embodiments, memory bridge 505 may be a northbridge chip, and I / O bridge 507 may be a southbridge chip. Additionally, communication paths 506 and 513, as well as other communication paths within computer system 500, may be implemented using any technically suitable protocol, including but not limited to AGP (Accelerated Graphics Port), HyperTransport, or any other bus or point communication protocol known in the art.

[0053] In some embodiments, the parallel processing subsystem 512 includes a graphics subsystem that transmits pixels to an optional display device 510, which can be any conventional cathode ray tube, liquid crystal display, light emitting diode display, etc. In such embodiments, the parallel processing subsystem 512 includes circuits optimized for graphics and video processing, including, for example, video output circuitry. Figure 6 and Figure 7 As described in more detail, such circuitry may be incorporated across one or more parallel processing units (PPUs) (also referred to herein as parallel processors) included within parallel processing subsystem 512 .

[0054] In other embodiments, parallel processing subsystem 512 includes circuitry optimized for general-purpose and / or computational processing. Again, such circuitry may be incorporated into one or more PPUs included within parallel processing subsystem 512, which are configured to perform such general-purpose and / or computational operations. In other embodiments, one or more PPUs included within parallel processing subsystem 512 may be configured to perform graphics processing, general-purpose processing, and computational processing operations. System memory 504 includes at least one device driver configured to manage processing operations of one or more PPUs within parallel processing subsystem 512.

[0055] In various embodiments, the parallel processing subsystem 512 may be coupled with Figure 5 One or more other elements of the CPU 502 may be integrated together to form a single system. For example, the parallel processing subsystem 512 may be integrated with the CPU 502 and other connecting circuits on a single chip to form a system on a chip (SoC).

[0056] In one embodiment, CPU 502 is the main processor of computer system 500, which controls and coordinates the operation of other system components. In one embodiment, CPU 502 issues commands that control the operation of the PPUs. In some embodiments, communication path 513 is a PCI Express link, as is known in the art, in which a dedicated channel is allocated to each PPU. Other communication paths may also be used. The PPUs advantageously implement a highly parallel processing architecture. The PPUs can be equipped with any amount of local parallel processing memory (PP memory).

[0057] It should be understood that the system shown herein is illustrative and that variations and modifications are possible. The connection topology, including the number and arrangement of bridges, the number of CPUs 502, and the number of parallel processing subsystems 512, can be modified as needed. For example, in some embodiments, the system memory 504 may be connected to the CPU 502 directly rather than through the memory bridge 505, and other devices will communicate with the system memory 504 through the memory bridge 505 and the CPU 502. In other embodiments, the parallel processing subsystem 512 may connect to the I / O bridge 507 or directly to the CPU 502 rather than to the memory bridge 505. In other embodiments, the I / O bridge 507 and the memory bridge 505 may be integrated into a single chip rather than existing in the form of one or more discrete devices. Finally, in some embodiments, there may be no Figure 5 For example, switch 516 could be omitted, and network adapter 518 and add-in cards 520 , 521 would connect directly to I / O bridge 507 .

[0058] Figure 6 According to various embodiments, Figure 5 A block diagram of a parallel processing unit (PPU) 602 in the parallel processing subsystem 512 of FIG. Figure 6 One PPU 602 is depicted, but as described above, parallel processing subsystem 512 may include any number of PPUs 602. As shown, PPU 602 is coupled to a local parallel processing (PP) memory 604. PPU 602 and PP memory 604 may be implemented using one or more integrated circuit devices (e.g., a programmable processor, an application-specific integrated circuit (ASIC), or a memory device) or in any other technically feasible manner.

[0059] In some embodiments, PPU 602 comprises a graphics processing unit (GPU), which can be configured to implement a graphics rendering pipeline to perform various operations related to generating pixel data based on graphics data provided by CPU 502 and / or system memory 504. When processing graphics data, PP memory 604 can be used as graphics memory to store one or more conventional frame buffers and, if necessary, one or more other rendering targets. PP memory 604 can be used to store and update pixel data and transmit final pixel data or display frames to optional display device 510 for display. In some embodiments, PPU 602 can also be configured for general processing and computing operations. In some embodiments, computer system 500 can be a server machine in a cloud computing environment. In such an embodiment, computer system 500 may not have display device 510. Instead, computer system 500 can generate equivalent output information by sending commands in the form of messages over a network via network adapter 518.

[0060] In some embodiments, the CPU 502 is the main processor of the computer system 500, which controls and coordinates the operation of the other system components. In one embodiment, the CPU 502 issues commands that control the operation of the PPU 602. In some embodiments, the CPU 502 writes a command stream for the PPU 602 into a data structure (in a memory) that can be located in the system memory 504, the PP memory 604, or other storage location accessible to both the CPU 502 and the PPU 602. Figure 5 or Figure 6 (not explicitly shown in the figure). A pointer to the data structure is written to a command queue (also referred to herein as a push buffer) to initiate processing of the command stream in the data structure. In one embodiment, the PPU 602 reads the command stream from the command queue and then executes the commands asynchronously with respect to the operation of the CPU 502. In embodiments where multiple push buffers are generated, the application can specify an execution priority for each push buffer through the device driver to control the scheduling of different push buffers.

[0061] In one embodiment, PPU 602 includes an I / O (input / output) unit 605 that communicates with the rest of computer system 500 via communication path 513 and memory bridge 505. In one embodiment, I / O unit 605 generates data packets (or other signals) for transmission on communication path 513 and also receives all incoming data packets (or other signals) from communication path 513, directing the incoming packets to the appropriate components of PPU 602. For example, commands related to processing tasks can be directed to host interface 606, while commands related to memory operations (e.g., reading from or writing to PP memory 604) can be directed to crossbar unit 610. In one embodiment, host interface 606 reads each command queue and sends the command stream stored in the command queue to front end 612.

[0062] As above combined Figure 5 As described above, the connection of PPU 602 to the rest of computer system 500 can vary. In some embodiments, parallel processing subsystem 512, including at least one PPU 602, is implemented as an add-in card that can be inserted into an expansion slot of computer system 500. In other embodiments, PPU 602 can be integrated on a single chip with a bus bridge (e.g., memory bridge 505 or I / O bridge 507). Likewise, in other embodiments, some or all elements of PPU 602 can be included with CPU 502 in a single integrated circuit or system on a chip (SoC).

[0063] In one embodiment, the front end 612 sends processing tasks received from the host interface 606 to a work distribution unit (not shown) within the task / work unit 607. In one embodiment, the work distribution unit receives a pointer to the encoded processing task as task metadata (TMD) and stores it in memory. The pointer to the TMD is contained in a command stream, which is stored as a command queue and received by the front end unit 612 from the host interface 606. The processing task, which can be encoded as a TMD, also includes an index associated with the data to be processed as state parameters and commands that define how to process the data. For example, the state parameters and commands can define a program to be executed on the data. Similarly, for example, a TMD can specify the number and configuration of a set of CTAs. Typically, each TMD corresponds to a task. The task / work unit 607 receives tasks from the front end 612 and ensures that the GPC 608 is configured to a valid state before starting the processing task specified by each TMD. Each TMD can be assigned a priority, which is used to schedule the execution of the processing task. Processing tasks can also be received from the processing cluster array 630. Optionally, the TMD may include a parameter that controls whether the TMD is added to the head or tail of a list of processing tasks (or a list of pointers to processing tasks), thereby providing another level of control over execution priority.

[0064] In one embodiment, PPU 602 implements a highly parallel processing architecture based on a processing cluster array 630 that includes a set of C general processing clusters (GPCs) 608, where C ≥ 1. Each GPC 608 is capable of executing a large number of concurrent threads (e.g., hundreds or thousands), where each thread is an instance of a program. In various applications, different GPCs 608 can be assigned to process different types of programs or to perform different types of computations. The assignment of GPCs 608 can vary depending on the workload generated by each type of program or computation.

[0065] In one embodiment, the memory interface 614 includes a set of D partition units 615, where D ≥ 1. Each partition unit 615 is coupled to one or more dynamic random access memories (DRAMs) 620 residing within the PPM memory 604. In some embodiments, the number of partition units 615 is equal to the number of DRAMs 620, and each partition unit 615 is coupled to a different DRAM 620. In other embodiments, the number of partition units 615 may be different from the number of DRAMs 620. Those skilled in the art will appreciate that any other technically suitable memory device may be used in place of the DRAMs 620. In operation, various render targets, such as texture maps and frame buffers, may be stored on the DRAMs 620, allowing the partition units 615 to write portions of each render target in parallel to efficiently use the available bandwidth of the PPM memory 604.

[0066] In one embodiment, a given GPC 608 can process data to be written to any DRAM 620 in PP memory 604. In one embodiment, the crossbar unit 610 is configured to route the output of each GPC 608 to the input partition unit 615 of any GPC 608 or to any other GPC 608 for further processing. The GPCs 608 communicate with the memory interface 614 via the crossbar unit 610 to read from or write to the various DRAMs 620. In some embodiments, the crossbar unit 610 has a connection to the I / O unit 605 in addition to the connection to the PP memory 604 via the memory interface 614, thereby enabling the processing cores in different GPCs 608 to communicate with the system memory 504 or other memory that is not local to the PPU 602. Figure 6 In an embodiment of the present invention, the crossbar unit 610 is directly connected to the I / O unit 605. The crossbar unit 610 can use virtual channels to separate the traffic flow between the GPCs 608 and the partition units 615.

[0067] In one embodiment, the GPCs 608 can be programmed to perform processing tasks associated with a wide variety of applications, including, but not limited to, linear and nonlinear data transformations, filtering of video and / or audio data, modeling operations (e.g., applying physical laws that determine the position, velocity, and other properties of an object), image rendering operations (e.g., tessellation shader, vertex shader, geometry shader, and / or pixel / fragment shader programs), general computational operations, etc. In operation, the PPU 602 is configured to transfer data from the system memory 504 and / or the PP memory 604 to one or more on-chip memory units, process the data, and write the resulting data back to the system memory 504 and / or the PP memory 604. The resulting data can then be accessed by other system components, including the CPU 502, another PPU 602 in the parallel processing subsystem 512, or another parallel processing subsystem 512 in the computer system 500.

[0068] In one embodiment, any number of PPUs 602 may be included in the parallel processing subsystem 512. For example, multiple PPUs 602 may be provided on a single add-in card, or multiple add-in cards may be connected to the communication path 513, or one or more of the PPUs 602 may be integrated into a bridge chip. The PPUs 602 in a multi-PPU system may be identical to or different from one another. For example, different PPUs 602 may have different numbers of processing cores and / or different amounts of PP memories 604. In implementations where multiple PPUs 602 are present, these PPUs may operate in parallel to process data at a higher throughput than a single PPU could. Systems containing one or more PPUs 602 may be implemented in a variety of configurations and forms, including but not limited to desktops, laptops, handheld personal computers or other handheld devices, servers, workstations, gaming consoles, embedded systems, and the like.

[0069] Figure 7 According to various embodiments, Figure 6 6. As shown, the GPC 608 includes, but is not limited to, a pipeline manager 705, one or more texture units 715, a preROP unit 725, a work distribution crossbar 730, and an L1.5 cache 735.

[0070] In one embodiment, GPC 608 can be configured to execute a large number of threads in parallel to perform graphics, general processing and / or computing operations. As used herein, a "thread" refers to an instance of a specific program executed on a specific set of input data. In some embodiments, single instruction multiple data (SIMD) instruction issuance technology is used to support the parallel execution of a large number of threads without providing multiple independent instruction units. In other embodiments, single instruction multiple thread (SIMT) technology is used to support the parallel execution of a large number of generally synchronized threads by using a general instruction unit configured to issue instructions to a group of processing engines in GPC 608. Under an execution mechanism in which all processing engines typically execute the same instructions, SIMT execution allows different threads to more easily follow different execution paths through a given program. It will be understood by those skilled in the art that the SIMD processing scheme represents a functional subset of the SIMT processing scheme.

[0071] In one embodiment, the operation of GPC 608 is controlled by pipeline manager 705, which distributes processing tasks received from a work distribution unit (not shown) within task / work unit 607 to one or more streaming multiprocessors (SMs) 710. Pipeline manager 705 may also be configured to control work distribution crossbar 730 by specifying the destination of processed data output by SM 710.

[0072] In various embodiments, the GPC 608 includes a group of M SMs 710, where M ≥ 1. Furthermore, each SM 710 includes a group of function execution units (not shown), such as execution units and load-store units. Processing operations for any function execution unit can be pipelined, which enables a new instruction to be issued for execution before the previous instruction has completed execution. Any combination of function execution units within a given SM 710 can be provided. In various embodiments, the function execution units can be configured to support a variety of different operations, including integer and floating-point arithmetic (e.g., addition and multiplication), comparison operations, Boolean operations (AND, OR, 5OR), bit-shifting, and various algebraic functions (e.g., planar interpolation and trigonometric functions, exponential and logarithmic functions, etc.). Advantageously, the same function execution unit can be configured to perform different operations.

[0073] In various embodiments, each SM 710 includes multiple processing cores. In one embodiment, SM 710 includes a large number (e.g., 128, etc.) of different processing cores. Each core may include a fully pipelined, single-precision, double-precision, and / or mixed-precision processing unit, including a floating-point arithmetic logic unit (ALU) and an integer ALU. In one embodiment, the ALU implements the IEEE 754-2008 standard for floating-point arithmetic. In one embodiment, the core includes 64 single-precision (32-bit) floating-point cores, 64 integer cores, 32 double-precision (64-bit) floating-point cores, and 8 tensor cores.

[0074] In one embodiment, the Tensor Cores are configured to perform matrix operations, and, in one embodiment, one or more Tensor Cores are included in these cores. In particular, the Tensor Cores are configured to perform deep learning matrix arithmetic, such as convolution operations for neural network training and inference. In one embodiment, each Tensor Core operates on a 4×4 matrix and performs a matrix multiplication and accumulation operation D=A×B+C, where A, B, C, and D are 4×4 matrices.

[0075] In one embodiment, the matrix multiplication inputs A and B are 16-bit floating-point matrices, while the accumulation matrices C and D can be 16-bit floating-point or 32-bit floating-point matrices. The Tensor Cores operate on 16-bit floating-point input data with 32-bit floating-point accumulation. The 16-bit floating-point multiplication requires 64 operations and produces a full-precision product, which is then accumulated with other intermediate products using 32-bit floating-point addition to perform a 4×4×4 matrix multiplication. In practice, the Tensor Cores are used to perform larger two-dimensional or higher-dimensional matrix operations that are constructed from these smaller elements. APIs such as the CUDA 9 C++ API expose specialized matrix load, matrix multiplication and accumulation, and matrix store operations to efficiently use Tensor Cores in CUDA-C++ programs. At the CUDA level, the warp-level interface assumes that the 16×16 size matrix spans all 32 threads of the warp.

[0076] Neural networks rely heavily on matrix math operations, and complex, multi-layer networks require significant floating-point performance and bandwidth for efficiency and speed. In various embodiments, the SM 710 features thousands of processing cores optimized for matrix math operations and delivers tens to hundreds of TFLOPS of performance. The SM 710 provides a computing platform capable of delivering the performance required for deep neural network-based artificial intelligence and machine learning applications.

[0077] In various embodiments, each SM 710 may also include a plurality of special function units (SFUs) that perform special functions (e.g., attribute evaluation, reciprocal square root, etc.). In one embodiment, the SFUs may include a tree traversal unit configured to traverse a hierarchical tree data structure. In one embodiment, the SFUs may include a texture unit configured to perform texture map filtering operations. In one embodiment, the texture unit is configured to load a texture map (e.g., a 2D array of texels) from memory and sample the texture map to generate sampled texture values ​​for use by a shader program executed by the SM. In various embodiments, each SM 710 may also include a plurality of load / store units (LSUs) that implement load and store operations between a shared memory / L1 cache and a register file internal to the SM 710.

[0078] In one embodiment, each SM 710 is configured to process one or more thread groups. As used herein, a "thread group" or "thread bundle" refers to a group of threads that execute the same program simultaneously on different input data, with a thread in the group being assigned to a different execution unit within the SM 710. The number of threads included may be less than the number of execution units in the SM 710, in which case some executions may be idle during the cycle in which the thread group is processed. A thread group may also include more threads than the number of execution units in the SM 710, in which case processing can proceed in consecutive clock cycles. Since each SM 710 can support up to G thread groups simultaneously, up to G*M thread groups can be executed in a GPC 608 at any given time.

[0079] Additionally, in one embodiment, multiple related thread groups can be simultaneously active (in different stages of execution) within an SM 710. Such a collection of thread groups is referred to herein as a "cooperative thread array" ("CTA") or "thread array." The size of a particular CTA is equal to m*k, where k is the number of concurrently executing threads within the thread group, typically an integer multiple of the number of execution units within an SM 710, and m is the number of concurrently active thread groups within the SM 710. In some embodiments, a single SM 710 can simultaneously support multiple CTAs, with such CTAs having a granularity for allocating work to the SM 710.

[0080] In one embodiment, each SM 710 contains a level 1 (L1) cache, or uses space in a corresponding L1 cache external to the SM 710 to support, among other things, load and store operations performed by the execution units. Each SM 710 may also access a level 2 (L2) cache (not shown) that is shared between all GPCs 608 in the PPU 602. The L2 cache may be used to transfer data between threads. Finally, the SM 710 may also access off-chip "global" memory, which may include PP memory 604 and / or system memory 504. It will be appreciated that any memory external to the PPU 602 may be used as global memory. Additionally, as Figure 7 As shown, a level 1.5 (L1.5) cache (735) can be included in GPC 608 and is configured to receive and store data requested from memory by SM 710 via memory interface 614. This data includes, but is not limited to, instructions, uniform data, and constant data. In embodiments having multiple SMs 710 within GPC 608, SMs 710 can advantageously share common instructions and data cached in L1.5 cache 735.

[0081] In one embodiment, each GPC 608 may have an associated memory management unit (MMU) 720 that is configured to map virtual addresses to physical addresses. In various embodiments, the MMU 720 may reside within the GPC 608 or within the memory interface 614. The MMU 720 includes a set of page table entries (PTEs) that are used to map virtual addresses to physical addresses of tiles or memory pages and, optionally, cache line indices. The MMU 720 may include an address translation lookaside buffer (TLB) or may reside within the SM 710, within one or more L1 caches, or within the GPC 608.

[0082] In one embodiment, in graphics and compute applications, GPC 608 may be configured such that each SM 710 is coupled to a texture unit 715 to perform texture mapping operations, such as determining texture sample locations, reading texture data, and filtering texture data.

[0083] In one embodiment, each SM 710 sends the processed tasks to the work distribution crossbar 730 to provide the processed tasks to another GPC 608 for further processing or to store the processed tasks in an L2 cache (not shown), parallel processing memory 604, or system memory 504 via the crossbar unit 610. In addition, a pre-raster operation (preROP) unit 725 is configured to receive data from the SM 710 and direct the data to one or more raster operation (ROP) units within the partition unit 615 to perform color blending optimization, organize pixel color data, and perform address translation.

[0084] It will be appreciated that the architecture described herein is illustrative and that variations and modifications are possible. Among other things, any number of processing units, such as SMs 710, texture units 715, or preROP units 725, may be included within a GPC 608. Furthermore, as described above in conjunction with Figure 6 As described, a PPU 602 may include any number of GPCs 608 configured to be functionally similar to one another, such that execution behavior does not depend on which GPC 608 receives a particular processing task. Furthermore, each GPC 608 operates independently of the other GPCs 608 in the PPU 602 to execute tasks for one or more application programs.

[0085] Figure 88 is a block diagram of an exemplary system-on-chip (SoC) integrated circuit 800 according to various embodiments. The SoC integrated circuit 800 includes one or more application processors 802 (e.g., CPUs), one or more graphics processors 804 (e.g., GPUs), one or more image processors 806, and / or one or more video processors 808. The SoC integrated circuit 800 also includes peripheral devices or bus components, such as a serial interface controller 814 that implements a universal serial bus (USB), a universal asynchronous receiver / transmitter (UART), a serial peripheral interface (SPI), a secure digital input output (SDIO), an inter-IC sound (IC sound), and a serial interface controller 814 that implements a universal serial bus (USB). 2 The SoC integrated circuit 800 also includes a display device 818 coupled to a display interface 820, such as a High-Definition Multimedia Interface (HDMI) and / or a Mobile Industry Processor Interface (MIPI). The SoC integrated circuit 800 also includes a flash memory subsystem 824 that provides storage on the integrated circuit, and a memory controller 822 that provides a storage interface for accessing the storage device.

[0086] In one or more embodiments, one or more types of integrated circuit components are used to implement the SoC integrated circuit 800. For example, the SoC integrated circuit 800 may include one or more processor cores for the application processor 802 and / or the graphics processor 804. Additional functionality associated with the serial interface controller 814, the display device 818, the display interface 820, the image processor 806, the video processor 808, AI acceleration, machine vision, and / or other specialized tasks may be provided through application specific integrated circuits (ASICs), application specific standard parts (ASSPs), field programmable gate arrays (FPGAs), and / or other types of custom components.

[0087] In summary, the disclosed technology performs rendering of synthetic infrared images for machine learning of gaze detection. The rendering can be performed using a geometric representation of the eyes, pupils, and / or other objects in the face that reflects gaze, anatomical axis parallax, and / or pupil constriction in the face. The rendering can also or alternatively be performed using texture, refractive index, and / or subsurface scattering of infrared illumination that simulates the face. The rendering can also or alternatively be performed using random properties such as blur, intensity, contrast, distance from the face to the camera, exposure, sensor noise, skin color, iris hue, skin texture, and / or reflections.

[0088] One technical advantage of the disclosed technology includes high accuracy and / or better performance in machine learning models that are trained using images to perform near-eye gaze estimation, long-range gaze tracking, pupil detection, and / or other types of reasoning. Another technical advantage of the disclosed technology includes improved interaction between humans and virtual reality systems, augmented reality systems, and / or other human-computer interaction technologies that utilize the output of machine learning models. Thus, the disclosed technology provides technical improvements in the training, execution, and performance of machine learning models, computer systems, applications, and / or technologies for performing gaze tracking, generating facial simulation images, and / or interacting with humans.

[0089] 1. In some embodiments, the processor includes one or more arithmetic logic units (ALUs) to calculate one or more activation values ​​for one or more neural networks trained to infer eye gaze information based at least in part on an eye position in one or more images of one or more faces indicated by infrared light reflections from the one or more images.

[0090] 2. The processor of clause 1, wherein the one or more neural networks are further trained to infer eye gaze information based at least on labels associated with the one or more images.

[0091] 3. The processor of clauses 1-2, wherein the labels comprise a region mapping of pixels in the image to semantic labels.

[0092] 4. The processor of clauses 1-3, wherein the semantic label comprises at least one of pupil, iris, sclera, skin, and corneal reflection.

[0093] 5. The processor of clauses 1-4, wherein the label comprises at least one of a gaze vector, an eye position, and a pupil position.

[0094] 6. The processor of clauses 1-5, wherein the one or more images are generated by generating a geometric representation of at least a portion of a face; and rendering the image of the geometry using a rendering setting that simulates illuminating the face with infrared light.

[0095] 7. A processor of clauses 1-6, wherein generating a geometric representation of at least a portion of a face comprises placing the eyes in the face based on a line of sight from the center of the eyes to a randomly selected gaze point, and rotating the eyes in a temporal direction to model the difference between the line of sight and the pupil axis of the eyes.

[0096] 8. The processor of clauses 1-7, wherein generating the geometric representation of at least a portion of the face further comprises positioning a pupil in the eye to model constrictive movement of the pupil.

[0097] 9. The processor of clauses 1-8, wherein rendering the image comprises rendering the surface in the image using a texture that matches observed characteristics of the surface under monochromatic infrared imaging.

[0098] 10. The processor of clauses 1-9, wherein rendering the images comprises rendering the surface in the one or more images using a refractive index that matches observed properties of the surface under monochromatic infrared imaging; and surface scattering infrared light in the one or more images based on the refractive index.

[0099] 11. In some embodiments, a method includes generating neural network weight information to configure one or more processors to infer eye gaze information based at least in part on eye positions of one or more images of one or more faces indicated by infrared light reflections from the one or more images.

[0100] 12. The method of clause 11, wherein the one or more images are rendered using one or more random image attributes.

[0101] 13. The method of clauses 11-12, wherein the one or more random image properties include blur, intensity, contrast, distance of the face to the camera, exposure, and at least one noise in the sensor.

[0102] 14. The method of clauses 11-13, wherein the one or more random image attributes include at least one of skin color, iris tint, skin texture, and reflection.

[0103] 15. The method of clauses 11-14, wherein generating the neural network weight information to configure the one or more processors to infer the eye gaze information includes updating the neural network weight information based on a loss term of activation of the neural network on skin areas in one or more images representing one or more faces.

[0104] 16. In some embodiments, a machine-readable medium has stored thereon a set of instructions that, if executed by one or more processors, cause the one or more processors to at least train one or more neural networks to infer eye gaze information based at least in part on eye positions of one or more images of one or more faces indicated by infrared light reflections from the one or more images.

[0105] 17. The non-transitory computer-readable medium of clause 16, wherein the one or more images are rendered using one or more random image attributes.

[0106] 18. The non-transitory computer-readable medium of clauses 16-17, wherein the one or more random attributes include at least one of blur, intensity, contrast, distance of the face from the camera, exposure, sensor noise, skin color, iris tint, skin texture, and reflection.

[0107] 19. The non-transitory computer-readable medium of clauses 16-18, wherein the eye gaze information comprises at least one of a line of sight and a pupil position.

[0108] 20. The non-transitory computer-readable medium of clauses 16-19, wherein the one or more neural networks comprise a convolutional neural network.

[0109] Any and all combinations of any claim elements recited in any claim and / or any elements described in this application, in any manner, are within the intended scope of this disclosure and protection.

[0110] The descriptions of the various embodiments have been given for purposes of illustration, but are not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments.

[0111] Aspects of the present embodiments may be embodied as systems, methods, or computer program products. Thus, aspects of the present disclosure may take the form of entirely hardware embodiments, entirely software embodiments (including firmware, resident software, microcode, etc.), or embodiments combining software and hardware aspects that may generally be referred to as such. Furthermore, aspects of the present disclosure may take the form of a computer program product embodied in one or more computer-readable media having computer-readable program code embodied thereon.

[0112] Any combination of one or more computer-readable media can be utilized. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared or semiconductor system, device or equipment, or any suitable combination of the foregoing. More specific examples (non-exhaustive list) of computer-readable storage media will include the following: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), only internal memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any other suitable combination of the foregoing. In the context of this article, a computer-readable storage medium can be any tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment.

[0113] Aspects of the present disclosure are described above with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present disclosure. It will be understood that each block of the flowchart illustration and / or block diagram and the combination of blocks in the flowchart illustration and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer or other programmable data processing device to produce a machine. When the instructions are executed by the processor of a computer or other programmable data processing device, the functions / actions specified in the flowchart and / or block diagram blocks can be implemented. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor or a field programmable gate array.

[0114] The flow charts and block diagrams in the accompanying drawings illustrate the architecture, functions and operations of possible implementations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a module, segment or portion of a code, which includes one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative embodiments, the functions indicated in the box may not occur in the order indicated in the figure. For example, depending on the functions involved, the two boxes shown in succession can actually be executed substantially simultaneously, or these boxes can sometimes be executed in the opposite order. It should also be noted that each box in the block diagram and / or flow chart illustration and the combination of the boxes in the block diagram and / or flow chart illustration can be implemented by a system based on dedicated hardware that performs a specified function or action, or a combination of dedicated hardware and computer instructions.

[0115] While the foregoing is directed to embodiments of the present disclosure, other and further embodiments of the disclosure may be devised without departing from the basic scope thereof, and the scope of the disclosure is determined by the claims that follow.

Claims

1. A processor, comprising: One or more circuits generate a geometric representation of at least a portion of a simulated face, simulate real-world conditions, render an image of the geometric representation to simulate capture of an eye region under the simulated real-world conditions, and cause one or more neural networks to be trained to infer eye gaze information based at least in part on one or more synthetically generated images that include one or more labels indicating one or more eye positions.

2. The processor according to claim 1, wherein: The one or more neural networks are also trained to infer the eye gaze information based at least on labels generated for the one or more synthetically generated images.

3. The processor according to claim 2, wherein: The labeling includes a region mapping of pixels in the image to semantic labels.

4. The processor of claim 3, wherein the semantic label comprises at least one of pupil, iris, sclera, skin, and corneal reflection.

5. The processor according to claim 2, wherein: The label includes at least one of a gaze vector, an eye position, and a pupil position.

6. The processor of claim 1 , wherein the one or more synthetically generated images are produced by: An image of the geometric representation is rendered using render settings that simulate infrared light illuminating the face.

7. The processor of claim 1 , wherein generating the geometric representation of at least a portion of the face comprises: Place the eyes on the face based on the line of sight from the eye center to a randomly chosen viewpoint; and The eye is rotated in its current orientation to simulate a difference between the line of sight and the pupil axis of the eye.

8. The processor of claim 7, wherein generating the geometric representation of at least a portion of the face further comprises positioning a pupil in the eye to simulate a constrictive displacement of the pupil.

9. The processor of claim 1, wherein rendering the image comprises rendering the surface in the image using a texture that matches characteristics of the surface as observed under monochromatic infrared imaging.

10. The processor of claim 6, wherein rendering the image comprises: rendering the surface in the one or more synthetically generated images using a refractive index that matches properties of the surface as observed under monochromatic infrared imaging; and Subsurface scattering of the infrared light in the one or more synthetically generated images is performed based on the refractive index.

11. A method comprising: Generate a geometric representation of at least a portion of a simulated face, simulate real-world conditions, render an image of the geometric representation to simulate capture of an eye region under the simulated real-world conditions, and train one or more neural networks to infer eye gaze information based at least in part on one or more synthetically generated images that include one or more labels indicating one or more eye positions.

12. The method of claim 11, wherein the one or more synthetically generated images are rendered using one or more random image attributes.

13. The method of claim 12, wherein the one or more random image attributes include at least one of blur, intensity, contrast, face-to-camera distance, exposure, and sensor noise.

14. The method of claim 12, wherein the one or more random image attributes include at least one of skin color, iris hue, skin texture, and reflection.

15. The method according to claim 11, further comprising: The neural network is updated based on a loss term of activations of the one or more neural networks over skin areas in the one or more synthetically generated images representing one or more faces.

16. A non-transitory computer-readable medium having stored thereon a set of instructions that, if executed by one or more processors, cause the one or more processors to at least: Generate a geometric representation of at least a portion of a simulated face, simulate real-world conditions, render an image of the geometric representation to simulate capture of an eye region under the simulated real-world conditions, and cause one or more neural networks to be trained to infer eye gaze information based at least in part on the one or more synthetically generated images of one or more faces including infrared light reflection indications of one or more images of one or more labels indicating one or more eye locations.

17. The non-transitory computer-readable medium of claim 16, wherein the one or more synthetically generated images are rendered using one or more random image attributes.

18. The non-transitory computer-readable medium of claim 17, wherein the one or more random attributes include at least one of blur, intensity, contrast, distance of the face from the camera, exposure, sensor noise, skin color, iris tint, skin texture, and reflection.

19. The non-transitory computer-readable medium of claim 16, wherein the eye gaze information comprises at least one of a gaze and a pupil position.

20. The non-transitory computer-readable medium of claim 16, wherein the one or more neural networks comprise a convolutional neural network.