Fixation determination using glare as an input
A machine learning system addresses the challenge of determining gaze direction in the presence of glare by explicitly learning to handle glare points, resulting in more accurate and robust gaze estimation.
Patent Information
- Application Number
- JP2021100474
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-06-16
- Filing Date
- 2021-06-16
- Publication Date
- 2025-06-23
- Estimated Expiration
- 2041-06-16
AI Technical Summary
Conventional gaze determination systems struggle to accurately determine the gaze direction of a subject when one or both eyes are occluded or obscured by glare, as light reflections can obscure the eyes and reduce system accuracy.
A machine learning system that explicitly learns to handle glare by inputting a separate representation of glare points, along with representations of the subject's face and eyes, into a network architecture. This allows the system to continue estimating gaze direction reliably even in the presence of glare.
The system achieves more accurate and robust gaze estimation by explicitly considering glare, leading to improved performance even in images with bright spots that might obscure the eyes.
Smart Images

Figure 0007696764000001 
Figure 0007696764000002 
Figure 0007696764000003
Abstract
Description
Technical Field
[0001] This application claims the benefit of U.S. Provisional Patent Application No. 62 / 948,793, filed on December 16, 2019, which is hereby incorporated by reference in its entirety.
Background Art
[0002] Embodiments of the present disclosure generally relate to machine learning systems. More particularly, embodiments of the present disclosure relate to gaze determination performed using a machine learning system having glare as an input.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] Recent convolutional neural networks (CNNs) have been developed to estimate the gaze direction of a subject represented in an image. Such a CNN can estimate the direction in which a subject is looking from an input image of the subject, for example, by determining certain features regarding the subject's eyes. This enables a system using such a CNN to automatically determine the direction in which the subject is looking and react in real time accordingly.
[0005] However, conventional gaze determination systems are not without drawbacks. When one or both eyes are occluded or not clearly represented in the input image in some other way, such systems often have difficulty determining the gaze direction of the subject. Conventional CNN systems often have particular difficulty determining the gaze direction in the presence of glare. Light from various sources often reflects off the eyes, glasses, or other nearby surfaces, resulting in bright spots in the image of the subject that can at least partially obscure his or her eyes from the image sensor (e.g., camera), thereby reducing the accuracy of conventional gaze determination systems.
[0006] The prior systems have attempted to compensate for or reduce the effects of glare using several methods, such as modifying the light source used in illumination, using various polarization techniques, identifying and removing glare pixels, or using other information such as head pose to compensate for the lack of eye information. However, each of these methods has proven to be of limited effectiveness. Accordingly, a system and method for a more robust gaze estimation system that incorporates glare points as explicit inputs into a gaze estimation machine learning network architecture are described herein. An exemplary gaze estimation system may use a camera or other image determination device and a processor such as a parallel processor having the ability to perform the inference operation of a machine learning network such as a CNN. In certain embodiments of the present disclosure, the system may receive an image of a subject captured by a camera. One or more machine learning networks are constructed to take as input: a separate representation of glare in the image of the subject, one or more representations of at least a portion of the subject's face, and a portion of the image corresponding to at least one eye of the subject. From these inputs, the machine learning network determines and outputs an estimated value of the gaze direction of the subject as shown in the image. This gaze direction may be transmitted to various systems that initiate operations based on the determined gaze.
[0007] The glare representation, the representation of the face of the subject, and the image portion corresponding to the eyes of the subject can be determined in any manner. For example, prior to the processing by the aforementioned machine learning network, the system can separately determine each of the glare representation, the face of the subject, and the image portion corresponding to the eyes of the subject. The glare representation may be a binary mask of glare points in the input image determined in any manner, while the representation of the face of the subject may be a similarly determined binary mask of a portion of the face of the subject including at least his or her eyes. The image portion corresponding to the eyes may be an eye crop, where the eyes are identified using any shape or object recognition technique, whether based on machine learning or not. **Means for Solving the Problem**
[0008] In one embodiment of the present disclosure, one machine learning model may have a separate representation of glare, for example, a binary mask of glare points, and a representation of at least a portion of the face of the subject, for example, a binary mask of a portion of the face of the subject including at least his or her eyes.
[0009] In another embodiment of the present disclosure, the separate representation of glare can be input into one machine learning model, while the representation of at least a portion of the face of the subject can be input into another machine learning model. Specifically, the glare representation can be input into one set of Fully connected layer while the representation of at least a portion of the face can be input into a feature extraction model.
[0010] In one embodiment of the present disclosure, one machine learning model has, as an input, one of the portions of the image corresponding to one of the eyes of the subject, while another machine learning model has, as an input, another portion of the image corresponding to the other eye of the subject. That is, two different machine learning models have, as their respective inputs, two eye crops obtained from the image of the subject.
[0011] In another embodiment of the present disclosure, one machine learning model has as input one representation of a person's face, e.g., the aforementioned binary mask of the person's face determined from an input image, and another machine learning model has as input another representation of the person's face, e.g., a coarse face grid also determined from the input image.
[0012] The final output of the machine learning model is the estimated gaze direction or focus point. Any action can be initiated in response to this output. For example, the output of the machine learning model may be the gaze direction of a vehicle driver, and the action initiated may be a vehicle action executed in response to the determined gaze direction. For example, when it is determined that the driver is not looking in the direction in which the vehicle is moving, the vehicle may initiate an alert to the driver to draw his or her attention back to the road.
[0013] Accordingly, embodiments of the present disclosure may be considered to provide a machine learning system for determining a person's gaze direction by explicitly learning glare. More specifically, the position of the glare point in the image is used as a separate input to the machine learning model. Accordingly, the machine learning model of embodiments of the present disclosure receives as input both an image of a person and an explicit representation of the glare point extracted from the image. From these (and optionally other) inputs, the gaze direction is determined. Various actions can then be initiated based on the determined gaze direction.
[0014] When considering the following detailed description in conjunction with the accompanying drawings, the foregoing and other objects and advantages of the present disclosure will become apparent, and like reference characters in the figures refer to like parts.
Brief Description of the Drawings
[0015]
Figure 1A
Figure 1B
Figure 2
Figure 3
Figure 4A
Figure 4B
Figure 4C
Figure 4D
Figure 5
Figure 6
Figure 7
[0016] In one embodiment, the present disclosure relates to a machine learning system and method for learning glare and thereby determining a fixation direction in a manner that has greater resistance to the effects of glare. The machine learning system has, in addition to the image itself, as explicit input, a separate representation of glare, for example, information on the position of glare points in the image. In this manner, the machine learning system explicitly considers glare during the determination of the fixation direction, thereby producing more accurate results for images containing glare.
[0017] Figure 1A conceptually shows the function of a conventional gaze determination process in the presence of glare. Usually, conceptually, an image 100 is input into a CNN 110 trained to estimate the gaze direction of a subject in the input image 100. The presence of glare often prevents this process from producing satisfactory results. Here, for example, a glare spot 130 on the subject's glasses makes it difficult for the CNN 110 to detect the subject's eyes well enough to determine their gaze direction with high confidence, obscuring the subject's eyes. Since the CNN 110 has insufficient information to determine the gaze, the output of the CNN 110 can be inaccurate and meaningless, or the affected input image can be completely discarded.
[0018] Figure 1B conceptually shows the function of the glare-resistant gaze determination process of an embodiment of the present disclosure. In contrast to the conventional CNN 110 of Figure 1A, an embodiment of the present disclosure describes a machine learning model that specifically learns glare and is more resistant to its effects. Thus, when a separate representation of the image 100 and the glare points 120 extracted from the image 100 are input into the machine learning model 140 of an embodiment of the present disclosure, the model 140 can continue to infer the gaze direction of the subject with sufficient reliability, resulting in an output gaze vector that is accurate even in the presence of glare spots 130 that may partially obscure the subject's eyes. Figures 1A and 1B show the image 100 input into the machine learning model 140 to conceptually show that the machine learning model 140 takes as input information extracted from or related to a captured image of the subject, but it should be noted that in practice, the input to the machine learning model 140 may also be different representations or sets of information other than images. For example, as will be described further below, the input to the model 140 may also be a face grid including glare points.
[0019] Figure 2 is a block diagram representation of an exemplary glare-resistant machine learning model of an embodiment of the present disclosure. The machine learning model of Figure 2 includes feature extraction layers 230, 235, and 240 arranged as shown, and Fully connected layerIt includes 245 and 255. Specifically, the feature extraction layer 230 has two different inputs, the glare representation 200 and the face representation 205, which are concatenated by the concatenation block 210. The concatenated inputs are then sent to the feature extraction layer 230. The feature extraction block 230 can extract the features of the input image 100 according to any method. For example, the feature extraction block 230 may be the feature learning part of a CNN that uses any convolutional kernel and pooling layer structured in any suitable manner and suitable for extracting the features of the subject expected to be captured in the input image. Therefore, the feature extraction block 230 outputs features that can represent the orientation of the face of the subject represented in the input image.
[0020] The feature extraction layers 235 and 240 may each have a single input, the left eye representation 215 and the right eye representation 220. For example, the left eye representation 215 may be a crop of the left eye of the subject in the input image 100, and the right eye representation 220 may be a crop of the right eye of the subject in the input image 100. Similar to the feature extraction layer 230, the feature extraction layers 235 and 240 may each be any feature learning part of a CNN that uses any convolutional kernel and pooling layer structured in any suitable manner and suitable for extracting the features of the subject's eyes. The output of each feature extraction layer 235, 240 is a set of eye features that can represent the gaze direction of the pupils of the inputs 215 and 220, respectively. Since the feature extraction layers 235 and 240 are each configured to analyze the eyes as inputs, their corresponding weight values may be shared or the same for efficiency. However, the embodiments of the present disclosure contemplate any weight values for each of the feature extraction layers 235 and 240, whether shared or not.
[0021] Fully connected layer 245 has the face grid 225 as its single input. Fully connected layer 245 may be any classifier suitable for classifying the input face image into a position classification indicating the position and / or location of the subject's face in the input image 100 where there is a face. For example, Fully connected layer245 may be a multi - layer perceptron or any other CNN configured and trained to classify an input object into one of several individual positions. Fully connected layer Thus, Fully connected layer 245 outputs the likelihood of the position of a face within the input image 100.
[0022] The outputs of the feature extraction layers 230, 235, and 240, and Fully connected layer the output of 245 are concatenated by the concatenation block 250 and then output an estimated value of the gaze direction of the subject in the input image 100. Fully connected layer This is input into 255. Fully connected layer 255 may be any classifier suitable for classifying the input features of the subject's face into a direction classification indicating the direction the subject is looking. For example, Fully connected layer 255 may be a multi - layer perceptron or any other CNN configured and trained to classify the input face and eye features and positions into one of several individual gaze directions. Fully connected layer Or it may be other. Fully connected layer The output classification of 255 may be any representation of the gaze direction, such as a vector, a focus point on an arbitrary predefined virtual plane, or the like.
[0023] The glare representation 200 may be any separate representation of glare within the input image 100, such as the glare representation 120. That is, the glare representation 200 may be any input that conveys the position and size of any glare present within the input image 100. As one example, the glare representation 200 may be a binary mask of glare points extracted from the input image 100, that is, an image containing only those pixels of the input image 100 that are determined to represent glare. Thus, all pixels of the binary mask, except those pixels whose corresponding pixels in the input image 100 are determined to be glare pixels, are black pixels (e.g., pixels having a value of 0 in the binary representation). As another example, the glare representation 200 may be a vector of the positions of those pixels within the input image 100 that are determined to contain glare.
[0024] The face representation 205 may be any separate representation of the face within the input image 100. That is, the face representation 205 may be any input that conveys only the face of the subject present within the input image 100. For example, the face representation 205 may be a binary mask of face pixels extracted from the input image 100, where all the face pixels, or those pixels of the input image 100 determined to represent the face of the subject, are of one color (e.g., have one value), while all other pixels are black pixels (e.g., have a different value). As another example, the face representation 200 may be a vector of the positions of those pixels within the input image 100 determined to represent the face of the subject.
[0025] The left-eye representation 215 and the right-eye representation 220 may each be any representation of the corresponding eye of the subject within the input image 100 that can indicate the gaze direction of that eye. For example, the left-eye representation 215 may be a cropped portion of the input image 100 that includes only the left eye of the subject, while the right-eye representation 220 may be a cropped portion of the input image 100 that includes only the right eye of the subject.
[0026] The face grid 225 may be any input that conveys the position of the face within the input image 100. As one example, the face grid 225 may be a binary mask of the face of the subject, similar to the face representation 205. This face grid 225 may be of any resolution, for example, the same resolution as the input image 100, or a coarser representation such as a 25×25 square image binary mask where the individual boxes corresponding to the face within the input image 100 are represented using one or more non-black colors, while the remaining boxes are black. The face grid 225 may be of any other resolution, for example, 12×12 or the like. The face represented in the face grid 225 may be scaled to any particular size to assist in calculations or kept at the same scale as the input image 100. Further, the face grid 225 may contain any information, whether an image or not, that conveys the spatial position of the face of the subject within the input image 100. For example, the face grid 225 may be a set of landmark points of the face of the subject determined in any manner.
[0027] Figure 3 is a block diagram representation of one exemplary gaze determination system of an embodiment of the present disclosure. Here, any electronic computing device that includes a processing circuit capable of performing the gaze determination operation of an embodiment of the present disclosure may be used. Computing device 300 is in electronic communication with both camera 310 and gaze assistance system 320. During operation, camera 310, which may correspond to the cabin camera 441 of FIGS. 4A and 4C below, captures and transmits an image of the subject to computing device 300, which then implements, for example, the machine learning model of FIG. 2, to determine the input shown in FIG. 2 from the image of camera 310 and calculate the output gaze direction of the subject. Computing device 300 transmits this gaze direction to gaze assistance system 320, which takes an action or executes one or more operations in response. Computing device 300 may be any one or more electronic computing devices suitable for implementing the machine learning model of an embodiment of the present disclosure, for example, computing device 500, which will be described in more detail later.
[0028] The gaze assistance system 320 may be any system having the ability to perform one or more actions based on the gaze direction it receives from the computing device 300. Any configuration of the camera 310, the computing device 300, and the gaze assistance system 320 is intended. As one example, the gaze assistance system 320 may be an autonomous vehicle having the ability to determine and respond to the gaze direction of a driver or another passenger, such as the autonomous vehicle 400 described in more detail below. In this example, the camera 310 and the computing device 300 may be disposed within the vehicle, while the gaze assistance system 320 may represent the vehicle itself. The camera 310 may be placed at any position within the vehicle that permits it to view the driver or passenger. Thus, the camera 310 can capture images of the driver or passenger and transmit them to the computing device 300 that calculates the inputs 200, 205, 215, 220, and 225 and determines the driver's gaze direction. This gaze direction can then be transmitted to another software module that determines, for example, an action that the vehicle can take in response. For example, the vehicle can determine that the gaze direction represents a distracted driver or a driver not paying attention to the road and can initiate any type of action in response. Such actions can include any type of alert issued to the driver (e.g., visual or audible alerts, alerts on a head-up display, or the like), autonomous driving initiation, braking or turning maneuvers, or any other action. The computing device 300 may be disposed within the vehicle of the gaze assistance system 320 as a local processor, or may be a remote processor that receives images from the camera 310 and wirelessly transmits the gaze direction or commands to the vehicle of the gaze assistance system 320.
[0029] As another example, the gaze assistance system 320 may be a virtual reality or augmented reality system having the ability to display images in response to a user's movement and gaze. In this example, the gaze assistance system 320 includes a virtual reality or augmented reality display, such as a headset configured to be worn by the user and project images thereon. The camera 310 and the computing device 300 may be placed within the headset such that the camera 310 captures an image of the user's eyes and the computing device 300 determines his or her gaze direction. This gaze direction may then be transmitted to the virtual reality or augmented reality display, which may then perform any action in response. For example, to conserve computing resources, the gaze assistance system 320 may render only those virtual reality or augmented reality elements that are within the user's field of view as determined using the determined gaze direction. Similarly, the gaze assistance system 320 may be able to warn the user of an object or event that has been determined to be outside of the user's field of view but that the user may wish to avoid or that may be of interest. As in the example of the autonomous vehicle described above, the computing device 300 of the virtual reality or augmented reality system may be located within the system 320, for example, within the headset itself, or may be located remotely such that the images are wirelessly transmitted to the computing device 300 and the computed gaze direction is then wirelessly sent back to the headset where it may then perform various actions in response.
[0030] As yet another example, the gaze assistance system 320 may be a computer-based advertising system that determines which visual stimuli - for example, and without limitation, advertisements, alerts, objects, people, or other visible areas or points of interest - the user is looking at. More specifically, the gaze assistance system may be any electronic computing system or device, such as, for example, a desktop computer, laptop computer, smartphone, server computer, or the like. The camera 310 and the computing device 300 may be incorporated into this computing device to point towards the user, for example, within or closest to the display of the computing device. The camera 310 can capture an image of the user, and the computing device 300 can determine his or her gaze direction. The determined gaze direction can then be transmitted to the gaze assistance system 320, such as, for example, the computing device that displays advertisements for the user, a remote computing device, or the like. The computing device can then use the calculated gaze direction to determine the target of the user's focus while providing information on the effectiveness of various advertisements, alerts, or other visual stimuli.
[0031] Figure 4A is a diagram of an exemplary autonomous vehicle 400 according to some embodiments of the present disclosure. The autonomous vehicle 400 (or referred to herein as "vehicle 400") can include a passenger vehicle, such as a car, truck, bus, first responder vehicle, shuttle, electric or motorized bicycle, motorcycle, fire truck, police vehicle, ambulance, boat, construction vehicle, submarine, drone, and / or another type of vehicle (e.g., unmanned and / or carrying one or more passengers), but is not limited thereto. Autonomous vehicles are generally described in terms of the level of automation defined by the National Highway Traffic Safety Administration (NHTSA), a department of the United States Department of Transportation, and the Society of Automotive Engineers (SAE) "Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicle" (Standard No. J3016-201806 published on June 15, 2018, Standard No. J3016-201609 published on September 30, 2016, and previous and future versions of this standard). The moving vehicle 400 can have the ability to function according to one or more of automation levels 3 to 5 of the autonomous driving level. For example, the moving vehicle 400 can have the ability of conditional automation (level 3), highly automated (level 4), and / or fully automated (level 5) according to the embodiment.
[0032] The moving vehicle 400 can include components such as the chassis, body, wheels (e.g., 2, 4, 6, 8, 18, etc.), tires, axles, and other components of the moving vehicle. The moving vehicle 400 can include a propulsion system 450, such as an internal combustion engine, a hybrid power plant, a fully electric engine, and / or another type of propulsion system. The propulsion system 450 can be connected to the drive train of the moving vehicle 400 and can include a transmission to enable the propulsion force of the moving vehicle 400. The propulsion system 450 can be controlled in response to receiving a signal from the throttle / acceleration device 452.
[0033] A steering system 454, which may include a steering wheel, can be used to steer a moving vehicle 400 (e.g., along a desired path or route) when the propulsion system 450 is operating (e.g., when the moving vehicle is in motion). The steering system 454 can receive a signal from a steering actuator 456. The steering wheel may be an option for a fully automated (level 5) function.
[0034] A brake sensor system 446 can be used to operate the vehicle brakes in response to receiving a signal from a brake actuator 448 and / or a brake sensor.
[0035] The controller 436, which may include one or more CPUs, a system on chip (SoC) 404 (FIG. 4C), and / or a GPU, can provide signals (e.g., representations of commands) to one or more components and / or systems of the moving vehicle 400. For example, the controller can send signals to operate the steering system 454 via one or more steering actuators 456, to operate the moving vehicle brakes via one or more brake actuators 448, and / or to operate the propulsion system 450 via one or more throttle / acceleration devices 452. The controller 436 can include one or more onboard (e.g., integrated) computing devices (e.g., a supercomputer) that process sensor signals and output operational commands (e.g., signals representing commands) to enable autonomous driving and / or to assist the driver in operating the moving vehicle 400. The controller 436 can include a first controller 436 for autonomous driving functions, a second controller 436 for functional safety functions, a third controller 436 for artificial intelligence functions (e.g., computer vision), a fourth controller 436 for infotainment functions, a fifth controller 436 for redundancy in emergency situations, and / or other controllers. In some examples, a single controller 436 can process two or more of the aforementioned functions, and two or more controllers 436 can process a single function and / or any combination thereof.
[0036] Controller 436 can provide signals for controlling one or more components and / or systems of the moving vehicle 400 in response to sensor data (e.g., sensor inputs) received from one or more sensors. The sensor data can be received from, for example and without limitation, a global navigation satellite system sensor 458 (e.g., a global positioning system sensor), a RADAR sensor 460, an ultrasonic sensor 462, a LIDAR sensor 464, an inertial measurement unit (IMU) sensor 466 (e.g., an accelerometer, a gyroscope, a magnetic compass, a magnetometer, etc.), a microphone 496, a stereo camera 468, a wide view camera 470 (e.g., a fish-eye camera), an infrared camera 472, a surround camera 474 (e.g., a 360-degree camera), a long-range and / or mid-range camera 498, a speed sensor 444 (e.g., for measuring the speed of the moving vehicle 400), a vibration sensor 442, a steering sensor 440, a brake sensor 446 (e.g., as part of a brake sensor system 446), and / or other sensor types.
[0037] One or more of the controllers 436 of the vehicle 400 can receive inputs (e.g., represented by input data) from the instrument cluster 432 of the vehicle 400 and provide outputs (e.g., represented by output data, display data, etc.) via a human-machine interface (HMI) display 434, an audible annunciator, a loudspeaker, and / or other components of the vehicle 400. The outputs can include information such as vehicle velocity, speed, time, map data (e.g., the HD map 422 of FIG. 4C), position data (e.g., the position of the vehicle 400 on a map, etc.), direction, the positions of other vehicles (e.g., occupancy grid), information regarding objects and the situation of objects as perceived by the controller 436, etc. For example, the HMI display 434 can display information regarding the presence of one or more objects (e.g., road signs, warning signs, changes in traffic signals, etc.) and / or driving operations that the vehicle has performed, is performing, or will perform (e.g., currently changing lanes, exiting at Exit 34B within 3.22 km (2 miles), etc.).
[0038] The vehicle 400 further includes a network interface 424 that can communicate via one or more networks using one or more wireless antennas 426 and / or a modem. For example, the network interface 424 can have the ability to communicate via LTE, WCDMA (registered trademark), UMTS, GSM, CDMA2000, etc. The wireless antenna 426 can also use local area networks such as Bluetooth (registered trademark), Bluetooth (registered trademark) LE, Z-Wave, ZigBee, and / or low power wide-area networks (LPWAN) such as LoRaWAN, SigFox to enable communication between objects (e.g., vehicles, mobile devices, etc.) in the environment.
[0039] Figure 4B is an example of the camera positions and fields of view of the exemplary autonomous vehicle 400 of FIG. 4A according to some embodiments of the present disclosure. The cameras and their respective fields of view are one exemplary embodiment and are not intended to be limiting. For example, additional and / or alternative cameras may be included and / or the cameras may be placed at different positions on the moving vehicle 400.
[0040] The camera type of the camera may include, but is not limited to, a digital camera adapted to be used with components and / or systems of the moving vehicle 400. The camera can operate at automotive safety integrity level (ASIL) B and / or at another ASIL. The camera type may have the ability to capture images at any rate, such as 60 frames per second (fps), 120 fps, 240 fps, etc., depending on the embodiment. The camera may have the ability to use a rolling shutter, a global shutter, another type of shutter, or a combination thereof. In some examples, the color filter array may include an RCCC (red clear clear clear) color filter array, an RCCB (red clear clear blue) color filter array, an RBGC (red blue green clear) color filter array, a Foveon X3 color filter array, a Bayer sensor (RGGB) color filter array, a monochrome sensor color filter array, and / or another type of color filter array. In some embodiments, clear pixel cameras, such as cameras with RCCC, RCCB, and / or RBGC color filter arrays, may be used in efforts to increase light sensitivity.
[0041] In some examples, one or more of the cameras can be used to perform advanced driver assistance system (ADAS) functions (e.g., as part of a redundant or fail-safe design). For example, a multi-functional mono-camera can be installed to provide functions including lane departure warning, traffic sign assist, and intelligent headlamp control. One or more of the cameras (e.g., all of the cameras) can record and provide image data (e.g., video) simultaneously.
[0042] One or more of the cameras can be mounted on mounting components such as custom-designed (3D printed) parts to remove stray light and reflections from inside the vehicle that can interfere with the camera's image data capture ability (e.g., reflections from the dashboard reflected in the front windshield mirror). Referring to the side mirror mounting component, the side mirror component can be custom 3D printed such that the camera mounting plate conforms to the shape of the side mirror. In some examples, the camera can be integrated within the side mirror. For side view cameras, the camera can also be integrated within four struts located at each corner of the cabin.
[0043] A camera having a field of view that includes a portion of the environment in front of the moving vehicle 400 (e.g., a forward-facing camera) can be used for surround view to assist in identifying the forward path and obstacles and to assist in providing information essential for the generation of an occupancy grid and / or determination of a preferred moving vehicle path with the aid of one or more controllers 436 and / or a control SoC. The forward-facing camera can be used to perform many of the same ADAS functions as LIDAR, including emergency braking, pedestrian detection, and collision avoidance. The forward-facing camera can also be used for ADAS functions and systems including other functions such as lane departure warning (LDW), autonomous cruise control (ACC), and / or traffic sign recognition.
[0044] A variety of cameras can be used in a forward-facing configuration, including, for example, a monocular camera platform that includes a CMOS (complementary metal oxide semiconductor) color imaging device. Another example may be a wide-view camera 470 that can be used to capture objects entering the view from the surroundings (e.g., pedestrians, intersecting traffic, or bicycles). Although only one wide-view camera is shown in FIG. 4B, any number of wide-view cameras 470 may be present on the moving vehicle 400. Additionally, a long-range camera 498 (e.g., a long-view stereo camera pair) can be used for depth-based object detection, particularly for objects for which the neural network has not yet been trained. The long-range camera 498 can also be used for object detection and classification, as well as for basic object tracking.
[0045] One or more stereo cameras 468 can also be included in the forward-facing configuration. The stereo camera 468 can include an integrated control unit with an expandable processing unit that can provide a programmable logic (e.g., FPGA) and a multi-core microprocessor with a CAN or Ethernet® interface integrated on a single chip. Such a unit can be used to generate a 3D map of the environment of the moving vehicle, including distance estimates for all points in the image. An alternative stereo camera 468 can include a compact stereo vision sensor that includes two camera lenses (one each on the left and right) and an image processing chip that can measure the distance from the moving vehicle to a target object and use the generated information (e.g., metadata) to activate autonomous emergency braking and lane departure warning functions. Other types of stereo cameras 468 may be used in addition to, or in place of, those described herein.
[0046] A camera (e.g., a side-view camera) having a field of view that includes a portion of the environment relative to the side of the moving vehicle 400 can be used for surround view to provide information for creating and updating an occupancy grid and for generating a side-impact collision warning. For example, surround cameras 474 (e.g., four surround cameras 474 as shown in FIG. 4B) can be positioned around the moving vehicle 400. The surround cameras 474 can include wide-view cameras 470, fisheye cameras, 360-degree cameras, and / or the like. For example, four fisheye cameras can be disposed in front of, behind, and on the sides of the moving vehicle. In an alternative arrangement, the moving vehicle may use three surround cameras 474 (e.g., left, right, and rear), and one or more other cameras (e.g., a forward-facing camera) may be utilized as a fourth surround-view camera.
[0047] A camera (e.g., a rear-view camera) having a field of view that includes a portion of the environment relative to the rear of the moving vehicle 400 can be used for parking assistance, surround view, rear collision warning, and for creating and updating an occupancy grid. As described herein, a wide variety of cameras can be used, including but not limited to cameras suitable as forward-facing cameras (e.g., long-range and / or mid-range cameras 498, stereo cameras 468), infrared cameras 472, etc.
[0048] A camera having a field of view that includes a portion of the interior or cabin of the vehicle 400 can be used to monitor one or more states of a driver, passenger, or object within the cabin. Any type of camera described herein that provides a view of the cabin or its interior may be used, and any type of camera, including but not limited to cabin camera 441, can be disposed anywhere on or within the vehicle 400. For example, the cabin camera 441 can be disposed within or on some portion of the vehicle 400 dashboard, rear-view mirror, side-view mirror, seat, or door, and can be oriented to capture an image of any driver, passenger, or any other object or portion of the vehicle 400.
[0049] Figure 4C is a block diagram of an exemplary system architecture of the exemplary autonomous vehicle 400 of FIG. 4A, according to some embodiments of the present disclosure. It should be understood that this and other arrangements described herein are presented by way of example only. Other arrangements and elements (e.g., machines, interfaces, functions, orders, groupings of functions, etc.) may be used in addition to or instead of those shown, and some elements may be excluded altogether. Further, many of the elements described herein are functional entities that may be implemented as individual or distributed components or in combination with other components, and in any suitable combination and location. The various functions described herein as being performed by an entity may be implemented by hardware, firmware, and / or software. For example, the various functions may be implemented by a processor that executes instructions stored in a memory.
[0050] Each of the components, features, and systems of the moving vehicle 400 of FIG. 4C is illustrated as being connected via a bus 402. The bus 402 may include a Controller Area Network (CAN) data interface (or referred to as a "CAN bus"). CAN may be a network within the moving vehicle 400 used to assist in the control of various features and functions of the moving vehicle 400, such as the operation of brakes, acceleration, brakes, steering, front windshield wipers, etc. The CAN bus may be configured to have dozens or hundreds of nodes, each having its own unique identifier (e.g., CAN ID). The CAN bus may be read to find steering wheel angle, ground speed, engine revolutions per minute (RPM), button position, and / or other moving vehicle status indicators. The CAN bus may be ASIL B compliant.
[0051] Bus 402 is described herein as being a CAN bus, but this is not intended to be limiting. For example, in addition to, or as an alternative to, the CAN bus, FlexRay and / or Ethernet® may be used. Additionally, a single line is used to represent bus 402, but this is not intended to be limiting. For example, any number of buses 402 may exist that include one or more CAN buses, one or more FlexRay buses, one or more Ethernet® buses, and / or one or more other types of buses that use different protocols. In some examples, two or more buses 402 may be used to perform different functions and / or may be used for redundancy. For example, a first bus 402 may be used for a collision avoidance function and a second bus 402 may be used for actuation control. In any example, each bus 402 may communicate with any of the components of the vehicle 400, and two or more buses 402 may communicate with the same component. In some examples, each SoC 404, each controller 436, and / or each computer within the vehicle may have access to the same input data (e.g., input from the sensors of the vehicle 400) and may be connected to a common bus such as a CAN bus.
[0052] Vehicle 400 may include one or more controllers 436, such as those described herein with respect to FIG. 4A. Controller 436 may be used for various functions. Controller 436 may be coupled to any of the various other components and systems of vehicle 400 and may be used for the control of vehicle 400, the artificial intelligence of vehicle 400, the infotainment for vehicle 400, and / or the like.
[0053] The mobile vehicle 400 may include a system-on-chip (SoC) 404. The SoC 404 may include a CPU 406, a GPU 408, a processor 410, a cache 412, an accelerator 414, a data store 416, and / or other components and features not shown. The SoC 404 may be used to control the mobile vehicle 400 within various platforms and systems. For example, the SoC 404 may be coupled in a system (e.g., the system of the mobile vehicle 400) having an HD map 422 that can obtain map refreshes and / or updates via a network interface 424 from one or more servers (e.g., server 478 of FIG. 4D).
[0054] The CPU 406 may include a CPU cluster or CPU complex (or also referred to as a "CCPLEX"). The CPU 406 may include multiple cores and / or an L2 cache. For example, in some embodiments, the CPU 406 may include 8 cores within a coherent multiprocessor configuration. In some embodiments, the CPU 406 may include 4 dual-core clusters, each cluster having a dedicated L2 cache (e.g., a 2MB L2 cache). The CPU 406 (e.g., CCPLEX) may be configured to support simultaneous cluster operation that allows any combination of clusters of the CPU 406 to become active at any given time.
[0055] The CPU 406 can implement a power management capability that includes one or more of the following features: individual hardware blocks can be automatically clock-gated when in an idle state to conserve dynamic power, each core clock can be gated when the core is not actively executing instructions by the execution of WFI / WFE instructions, each core can be independently power-gated, each core cluster can be independently clock-gated when all cores are clock-gated or power-gated, and / or each core cluster can be independently power-gated when all cores are power-gated. The CPU 406 can further implement an enhanced algorithm for managing power states, where the allowed power states and the expected wake-up times are specified and the hardware / microcode determines the best power state to input to the cores, clusters, and CCPLEX. The processing cores can support a simplified power state input sequence in software where the work is offloaded to microcode.
[0056] The GPU 408 may include an integrated GPU (or referred to herein as "iGPU"). The GPU 408 can be programmable and can be efficient for parallel workloads. In some examples, the GPU 408 can use an enhanced tensor instruction set. The GPU 408 may include one or more streaming microprocessors, where each streaming microprocessor may include an L1 cache (e.g., an L1 cache having at least 96 KB of storage capacity), and two or more of the streaming microprocessors may share a cache (e.g., an L2 cache having 512 KB of storage capacity). In some embodiments, the GPU 408 may include at least 8 streaming microprocessors. The GPU 408 can use a computer-based application programming interface (API). Additionally, the GPU 408 can use one or more parallel computing platforms and / or programming models (e.g., NVIDIA's CUDA).
[0057] The GPU 408 can be power-optimized for the best performance in automotive and embedded use cases. For example, the GPU 408 can be manufactured on FinFET (Fin field-effect transistor). However, this is not intended to be limiting, and the GPU 408 can be manufactured using other semiconductor manufacturing processes. Each streaming microprocessor can incorporate several mixed-precision processing cores partitioned into multiple blocks. By way of example and not limitation, for instance, 64 PF32 cores and 32 PF64 cores may be partitioned into 4 processing blocks. In such an example, each processing block may be allocated 16 FP32 cores, 8 FP64 cores, 16 INT32 cores, 2 mixed-precision NVIDIA tensor cores for deep learning matrix operations, an L0 instruction cache, a warp scheduler, a dispatch unit, and / or a 64KB register file. Additionally, the streaming microprocessor may include independent parallel integer and floating-point data paths to provide efficient execution of workloads having a mix of compute and addressing operations. The streaming microprocessor may include independent thread scheduling capabilities to enable higher fine-grained synchronization and cooperation between parallel threads. The streaming microprocessor may include a combined L1 data cache and shared memory unit to simplify programming while improving performance.
[0058] In some examples, GPU 408 may include high bandwidth memory (HBM) and / or a 16 GB HBM2 memory subsystem to provide for a peak memory bandwidth of 900 GB / second. In some examples, in addition to, or instead of, HBM memory, synchronous graphics random-access memory (SGRAM), such as graphics double data rate type five synchronous random-access memory (GDDR5), may be used.
[0059] GPU 408 can include unified memory technology that includes access counters to enable more accurate movement of those memory pages to the processor that most frequently accesses the memory pages, thereby improving the efficiency of the storage ranges shared among the processors. In some examples, address translation service (ATS) support may be used to enable the GPU 408 to directly access the CPU 406 page table. In such examples, when the GPU 408 memory management unit (MMU) experiences a miss, an address translation request may be sent to the CPU 406. In response, the CPU 406 can examine its page table for the virtual-to-physical mapping of the address and send the translation back to the GPU 408. As such, unified memory technology can enable a single unified virtual address space for the memory of both the CPU 406 and the GPU 408, thereby simplifying GPU 408 programming and porting of applications to the GPU 408.
[0060] In addition, GPU 408 may include an access counter that can record the frequency of access of GPU 408 to the memory of other processors. The access counter can help ensure that memory pages are moved to the physical memory of the processor that most frequently accesses that page.
[0061] SoC 404 may include any number of caches 412, including those described herein. For example, cache 412 may include an L3 cache that is available to both CPU 406 and GPU 408 (e.g., connected to both CPU 406 and GPU 408). Cache 412 may include a write-back cache that can record the state of lines, such as by using a cache coherence protocol (e.g., MEI, MESI, MSI, etc.). The L3 cache may include more than 4 MB, although smaller cache sizes may be used, depending on the embodiment.
[0062] SoC 404 may include an arithmetic logic unit (ALU) that can be utilized when executing processing for any of the various tasks or operations of vehicle 400 (e.g., processing DNN). In addition, SoC 404 may include a floating point unit (FPU) (or other math coprocessor or numeric coprocessor type) for performing mathematical operations within the system. For example, SoC 104 may include one or more FPUs integrated as execution units within CPU 406 and / or GPU 408.
[0063] The SoC 404 may include one or more accelerators 414 (e.g., a hardware accelerator, a software accelerator, or a combination thereof). For example, the SoC 404 may include a hardware acceleration cluster that may include an optimized hardware accelerator and / or a large on-chip memory. The large on-chip memory (e.g., 4MB of SRAM) may enable the hardware acceleration cluster to accelerate neural networks and other operations. The hardware acceleration cluster may be used to complement the GPU 408 and to offload some of the tasks of the GPU 408 (e.g., to free up more cycles of the GPU 408 for performing other tasks). As an example, the accelerator 414 may be used for target workloads that are stable enough to be suitable for acceleration (e.g., perception, convolutional neural networks (CNNs)). As used herein, the term "CNN" may include all types of CNNs, including region-based or regional convolutional neural networks (RCNNs) and fast RCNNs (e.g., as used for object detection).
[0064] The acceleration device 414 (e.g., a hardware acceleration cluster) may include a deep learning accelerator (DLA). The DLA may include one or more tensor processing units (TPUs) that can be configured to provide an additional 10 trillion operations per second for deep learning applications and inferences. The TPU may be an accelerator configured and optimized to execute image processing functions (e.g., those of CNN, RCNN, etc.). The DLA can further be optimized for a specific set of neural network types and floating-point operations, as well as for inferences. The design of the DLA can provide more performance per millimeter than a general-purpose GPU and greatly exceed the performance of a CPU. The TPU can execute several functions, including, for example, a single-instance convolution function that supports INT8, INT16, and FP16 data types for both features and weights, and a post-processor function.
[0065] The DLA can rapidly and efficiently execute neural networks, particularly CNNs, with processed or unprocessed data for any of a variety of functions, including but not limited to: CNNs for object identification and detection using data from a camera sensor, CNNs for distance estimation using data from a camera sensor, CNNs for emergency vehicle detection and identification and detection using data from a microphone, CNNs for face recognition and moving vehicle owner identification using data from a camera sensor, and / or CNNs for security and / or safety-related events.
[0066] The DLA can execute any function of the GPU 408, and by using the inference accelerator, for example, a designer can target either the DLA or the GPU 408 for any function. For example, the designer can focus on processing CNNs and floating-point operations on the DLA and leave other functions to the GPU 408 and / or other acceleration devices 414.
[0067] Acceleration device 414 (e.g., a hardware acceleration cluster) may include, or may be referred to herein as, a programmable vision accelerator (PVA) and a computer vision acceleration device. The PVA may be designed and configured to accelerate computer vision algorithms for advanced driver assistance systems (ADAS), autonomous driving, and / or augmented reality (AR) and / or virtual reality (VR) applications. The PVA can provide a balance between performance and flexibility. For example, each PVA may include, but is not limited to, any number of reduced instruction set computer (RISC) cores, direct memory access (DMA), and / or any number of vector processors.
[0068] The RISC core can interact with an image sensor (e.g., the image sensor of any of the cameras described herein), an image signal processor, and / or the like. Each RISC core may include any amount of memory. The RISC core can use any of several protocols, depending on the embodiment. In some examples, the RISC core can execute a real-time operating system (RTOS). The RISC core can be implemented using one or more integrated circuit devices, application-specific integrated circuits (ASICs), and / or memory devices. For example, the RISC core may include an instruction cache and / or tightly coupled RAM.
[0069] The DMA can enable the components of the PVA to access the system memory independent of the CPU406. The DMA can support any number of features used to provide optimization for the PVA, including but not limited to supporting multidimensional addressing and / or circular addressing. In some examples, the DMA can support addressing up to six dimensions or more, which may include block width, block height, block depth, horizontal block stepping, vertical block stepping, and / or depth stepping.
[0070] The vector processor may be a programmable processor designed to efficiently and flexibly execute the programming of computer vision algorithms and provide signal processing capabilities. In some examples, the PVA may include a PVA core and two vector processing subsystem partitions. The PVA core may include a processor subsystem, a DMA engine (e.g., two DMA engines), and / or other peripheral devices. The vector processing subsystem can operate as the primary processing engine of the PVA and may include a vector processing unit (VPU), an instruction cache, and / or vector memory (e.g., VMEM). The VPU core may include a digital signal processor, such as a single instruction, multiple data (SIMD), very long instruction word (VLIW) digital signal processor. The combination of SIMD and VLIW can increase throughput and speed.
[0071] Each vector processor may include an instruction cache and may be coupled to dedicated memory. As a result, in some examples, each vector processor may be configured to execute independently from other vector processors. In other examples, the vector processors included in a particular PVA may be configured to use data parallel processing. For example, in some embodiments, multiple vector processors included in a single PVA may be able to execute the same computer vision algorithm, but on different regions of an image. In other examples, the vector processors included in a particular PVA may be able to execute different computer vision algorithms simultaneously on the same image, or even execute different algorithms sequentially on an image or portions of an image. In particular, any number of PVAs may be included in a hardware acceleration cluster, and any number of vector processors may be included in each PVA. Additionally, the PVA may include additional error correcting code (ECC) memory to enhance overall system safety.
[0072] Accelerator 414 (e.g., a hardware acceleration cluster) may include a computer vision network on chip and SRAM to provide high bandwidth, low latency SRAM for the accelerator 414. In some examples, the on-chip memory may include at least 4MB of SRAM consisting of, for example and without limitation, 8 field-configurable memory blocks that may be accessible by both the PVA and the DLA. Each pair of memory blocks may include an advanced peripheral bus (APB) interface, configuration circuitry, a controller, and a multiplexer. Any type of memory may be used. The PVA and the DLA may access the memory via a backbone that provides high-speed access to the PVA and the DLA to the memory. The backbone may include a computer vision network on chip that interconnects the PVA and the DLA to the memory (e.g., using an APB).
[0073] The computer vision network-on-chip may include an interface that determines that both the PVA and the DLA are operable and provide valid signals before any control signal / address / data transmission. Such an interface can provide separate phases and separate channels for transmitting control signals / addresses / data, as well as burst-type communication for continuous data transfer. This type of interface can comply with the ISO26262 or IEC61508 standard, although other standards and protocols may be used.
[0074] In some examples, the SoC404 may include a real-time ray tracing hardware accelerator as described in U.S. Patent Application No. 16 / 101,232, filed on August 10, 2018. The real-time ray tracing hardware accelerator can be used to quickly and efficiently determine the position and scale of objects (e.g., within a world model) for generating real-time visualization simulations for RADAR signal interpretation, for acoustic propagation synthesis and / or analysis, for simulation of SONAR systems, for general wave propagation simulations, for comparison against LIDAR data for localization and / or other functions, and / or for other uses. In some embodiments, one or more tree traversal units (TTUs) may be used to perform one or more ray tracing related operations.
[0075] The acceleration device 414 (e.g., a hardware acceleration device cluster) has various uses for autonomous driving. The PVA may be a programmable vision acceleration device that can be used in extremely important processing stages in ADAS and autonomous vehicles. The capabilities of the PVA are suitable for areas of algorithms that require predictable processing at low power and low latency. In other words, the PVA functions well with low latency, low power, and predictable execution time, even on small data sets and with semi-dense or dense normal calculations. Therefore, since the PVA is efficient in object detection and integer calculations, in the context of a platform for autonomous vehicles, the PVA is designed to execute classic computer vision algorithms.
[0076] For example, according to one embodiment of the present technology, the PVA is used to execute computer stereo vision. A semi-global matching-based algorithm may be used in some examples, but this is not intended to be limiting. Many applications for level 3-5 autonomous driving require motion estimation / stereo matching on the fly (e.g., SFM (structure from motion), pedestrian recognition, lane detection, etc.). The PVA can execute computer stereo vision functions with inputs from two monocular cameras.
[0077] In some examples, the PVA can be used to execute high-density optical flow. For example, the PVA can be used to process raw RADAR data (e.g., using 4D fast Fourier transform) to provide processed RADAR signals before emitting the next RADAR pulse. In other examples, the PVA is used for flight depth processing time, for example, by processing the raw time of flight data to provide the processed time of flight data.
[0078] DLA can be used to execute any type of network to enhance control and driving safety, including, for example, a neural network that outputs a measure of the reliability of each object detection. Such reliability values can be interpreted as probabilities or as providing the relative "weight" of each detection compared to other detections. This reliability value enables the system to make further decisions regarding which detections should be considered true positive detections rather than false positive detections. For example, the system can set a reliability threshold and consider only detections that exceed the threshold as true positive detections. In an automatic emergency braking (AEB) system, a false positive detection would cause the moving vehicle to automatically execute emergency braking, which is clearly undesirable. Therefore, only the most confident detections should be considered as triggers for AEB. DLA can execute a neural network that regresses reliability values. The neural network can receive as its inputs at least some subsets of parameters such as bounding box dimensions, ground plane estimation obtained (e.g., from another subsystem), the orientation, distance, 3D position estimation of an object obtained from the neural network and / or other sensors (e.g., LIDAR sensor 464 or RADAR sensor 460), and the output of an inertial measurement unit (IMU) sensor 466 that correlates with the above, among others.
[0079] SoC 404 can include a data store 416 (e.g., memory). The data store 416 can be the on-chip memory of the SoC 404 and can store neural networks to be executed by the GPU and / or DLA. In some examples, the data store 416 can have a capacity large enough to store multiple instances of neural networks for redundancy and safety. The data store 416 can include an L2 or L3 cache 412. References to the data store 416 can include references to memory related to PVA, DLA, and / or other accelerators 414 as described herein.
[0080] SoC 404 may include one or more processors 410 (e.g., an embedded processor). The processor 410 may include a boot and power management processor, which may be a dedicated processor and subsystem for handling boot power and management capabilities and related security enforcement. The boot and power management processor may be part of the SoC 404 boot sequence and can provide runtime power management services. The boot power and management processor can provide clock and voltage programming, assistance with system low-power state transitions, management of the SoC 404 heat and temperature sensors, and / or management of the SoC 404 power state. Each temperature sensor may be implemented as a ring oscillator whose output frequency is proportional to temperature, and the SoC 404 can use the ring oscillator to detect the temperature of the CPU 406, GPU 408, and / or accelerator 414. If the temperature is determined to exceed a threshold, the boot and power management processor can enter a temperature fault routine and place the SoC 404 in a lower power state and / or put the vehicle 400 in a chauffeur safety stop mode (e.g., safely stop the vehicle 400).
[0081] The processor 410 may further include a set of embedded processors that can perform the functions of an audio processing engine. The audio processing engine may be an audio subsystem that enables full hardware support for multi-channel audio via multiple interfaces and a wide and flexible range of audio I / O interfaces. In some examples, the audio processing engine is a dedicated processor core with a digital signal processor having dedicated RAM.
[0082] Processor 410 may further include an always-on processor engine that can provide the necessary hardware features to support low-power sensor management and wake usage scenarios. The always-on processor engine may include a processor core, tightly coupled RAM, support peripherals (e.g., timers and interrupt controllers), various I / O controller peripherals, and routing logic.
[0083] Processor 410 may further include a safety cluster engine that includes a dedicated processor subsystem for processing the safety management of automotive applications. The safety cluster engine may include two or more processor cores, tightly coupled RAM, support peripherals (e.g., timers, interrupt controllers, etc.), and / or routing logic. In the safety mode, two or more cores can operate in a lockstep mode and function as a single core with comparison logic for detecting any differences during their operations.
[0084] Processor 410 may further include a real-time camera engine that may include a dedicated processor subsystem for processing real-time camera management.
[0085] Processor 410 may further include a high dynamic range signal processor that includes an image signal processor, which is a hardware engine that is part of the camera processing pipeline.
[0086] The processor 410 may include a video image synthesizer, which is a processing block (e.g., implemented in a microprocessor) that implements the post-video processing functions required by the video playback application to produce the final image for the player window. The video image synthesizer can perform lens distortion correction with the wide-view camera 470, with the surround camera 474, and / or with the in-cabin monitoring camera sensor. The in-cabin monitoring camera sensor is preferably monitored by a neural network running on another instance of a high-level SoC that is configured to identify in-cabin events and respond appropriately. The in-cabin system can perform lip reading to activate cellular service and make calls, compose emails, change the destination of the moving vehicle, activate or change the infotainment system and settings of the moving vehicle, or provide voice-activated web surfing. Certain functions are available to the driver only when operating in autonomous mode and are otherwise disabled.
[0087] The video image synthesizer may include enhanced temporal noise reduction for both spatial and temporal noise reduction. For example, when motion occurs in the video, the noise reduction reduces the weight of the information provided by adjacent frames and appropriately weights the spatial information. When an image or a part of an image does not contain motion, the temporal noise reduction performed by the video image synthesizer can reduce the noise in the current image using information from the previous image.
[0088] The video image synthesizer may also be configured to perform stereo rectification on the input stereo lens frame. The video image synthesizer can further be used for user interface synthesis when the operating system desktop is in use, and the GPU 408 is not required to continuously render new surfaces. Even when the power of the GPU 408 is turned on and 3D rendering is actively performed, the video image synthesizer can be used to offload the GPU 408 to improve performance and responsiveness.
[0089] SoC 404 may further include a mobile industry processor interface (MIPI) camera serial interface, a high-speed interface, and / or a video input block that can be used for the camera and related pixel input functions to receive video and inputs from the camera. SoC 404 may further include an input / output controller that can be controlled by software and can be used to receive I / O signals that are not committed to a specific role. SoC 404 may further include a wide range of peripheral interfaces to enable communication with peripheral devices, audio codecs, power management, and / or other devices. SoC 404 can be used to process data from cameras (e.g., connected via gigabit multimedia serial link and Ethernet (R)), sensors (e.g., LIDAR sensor 464, RADAR sensor 460 that can be connected via Ethernet (R)), data from the bus 402 (e.g., the speed, steering wheel position, etc. of the moving vehicle 400), and data from GNSS sensors 458 (e.g., connected via Ethernet (R) or CAN bus). SoC 404 may include its own DMA engine and may further include a dedicated high-performance large-capacity storage controller that can be used to free the CPU 406 from routine data management tasks.
[0090] The SoC404 may be an end-to-end platform with a flexible architecture that extends to automation levels 3-5, thereby leveraging and efficiently using computer vision and ADAS techniques for diversity and redundancy, and providing a platform for a flexible and reliable driving software stack together with deep learning tools, providing an integrated functional safety architecture. The SoC404 can be faster, more reliable, more energy-efficient, and more space-efficient than conventional systems. For example, when the accelerator 414 is coupled to the CPU 406, the GPU 408, and the data store 416 can provide a fast and efficient platform for level 3-5 autonomous vehicles.
[0091] Accordingly, the present technology provides capabilities and functionality that cannot be achieved by conventional systems. For example, computer vision algorithms can be implemented using a high-level programming language such as the C programming language to execute a variety of processing algorithms over a variety of visual data and can be executed on a CPU. However, the CPU often cannot meet the performance requirements of many computer vision applications, such as those related to execution time and power consumption. Specifically, many CPUs cannot execute real-time composite object detection algorithms, which are requirements for in-vehicle ADAS applications and actual level 3-5 autonomous vehicles.
[0092] In contrast to conventional systems, by providing a CPU complex, a GPU complex, and a hardware acceleration cluster, the techniques described herein enable multiple neural networks to be executed simultaneously and / or sequentially, and results to be combined to enable level 3-5 autonomous driving functions. For example, a CNN executed on a DLA or a dGPU (e.g., GPU 420) can include text and word recognition that enables a supercomputer to read and understand traffic signs, including signs that the neural network has not been specifically trained on. The DLA can further include a neural network that can identify, interpret, and provide a semantic understanding of the signs and pass the semantic understanding to a route planning module executed on the CPU complex.
[0093] As another example, multiple neural networks can be executed simultaneously as required for level 3, 4, or 5 driving. For example, along with the electro-optical, a warning sign consisting of "Caution: Flashing light indicates frozen condition" can be interpreted independently or collectively by several neural networks. The sign itself can be identified as a traffic sign by a first deployed neural network (e.g., a neural network that has been trained), and the text "Flashing light indicates frozen condition" can be interpreted by a second deployed neural network that informs the route planning software of the moving vehicle (preferably executed on the CPU complex) that a frozen condition exists when the flashing light is detected. The flashing light can be identified by informing the route planning software of the moving vehicle of the presence (or absence) of the flashing light and operating a third deployed neural network over multiple frames. All three neural networks can be executed simultaneously, such as within the DLA and / or on the GPU 408.
[0094] In some examples, a CNN for face recognition and mobile vehicle owner identification can use data from a camera sensor to identify the presence of the regular driver and / or owner of the mobile vehicle 400. The always-on sensor processing engine can be used to unlock and light up the mobile vehicle when the owner approaches the driver's side door, and also, in security mode, to stop the operation of the mobile vehicle when the owner leaves the mobile vehicle. In this way, the SoC 404 provides security against theft and / or carjacking.
[0095] In another example, a CNN for emergency vehicle detection and identification can use data from the microphone 496 to detect and identify emergency vehicle sirens. In contrast to conventional systems that use a general classifier to detect sirens and manually extract features, the SoC 404 uses a CNN for classification of environmental and urban sounds, as well as for classification of visual data. In a preferred embodiment, the CNN executed on the DLA is trained to identify the relative end speed of an emergency vehicle (e.g., by using the Doppler effect). The CNN can also be trained to identify emergency vehicles specific to the local area where the mobile vehicle is operating, as identified by the GNSS sensor 458. Therefore, for example, when operating in Europe, the CNN will attempt to detect European sirens, and when in the United States, the CNN will attempt to identify only North American sirens. After an emergency vehicle is detected, a control program can be used to execute an emergency vehicle safety routine to slow down the mobile vehicle, stop it at the side of the road, park the mobile vehicle, and / or idle the mobile vehicle, with the assistance of the ultrasonic sensor 462 until the emergency vehicle has passed.
[0096] The mobile vehicle may include a CPU 418 (e.g., a discrete CPU or a dCPU) that can be connected to the SoC 404 via a high-speed interconnect (e.g., PCIe). The CPU 418 may include, for example, an X86 processor. The CPU 418 may be used to perform any of a variety of functions, including, for example, mediating potential inconsistencies between the ADAS sensors and the SoC 404 and / or monitoring the status and health of the controller 436 and / or the infotainment SoC 430.
[0097] The mobile vehicle 400 may include a GPU 420 (e.g., a discrete GPU or a dGPU) that can be connected to the SoC 404 via a high-speed interconnect (e.g., NVIDIA's NVLINK). The GPU 420 can provide additional artificial intelligence capabilities, such as by executing redundant and / or different neural networks, and can be used to train and / or update neural networks based on inputs (e.g., sensor data) from the sensors of the mobile vehicle 400.
[0098] The mobile vehicle 400 may further include a network interface 424 that may include one or more wireless antennas 426 (e.g., one or more wireless antennas for different communication protocols such as cellular antennas, Bluetooth® antennas, etc.). The network interface 424 may be used to enable wireless connections with the cloud via the Internet (e.g., with the server 478 and / or other network devices), with other mobile vehicles, and / or with computing devices (e.g., a passenger's client device). To communicate with other mobile vehicles, a direct link may be established between two mobile vehicles and / or an indirect link may be established (e.g., through a network and via the Internet). The direct link may use and provide a vehicle-to-vehicle communication link. The vehicle-to-vehicle communication link may provide vehicle 400 information regarding mobile vehicles in proximity to the mobile vehicle 400 (e.g., mobile vehicles in front of, beside, and / or behind the mobile vehicle 400). This functionality may be part of the cooperative adaptive cruise control function of the mobile vehicle 400.
[0099] The network interface 424 may include a system-on-chip (SoC) that provides modulation and demodulation functions and enables the controller 436 to communicate via a wireless network. The network interface 424 may include a radio frequency front end for upconversion from baseband to radio frequency and downconversion from radio frequency to baseband. The frequency conversion can be performed through well-known processes and / or using a superheterodyne process. In some examples, the radio frequency front end functionality may be provided by a separate chip. The network interface may include wireless functionality for communicating via LTE, WCDMA®, UMTS, GSM, CDMA2000, Bluetooth®, Bluetooth® LE, Wi-Fi, Z-Wave, ZigBee, LoRaWAN, and / or other wireless protocols. The mobile vehicle 400 may further include a data store 428 that may include storage outside of the chip (e.g., outside of the SoC 404). The data store 428 may include one or more storage elements including RAM, SRAM, DRAM, VRAM, flash, hard disk, and / or other components and / or devices capable of storing at least 1 bit of data.
[0100] The vehicle 400 may further include a GNSS sensor 458 (e.g., GPS and / or assisted GPS sensor) to assist with mapping, perception, occupancy grid generation, and / or path planning functions. For example, any number of GNSS sensors 458 may be used, including but not limited to GPS with an Ethernet® to serial (RS-232) bridge USB connector. The mobile vehicle 400 may further include a RADAR sensor 460. The RADAR sensor 460 may be used by the mobile vehicle 400 for long-range mobile vehicle detection even in darkness and / or harsh weather conditions. The RADAR functional safety level may be ASIL B. In some examples, the RADAR sensor 460 may use CAN and / or bus 402 for control and to access object tracking data (e.g., to send data generated by the RADAR sensor 460) using access to Ethernet® for accessing raw data. A wide variety of RADAR sensor types may be used. For example, and without limitation, the RADAR sensor 460 may be suitable for front, rear, and side RADAR use. In some examples, a pulse Doppler RADAR sensor is used.
[0101] The RADAR sensor 460 can include different configurations, such as long range with a narrow field of view, short range with a wide field of view, and short-range side coverage. In some examples, the long-range RADAR can be used for an adaptive cruise control function. The long-range RADAR system can provide a wide field of view realized by two or more independent scans, such as within a range of 250 m. The RADAR sensor 460 can help distinguish between static and moving objects and can be used by an ADAS system for emergency brake assist and forward collision warning. The long-range RADAR sensor can include a monostatic multi-modal RADAR with a plurality of (e.g., six or more) fixed RADAR antennas and high-speed CAN and FlexRay interfaces. In one example with six antennas, the four central antennas can create a focused beam pattern designed to record the surroundings of the moving vehicle 400 at high speed while minimizing interference from traffic in adjacent lanes. The other two antennas can widen the field of view and enable rapid detection of moving vehicles entering or leaving the lane of the moving vehicle 400.
[0102] As an example, the mid-range RADAR system can include a range up to 460 m (front) or 80 m (rear), and a field of view up to 42 degrees (front) or 450 degrees (rear). The short-range RADAR system can include, but is not limited to, RADAR sensors designed to be installed at both ends of the rear bumper. When installed at both ends of the rear bumper, such a RADAR sensor system can create two beams that constantly monitor the blind spots behind and adjacent to the moving vehicle.
[0103] The short-range RADAR system can be used in an ADAS system for blind spot detection and / or lane change assist.
[0104] The moving vehicle 400 may further include an ultrasonic sensor 462. The ultrasonic sensor 462, which may be positioned at the front, rear, and / or sides of the moving vehicle 400, may be used for parking assist and / or for creating and updating an occupancy grid. A wide variety of ultrasonic sensors 462 may be used, and different ultrasonic sensors 462 may be used for different ranges of detection (e.g., 2.5 m, 4 m). The ultrasonic sensor 462 can operate at a functional safety level of ASIL B.
[0105] The moving vehicle 400 may include a LIDAR sensor 464. The LIDAR sensor 464 may be used for object and pedestrian detection, emergency braking, collision avoidance, and / or other functions. The LIDAR sensor 464 may also be at a functional safety level of ASIL B. In some examples, the moving vehicle 400 may include a plurality (e.g., 2, 4, 6, etc.) of LIDAR sensors 464 that can use Ethernet (registered trademark) (e.g., to provide data to a gigabit Ethernet (registered trademark) switch).
[0106] In some examples, the LIDAR sensor 464 may have the ability to provide a list of objects and their distances in a 360-degree field of view. Commercially available LIDAR sensors 464 may have an accuracy of, for example, 2 cm to 3 cm, support a 100 Mbps Ethernet (registered trademark) connection, and may have an advertised range of about 100 m. In some examples, one or more non-protruding LIDAR sensors 464 may be used. In such examples, the LIDAR sensor 464 may be implemented as a small device that can be incorporated into the front, rear, sides, and / or corners of the moving vehicle 400. In such examples, the LIDAR sensor 464 may have a range of 200 m even for low-reflectivity objects and can provide up to a 120-degree horizontal and 35-degree vertical field of view. The LIDAR sensor 464 attached to the front may be configured for a horizontal field of view between 45 degrees and 135 degrees.
[0107] In some examples, LIDAR technologies such as 3D flash LIDAR may also be used. 3D flash LIDAR uses a flash of laser as a transmitter to illuminate the area around a moving vehicle up to about 200 m. The flash LIDAR unit includes a receptor that records the laser pulse travel time and the reflected light on each pixel, corresponding in sequence to the range from the moving vehicle to the object. Flash LIDAR can enable high-precision and distortion-free images of the surroundings to be generated with every laser flash. In some examples, four flash LIDAR sensors may be deployed, one on each side of the moving vehicle 400. Available 3D flash LIDAR systems include a solid-state 3D steering array LIDAR camera (e.g., a non-scanning LIDAR device) that has no moving parts other than a blower. The flash LIDAR device can use class I (eye-safe) laser pulses of 5 nanoseconds per frame and can capture the reflected laser light in the form of 3D range point clouds and co-recorded intensity data. By using flash LIDAR, also, since the flash LIDAR is a solid-state device with no moving parts, the LIDAR sensor 464 can be less susceptible to the effects of motion blur, vibration, and / or shock.
[0108] The moving vehicle may further include an IMU sensor 466. In some examples, the IMU sensor 466 may be positioned at the center of the rear axle of the moving vehicle 400. The IMU sensor 466 can include, for example, but is not limited to, an accelerometer, a magnetometer, a gyroscope, a magnetic compass, and / or other sensor types. In some examples, in a 6-axis application, the IMU sensor 466 can include an accelerometer and a gyroscope, while in a 9-axis application, the IMU sensor 466 can include an accelerometer, a gyroscope, and a magnetometer.
[0109] In some embodiments, the IMU sensor 466 may be implemented as a miniature, high-performance GPS-aided inertial navigation system (GPS / INS) that combines a micro-electro-mechanical system (MEMS) inertial sensor, a high-sensitivity GPS receiver, and an advanced Kalman filtering algorithm to provide estimates of position, velocity, and attitude. As such, in some examples, the IMU sensor 466 may enable the moving vehicle 400 to estimate its direction of travel without the need for input from a magnetic sensor by directly observing and correlating changes in velocity from the GPS to the IMU sensor 466. In some examples, the IMU sensor 466 and the GNSS sensor 458 may be combined in a single integrated unit. The moving vehicle may include a microphone 496 disposed within and / or around the moving vehicle 400. The microphone 496 may be used, among other things, for emergency vehicle detection and identification.
[0110] The moving vehicle may further include any number of camera types, including a stereo camera 468, a wide-view camera 470, an infrared camera 472, a surround camera 474, a long-range and / or mid-range camera 498, and / or other camera types. The cameras may be used to capture image data around the entire outer surface of the moving vehicle 400. The type of cameras used depends on the embodiment and requirements of the moving vehicle 400, and any combination of camera types may be used to achieve the necessary coverage around the moving vehicle 400. Additionally, the number of cameras may vary depending on the embodiment. For example, the moving vehicle may include six cameras, seven cameras, ten cameras, twelve cameras, and / or another number of cameras. The cameras may, as one example, support but are not limited to Gigabit Multimedia Serial Link (GMSL) and / or Gigabit Ethernet (registered trademark). Each camera is described in more detail herein with reference to FIGS. 4A and 4B.
[0111] The moving vehicle 400 may further include a vibration sensor 442. The vibration sensor 442 can measure the vibration of components of the moving vehicle, such as an axle. For example, a change in vibration may indicate a change in the road surface. In another example, when two or more vibration sensors 442 are used, the difference in vibration can be used to determine the friction or slip of the road surface (for example, when the difference in vibration is between a power-driven axle and a freely rotating axle).
[0112] The moving vehicle 400 may include an ADAS system 438. In some examples, the ADAS system 438 may include a SoC. The ADAS system 438 may include autonomous / adaptive / automatic cruise control (ACC), cooperative adaptive cruise control (CACC), forward crash warning (FCW), automatic emergency braking (AEB), lane departure warning (LDW), lane keep assist (LKA), blind spot warning (BSW), rear cross-traffic warning (RCTW), collision warning system (CWS), lane centering (LC), and / or other features and functions.
[0113] The ACC system may use a RADAR sensor 460, a LIDAR sensor 464, and / or a camera. The ACC system may include longitudinal ACC and / or lateral ACC. Longitudinal ACC monitors and controls the distance to the vehicle immediately in front of the moving vehicle 400 and automatically adjusts the vehicle speed to maintain a safe distance from the vehicle in front. Lateral ACC performs distance keeping and advises the moving vehicle 400 to change lanes when necessary. Lateral ACC is related to other ADAS applications such as LCA and CWS.
[0114] CACC uses information from other vehicles that can be received via a wireless link through the network interface 424 and / or the wireless antenna 426 from other vehicles, or indirectly through a network connection (e.g., via the Internet). The direct link may be provided by a vehicle-to-vehicle (V2V) communication link, while the indirect link may be an infrastructure-to-vehicle (I2V) communication link. Generally, the V2V communication concept provides information about the vehicle immediately in front (e.g., the vehicle immediately in front of the moving vehicle 400 in the same lane as the moving vehicle 400), while the I2V communication concept provides information about traffic further ahead. The CACC system may include either or both of an I2V information source and a V2V information source. Given the information of the vehicle in front of the moving vehicle 400, CACC can be more reliable, and CACC has the potential to make the traffic flow smoother and reduce road congestion.
[0115] The FCW system is designed to warn the driver of danger so that the driver can take corrective action. The FCW system uses a forward-facing camera and / or RADAR sensor 460 connected to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically connected to driver feedback such as a display, speaker, and / or vibrating component. The FCW system can provide warnings in the form of acoustic, visual alerts, vibrations, and / or quick brake pulses.
[0116] The AEB system can detect an impending forward collision with another moving vehicle or other object and automatically apply the brakes if the driver does not take corrective action within a specified time or distance parameter. The AEB system can use a forward-facing camera and / or RADAR sensor 460 connected to a dedicated processor, DSP, FPGA, and / or ASIC. When the AEB system detects danger, it usually first warns the driver to take corrective action to avoid the collision. If the driver does not take corrective action, the AEB system can automatically apply the brakes as part of an effort to prevent or at least mitigate the impact of the predicted collision. The AEB system may include techniques such as dynamic brake support and / or collision imminent braking.
[0117] The LDW system provides visual, audible, and / or tactile warnings, such as vibrations of the steering wheel or seat, to warn the driver when the moving vehicle 400 crosses a lane dividing line. The LDW system does not activate when the driver indicates an intentional lane departure by activating the turn indicator. The LDW system can use a front-facing camera connected to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically connected to driver feedback such as a display, speaker, and / or vibrating component.
[0118] The LKA system is a modified form of the LDW system. The LKA system provides a steering input or brake to correct the moving vehicle 400 when the moving vehicle 400 begins to deviate from the lane. The BSW system detects and warns the driver of a moving vehicle in the blind spot of a motor vehicle. The BSW system can provide visual, audible, and / or tactile warnings to indicate that a merge or lane change is not safe. The system can provide additional warnings when the driver uses the turn indicator. The BSW system uses a rear-facing camera and / or RADAR sensor 460 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback, such as a display, speaker, and / or vibrating component.
[0119] The RCTW system can provide visual, audible, and / or tactile notifications when an object is detected outside the range of the rear camera while the vehicle 400 is backing up. Some RCTW systems include AEB to ensure that the vehicle brakes are applied to avoid a collision. The RCTW system can use one or more rear-facing RADAR sensors 460 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback, such as a display, speaker, and / or vibrating component.
[0120] Conventional ADAS systems allow the driver to be warned and to determine whether a safety condition actually exists and act accordingly. Thus, conventional ADAS systems, while not usually catastrophic, tend to produce false positive results that can annoy and distract the driver. However, in the case of an autonomous vehicle 400, if the results conflict, the moving vehicle 400 itself must decide whether to accept the results from the primary or secondary computer (e.g., the first controller 436 or the second controller 436). For example, in some embodiments, the ADAS system 438 may be a backup and / or secondary computer for providing perception information to a backup computer rationality module. The backup computer rationality monitor can execute diverse software redundant in hardware components to detect failures in perception and dynamic driving tasks. The output from the ADAS system 438 can be provided to the supervisory MCU. If the outputs from the primary and secondary computers conflict, the supervisory MCU needs to determine how to resolve that conflict to ensure safe operation.
[0121] In some examples, the primary computer can be configured to provide a reliability score to the supervisory MCU indicating the reliability of the primary computer in the selected result. If the reliability score exceeds a threshold, the supervisory MCU can follow the instructions of the primary computer regardless of whether the secondary computer gives conflicting or inconsistent results. If the reliability score does not meet the threshold and the primary and secondary computers indicate different (e.g., conflicting) results, the supervisory MCU can mediate between the computers to determine an appropriate result.
[0122] The supervisory MCU may be configured to execute a neural network trained and configured to determine, based on outputs from the primary computer and the secondary computer, a state in which the secondary computer provides a false alarm. Thus, the neural network within the supervisory MCU can learn when the output of the secondary computer is reliable and when it is not. For example, when the secondary computer is a RADAR-based FCW system, the neural network within the supervisory MCU can learn when the FCW identifies metallic objects such as manhole covers or gratings in a drain that are not actually dangerous but trigger an alarm. Similarly, when the secondary computer is a camera-based LDW system, the neural network within the supervisory MCU can learn to ignore the LDW when a person on a bicycle or a pedestrian is present and lane departure is actually the safest operation. In an embodiment that includes a neural network running on the supervisory MCU, the supervisory MCU may include at least one of a DLA or a GPU suitable for executing the neural network with associated memory. In a preferred embodiment, the supervisory MCU may comprise components of the SoC404 and / or be included as components of the SoC404.
[0123] In other examples, the ADAS system 438 may include a secondary computer that executes ADAS functions using conventional rules of computer vision. As such, the secondary computer can use classical computer vision rules (if-then), and the presence of the neural network within the supervisory MCU can improve reliability, safety, and performance. For example, the various implementation forms and intentional non-identities make the overall system more fault-tolerant, especially against faults caused by software (or software-hardware interface) functions. For example, if there is a software bug or error in the software running on the primary computer and the non-identical software code running on the secondary computer provides the same overall result, the supervisory MCU may have greater confidence that the overall result is correct and that the bug in the software or hardware used by the primary computer is not causing a critical error.
[0124] In some examples, the output of the ADAS system 438 may be supplied to the perception block of the primary computer and / or the dynamic driving task block of the primary computer. For example, if the ADAS system 438 indicates a forward collision warning due to an object directly ahead, the perception block can use this information when identifying the object. In other examples, the secondary computer may have its own neural network that is trained as described herein and thus reduces the risk of misjudgment.
[0125] The mobile vehicle 400 may further include an infotainment SoC 430 (e.g., an in-vehicle infotainment system (IVI) in the mobile vehicle). Although illustrated and described as an SoC, the infotainment system may not be an SoC and may include two or more individual components. The infotainment SoC 430 may include a combination of hardware and software used to provide the mobile vehicle 400 with audio (e.g., music, mobile devices, navigation instructions, news, radio, etc.), video (e.g., TV, movies, streaming, etc.), telephony (e.g., hands-free calling), network connectivity (e.g., LTE, Wi-Fi, etc.), and / or information services (e.g., navigation systems, rear parking assistance, wireless data systems, fuel level, total mileage, brake fuel level, oil level, opening / closing doors, air filter information, and other mobile vehicle-related information). For example, the infotainment SoC 430 may include radio, disk player, navigation system, video player, USB and Bluetooth® connection, car computer, in-vehicle entertainment, Wi-Fi, steering wheel audio control, hands-free voice control, heads-up display (HUD), HMI display 434, telematics device, control panel (e.g., for controlling and / or interacting with various components, features, and / or systems), and / or other components. The infotainment SoC 430 may be further used to provide information (e.g., visual and / or audible) to the user of the mobile vehicle, such as information from the ADAS system 438, autonomous driving information such as planned mobile vehicle operations, trajectories, surrounding environment information (e.g., intersection information, mobile vehicle information, road information, etc.), and / or other information.
[0126] The infotainment SoC 430 may include GPU functionality. The infotainment SoC 430 can communicate with other devices, systems, and / or components of the moving vehicle 400 via a bus 402 (e.g., CAN bus, Ethernet®, etc.). In some examples, the GPU of the infotainment system can execute some self-driving functions in the event that the primary controller 436 (e.g., the primary and / or backup computer of the moving vehicle 400) fails, and the infotainment SoC 430 can be coupled to the supervisory MCU. In such examples, the infotainment SoC 430 can put the moving vehicle 400 into the chauffeur's safe stop mode as described herein.
[0127] The moving vehicle 400 may further include an instrument cluster 432 (e.g., digital dash, electronic instrument cluster, digital instrument panel, etc.). The instrument cluster 432 may include a controller and / or a supercomputer (e.g., an individual controller or supercomputer). The instrument cluster 432 may include a set of instruments such as a speedometer, fuel level, oil pressure, tachometer, odometer, turn indicator, gear shift position indicator, seat belt warning light, parking brake warning light, engine malfunction light, airbag (SRS) system information, lighting control device, safety system control device, navigation information, etc. In some examples, information can be displayed and / or shared between the infotainment SoC 430 and the instrument cluster 432. In other words, the instrument cluster 432 may be included as part of the infotainment SoC 430, and vice versa.
[0128] FIG. 4D is a system diagram of communication between a cloud-based server of FIG. 4A and an exemplary autonomous vehicle 400 according to some embodiments of the present disclosure. The system 476 may include a server 478, a network 490, and a moving vehicle including the moving vehicle 400. The server 478 may include a plurality of GPUs 484(A)-484(H) (collectively referred to herein as GPU 484), PCIe switches 482(A)-482(H) (collectively referred to herein as PCIe switch 482), and / or CPUs 480(A)-480(B) (collectively referred to herein as CPU 480). The GPUs 484, CPUs 480, and PCIe switches may be interconnected with each other by high-speed interconnects, such as, but not limited to, an NVLink interface 488 and / or a PCIe connection 486 developed by NVIDIA, for example. In some examples, the GPUs 484 are connected via NVLink and / or an NVSwitch SoC, and the GPUs 484 and the PCIe switches 482 are connected via a PCIe interconnect. Eight GPUs 484, two CPUs 480, and two PCIe switches are illustrated, but this is not intended to be limiting. Depending on the embodiment, each server 478 may include any number of GPUs 484, CPUs 480, and / or PCIe switches. For example, each server 478 may include eight, sixteen, thirty-two, and / or more GPUs 484, respectively.
[0129] Server 478 can receive, via network 490, image data representing an image indicating an unexpected or changed road condition, such as a recently started road construction, from a moving vehicle. Server 478 can transmit, via network 490, neural network 492, an updated neural network 492, and / or map information 494 including information regarding traffic and road conditions to the moving vehicle. The update of the map information 494 may include an update of the HD map 422, such as information regarding a construction site, a depression, a detour, a flood, and / or other obstacles. In some examples, the neural network 492, the updated neural network 492, and / or the map information 494 may result from new training and / or experience represented in data received from any number of moving vehicles in the environment and / or based on training performed in a data center (e.g., using server 478 and / or other servers).
[0130] Server 478 can be used to train a machine learning model (e.g., a neural network) based on training data. The training data can be generated by a moving vehicle and / or (e.g., using a game engine) generated in a simulation. In some examples, the training data is tagged (e.g., when the neural network benefits from supervised learning) and / or undergoes other preprocessing, while in other examples, the training data is not tagged and / or preprocessed (e.g., when the neural network does not require supervised learning). The training can be performed according to any one or more classes of machine learning techniques, including but not limited to, for example, the following classes: supervised training, semi-supervised training, unsupervised training, self-learning, reinforcement learning, associative learning, transfer learning, feature learning (including principal component and cluster analysis), multi-linear subspace learning, manifold learning, representation learning (including pre-dictionary learning), rule-based machine learning, anomaly detection, and their variants or combinations. After the machine learning model is trained, the machine learning model can be used by the moving vehicle (e.g., transmitted to the moving vehicle via network 490), and / or the machine learning model can be used by server 478 to remotely monitor the moving vehicle.
[0131] In some examples, server 478 can receive data from a moving vehicle and apply the data to a latest real-time neural network for real-time intelligent inference. Server 478 can include a deep learning supercomputer and / or a dedicated AI computer powered by GPU 484, such as DGX and DGX Station machines developed by NVIDIA. However, in some examples, server 478 can include a deep learning infrastructure that uses only CPU-powered data centers.
[0132] The deep learning infrastructure of server 478 can have the ability of high-speed real-time inference, and use that ability to evaluate and verify the condition of the processors, software, and / or related hardware within moving vehicle 400. For example, the deep learning infrastructure can receive periodic updates from moving vehicle 400, such as a sequence of images and / or objects in which moving vehicle 400 is located within that sequence of images (e.g., via computer vision and / or other machine learning object classification techniques). The deep learning infrastructure can execute its own neural network to identify objects and compare them with the objects identified by moving vehicle 400. If the results do not match and the infrastructure concludes that the AI within moving vehicle 400 is not functioning properly, server 478 can send a signal to moving vehicle 400 that infers control, notifies the passengers, and commands the fail-safe computer of moving vehicle 400 to complete a safe parking operation.
[0133] For inference, server 478 can include GPU 484 and one or more programmable inference acceleration devices (e.g., NVIDIA's TensorRT). The combination of a GPU-powered server and inference acceleration can enable real-time responsiveness. In other examples, such as when performance is not as highly required, a server powered by a CPU, FPGA, and other processors can be used for inference.
[0134] FIG. 5 is a block diagram of an example of a computing device 500 suitable for use in the implementation of some embodiments of the present disclosure. Computing device 500 can include an interconnect system 502 that indirectly or directly connects the following devices: memory 504, one or more central processing units (CPUs) 506, one or more graphics processing units (GPUs) 508, a communication interface 510, an I / O port 512, input / output components 514, a power supply 516, one or more presentation components 518 (e.g., a display), and one or more logic units 520.
[0135] Although the various blocks of FIG. 5 are shown as being connected via an interconnect system 502 with lines, this is not intended to be limiting and is merely for clarity. For example, in some embodiments, a presentation component 518 such as a display device may be considered an I / O component 514 (e.g., if the display is a touch screen). As another example, the CPU 506 and / or GPU 508 may include memory (e.g., memory 504 may represent a storage device in addition to the memory of the GPU 508, CPU 506, and / or other components). In other words, the computing device of FIG. 5 is merely illustrative. Categories such as "workstation", "server", "laptop", "desktop", "tablet", "client device", "mobile device", "handheld device", "gaming console", "electronic control unit (ECU)", "virtual reality system", "augmented reality system", and / or other device or system types are all intended to be within the scope of the computing device of FIG. 5 and thus are not distinguished.
[0136] The interconnect system 502 may represent one or more links or buses, such as an address bus, a data bus, a control bus, or a combination thereof. The interconnect system 502 may include one or more bus or link types, such as an Industry Standard Architecture (ISA) bus, an Extended Industry Standard Architecture (EISA) bus, a VESA (Video Electronics Standards Association) bus, a Peripheral Component Interconnect (PCI) bus, a Peripheral Component Interconnect Express (PCIe) bus, and / or another type of bus or link. In some embodiments, there are direct connections between components. As an example, the CPU 506 may be directly connected to the memory 504. Further, the CPU 506 may be directly connected to the GPU 508. If there are direct or point-to-point connections between components, the interconnect system 502 may include a PCIe link for making the connections. In these examples, the PCI bus need not be included in the computing device 500.
[0137] The memory 504 may include any of a variety of computer-readable media. The computer-readable media may be any available media that can be accessed by the computing device 500. The computer-readable media may include both volatile and non-volatile media, and removable and non-removable media. By way of example, and not limitation, the computer-readable media may comprise computer storage media and communication media.
[0138] A computer storage medium can include both volatile and non-volatile media and / or removable and non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules, and / or other data types. For example, memory 504 can store computer-readable instructions (such as those representing a program and / or program elements) such as an operating system. A computer storage medium can include, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by computing device 500. In this specification, a computer storage medium does not include a signal itself.
[0139] A computer storage medium can implement computer-readable instructions, data structures, program modules, and / or other data types in a modulated data signal such as a carrier wave or other transport mechanism, and includes any information delivery medium. The term "modulated data signal" can refer to a signal that has one or more of its characteristic sets or that changes in such a way as to encode information in the signal. By way of example, and without limitation, a computer storage medium can include wired media such as a wired network or direct wired connection, and wireless media such as acoustic, RF, infrared and other wireless media. Any combination of the foregoing should also be included within the scope of computer-readable media.
[0140] The CPU 506 may be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 500 to execute one or more of the methods and / or processes described herein. The CPU 506 may each include one or more (e.g., 1, 2, 4, 8, 28, 72, etc.) cores having the ability to process multiple software threads simultaneously. The CPU 506 may include any type of processor and, depending on the type of computing device 500 implemented, may include different types of processors (e.g., a processor having fewer cores for a mobile device and a processor having more cores for a server). For example, depending on the type of computing device 500, the processor may be an Advanced RISC Machines (ARM) processor implemented using reduced instruction set computing (RISC), or an x86 processor implemented using complex instruction set computing (CISC). The computing device 500 may include one or more CPUs 506 within one or more microprocessors or auxiliary coprocessors, such as a compute coprocessor.
[0141] In addition to or instead of the CPU 506, the GPU 508 may be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 500 to execute one or more of the methods and / or processes described herein. One or more of the GPUs 508 may be an integrated GPU (e.g., may be with one or more of the CPUs 506), and / or one or more of the GPUs 508 may be a discrete GPU. In an embodiment, one or more of the GPUs 508 may be one or more coprocessors of the CPU 506. The GPU 508 may be used by the computing device 500 to render graphics (e.g., 3D graphics) or perform general-purpose computing. For example, the GPU 508 may be used for general-purpose computing on the GPU (GPGPU). The GPU 508 may include hundreds or thousands of cores having the ability to process hundreds or thousands of software threads simultaneously. The GPU 508 can generate pixel data for an output image in response to rendering commands (e.g., rendering commands from the CPU 506 received via a host interface). The GPU 508 may include graphics memory, such as display memory, for storing pixel data or any other suitable data, such as GPGPU data. The display memory may be included as part of the memory 504. The GPU 508 may include two or more GPUs operating in parallel (e.g., via a link). The link can be directly connected to the GPUs (e.g., using NVLINK), or the GPUs can be connected via a switch (e.g., using NVSwitch). When coupled together, each GPU 508 can generate pixel data or GPGPU data for different portions of the output or different outputs (e.g., the first GPU for a first image and the second GPU for a second image). Each GPU can include its own memory or can share memory with other GPUs.
[0142] In addition to and / or instead of the CPU 506 and / or the GPU 508, the logic unit 520 may be configured to execute at least some of the computer-readable instructions to control one or more of the computing devices 500 to execute one or more of the methods and / or processes described herein. In an example, the CPU 506, the GPU 508, and / or the logic unit 520 may execute any combination of methods, processes, and / or portions thereof discretely or in parallel. One or more of the logic units 520 may be part of and / or integrated with one or more of the CPU 506 and / or the GPU 508, and / or one or more of the logic units 520 may be discrete components with respect to the CPU 506 and / or the GPU 508 or otherwise external to them. In an example, one or more of the logic units 520 may be one or more coprocessors of one or more of the CPU 506 and / or one or more of the GPU 508.
[0143] Examples of the logic unit 520 include one or more processing cores and / or their components, such as tensor cores (TC), tensor processing units (TPU), pixel visual cores (PVC), vision processing units (VPU), graphics processing clusters (GPC), texture processing clusters (TPC), streaming multiprocessors (SM), tree traversal units (TTU), artificial intelligence accelerators (AIA), deep learning accelerators (DLA), arithmetic logic units (ALU), application-specific integrated circuits (ASIC), floating-point units (FPU), I / O elements, peripheral component interconnect (PCI) or peripheral component interconnect express (PCIe) elements, and / or the like.
[0144] The communication interface 510 can include one or more receivers, transmitters, and / or transceivers that enable the computing device 500 to communicate with other computing devices via an electronic communication network, including wired and / or wireless communication. The communication interface 510 can include components and functions for enabling communication via any of several different networks, such as wireless networks (e.g., Wi-Fi, Z-Wave, Bluetooth®, Bluetooth® LE, ZigBee, etc.), wired networks (e.g., communicating via Ethernet® or InfiniBand), low-power wide-area networks (e.g., LoRaWAN, SigFox, etc.), and / or the Internet.
[0145] The I / O port 512 can enable the computing device 500 to be logically connected to other devices, including some of which may be built in (e.g., integrated) into the computing device 500, such as I / O components 514, presentation components 518, and / or other components. Exemplary I / O components 514 include microphones, mice, keyboards, joysticks, game pads, game controllers, satellite dishes, scanners, printers, wireless devices, etc. The I / O components 514 can provide a natural user interface (NUI) that processes air gestures, voice, or other physiological inputs generated by the user. In some cases, the input can be sent to an appropriate network element for further processing. The NUI can implement any combination of voice recognition, stylus recognition, face recognition, biometric recognition, gesture recognition on and adjacent to the screen, air gestures, head and gaze tracking, and touch recognition related to the display of the computing device 500 (as will be described in more detail later). The computing device 500 can include a depth camera, such as a stereoscopic camera system, an infrared camera system, an RGB camera system, touch screen technology, and combinations thereof, for gesture detection and recognition. Additionally, the computing device 500 can include an accelerometer or gyroscope that enables motion detection (e.g., as part of an inertia measurement unit (IMU)). In some examples, the output of the accelerometer or gyroscope can be used by the computing device 500 to render immersive augmented reality or virtual reality. * The exemplary I / O components 514 include microphones, mice, keyboards, joysticks, game pads, game controllers, satellite dishes, scanners, printers, wireless devices, etc. The I / O components 514 can provide a natural user interface (NUI) that processes air gestures, voice, or other physiological inputs generated by the user. In some cases, the input can be sent to an appropriate network element for further processing. The NUI can implement any combination of voice recognition, stylus recognition, face recognition, biometric recognition, gesture recognition on and adjacent to the screen, air gestures, head and gaze tracking, and touch recognition related to the display of the computing device 500 (as will be described in more detail later). The computing device 500 can include a depth camera, such as a stereoscopic camera system, an infrared camera system, an RGB camera system, touch screen technology, and combinations thereof, for gesture detection and recognition. Additionally, the computing device 500 can include an accelerometer or gyroscope that enables motion detection (e.g., as part of an inertia measurement unit (IMU)). In some examples, the output of the accelerometer or gyroscope can be used by the computing device 500 to render immersive augmented reality or virtual reality.
[0146] The power supply device 516 can include a hard-wired power supply device, a battery power supply device, or a combination thereof. The power supply device 516 can provide power to the computing device 500 to enable the components of the computing device 500 to operate. The presentation component 518 may include a display (e.g., a monitor, a touch screen, a television screen, a head-up display (HUD), other display types, or a combination thereof), a speaker, and / or other presentation components. The presentation component 518 can receive data from other components (e.g., GPU 508, CPU 506, etc.) and output the data (e.g., as an image, video, sound, etc.).
[0147] The present disclosure may be described in the general context of computer-executable instructions, such as program modules, being executed by a computer or other machine, such as a portable information terminal or other handheld device. Generally, program modules include routines, programs, objects, components, data structures, etc., which refer to code that performs specific tasks or implements specific abstract data types. The present disclosure may be implemented in a variety of configurations, including handheld devices, household appliances, general-purpose computers, more specialized computing devices, etc. The present disclosure may also be implemented in a distributed computing environment where tasks are performed by remote processing devices linked via a communication network.
[0148] As used herein, the description of "and / or" with respect to two or more elements should be construed to mean either only one element, or a combination of elements. For example, "element A, element B, and / or element C" can include only element A, only element B, only element C, element A and element B, element A and element C, element B and element C, or elements A, B, and C. Additionally, "at least one of element A or element B" can include at least one of element A, at least one of element B, or at least one of element A and at least one of element B. Further, "at least one of element A and element B" can include at least one of element A, at least one of element B, or at least one of element A and at least one of element B.
[0149] The subject matter of this disclosure has been described with specificity in order to meet statutory requirements. However, the description itself is not intended to limit the scope of the disclosure. Rather, the inventors intend that the claimed subject matter be implemented in other forms, including in combinations of different steps or steps similar to those described in this document, in conjunction with other current or future technologies. Further, the terms "step" and / or "block" may be used herein to imply different elements of a method of use, but these terms should not be construed to imply any particular order among the various steps disclosed herein except where the order of individual steps is explicitly recited and so recited.
[0150] Each block of the method described below in connection with FIG. 6 can include a computational process that can be executed using any combination of hardware, firmware, and / or software. For example, various functions can be implemented by a processor executing instructions stored in memory. The method can also be implemented as computer-usable instructions stored on a computer storage medium. The method can be provided, by way of example, as a stand-alone application, service, or hosted service (stand-alone or in combination with another hosted service), or as a plug-in to another product. Additionally, the method of FIG. 6 is described, by way of example, with respect to the exemplary autonomous vehicle system of FIGS. 4A-4D. However, these methods can be executed additionally or alternatively by any one system, or any combination of systems, including but not limited to those described herein.
[0151] FIG. 6 is a flowchart showing process steps for determining a gaze direction in the presence of glare according to an embodiment of the present disclosure. The process of FIG. 6 can begin when computing device 300 receives an image of a subject captured by camera 310 (step 600). The computing device then identifies the face and eyes of the subject in the received image (step 610). The face of the subject can be located within the image using any method or process, including known computer vision-based face detection processes that detect faces without using a neural network, such as edge detection methods, feature search methods, probabilistic face models, graph matching, classifiers such as histograms of oriented gradient (HOG) supplied to support vector machines, HaarCascade classifiers, and the like. The determination of the face position can also be performed using neural network-based face recognition methods, such as those using a deep neural network (DNN) face recognition scheme, and any others. Embodiments of the present disclosure also contemplate locating the eyes of the subject from the same image. The eye location can be performed in any manner by a known computer vision-based eye detection process, including, for example, any of the aforementioned non-neural network-based techniques, neural network-based eye recognition methods, and the like.
[0152] After the face and eyes of the subject have been located within the received image, computing device 300 extracts glare points from the received image and generates a face representation (step 620). Glare can be identified using any method or process. For example, glare pixels can be identified as those pixels of the received image having a luminance value exceeding a predetermined threshold. As another example, glare pixels can be identified according to a known or characteristic glare shape (e.g., circular or star-shaped spot, or the like) within the captured image. A glare mask can be formed, for example, by setting glare pixels as white or saturated pixels and the remaining pixels as black pixels.
[0153] The face expression, for example, the binary mask of the identified face, can be determined using any method or process. Such a binary mask can be formed, for example, by setting the pixels corresponding to the identified face to white or another uniform color and setting the remaining pixels to black pixels.
[0154] Eye crops are selected (step 630) by, for example, drawing a bounding box around each eye identified in step 610 in a known manner and cropping the image accordingly. Similarly, a face grid is then calculated (step 640). The face grid 225 can be formed by cropping the identified face according to a bounding box that can be drawn in any known manner, and by assigning one or more colors (or a first value) to grid elements that contain face pixels and setting grid elements that do not contain face pixels to black (or a value different from the first value), and projecting the cropped image onto a predetermined coarse grid.
[0155] The mask of step 620, the eye crops of step 630, and the face grid of step 640 are input into the machine learning models (such as feature extraction layers 230, 235, and 240, and Fully connected layer 245 and 255) of the embodiments of the present disclosure in order to calculate the gaze direction as described above (step 650). This gaze direction is then output to the gaze assistance system 320 (step 660), and the gaze assistance system 320 can perform any one or more actions based on the gaze direction it receives (step 670).
[0156] The embodiments of the present disclosure contemplate machine learning models configured in other ways in addition to those represented in FIG. 2. FIG. 7 is a block diagram representation of one such alternative example. Here, the machine learning model of FIG. 7 includes Fully connected layer 710, 740, and 755, and feature extraction layers 715, 745, and 750. The feature extraction layers 745 and 750, and Fully connected layer755 has, as inputs, a left-eye expression 725, a right-eye expression 730, and a face expression 735. The three branches from the right end of the model in FIG. 7 may be substantially the same as those in FIG. 2. That is, the feature extraction layers 745 and 750, and Fully connected layer 755 may be substantially the same as the feature extraction layers 235 and 240, and Fully connected layer 245, respectively. Similarly, the inputs 725, 730, and 735 may be substantially the same as the inputs 215, 220, and 225, respectively.
[0157] The glare expression 700 and the face expression 705 may be substantially the same as the inputs 200 and 205 in FIG. 2. Thus, the glare expression 700 may be a binary mask of the glare points identified in the input image 100, and the face expression 705 may be a binary mask of the face pixels extracted from the input image 100. However, unlike FIG. 2, the glare expression 700 is Fully connected layer input to 710, and the face expression 705 is input to the feature extraction layer 715. Fully connected layer 710 outputs features that explain the spatial position of the glare in the input image 100, and the feature extraction layer 715 outputs features that explain the orientation of the face in the input image 100. These two outputs are concatenated by the concatenation block 720 and output a set of features that explain the orientation of the face in the input image 100 Fully connected layer which is input to 740. Fully connected layer The output of 740, the feature extraction layers 745 and 750, and Fully connected layer 755 is then processed as described above with respect to FIG. 2. Specifically, the outputs are concatenated by the concatenation block 760 and input to Fully connected layer 765 which outputs the gaze direction. In this manner, the spatial position of the glare is used to determine the orientation of the face, rather than the glare points themselves.
[0158] In the embodiments of the present disclosure Fully connected layerAnd the training of any of the feature extraction layers can be performed in any suitable manner. In some embodiments, the machine learning models of the embodiments of the present disclosure, such as the architectures of FIGS. 2 and 7, can be trained end-to-end. In certain embodiments, the Fully connected layer and the feature extraction layers can be trained in a supervised manner using labeled input images of a subject whose gaze direction is known. In some embodiments, glare spots can be added to these input images to simulate the presence of glare. The added glare spots can be added to each image in any manner, such as random placement within various images, placement at predetermined positions within the input image, or the like.
[0159] Those skilled in the art will appreciate that the embodiments of the present disclosure are not limited to gaze estimation. Specifically, the methods and systems of the embodiments of the present disclosure can be used to account for glare in the determination of any eye movement variable. For example, the machine learning models of FIGS. 2 and / or 7 can be constructed and trained such that the outputs of FC layers 255, 765 are estimates of any other eye movement variable in addition to gaze, such as pupil size, pupil movement, cognitive load, fatigue, such as due to drugs or alcohol, incapacitation, person or object recognition through the iris, any other physiological response, or the like. This can be accomplished, for example, by training the machine learning models of the embodiments of the present disclosure using an input data set (such as a glare mask, face mask, eye crop, face grid, etc.) labeled with the value of any eye movement variable of the subject. The machine learning model can also be constructed to extract any features useful in the determination of such eye movement variables. In this manner, the machine learning models of the embodiments of the present disclosure can be constructed and trained to estimate any eye movement variable in a glare-tolerant manner, i.e., to account for glare in the determination of any eye movement variable.
[0160] The foregoing has been presented for purposes of illustration and description to enable a complete understanding of the disclosure. However, it will be apparent to those skilled in the art that specific details are not required in order to practice the methods and systems of the disclosure. Accordingly, the foregoing description of specific embodiments of the invention has been presented for purposes of illustration and description. They are not intended to be exhaustive or to limit the invention to the precise forms disclosed. Numerous modifications and variations are possible in light of the above teachings. For example, the glare point can be input into various machine learning models in any manner, such as by input of a binary mask of the glare point, input of a table or set of glare positions, or any other representation of glare. The embodiments were chosen and described in order to best illustrate the principles of the invention and its practical applications, thereby enabling those skilled in the art to best utilize the disclosure and various embodiments with various modifications as are suited to the particular use contemplated. Additionally, different features of the various embodiments, whether disclosed or not, may be combined or matched in other ways to produce further embodiments contemplated by the disclosure.
Claims
1. A method for determining a gaze direction in the presence of glare, comprising: Determining the gaze direction of a subject in an image using a parallel processing circuit, wherein the gaze direction is determined at least in part from the outputs of a plurality of machine learning models having as inputs a separated representation of glare in the image, a plurality of representations of at least a portion of the face of the subject in the image, and a portion of the image corresponding to at least one eye of the subject; Initiating an operation based on the determined gaze direction; and wherein the plurality of representations of at least a portion of the face of the subject further includes a first representation and a second representation of the face of the subject, the first representation of the face of the subject includes a binary mask of the face of the subject, the second representation of the face of the subject includes a face grid of the face of the subject, and the plurality of machine learning models includes a first machine learning model having as an input the first representation of the face of the subject and a second machine learning model having as an input the second representation of the face of the subject.
2. The method of claim 1, further comprising generating, based on the image using a processing circuit, the separated representation of glare, the plurality of representations of at least a portion of the face of the subject, and the portion of the image corresponding to at least one eye of the subject.
3. The method of claim 1, wherein one of the machine learning models has as inputs both the separated representation of glare and the representation of at least a portion of the face of the subject.
4. The method of claim 1, wherein one of the machine learning models has as an input the separated representation of glare and another one of the machine learning models has as an input the representation of at least a portion of the face of the subject.
5. The method according to claim 4, wherein one of the machine learning models includes a fully connected layer having the separated representation of the glare as an input.
6. The method according to claim 1, wherein the image is generated using sensor data obtained using a sensor device corresponding to a vehicle, and the starting step further includes starting an operation of the vehicle based on the determined gaze direction.
7. The portion of the image corresponding to at least one eye of the subject further includes a first portion of the image corresponding to a first eye of the subject and a second portion of the image corresponding to a second eye of the subject, The plurality of machine learning models includes a first machine learning model having the first portion of the image as an input and a second machine learning model having the second portion of the image as an input. The method according to claim 1.
8. The method according to claim 1, wherein the separated representation of the glare in the image further includes a binary mask of glare points extracted from the image.
9. A system for determining a gaze direction in the presence of glare, comprising: a memory; and a parallel processing circuit, the parallel processing circuit being configured to: determine a gaze direction of a subject in an image, the gaze direction being determined at least in part from an output of a plurality of machine learning models having a separated representation of glare in the image as an input, a plurality of representations of at least a portion of the face of the subject in the image, and a portion of the image corresponding to at least one eye of the subject; and start an operation based on the determined gaze direction and is configured to perform, the plurality of representations of at least a portion of the face of the subject further includes a first representation and a second representation of the face of the subject. The first representation of the face of the subject includes a binary mask of the face of the subject, The second representation of the face of the subject includes a face grid of the face of the subject, The plurality of machine learning models includes a first machine learning model having, as an input, the first representation of the face of the subject, and a second machine learning model having, as an input, the second representation of the face of the subject, a system.
10. The parallel processing circuit is further configured to generate the separated representation of the glare, the plurality of representations of at least a part of the face of the subject, and the part of the image corresponding to at least one eye of the subject, and the separated representation of the glare, the plurality of representations of at least a part of the face of the subject, and the part of the image corresponding to at least one eye of the subject are each generated based on the image, Claim The method according to 9.
11. One of the machine learning models has, as inputs, both the separated representation of the glare and the representation of at least a part of the face of the subject, the system according to Claim 9.
12. One of the machine learning models has, as an input, the separated representation of the glare, and another one of the machine learning models has, as an input, the representation of at least a part of the face of the subject, the system according to Claim 9.
13. The one of the machine learning models includes a fully connected layer having, as an input, the separated representation of the glare, the system according to Claim 12.
14. The starting further includes starting an operation of a vehicle based on the determined gaze direction, the system according to Claim 9.
15. The portion of the image corresponding to at least one eye of the subject further includes a first portion of the image corresponding to the first eye of the subject and a second portion of the image corresponding to the second eye of the subject. The plurality of machine learning models includes a first machine learning model having the first portion of the image as an input and a second machine learning model having the second portion of the image as an input. The system according to claim 9.
16. The plurality of representations of at least a portion of the face of the subject further includes a first representation and a second representation of the face of the subject. The plurality of machine learning models includes a first machine learning model having the first representation of the face of the subject as an input and a second machine learning model having the second representation of the face of the subject as an input. The system according to claim 9.
Citation Information
Patent Citations
Driving support apparatus and driving support method
JP2017129973A
Addressing glare in eye tracking
JP2017521738A
Information processing apparatus for estimating person's line of sight and estimation method, and learning device and learning method
JP2019028843A
Information processing apparatus, driver monitoring system, information processing method, and information processing program
JP2019091255A
Vehicle monitoring device
JP2020047087A