Monitoring the attention and cognitive load of occupants for autonomous and semi-autonomous driving applications
The system integrates interior and exterior vehicle perception to accurately assess driver attention and cognitive load, addressing inaccuracies in conventional systems by comparing estimated field-of-view with external perception, enhancing safety in autonomous and semi-autonomous driving.
Patent Information
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- NVIDIA CORP
- Filing Date
- 2021-10-11
- Publication Date
- 2026-05-07
AI Technical Summary
Conventional driver monitoring systems inaccurately assess driver attention and cognitive load due to independent measurement without considering external environmental conditions, leading to excessive and unreliable warnings.
A system that integrates interior and exterior vehicle perception by comparing the driver's estimated field-of-view with external vehicle perception using deep neural networks to determine the driver's attention and cognitive load, ensuring accurate and informed decision-making.
Enhances the accuracy and reliability of driver state assessment by considering both interior and exterior environmental factors, reducing unnecessary warnings and improving safety in autonomous and semi-autonomous driving scenarios.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
STATE OF THE ART
[0001] Cognitive and visual attention play a crucial role in a driver's ability to recognize safety-critical events and thus control the vehicle safely. For example, if a driver is distracted—e.g., by daydreaming, distraction, using mental capacity for tasks other than the current driving task, observing off-road activity, etc.—the driver may be unable to make appropriate planning and control decisions, or the correct decisions may be delayed. To address this, some vehicle systems—such as Advanced Driver Assistance Systems (ADAS)—are employed to generate audible, visual, and / or tactile warnings or cues for the driver to inform them about environmental and / or road conditions (e.g., vulnerable road users (VRUs), traffic lights, heavy traffic, potential collisions, etc.). With the increasing number of warning systems (e.g.,Automatic emergency braking (AEB), blind spot detection (BSD), forward collision warning (FCW), etc., can generate an overwhelming number of warnings for the driver. As a result, a driver might disable the safety systems for one or more of these systems, thereby negating their effectiveness in mitigating unsafe events.
[0002] In some conventional systems, driver monitoring systems and / or driver fatigue detection systems can be used to determine the driver's current state. However, these conventional systems measure the driver's cognitive load or attention independently. Once a driver's state has been determined, the driver may, for example, be warned or alerted to their detected inattention or elevated cognitive load. This not only results in a warning or alert being issued in addition to the existing warnings from other ADAS systems, but also means that the determination of attention or high cognitive load can be inaccurate or imprecise. Regarding attention, for example, a user's gaze within the vehicle can be measured—such as when the user looks into the vehicle cabin.A driver who is rated as attentive because their gaze is fixed on the windshield may not actually be paying attention because external environmental conditions—such as the position of static or dynamic objects, road conditions, waiting times, etc.—are not being taken into account. Regarding cognitive load, conventional systems use deep neural networks (DNNs) trained on simulated data, which correlate pupil size or other eye characteristics, eye movements, blink rate or other eye measurements, and / or other information with current cognitive load. However, these measurements are subjective and can be inaccurate or imprecise for certain users—for example, cognitive load may affect different drivers differently, and / or some drivers may perform better than others under high cognitive load.If driver inattention and cognitive load are considered independently, and environmental conditions outside the vehicle are ignored, this can lead to excessive warnings based on inaccurate, unreliable, or fragmentary information about the driver's current state. In this technical field, reference is also made to DE 10 2020 122 357 A1, DE 10 2020 113 712 A1, DE 10 2014 109 079 A1, DE 10 2015 101 239 A1, and DE 10 2015 101 358 A1. DE 10 2019 122 267 A1 describes a method for tracking the gaze direction of a vehicle occupant. SUMMARY
[0003] The present embodiments relate to monitoring occupant attention and cognitive load for semi-autonomous or autonomous driving applications. A system and a method according to the independent claims are disclosed which, in addition to monitoring attention and / or cognitive load, compare estimated field-of-view or gaze information of a user with vehicle perception information corresponding to an environment outside the vehicle. Consequently, the interior monitoring of a driver or vehicle occupant can be extended to the vehicle's exterior to determine whether the driver or occupant has processed or seen certain types of objects, environmental conditions, or other information outside the vehicle—e.g., dynamic actors, static objects, vulnerable road users (VRUs), waiting status information, signs, potholes, speed bumps, debris, etc.If it is determined that a projected representation – e.g., in a space coordinate system – of a user's field of view overlaps with a detected object, condition, and / or similar, the system may assume that the user has seen the object, condition, etc., and may refrain from taking any action (e.g., generating a notification, activating an AEB system, etc.).
[0004] In some embodiments, for a more comprehensive understanding of the user's state—e.g., according to the user's current ability to process seen or visualized information—the user's attention and / or cognitive load can be monitored to determine whether one or more actions should be taken (e.g., to generate a visual, audible, tactile, or other type of notification, to take control of the vehicle, to activate one or more AEB systems, etc.). Thus, even if it is determined that the object, the road condition, etc.,Once an object has entered the user's field of vision, an action appropriate to the object, road conditions, and / or the like may be performed if the user's assessed attention and / or cognitive load indicates that the user may not have fully processed the information to make an informed driving decision without the action. Consequently, and unlike conventional systems, notifications, AEB activations, and / or other actions can be determined based on a more objective and complete user state, as determined by cognitive load, attention, and / or a comparison between the external perception of the vehicle and the user's estimated perception as projected onto the vehicle from the outside. Furthermore, simulations of real-world scenarios—e.g.,in a virtual simulated environment - by using more objective measures of attention and / or cognitive load, they can be more accurate and reliable and thus better suited for development, testing and ultimately use in a real system. BRIEF DESCRIPTION OF THE DRAWINGS
[0005] The systems and methods presented for monitoring occupant attention and cognitive load for semi-autonomous or autonomous driving applications are described in detail below with reference to the accompanying figures, whereby the following applies: Fig. Figure 1 shows a data flow diagram for a method for monitoring attention and / or cognitive load according to some embodiments of the present disclosure; Fig. 2A shows an example graphic that was generated using eye movement data according to some embodiments of the present disclosure; Fig. Figure 2B shows an example graphic that includes visualizations of vehicle regions generated using eye movement data according to some embodiments of the present disclosure; Fig. Figures 2C-2D show example images of eye positions in a time step or frame used to determine eye movement information according to some embodiments of the present disclosure; Fig. 2E shows an example image of a heat map corresponding to eye movement data collected over a certain period of time, according to some embodiments of the present disclosure; Fig. Figure 3 shows an example visualization of a view or field of vision representation outside a vehicle for comparing the vehicle perception with the estimated occupant perception according to some embodiments of the present disclosure; Fig. Figure 4 is a flowchart showing a method for determining actions based on the driver's attention and / or cognitive load according to some embodiments of the present disclosure; Fig. Figure 5A is an illustration of an example of an autonomous vehicle according to some embodiments of the present disclosure; Fig. 5B is an example of camera positions and fields of view for the autonomous vehicle from Fig. 5A is represented according to some embodiments of the present disclosure; Fig. 5C is a block diagram of an example system architecture of the exemplary autonomous vehicle from Fig. 5A is represented according to some embodiments of the present disclosure; Fig. 5D is a system diagram for the communication between the cloud-based server(s) and the example autonomous vehicle. Fig. 5A is represented according to some embodiments of the present disclosure; Fig. Figure 6 is a block diagram of an exemplary computing device suitable for use in the implementation of some embodiments of the present disclosure; and Fig. Figure 7 is a block diagram of an exemplary data center suitable for use in the implementation of some embodiments of the present disclosure. DETAILED DESCRIPTION
[0006] Systems and methods relating to monitoring occupant attention and cognitive load for semi-autonomous or autonomous driving applications are disclosed. Although the present disclosure can be described in relation to an example of an autonomous vehicle 500 (alternatively referred to herein as "Vehicle 500" or "Ego-Vehicle 500"), an example of which is given here in relation to the Fig. The fact that the present disclosure can be described in sections 5A-5D is not intended to be limiting. The systems and methods described herein can be used, for example, by non-autonomous vehicles, semi-autonomous vehicles (e.g., in one or more advanced driver assistance systems (ADAS)), robots, warehouse vehicles, off-road vehicles, flying ships, boats, and / or other types of vehicles. Even though the present disclosure can be described in relation to autonomous driving, this is not to be understood as a limitation. The systems and methods described herein can be used, for example, in robotics, in flight systems (e.g., for determining attention and / or cognitive load), in boat systems, in simulation environments (e.g.,to simulate actions based on the attention and / or cognitive load of a human operator of virtual vehicles within a virtual simulation environment) and / or in other technology areas.
[0007] With reference to Fig. 1, shows Fig. 1 A data flow diagram for a method 100 for monitoring attention and / or cognitive load according to some embodiments of the present disclosure. It is understood that this and other arrangements described herein are given only as examples. Other arrangements and elements (e.g., machines, interfaces, functions, commands, groupings of functions, etc.) may be used in addition to or instead of those shown, and some elements may be omitted entirely. Furthermore, many of the elements described herein are functional units that may be implemented as discrete or distributed components or in conjunction with other components and in any suitable combination and position. Various functions described herein as being performed by units may be performed by hardware, firmware, software, and / or any combination thereof.For example, various functions can be performed by a processor that executes instructions stored in memory.
[0008] The method 100 may involve generating and / or receiving sensor data 102A and / or 102B (here collectively referred to as "sensor data 102") from one or more sensors of a vehicle 500 (which may be similar to the vehicle 500 or may include non-autonomous or semi-autonomous vehicles).The sensor data 102 can be used within the procedure 100 to track body movements or posture of one or more occupants of the vehicle 500, to track eye movements of one or more occupants of the vehicle 500, to determine the attention and / or cognitive load of one or more occupants, to project a representation of a gaze or field of view of one or more occupants outside the vehicle 500, to generate outputs 114 using one or more deep neural networks (DNNs) 112, to compare the field of view or gaze representation with the outputs 114, to determine a state of one or more occupants, to determine one or more actions to be taken based on the state, and / or to perform other tasks or operations. The sensor data 102 can, without limitation, include sensor data 102 from any type of sensor, such as..., but not limited to the sensors described herein in relation to the vehicle 500 and / or other vehicles or objects – such as robotic devices, VR systems, AR systems, etc., in some examples. As a non-limiting example and with reference to . Fig. 5A-5C, the sensor data 102 can include data generated by, among others, the following sensors: GNSS sensors 558 (e.g., GPS sensors, DGPS sensors, etc.), RADAR sensors 560, ultrasonic sensors 562, LIDAR sensors 564, IMU sensors 566 (e.g., accelerometer(s), gyroscope(s), magnetic compass(s), magnetometer(s), etc.), microphone(s) 596, stereo camera(s) 568, wide-angle camera(s) 570 (e.g., fisheye cameras), infrared camera(s) 572, surround-view camera(s) 574 (e.g., 360-degree cameras), long-range and / or medium-range camera(s) 598, in-cabin cameras, thermal, pressure, or Touch sensors in the cabin, motion sensors in the cabin, microphones in the cabin, speed sensor(s) 544 (e.g. to measure the speed of the vehicle 500 and / or the distance traveled) and / or other sensor types.
[0009] In some embodiments, the sensor data 102A may correspond to sensor data generated by one or more in-cabin sensors, such as one or more in-cabin cameras, in-cabin near-infrared (NIR) sensors, in-cabin microphones and / or the like, and the sensor data 102B may correspond to sensor data generated by one or more external sensors of the vehicle 500, such as one or more cameras, RADAR sensor(s) 560, ultrasonic sensor(s) 562, LIDAR sensor(s) 564, and / or the like. Thus, sensor data 102A can correspond to sensors with a sensory field or field of view inside the vehicle 500 (e.g. cameras with the occupant(s), such as the driver, in their field of view) and sensor data 102B to sensors with a sensory field or field of view outside the vehicle 500 (e.g. cameras, LiDAR sensors, etc. with sensory fields that include the environment outside the vehicle 500).In some embodiments, however, the sensor data 102A and the sensor data 102B may also include sensor data from any sensors with sensor fields inside and / or outside the vehicle 500.
[0010] The sensor data 102A can be used by a body tracker 104 and / or an eye tracker 106 to determine gestures, postures, activities, eye movements (e.g., saccade velocity, smooth tracking, gaze positions, directions or vectors, pupil size, blink rate, range and distribution of the road scan, etc.) and / or other information about an occupant—e.g., a driver—of the vehicle 500. This information can then be used by an attention determiner 108 to determine the attention of one or more occupants, by a cognitive load determiner 110 to determine the cognitive load of the occupant(s), and / or by a field-of-view projector 116 (FOV) to compare it, via a comparator 120, with the external perceptual outputs 114 from one or more deep neural networks (DNNs) 112.Information representative of the attention, cognitive load and / or outputs of the comparator 120 can be analyzed by a state machine 122 to determine a state of the occupant(s), and the state can be used by an action determiner 124 to determine one or more actions or operations to be performed (e.g., to issue a visual, auditory and / or tactile alert, to suppress an alert, to activate an ADAS system, to take over autonomous control of the vehicle 500, etc.).
[0011] The Body Tracker 104 can use sensor data 102A – e.g., sensor data from one or more in-cabin cameras, microphones, pressure sensors, temperature sensors, etc. – to determine body posture, pose, activity, or condition (e.g., two hands on the steering wheel, one hand on the steering wheel, texting, reading, slouching, sudden illness, incapacitated, distracted, etc.) and / or other information about one or more occupants. The Body Tracker 104 can, for example, execute one or more machine learning algorithms, deep neural networks, computer visualization algorithms, image processing algorithms, mathematical algorithms, and / or similar processes to determine the body tracking information. In some non-restrictive embodiments, the Body Tracker 104 may include similar features, functions and / or components as described in the preliminary U.S. application No. 16 / 915,577 filed on June 29, 2020, and / or in the preliminary U.S. application No. 16 / 915,577 filed on June 19, 2020.The preliminary US application No. 16 / 907,125, filed in June 2020, is described and is incorporated herein in full by reference.
[0012] The Eye Tracker 106 can use sensor data 102A—e.g., sensor data from one or more cameras in the cabin, NIR cameras or sensors, and / or other eye-tracking sensor types—to determine gaze directions and movements, fixations, road scanning behavior (e.g., road scanning patterns, distribution, and range), saccade information (e.g., speed, direction, etc.), blink rate, smooth tracking information (e.g., speed, direction, etc.), and / or other information. The Eye Tracker 106 can determine time periods corresponding to specific states, such as the duration of a fixation, and / or track how often certain states are determined—e.g., the number of fixations, saccades, smooth tracking, etc. The Eye Tracker 106 can monitor or analyze each eye individually and / or monitor or analyze both eyes together.For example, both eyes can be monitored to measure the depth of an occupant's gaze using triangulation. In some embodiments, the Eye Tracker 106 can execute one or more machine learning algorithms, deep neural networks, computer visualization algorithms, image processing algorithms, mathematical algorithms, and / or similar algorithms to determine the eye-tracking information. In some non-restrictive embodiments, the Eye Tracker 106 may include similar features, functions and / or components as described in U.S. Preliminary Application No. 16 / 363,648, filed on October 8, 2018; U.S. Preliminary Application No. 16 / 544,442, filed on August 19, 2019; U.S. Preliminary Application No. 17 / 010,205, filed on September 2, 2020; U.S. Preliminary Application No. 16 / 859,741, filed on April 27, 2020; and U.S. Preliminary Application No. 17 / 004,252, filed on April 27, 2020.August 2020, and / or the preliminary US application No. 17 / 005,914, filed on August 28, 2020, each of which is incorporated herein in full by reference.
[0013] The Attention Assessor 108 can be used to determine the attention level of the occupant(s). For example, data from the Body Tracker 104 and / or the Eye Tracker 106 can be processed or analyzed by the Attention Assessor 108 to determine an attention score, a mark, or a level. The Attention Assessor 108 can execute one or more machine learning algorithms, deep neural networks, computer visualization algorithms, image processing algorithms, mathematical algorithms, and / or similar processes to determine attention. For example, with regard to Fig. 2A-2E, includes Fig. 2A a diagram 202 that corresponds to a current (e.g., a current point in time or a period of time - such as one second, three seconds, five seconds, etc.) gaze direction and gaze information. For example, the gaze direction can be represented by points 212, where the (x, y) positions in graph 202 can have corresponding positions with respect to the vehicle 500. Graph 202 can be used to determine the type of eye movement - e.g., the most recent type of eye movement - such as a saccade (see Fig. 2A) or a uniform tracking, fixation, etc. As another example of Fig. 2B, the points 212 from diagram 202 can be transferred to diagram 204, which represents the current and / or most recent (e.g., within the last second, three seconds, etc.) gaze regions of the occupant(s). For example, any number of gaze regions 214 (e.g., gaze regions 214A-214F) can be used to determine behavior while scanning the road, fixations, and / or other information that can be used by the attention assessor 108. Non-restrictive examples of the viewing regions 214 include a left side-viewing region 214A (e.g., corresponding to a window on the driver's side, a driver's side mirror, etc.), a left front viewing region 214B (e.g., corresponding to the left half or part of a windshield), a right front viewing region 214C (e.g., corresponding to the right half or part of a windshield), a right side-viewing region 214D (e.g., corresponding to the, for example,a passenger window, a passenger mirror, etc.), a viewing region 214E for the instrument cluster (e.g., corresponding to an instrument cluster 532 or an instrument panel behind, below, and / or above the steering wheel) and / or a viewing region 214F for the center console (e.g., corresponding to controls, displays, touchscreen interfaces, radio controls, climate control controls, hazard warning light controls, in-vehicle infotainment (IVI), in-car entertainment (ICE), and / or other center console functions).
[0014] With reference to Fig. Diagrams 206 and 208, which include visualizations of an occupant—for example, with a stronger focus on the occupant's eyes in Diagram 206 and a broader focus on the occupant in Diagram 208—can be used to create Diagrams 202, 204, and / or 210. For example, the determination that the occupant is looking toward the right front of vehicle 500—for example, toward the right front gaze region 214C—and that the occupant's last movement was a saccade can be made using one or more instances of charts 206 and 208. The (x, y) position of the occupant's head and / or eyes may have a known correlation with a gaze region 214 and / or with an (x, y) position on a gaze region chart—such as charts 202 and / or 204.In this way, the orientation of the user's head and / or eyes can be determined and used to ascertain the gaze direction and / or position for a given image. Furthermore, the results can be used across any number of images (e.g., two seconds of frames captured at 30 frames per second, or 60 images) to track movement types—such as saccades, blink rate, smooth tracking, fixations, street searching, and / or similar behaviors. In some embodiments, such as in... Fig. As shown in Figure 2E, the occupant's gaze information can be tracked over a specific period and used to create a heat map 210 (e.g., with darker regions corresponding to more frequent gaze positions or directions than regions that are lighter or have less dense dot patterns) that reflects the occupant's gaze positions and directions over time (e.g., over a period of thirty seconds, one minute, three minutes, five minutes, etc.). For example, with reference to heat map 210—whose coordinate system is similar to that of map 202 and / or 204—heat map 210 can show that the occupant more frequently views the right front gaze region 214C and / or the right lateral gaze region 214D than other gaze regions 214. In some embodiments, heat map 210 can be updated with each image to include the information from the current image.For example, weighting can be applied to give more weight to the more recent gaze information when creating the heat map 210. The heat map 210 can then depict the behavior, patterns, and / or frequency of the occupant's scanning of the road, which can be used to determine the occupant's level of attention.
[0015] The attention assessor 108 can then use some or all of the eye-tracking information to determine the occupant's attention. For example, if the heat map shows that the occupant has frequently scanned the road over a wide area and their current gaze is directed at or immediately adjacent to the road surface, the current attention score may be classified as high. Another example: If the heat map shows that the occupant has been focusing more on places off the road—such as sidewalks, buildings, landscapes, etc.—and their current gaze is also directed toward one side of the vehicle 500 (for example, if the actual roadway is only at the periphery of the occupant's field of vision), the attention score may be low, at least for the eye-tracking-based information—for example,The occupant can be identified as dreaming or with a blank stare.
[0016] In some examples, which are described in more detail here, the attention determiner 108 can use the outputs of the comparator 120 to determine attention. For example, if a certain number of objects—e.g., VRUs, traffic signs, waiting conditions, etc.—have been identified using the external perception of the vehicle 500, the number of these objects seen and / or processed by the occupant over a specific period can be determined. For example, if 30 objects were detected and the occupant saw and / or processed 28 of them—e.g., as determined by the comparator 120 using the outputs of the FOV projector 116 and the outputs of the DNN(s) 112—the occupant's attention can be classified as high. However, if the occupant sees and / or processes only 15 of the 30 images, the occupant's attention can be classified as low (at least with respect to the comparator 120's calculations).In some embodiments, determining whether the occupant is processing the objects can be based on other attentional factors and / or the determination of cognitive load. For example, if a projection of the occupant's gaze (e.g., from a calculated gaze vector) or visual field overlaps—at least partially—with the detected objects, it can be determined that the occupant has seen the object. However, if other attentional information indicates that the occupant is actually fixated on a particular gaze or has briefly looked at their phone, or if the occupant is experiencing a high cognitive load, it can be determined that the occupant has not processed, registered, or paid attention to the object. In such cases, it can be determined that the occupant's attention is lower than it would have been if the occupant had processed or registered the object.
[0017] The Attention Determiner 108 can use the Body Tracker information from the Body Tracker 104 to determine attention, either in addition to or as an alternative to the Eye Tracker information. For example, pose, posture, activity, and / or other information about one or more occupants can be used to determine attention. Thus, if a driver has both hands on the steering wheel and maintains an upright posture, the driver's body-based contribution to attention may correspond to a high attention score. Conversely, a driver who has one hand on the steering wheel and is holding their phone in front of their face with the other may receive a low attention score, at least for the body-based contribution.
[0018] The Cognitive Load Assertion 110 can be used to determine the cognitive load of the inmate(s). For example, data from the Body Tracker 104 and / or the Eye Tracker 106 can be processed or analyzed by the Cognitive Load Assertion 110 to determine a cognitive load score, value, or level. The Cognitive Load Assertion 110 can execute one or more machine learning algorithms, deep neural networks, computer vision algorithms, image processing algorithms, mathematical algorithms, and / or similar processes to determine the cognitive load. For example, eye-based metrics such as pupil size (e.g., diameter) or other eye characteristics, eye movements, blink rate, or other eye measurements, and / or other information—e.g., from a computer vision algorithm and / or one or more DNNs—can be used to determine a value, score, or level of cognitive load (e.g.,to calculate (low, medium, high). For example, dilated pupils can indicate high cognitive load, and the Cognitive Load Determination System 110 can use pupil size, among other information, to determine the cognitive load score.
[0019] Behavior while scanning the road can also indicate cognitive load, such as whether the occupant is attentive or daydreaming. For example, if a heat map showing the occupant's pattern or behavior while scanning the road indicates that the occupant has scanned the entire scene from left to right, this information may indicate a lower cognitive load. However, if the occupant is not scanning the road or is fixating (e.g., with a blank stare while daydreaming, etc.) on non-essential parts of the environment—such as sidewalks, scenery, etc.—this may indicate a higher cognitive load. In some embodiments, environmental conditions (e.g., weather, time of day, etc.) and / or driving conditions (e.g., traffic, road conditions, etc.) may also play a role.) when determining cognitive load (and / or the attention determinants), so that snow, sleet, hail, rain, direct sunlight, darkness, heavy traffic, and / or other conditions can influence the determination of cognitive load. For example, the cognitive load for the same set of eye and / or body tracking information may be rated as low in clear weather and light traffic and as high in snowfall and heavy traffic.
[0020] In some embodiments, the cognitive load and / or attention of the occupant(s) can be based on a profile specific to the occupant(s). For example, during a current trip and / or one or more previous trips, eye-tracking information, body-tracking information, and / or other information about the occupant(s) can be monitored and used to determine an individual profile for the occupant(s), indicating when the respective occupant(s) is / are attentive, inattentive, cognitively more or less loaded, etc. For example, a first occupant might scan the road less but process objects and make the correct decisions faster or with a higher frequency, while a second occupant might scan the road less but process objects and make correct decisions less quickly or with a lower frequency.By customizing profiles for the first and second occupants, the first occupant can be prevented from receiving excessive warnings, while the second occupant can be warned more frequently to ensure safe driving if both exhibit similar road scanning behavior. An occupant's profile can also include a higher level of granularity, allowing the tracking of occupant behavior across specific roads, highways, routes, weather conditions, etc. This information can be used to determine the occupant's attention and / or cognitive load on future journeys on the same roads, highways, in the same weather, etc. In addition to determining cognitive load and / or attention, the occupant can customize their profile to include specific notifications or other types of actions or objects for which they wish to receive alerts (e.g.,(For example, a first occupant might not want zebra crossing notifications, while a second occupant might), etc.
[0021] In some embodiments, the sensor data 102B can be applied to one or more deep neural networks (DNNs) 112 that are trained to compute various different outputs 114. Before being applied to or inputted into the DNN(s) 112, the sensor data 102 can be preprocessed, for example, to convert, crop, upscale, reduce, enlarge, rotate, and / or otherwise modify the sensor data 102. For example, if the sensor data 102B corresponds to camera image data, the image data can be cropped, downscaled, upscaled, mirrored, rotated, and / or otherwise adapted to a suitable input format for the respective DNN(s) 112. In some embodiments, the sensor data 102B can include image data representing a picture or images, image data representing a video (e.g.,Snapshots of videos), and / or sensor data representing sensor fields (e.g., depth maps for LiDAR sensors, a value curve for ultrasonic sensors, etc.). In some examples, the sensor data 102B can be used without any preprocessing (e.g., in a raw or captured format), while in other examples, the sensor data 102 can be preprocessed (e.g., noise reduction, demosaicing, scaling, cropping, magnification, white balance, tone curve adjustment, etc., e.g., using a sensor data preprocessor (not shown)).
[0022] Although examples relating to the use of DNNs 112 (and / or the use of DNNs, computer vision algorithms, image processing algorithms, machine learning models, etc., in relation to the Body Tracker 104, the Eye Tracker 106, the Attention Determiner 108, and / or the Cognitive Load Determiner 110) are described here, this is not intended as a limitation. For example, and without limitation, the DNN(s) 112 and / or the computer vision algorithms, image processing algorithms, machine learning models, etc., described herein in relation to the Body Tracker 104, the Eye Tracker 106, the Attention Determiner 108, and / or the Cognitive Load Determiner 110, may include any type of machine learning model or algorithm, such as a machine learning model or...Machine learning models using linear regression, logistic regression, decision trees, support vector machines (SVMs), Naive Bayes, k-nearest neighbor (Knn), K-mean clustering, random forest, dimensionality reduction algorithms, gradient boosting algorithms, neural networks (e.g., autocoders, convolutional algorithms, recurrent algorithms, perceptrons, long / short memory / LSTM, Hopfield, Boltzmann, deep belief, deconvolutional algorithms, generative adversarial algorithms, liquid state machine, etc.), area of interest detection algorithms, computer vision algorithms, and / or other types of algorithms or machine learning models.
[0023] For example, the DNNs 112 can process the sensor data 102 to generate detections of lane markings, road boundaries, signs, poles, trees, static objects, vehicles and / or other dynamic objects, waiting conditions, intersections, distances, depths, object dimensions, etc. The detections can correspond to locations (e.g., in 2D image space, in 3D space, etc.), geometry, position, semantic information, and / or other information about the detection. For example, with lane markings, the positions of the lane markings and / or the types of lane markings (e.g., dashed, solid, yellow, white, zebra crossing, bike lane, etc.) can be detected by one or more DNNs 112 processing the sensor data 102. With regard to signs, the positions of signs or other information about waiting conditions and / or their types (e.g.,Right-of-way, stop signs, pedestrian crossings, traffic lights, priority, construction sites, speed limits, exits, etc.) can be detected using DNN(s) 112. For detected vehicles, motorcyclists, and / or other dynamic actors or road users, the locations and / or types of the dynamic actors can be identified and / or tracked and / or used to determine waiting conditions in a scene (e.g., if a vehicle behaves in a certain way in relation to an intersection, such as coming to a stop, the intersection or the corresponding waiting conditions can be detected as an intersection with a stop sign or traffic light).
[0024] The outputs of DNN(s) 112 can be subjected to post-processing in embodiments, such as converting raw outputs into useful outputs. For example, if a raw output corresponds to a confidence level for each point (e.g., in LiDAR, RADAR, etc.) or each pixel (e.g., in camera images) that the point or pixel represents a particular object type, post-processing can be performed to determine each of the points or pixels that represents a single instance of that object type. This post-processing can include temporal filtering, weighting, outlier removal (e.g., removing pixels or points that turn out to be outliers), upscaling (e.g., the outputs can be predicted at a lower resolution than an input data instance of the sensor, and the output can be upscaled back to the input resolution), reduction, curve fitting, and / or other post-processing techniques.The outputs 114 can – in embodiments after post-processing – be in either a 2D coordinate space (e.g., image space, LiDAR distance image space, etc.) and / or in a 3D coordinate system. In embodiments where the outputs 114 are in a 2D coordinate space and / or in a 3D coordinate space other than 3D space, the outputs 114 can be converted by the FOV projection 116 into the same coordinate system as the projected FOV output (e.g., into a 3D space coordinate system with an origin at a position on the vehicle).
[0025] In some non-restrictive examples, DNN(s) 112 and / or Issues 114 may be similar to those contained in U.S. Non-Preliminary Application No. 16 / 286,329, filed on February 26, 2019; U.S. Non-Preliminary Application No. 16 / 355,328, filed on March 15, 2019; U.S. Non-Preliminary Application No. 16 / 356,439, filed on March 18, 2019; U.S. Non-Preliminary Application No. 16 / 385,921, filed on April 16, 2019; U.S. Non-Preliminary Application No. 535,440, filed on August 8, 2019; U.S. Non-Preliminary Application No. 16 / 728,595, filed on December 27, 2019; and other non-preliminary applications. US application no. 16 / 728,598, filed on December 27, 2019, non-provisional US application no. 16 / 813,306, filed on March 9, 2020, non-provisional US application no. 16 / 848,102, filed on April 14, 2020, non-provisional US application no. 16 / 814,351, filed on March 10, 2020, non-provisional US application no. 16 / 911,007, filed on March 24, 2020.June 2020, and / or non-provisional U.S. application no. 16 / 514,230, filed on July 17, 2019, each of which is incorporated herein in full by reference.
[0026] The FOV projector 116 can use the eye tracker information from the eye tracker 106 to determine the occupant's gaze direction and projected field of view. For example, with regard to Fig. 3, if an occupant enters the front right field of vision 214C ( Fig. 2B) and in particular looks through the lower right part of the windshield 302, a two-dimensional (2D) or three-dimensional (3D) projection 306 of the occupant's field of vision (as shown) can be generated and extended into the environment outside the vehicle 500. Furthermore, the one or more DNNs 112, as described here, can process the sensor data 102B to generate outputs 114 that can be used—e.g., directly and / or after post-processing—to determine the positions (e.g., boundary shapes, boundary contours, boundary fields, points, pixels, etc.) of a person 304 (e.g., a VRU) and a sign 310 in the environment.Although a sign 310 and a person 304 are depicted, this is not to be understood as a limitation, and the objects or information of interest in the environment can include any roads, objects, environmental information, or other information, depending on the embodiment. For example, static objects, dynamic actors, intersections (or corresponding information), waiting conditions, lane lines, road boundaries, sidewalks, barriers, signs, poles, trees, landscapes, information obtained from a map (e.g., from an HD map), environmental conditions, and / or other information can be subjected to an overlap comparison with the comparator 120.
[0027] A non-restrictive example: A first DNN 112 can be used to detect sign 310, and a second DNN 112 can be used to detect person 304. In other examples, one and the same DNN 112 can be used to detect both sign 310 and person 304. The projection 306 (and / or one or more additional projections corresponding to earlier time steps or images) can then be compared by the comparator 120 with the detected objects—e.g., sign 310 and person 304—to determine whether the occupant saw the detected objects. If there is at least a partial overlap, it can be determined that the occupant has an object. In some embodiments, the determination of overlap may include a threshold value for the overlap, e.g., 50% overlap (e.g.,50% of the limiting shape is overlapped by part of the projection 306), 70% overlap, 90% overlap, etc. In other embodiments, any overlap may be sufficient to determine the overlap, or a complete overlap may be sufficient to determine the overlap.
[0028] To determine the overlap, the projection 306 and the outputs 114 can be calculated—or converted—in the same coordinate system. For example, the projection 306 and the outputs 114 can be determined relative to a common coordinate system (e.g., space)—e.g., with an origin at a location on the vehicle 500, such as on an axle (e.g., the center of the vehicle's rear axle), a front bumper, a windshield, etc. If the outputs 114 are calculated in 2D image space, in 3D space relative to a different origin, and / or in a different 2D or 3D coordinate space that is not the common coordinate space used by the comparator 120, the outputs 114 can be converted into the common coordinate space.
[0029] Although the projection 306 is very narrow (e.g., encompassing a focal area of the occupant's field of vision), this is not to be understood as a limitation. In some embodiments, the projection 306 may, for example, in addition to a portion or all of the estimated periphery of the occupant's field of vision, also include a focus of the gaze within the occupant's field of vision. In some embodiments, the determination of the overlap may be weighted based on the portion of the projection 306 that overlaps the object (e.g., the focus or center of a field of vision may be weighted more heavily for determining that the occupant saw the object than the peripheral area of the field of vision).For example, if a portion of projection 306 overlaps, corresponding only to the edge of the occupant's field of vision, the finding may be that the occupant did not see the object, or other criteria may need to be met to determine that the occupant did see the object (e.g., a threshold for time in the periphery must be met, a threshold for frequency - e.g., was the object in the periphery within the last second - etc.).
[0030] In some embodiments where it has been determined that the occupant saw the object(s), the comparator 120 can use one or more additional criteria to make a final decision. For example, duration (e.g., continuous, cumulative over a certain period, etc.) and / or frequency can be considered. In such examples, confirmation that the occupant saw the object can be provided if the occupant saw the object for a specific duration (e.g., a specific number of frames, such as 5 or 10 frames; a specific time interval, such as half a second, one second, two seconds, etc.) (e.g., using an overlap threshold). In some embodiments, the duration can be a cumulative duration over a specific period or a (sliding) time window.A non-restrictive example: If the cumulative duration is one second and the time period is five seconds, if an overlap of a quarter of a second occurs and three seconds later another overlap of three quarters of a second occurs, the cumulative duration may be satisfied within the time period, and it can be determined that the occupant saw the object. In some examples, an additional or alternative time frame determination may be used, requiring that the occupant see the object within a specific time window. For example, the occupant must have the object within three, five, ten seconds, etc., of the current image or time step for the final determination to be made that the occupant saw the object.
[0031] The comparator 120 can, at each time step, output a confidence level (e.g., based on the extent of the overlap, the duration of the overlap, the frequency of the overlap, etc.) that the occupant(s) saw the object(s), an overlap value (e.g., in percent), an overlap duration, an overlap frequency, and / or another value, score, or probability that the occupant(s) saw the object(s). This determination can be made for each object type and / or for each instance thereof. For example, if an action determination 124 includes a notification determination and the system is configured to issue a notification for unseen (and / or unprocessed) lane markings and a notification for unseen (and / or unprocessed) pedestrians, the comparator 120 can output predictions for each lane marking and predictions for each pedestrian (e.g.,at each time step or frame). Thus, the state machine (122) and / or the action determiner (124) can generate outputs and / or make decisions corresponding to lane markings, pedestrians, or a combination thereof. For example, if the action determiner 124 corresponds to one or more ADAS systems, state information corresponding to lane markings may correspond to different actions than state information corresponding to pedestrians. Separate outputs from the comparator 120 corresponding to different object types and / or instances thereof can be useful for the system to determine the occupant's state with respect to each object type or instance, and for the action determiner 124 to determine the correct action(s) to perform (or not perform).
[0032] The state machine 122 can receive as input the outputs of the comparator 120, the attention determiner 108, and / or the cognitive load determiner 110 and determine a state of the occupant(s)—for example, the driver. In some examples, the state machine 122 can determine a single state (e.g., attentive, distracted, focused, etc.) that can be used by the action determiner 124 when deciding which action(s) to perform or refrain from performing (e.g., suppressing a notification). In other examples, the state machine 122 can determine a state with respect to each object type and / or each instance thereof, which can be used by the action determiner 124 to make decisions about one or more different types of actions.In some embodiments, the state machine 122 can determine one or more states based on different levels or processing—e.g., a hierarchical state determination procedure. For example, an inattentive or disinterested state can be determined if an overlap threshold (e.g., overlap quantity, overlap duration, overlap frequency, etc.) has not been reached—e.g., an indication that the occupant(s) has not seen and / or processed the object(s). In such an example, the next stage of the procedure—e.g., processing the results of the attention determiner 108 and / or the cognitive load determination procedure 110—may not be executed. If an occupant has not been determined by the comparator 120 to have seen the object—such as, for example,If a VRU has been seen, the actions that would be performed if it did not see the object can be carried out (e.g., generating and / or issuing a notification or warning, activating an ADAS system such as an automatic emergency braking system (AEB), etc.).
[0033] In some embodiments, even if the overlap threshold(s) has not been reached, the state machine 122 can process the outputs of the attention determiner 108 and / or the cognitive load determiner 110 to determine a combined state—e.g., attentive (e.g., the behavior while scanning the road indicates attention) but inattentive (e.g., the overlap threshold(s) was not reached and the occupant(s) may not have seen the object(s)). In such an example, a different stage or type of action can be performed. In the above example with an AEB system that determines the occupant(s) is / are attentive (and / or has / have low cognitive load) but inattentive, the AEB system cannot be executed, but the notification or warning can be generated and issued to the occupant(s).B. acoustic, tactile, visual, and / or other means. In an example where the overlap threshold(s) has not been reached and attention is low and / or cognitive load is high, further preventive measures can be taken. For example, a notification can be issued, an ADAS system can be activated, and / or—if the vehicle is an autonomous vehicle—autonomous control can be assumed (e.g., to perform a safety maneuver, such as pulling over to the side of the road).
[0034] When an overlap threshold is reached—e.g., when the driver has seen the object—the state machine 122 can use the outputs of the attention determiner 108 and / or the cognitive load determiner 110 to determine a final state. If the occupant(s) is / are attentive—e.g., due to the overlap information from the comparator 120—the attention determiner, the value, level, etc., and / or the value, level, etc., of the cognitive load can be analyzed by the state machine 122 to determine the final state. Even if it is determined that the occupant(s) has / have seen the object(s) (e.g., pedestrian, vehicle, waiting state, intersection, pole, sign, etc.), the state machine 122 can be used to determine the ability (e.g., based on cognitive load) or the probability (e.g.,based on the occupant's attention, the occupant(s) may have processed the perception of the object(s). For example, if a driver sees a vehicle (e.g., based on a comparison of the driver's estimated field of view and the vehicle's position) that is some distance ahead of the ego vehicle 500, but has a high cognitive load and / or low attention (e.g., based on the heat map, fixations, etc.), the state may involve perceiving the vehicle but being inattentive (i.e., possibly not having processed the vehicle's presence). In such examples, the action determiner 124 may execute the actions as if the occupant(s) had not perceived the object(s), and / or trigger a lower-level action—such as issuing a notification or warning rather than executing an ADAS system (e.g., AEB).
[0035] In some embodiments, the state determination(s) of the state machine 122 can take into account the environmental conditions (e.g., weather, time of day, etc.) and / or the driving conditions (e.g., traffic, road condition, etc.) when determining the state. In some embodiments, for example, the determination of the attention determiners, cognitive load, and / or overlap can take the environmental and / or driving conditions into account, and the state machine 122 can determine the state based on this information. In other embodiments, the state machine 122 can additionally or alternatively take the environmental and / or driving conditions into account when determining the state, in addition to or as an alternative to the attention determiners 108, the cognitive load determiners 110, and / or the comparator 120, which take the environmental and / or driving conditions into account.For example, certain values, scores and / or attention levels may indicate an attentive driver when the weather is clear and it is bright outside (e.g., during the day), but an inattentive driver when it is snowing and / or dark outside (e.g., at night).
[0036] The action determiner 124 can include or be part of a warning and notification system, an ADAS system, a control system or control layer of an autonomous driving software stack, a planning system or layer of the autonomous driving stack, a collision avoidance system or layer of the autonomous driving stack, an actuator system or layer of the autonomous driving stack, etc. The action determiner 124 can perform various actions depending on the subject (or object) to which the state corresponds. As non-restrictive examples, the action determiner 124 can determine notification types, ADAS systems to execute or activate, autonomous systems to execute or activate, etc., for dynamic objects on the roadway (e.g., vehicles, pedestrians, cyclists, etc.).For objects off-road, such as poles, pedestrians, and / or similar, the action controller can specify notification types 124, but cannot activate ADAS or autonomous functions unless the vehicle 500 is also off-road or is leaving the road. For waiting conditions (e.g., streetlights, stop signs, etc.), the action controller 124 can specify notification types and / or ADAS systems to be executed or activated—for example, to stop the vehicle 500 at a red light if the light is red and the occupants have determined that they did not see it.
[0037] With reference to Fig. 4. Each block of the Method 400 described herein comprises a computational process that can be performed using any combination of hardware, firmware, and / or software. For example, various functions can be performed by a processor executing instructions stored in memory. The Method 400 can also be embodied as computer-usable instructions stored on computer storage media. The Method 400 can be provided by a standalone application, a service, a hosted service (alone or in combination with another hosted service), or a plug-in for another product, to name just a few. Additionally, Method 400 is illustrated by reference to Process 100 of Fig. 1 described. However, this procedure 400 can additionally or alternatively be performed by any process and / or any system or any combination of processes and / or systems, including but not limited to those described herein.
[0038] Fig. Figure 4 is a flowchart illustrating a method 400 for determining actions based on the driver's attention and / or cognitive load according to some embodiments of the present disclosure. Method 400 includes, in block B402, the determination of gaze information of a vehicle occupant. For example, the eye tracker 106 can determine gaze information (e.g., a gaze vector) and / or other information that can be used to determine the current gaze or field of view of an occupant of the vehicle 500 (e.g., a driver).
[0039] Procedure 400 includes in block B404 the projection of a representation of the gaze information outside the vehicle. For example, based on the gaze information, a projection (e.g., projection 306 in) can be generated. Fig. 3) are generated, whereby the projection can correspond to at least part of the occupant's field of vision.
[0040] Procedure 400, in block B406, involves the calculation of object position information corresponding to one or more objects outside the vehicle, using one or more deep neural networks (DNNs). For example, one or more DNNs can be used to calculate locations and / or semantic information corresponding to one or more objects, subjects, and / or other information in the environment outside the vehicle. Non-restrictive examples include calculating locations and / or semantic information about dynamic actors, static objects, waiting conditions, lane lines, road boundaries, poles, and / or signs using DNNs.
[0041] Procedure 400, in block B408, includes the determination of whether a threshold for overlap between the representation and one or more objects is met. For example, comparator 120 can compare the view's representation (e.g., projection 306) with one or more objects in the environment—e.g., within the same coordinate system—to determine whether an overlap threshold exists. The overlap threshold can also include determining whether the overlap occurred within a time window from the current time, for a cumulative time interval within a time window, and / or for a sufficiently long consecutive time interval.
[0042] If the threshold overlap in block B408 is not met, procedure 400 can proceed to block B410. Block B410 involves the execution of an initial action(s). For example, if the overlap threshold is not reached (and / or other thresholds such as frequency, duration, etc. are not met), which may indicate that the occupant did not see the object, an initial action or actions may be performed. For example, a warning or notification may be generated and issued, one or more ADAS systems may be activated, and / or similar actions may be taken.
[0043] If the threshold overlap in Block B408 is met, Procedure 400 can proceed to Block B412. Block B412 involves determining at least one cognitive load or attention measure of the inmate. For example, if the overlap threshold is met, meaning the inmate saw the object(s), the Cognitive Load Determiner 110 can determine the inmate's cognitive load and / or the Attention Determiner 108 can determine the inmate's attention. This information can then be used in Block B414 to determine whether the inmate was able to process the perception of the object(s).
[0044] Procedure 400 includes, in block B414, the determination of whether cognitive load is above a first threshold and / or attention is below a second threshold. For example, if cognitive load is above the first threshold—meaning the occupant is currently unable to process the perception of the object(s)—Procedure 400 can proceed to block B416. If attention is below a second threshold—meaning the occupant is currently unable to process the perception of the object(s)—Procedure 400 can proceed to block B416. In some embodiments, a combination of attention and cognitive load can be used to determine whether to proceed to block B416. Block B416 involves the execution of a second action(s). In embodiments, the second action(s) can be the same as the first action(s).Therefore, if block B408 detects that the occupant did not see the object and / or detects that the occupant saw the object but did not fully process it—for example, due to high cognitive load and / or low attention—the action(s) performed may be the same. However, in some embodiments, the second action(s) may differ from the first. For example, if it is detected that the occupant saw the object(s), but block B414 detects that the occupant may not have fully processed it, a notification or warning may be issued, but one or more ADAS systems may not be executed (at least initially).Another example: The first action(s) may include an audible and visual warning, and the second action(s) may only include a visual warning – e.g., a less intrusive warning, since the occupant has seen the object(s).
[0045] Another example: If the cognitive load is below the first threshold—meaning, for example, that the occupant is currently better able to process the perception of the object(s)—Procedure 400 can proceed to Block B418. If the attention is above the second threshold—indicating, for example, that the occupant is currently better able to process the perception of the object(s)—Procedure 400 can proceed to Block B418. Block B418 involves the execution of a third action(s). For example, if the cognitive load is low and / or attention is high, and the occupant has determined in Block B408 that they have seen the object(s), the third action(s) can be executed. In some embodiments, the third action(s) may involve no action being performed or an action being suppressed—for example, the suppression of a notification.Therefore, if it is determined that an occupant has seen one or more objects and exhibits low cognitive load and / or high attention, notifications, ADAS systems, and / or other actions can be suppressed to reduce the number of warnings, notifications, and / or autonomous or semi-autonomous vehicle activations. Thus, Procedure 100 and / or Procedure 400 can execute a hierarchical decision tree to determine the actions to be taken based on the occupant's state. EXEMPLARY AUTONOMOUS VEHICLE
[0046] Fig. Figure 5A illustrates an example of an autonomous vehicle 500 according to some embodiments of the present disclosure. The autonomous vehicle 500 (here alternatively referred to as "vehicle 500") can, without limitation, be a passenger vehicle, such as a car, truck, bus, rescue vehicle, shuttle, electric or motorized bicycle, motorcycle, fire engine, police vehicle, ambulance, boat, construction vehicle, underwater vehicle, drone, and / or another type of vehicle (e.g., an unmanned vehicle and / or a vehicle that carries one or more passengers).Autonomous vehicles are generally described in terms of automation levels defined by the National Highway Traffic Safety Administration (NHTSA), a division of the U.S. Department of Transportation, and the Society of Automotive Engineers (SAE) in their "Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles" (Standard No. J3016-201806, published June 15, 2018, Standard No. J3016-201609, published September 30, 2016, and earlier and future versions of this standard). The Vehicle 500 may be capable of functionality corresponding to one or more of Levels 3 through 5 of the autonomous driving levels. For example, depending on its configuration, the Vehicle 500 may be capable of conditional automation (Level 3), high automation (Level 4), and / or full automation (Level 5).
[0047] The vehicle 500 can include components such as a chassis, a vehicle body, wheels (e.g., 2, 4, 6, 8, 18, etc.), tires, axles, and other vehicle components. The vehicle 500 can include a propulsion system 550, such as an internal combustion engine, a hybrid electric drive, a fully electric motor, and / or another type of propulsion system. The propulsion system 550 can be connected to a drivetrain of the vehicle 500, which may include a transmission to enable the propulsion of the vehicle 500. The propulsion system 550 is controlled in response to signals received from the throttle / accelerator pedal (552).
[0048] A steering system 554, which may include a steering wheel, can be used to steer the vehicle 500 (e.g., along a desired path or route) when the drive system 550 is in operation (e.g., when the vehicle is in motion). The steering system 554 can receive signals from a steering actuator 556. For full automation functionality (Level 5), the steering wheel may be optional.
[0049] The brake sensor system 546 can be used to operate the vehicle brakes in response to receiving signals from the brake actuators 548 and / or brake sensors.
[0050] The controller(s) 536, which controls one or more system-on-chips (SoCs) 504 ( Fig. 5C) and / or GPU(s), can provide signals (e.g., in the form of commands) to one or more components and / or systems of the vehicle 500. For example, the controller(s) can send signals to actuate the vehicle brakes via one or more brake actuators 548, to actuate the steering system 554 via one or more steering actuators 556, and to actuate the propulsion system 550 via one or more throttle / acceleration controllers 552. The controller(s) 536 can include one or more onboard (e.g., integrated) computing devices (e.g., supercomputers) that process sensor signals and issue operating commands (e.g., signals representing commands) to enable autonomous driving and / or to assist a human driver in operating the vehicle 500.The controller(s) 536 can include a first controller 536 for autonomous driving functions, a second controller 536 for functional safety functions, a third controller 536 for an artificial intelligence function (e.g., machine vision), a fourth controller 536 for an infotainment function, a fifth controller 536 for emergency redundancy, and / or other controllers. In some examples, a single controller 536 can handle two or more of the above functionalities, two or more controllers 536 can handle a single functionality, and / or any combination thereof.
[0051] The controller(s) 536 can provide the signals to control one or more components and / or systems of the vehicle 500 in response to sensor data received from one or more sensors (e.g. sensor inputs). The sensor data can be received, for example, without restriction, from sensor(s) 558 of global navigation satellite systems (e.g., sensor(s) of the global positioning system), RADAR sensor(s) 560, ultrasonic sensor(s) 562, LiDAR sensor(s) 564, sensor(s) 566 of an inertial measurement unit (IMU) (e.g., accelerometer, gyroscope(s), magnetic compass(s), magnetometer, etc.), microphone(s) 596, stereo camera(s) 568, wide-view camera(s) 570 (e.g., fisheye cameras), infrared camera(s) 572, omnidirectional camera(s) 574 (e.g., 360-degree cameras), long-range cameras and / or medium-range camera(s) 598, speed sensor(s) 544 (e.g., radar). B.for measuring the speed of the vehicle 500), vibration sensor(s) 542, steering sensor(s) 540, brake sensor(s) (e.g. as part of the brake sensor system 546) and / or other types of sensors.
[0052] One or more of the controller(s) 536 can receive inputs (e.g., represented by input data) from an instrument cluster 532 of the vehicle 500 and provide outputs (e.g., represented by output data, display data, etc.) via a display 534, a human-machine interface (HMI), an acoustic alarm, a loudspeaker, and / or via other components of the vehicle 500. The outputs can include information such as vehicle speed, engine speed, time, map data (e.g., the HD map 522 of Fig. 5C), position data (e.g., the position of the vehicle 500, e.g., on a map), direction, position of other vehicles (e.g., an occupancy grid), information about objects and the status of objects as perceived by the controller(s) 536, etc. For example, the HMI display 534 can show information about the presence of one or more objects (e.g., a road sign, a warning sign, a changing traffic light, etc.) and / or information about driving maneuvers that the vehicle has performed, is currently performing, or will perform (e.g., changing lanes now, taking exit 34B in two miles, etc.).
[0053] The vehicle 500 also includes a network interface 524, which can use one or more wireless antenna(s) 526 and / or modem(s) for communication over one or more networks. The network interface 524 can, for example, enable communication via LTE, WCDMA, UMTS, GSM, CDMA2000, etc. The wireless antenna(s) 526 can also enable communication between objects in the environment (e.g., vehicles, mobile devices, etc.) via local area networks such as Bluetooth, Bluetooth LE, Z-Wave, ZigBee, etc., and / or Low Power Wide Area Network(s) (LPWANs) such as LoRaWAN, SigFox, etc.
[0054] Fig. 5B is an example of camera positions and fields of view for the autonomous vehicle 500 from Fig. 5A, according to some embodiments of the present disclosure. The cameras and the corresponding fields of view represent an exemplary embodiment and are not to be considered limiting. For example, additional and / or alternative cameras may be included and / or the cameras may be located at different positions on the vehicle 500.
[0055] The camera types may include, but are not limited to, digital cameras designed for use with components and / or systems of the Vehicle 500. The camera(s) may operate at Automotive Safety Integrity Level (ASIL) B and / or another ASIL. Depending on the configuration, the camera types may be capable of any frame rate, such as 60 frames per second (fps), 120 fps, 240 fps, etc. The cameras may be capable of using roller shutters, global shutters, another type of shutter, or a combination thereof.In some examples, the color filter array may include a red-clear-clear-clear (RCCC) color filter array, a red-clear-clear-blue (RCCB) color filter array, a red-blue-green-clear (RBGC) color filter array, a Foveon X3 color filter array, a Bayer sensor (RGGB) color filter array, a monochrome sensor color filter array, and / or another type of color filter array. In some embodiments, clear pixel cameras, such as cameras with an RCCC, RCCB, and / or RBGC color filter array, may be used in an effort to increase light sensitivity.
[0056] In some embodiments, one or more of the camera(s) can be used to perform functions of advanced driver assistance systems (ADAS) (e.g., as part of a redundant or fail-safe design). For example, a multi-function monocular camera can be installed that provides functions such as lane departure warning, traffic sign recognition, and intelligent headlight control. One or more of the camera(s) (e.g., all cameras) can simultaneously capture and provide image data (e.g., a video).
[0057] One or more cameras can be mounted in an assembly, such as a custom-designed (3D-printed) assembly, to eliminate stray light and reflections from inside the car (e.g., reflections from the dashboard reflected in the windshield mirrors) that could impair the camera's image acquisition capabilities. Referring to side mirror assemblies, the side mirror assemblies can be custom 3D-printed so that the camera mounting plate matches the shape of the side mirror. In some examples, the camera(s) can be integrated into the side mirror. For side-view cameras, the camera(s) can be integrated into the four pillars at each corner of the cab.
[0058] Cameras with a field of view that includes portions of the environment in front of the vehicle 500 (e.g., forward-facing cameras) are used for all-around vision to help identify forward paths and obstacles, and, with the aid of one or more controllers 536 and / or control SoCs, to provide information critical for generating an occupancy grid and / or determining preferred vehicle paths. Forward-facing cameras can be used to perform many of the same ADAS functions as LiDAR, including emergency braking, pedestrian detection, and collision avoidance. Forward-facing cameras can also be used for ADAS functions and systems, including lane departure warning (LDW), autonomous cruise control (ACC), and / or other functions such as traffic sign recognition.
[0059] A variety of cameras can be used in a forward-facing configuration, including, for example, a monocular camera platform incorporating a color image sensor with a CMOS (complementary metal oxide semiconductor). Another example could be a 570 wide-angle camera, which can be used to detect objects entering the field of view from the periphery (e.g., pedestrians, crossing traffic, or bicycles). Although in Fig. While only one wide-angle camera is shown in Figure 5B, any number of wide-angle cameras 570 can be present on the vehicle 500. Additionally, long-range camera(s) 598 (e.g., a pair of long-range stereo cameras) can be used for depth-based object detection, especially for objects for which a neural network has not yet been trained. The long-range camera(s) 598 can also be used for object detection and classification, as well as for basic object tracking.
[0060] One or more Stereo Cameras 568 can also be included in a forward-facing configuration. The Stereo Camera(s) 568 can include an integrated control unit comprising a scalable processing unit that can provide programmable logic (FPGA) and a multi-core microprocessor with an integrated CAN or Ethernet interface on a single chip. Such a unit can create a 3D map of the vehicle's surroundings, which also includes distance estimation for all points in the image. Alternatively, one or more Stereo Camera(s) 568 can include a compact stereo vision sensor, which can include two camera lenses (one on the left and one on the right) and an image processing chip that can measure the distance from the vehicle to the target object and use the generated information (e.g., metadata) to activate autonomous emergency braking and lane departure warning functions.Other types of stereo camera(s) 568 may be used in addition to or as an alternative to those described herein.
[0061] Cameras with a field of view that includes sections of the environment to the sides of the vehicle 500 (e.g., side-view cameras) can be used for all-around vision, providing information that is used to create and update the occupancy grid and to generate side-impact collision warnings. For example, the surrounding camera(s) 574 (e.g., four surrounding cameras 574, as in Fig. 5B) on the vehicle 500. The surround view camera(s) 574 can include wide-angle camera(s) 570, fisheye camera(s), 360-degree camera(s), and / or similar devices. For example, four fisheye cameras can be mounted on the front, rear, and sides of the vehicle. In an alternative configuration, the vehicle can use three surround view camera(s) 574 (e.g., left, right, and rear) and can use one or more other cameras (e.g., a forward-facing camera) as a fourth surround view camera.
[0062] Cameras with a field of view that includes sections of the area behind the vehicle 500 (e.g., reversing cameras) can be used as parking aids, for all-around visibility, rear collision warnings, and for creating and updating the occupancy grid. A variety of cameras can be used, including, but not limited to, cameras that are also suitable as forward-facing cameras (e.g., long-range and / or medium-range cameras 598, stereo cameras 568, infrared cameras 572, etc.), as described herein.
[0063] Fig. 5C is a block diagram of an example system architecture of the exemplary autonomous vehicle 500 by Fig. 5A, according to some embodiments of the present disclosure. It is understood that these and other arrangements described herein are given only as examples. Other arrangements and elements (e.g., machines, interfaces, functions, instructions, groupings of functions, etc.) may be used in addition to or instead of those shown, and some elements may be omitted entirely. Furthermore, many of the elements described herein are functional units that may be implemented as discrete or distributed components, or in conjunction with other components, and in any suitable combination and position. Various functions described herein as being performed by units may be performed by hardware, firmware, and / or software. For example, various functions may be performed by a processor executing instructions stored in memory.
[0064] Each of the components, features, and systems of the 500 vehicle are in Fig. 5C is shown to be connected via bus 502. Bus 502 may include a Controller Area Network (CAN) data interface (here alternatively referred to as the "CAN bus"). A CAN bus can be a network within the vehicle 500 used to support the control of various features and functions of the vehicle 500, such as actuation of the brakes, acceleration, braking, steering, windshield wipers, etc. A CAN bus may be configured to have dozens or even hundreds of nodes, each with its own unique identifier (e.g., a CAN ID). The CAN bus can be read to determine the steering wheel angle, base speed, engine revolutions per minute (RPM), button positions, and / or other vehicle status indicators. The CAN bus may be ASIL B compliant.
[0065] Although the 502 bus is described here as a CAN bus, this should not be interpreted as a limitation. For example, FlexRay and / or Ethernet can be used in addition to or as an alternative to the CAN bus. Although the 502 bus is represented by a single line, this should not be interpreted as a limitation. For example, any number of 502 buses can be present, which may include one or more CAN buses, one or more FlexRay buses, one or more Ethernet buses, and / or one or more other types of buses using a different protocol. In some examples, two or more 502 buses can be used to perform different functions and / or for redundancy. For example, a first 502 bus can be used for collision avoidance functionality, and a second 502 bus can be used for drive control.In each example, each bus 502 can communicate with any component of the vehicle 500, and two or more buses 502 can communicate with the same components. In some examples, each SoC 504, each controller 536, and / or each computer in the vehicle can have access to the same input data (e.g., inputs from sensors of the vehicle 500) and be connected to a common bus, such as the CAN bus.
[0066] The vehicle 500 can include one or more control units 536, such as those mentioned herein in relation to Fig. 5A. The controller(s) 536 can be used for a variety of functions. The controller(s) 536 can be coupled with any of the various other components and systems of the vehicle 500 and can be used to control the vehicle 500, the artificial intelligence of the vehicle 500, the infotainment system for the vehicle 500, and / or the like.
[0067] The vehicle 500 can include a system (or multiple systems) on a single chip (SoC) 504. The SoC 504 can include CPU(s) 506, GPU(s) 508, processor(s) 510, cache(s) 512, accelerator(s) 514, data storage 516, and / or other components and functions not shown. The SoC(s) 504 can be used to control the vehicle 500 in a variety of platforms and systems. For example, the SoC(s) 504 in a system (e.g., the system of the vehicle 500) can be combined with an HD card 522, which is accessed via a network interface 524 by one or more servers (e.g., the server(s) 578). Fig. 5D) can receive map updates.
[0068] The CPU(s) 506 can include a CPU cluster or CPU complex (hereafter referred to as "CCPLEX"). The CPU(s) 506 can include multiple cores and / or L2 caches. For example, in some embodiments, the CPU(s) 506 can include eight cores in a coherent multiprocessor configuration. In some embodiments, the CPU(s) 506 can include four dual-core clusters, each cluster having a dedicated L2 cache (e.g., a 2 MB L2 cache). The CPU(s) 506 (e.g., CCPLEX) can be configured to support simultaneous cluster operation, so that any combination of clusters from the CPU(s) 506 can be active at any given time.
[0069] The CPU(s) 506 can implement power management capabilities that include one or more of the following features: individual hardware blocks can be automatically clocked when idle to save dynamic power; each core clock can be controlled when the core is not actively executing instructions due to the execution of WFI / WFE instructions; each core can be independently power-controlled; each core cluster can be independently clocked when all cores are clocked or power-controlled; and / or each core cluster can be independently power-controlled when all cores are power-controlled. The CPU(s) 506 can further implement an advanced power state management algorithm in which permissible power states and expected wake-up times are specified, and the hardware / microcode determines the best power state to enter for the core, cluster, and CCPLEX.The processing kernels can support simplified performance status entry sequences in the software, with the work being outsourced to the microcode.
[0070] The GPU(s) 508 can include an integrated GPU (referred to herein alternatively as "1ıGPU"). The GPU(s) 508 can be programmable and efficient for parallel workloads. The GPU(s) 508 can use an extended tensor instruction set in some examples. The GPU(s) 508 can include one or more streaming microprocessors, each streaming microprocessor being able to include an L1 cache (e.g., an L1 cache with a minimum memory size of 96 KB), and two or more of the streaming microprocessors being able to share an L2 cache (e.g., an L2 cache with a memory size of 512 KB). In some embodiments, the GPU(s) 508 can include at least eight streaming microprocessors. The GPU(s) 508 can use a computational application programming interface (API). Additionally, the GPU(s) 508 can use one or more parallel computing platforms and / or programming models (e.g., NVIDIA's CUDA).
[0071] The GPU(s) 508 can be performance-optimized for the best computing performance in automotive and embedded applications. For example, the GPU(s) 508 can be manufactured using a FinFET field-effect transistor. However, this is not a limitation, and the GPU(s) 508 can also be manufactured using other semiconductor fabrication methods. Each streaming microprocessor can include a number of mixed-precision processing cores partitioned into multiple blocks. For example, and without limitation, 64 PF32 cores and 32 PF64 cores could be partitioned into four processing blocks. In such an example, each processing block could be allocated 16 FP32 cores, 8 FP64 cores, 16 INT32 cores, two mixed-precision NVIDIA TENSOR COREs for deep learning matrix arithmetic, an L0" instruction cache, a warp scheduler, a distribution unit and / or a 64 KB register file.Additionally, streaming microprocessors can include independent parallel integer and floating-point data paths to enable efficient execution of workloads with a mix of computation and addressing operations. These microprocessors can also include an independent thread scheduling function for finer-grained synchronization and cooperation between parallel threads. Furthermore, streaming microprocessors can incorporate a combined L1 data cache and shared memory unit to improve performance while simplifying programming.
[0072] The GPU(s) 508 can include high-bandwidth memory (HBM) and / or a 16 GB HBM2 memory subsystem to provide a peak memory bandwidth of approximately 900 GB / second in some examples. In some examples, synchronous graphics random-access memory (SGRAM), such as synchronous graphics double-data-rate type five random-access memory (GDDRS), can be used in addition to or as an alternative to HBM memory.
[0073] The GPU(s) 508 can incorporate a unified memory technology with access counters, enabling more accurate allocation of memory pages to the processor that accesses them most frequently, thus improving the efficiency of memory areas shared by multiple processors. In some examples, support for address translation services (ATS) can be used to allow the GPU(s) 508 to directly access the page tables of the CPU(s) 506. In such examples, if the memory management unit (MMU) of the GPU(s) 508 experiences an omission, an address translation request can be passed to the CPU(s) 506. In response, the CPU(s) 506 can search its page tables for a virtual-to-physical mapping for the address and pass the translation back to the GPU(s) 508.Therefore, the unified memory technology can enable a single unified virtual address space for the memory of both the CPU(s) 506 and the GPU(s) 508, thereby simplifying the programming of the GPU(s) 508 and the porting of applications to the GPU(s) 508.
[0074] Additionally, the GPU(s) 508 can include an access counter that tracks the frequency of GPU(s) accessing memory on other processors. This access counter can help ensure that memory pages are moved to the physical memory of the processor that accesses them most frequently.
[0075] The SoC(s) 504 can include any number of Cache(s) 512, including those described herein. The Cache(s) 512 can, for example, include an L3 cache available to both the CPU(s) 506 and the GPU(s) 508 (e.g., connected to both the CPU(s) 506 and the GPU(s) 508). The Cache(s) 512 can include a write-back cache capable of tracking the state of rows, such as by using a cache coherence protocol (e.g., MEI, MESI, MSI, etc.). The L3 cache can be 4 MB or larger, depending on the embodiment, although smaller cache sizes are also possible.
[0076] The SoC(s) 504 may include an arithmetic logic unit (ALU) that can be used to perform processing related to one of the many tasks or operations of the Vehicle 500, such as DNN processing. Furthermore, the SoC(s) 504 may include one or more floating-point units (FPUs) or other mathematical or numeric coprocessors for performing mathematical operations within the system. For example, the SoC(s) 104 may include one or more FPUs integrated as execution units in one or more CPU(s) 506 and / or GPU(s) 508.
[0077] The SoC(s) 504 can include one or more accelerators 514 (e.g., hardware accelerators, software accelerators, or a combination thereof). For example, the SoC(s) 504 can include a hardware acceleration cluster, which may contain optimized hardware accelerators and / or a large amount of on-chip memory. The large on-chip memory (e.g., 4 MB SRAM) can enable the hardware acceleration cluster to accelerate neural networks and other computations. The hardware acceleration cluster can be used to complement the GPU(s) 508 and offload some tasks from the GPU(s) 508 (e.g., to free up more GPU cycles for other tasks). For example, the Accelerator 514 can be used for targeted workloads (e.g., perception, convolutional neural networks (CNNs, etc.)) that are stable enough to be suitable for acceleration.The term “CNN”, as used here, can encompass all types of CNNs, including region-based or regional convolutional neural networks (RCNNs) and Fast RCNNs (e.g., for object detection).
[0078] The Accelerator(s) 514 (e.g., hardware acceleration cluster) can include a deep learning accelerator(s) (DLA(s)). The DLA(s) can include one or more tensor processing units (TPU(s)) that can be configured to provide an additional ten trillion operations per second for deep learning applications and derivation. The TPUs can be accelerators configured and optimized to perform image processing functions (e.g., for CNNs, RCNNs, etc.). The DLA(s) can further be optimized for a specific set of neural network types and floating-point operations, as well as for derivation. The design of the DLA(s) can provide more performance per millimeter than a typical general-purpose GPU and far surpasses the performance of a CPU. The TPU(s) can perform multiple functions, including a single-instance convolution function, which, for example,Supports INT8, INT16 and FP16 data types for both features and weights, as well as post-processor functions.
[0079] The DLA(s) can quickly and efficiently execute neural networks, especially CNNs, on processed or unprocessed data for a variety of functions, including, but not limited to: a CNN for object identification and recognition using camera sensor data; a CNN for distance estimation using camera sensor data; a CNN for emergency vehicle detection, identification, and recognition using microphone data; a CNN for facial recognition and vehicle owner identification using camera sensor data; and / or a CNN for safety-related events.
[0080] The DLA(s) can perform any function of the GPU(s) 508, and by using an inference accelerator, a designer can, for example, target either the DLA(s) or the GPU(s) 508 for any given function. The designer can, for instance, concentrate the processing of CNNs and floating-point operations on the DLA(s) and leave other functions to the GPU(s) 508 and / or other accelerator(s) 514.
[0081] The Accelerator(s) 514 (e.g., the hardware acceleration cluster) may include a programmable vision accelerator (“PVA”), which may alternatively be referred to herein as a machine vision accelerator. The PVA(s) may be designed and configured to accelerate machine vision algorithms for advanced driver-assistance systems (ADAS), autonomous driving, and / or augmented reality (AR) and / or virtual reality (VR) applications. The PVA(s) may provide a balance between performance and flexibility. For example, all PVA(s), without restriction, may include any number of reduced instruction set computer (RISC) cores, direct memory access (DMA), and / or any number of vector processors.
[0082] The RISC cores can interact with image sensors (e.g., the image sensors of any of the cameras described herein), image signal processor(s), and / or the like. Each RISC core can include any amount of memory. Depending on the embodiment, the RISC cores can use any number of protocols. In some examples, the RISC cores can run a real-time operating system (RTOS). The RISC cores can be implemented using one or more integrated circuit devices, application-specific integrated circuits (ASICs), and / or memory devices. For example, the RISC cores can include an instruction cache and / or tightly coupled RAM.
[0083] The DMA allows the components of the PVA(s) to access system memory independently of the CPU(s). The DMA can support any number of features used to optimize the PVA, including, but not limited to, support for multidimensional and / or circular addressing. In some examples, the DMA can support up to six or more dimensions of addressing, which may include block width, block height, block depth, horizontal block gradation, vertical block gradation, and / or depth gradation.
[0084] Vector processors can be programmable processors designed to efficiently and flexibly execute programming for machine vision algorithms and provide signal processing capabilities. In some examples, the PVA may include a PVA core and two vector processing subsystem partitions. The PVA core may include a processor subsystem, DMA engine(s) (e.g., two DMA engines), and / or other peripherals. The vector processing subsystem may act as the primary processing engine of the PVA and may include a vector processing unit (VPU), an instruction cache, and / or vector memory (e.g., VMEM). A VPU core may include a digital signal processor, such as a single-instruction multiple data (SIMD) very-long instruction word (VLIW) signal processor. The combination of SIMD and VLIW can increase throughput and speed.
[0085] Each vector processor can include an instruction cache and be coupled to dedicated memory. Therefore, in some examples, each vector processor can be configured to run independently of the others. In other examples, vector processors contained in a specific PVA can be configured to utilize data parallelism. For example, in some embodiments, the multitude of vector processors contained in a single PVA can execute the same machine vision algorithm, but on different regions of an image. In other examples, the vector processors contained in a specific PVA can simultaneously execute different machine vision algorithms on the same image, or even different algorithms on sequential images or sections of an image.Among other things, the hardware acceleration cluster can contain any number of PVAs, and each PVA can contain any number of vector processors. Additionally, the PVA(s) can include extra error correction code (ECC) memory to increase overall system security.
[0086] The Accelerator 514 (e.g., the hardware acceleration cluster) can include an on-chip machine vision network and SRAM to provide high-bandwidth, low-latency SRAM for the Accelerator 514. In some examples, the on-chip memory can include at least 4 MB of SRAM, consisting, for example, and without limitation, of eight field-configurable memory blocks accessible to both the PVA and the DLA. Each pair of memory blocks can include an advanced peripheral bus (APB) interface, configuration circuitry, a controller, and a multiplexer. Any type of memory can be used. The PVA and DLA can access the memory via a backbone that provides high-speed memory access for both the PVA and DLA.The backbone can include an on-chip network for machine vision that connects the PVA and DLA to the memory (e.g., using the APB).
[0087] The network can include an interface on the machine vision chip that, prior to the transmission of any control signal, address, or data, determines that both the PVA and the DLA provide ready and valid signals. Such an interface can provide separate phases and channels for the transmission of control signals, addresses, and data, as well as burst-type communication for continuous data transmission. This type of interface can conform to ISO 26262 or IEC 61508 standards, although other standards and protocols can also be used.
[0088] In some examples, the SoC(s) 504 can include a real-time ray-tracing hardware accelerator as described in US Patent Application No. 16 / 101,232, filed on August 10, 2018. The real-time ray-tracing hardware accelerator can be used to quickly and efficiently determine the positions and extents of objects (e.g., within a world model) for generating real-time visualization simulations, for radar signal interpretation, for sound propagation synthesis and / or analysis, for simulating SONAR systems, for general wave propagation simulation, for comparison with LiDAR data for localization purposes, and / or for other functions and / or applications. In some embodiments, one or more Tree Traversal Units (TTUs) can be used to perform one or more ray-tracing-related operations.
[0089] The Accelerator 514 (e.g., the hardware accelerator cluster) has a wide range of applications for autonomous driving. The PVA can be a programmable vision accelerator used for key processing stages in ADAS and autonomous vehicles. The PVA's capabilities are well-suited for algorithmic domains requiring predictable processing with low power consumption and low latency. In other words, the PVA demonstrates good performance for semi-dense or dense regular computations, even on small datasets, that require predictable runtimes with low latency and low power consumption. Consequently, PVAs in the context of autonomous vehicle platforms are designed to execute classical machine vision algorithms, as these are efficient at object recognition and operate with integer mathematics.
[0090] For example, according to one embodiment of the technology, the PVA is used to perform machine stereo vision. A semi-global matching-based algorithm can be used, although this should not be interpreted as a limitation. Many Level 3-5 autonomous driving applications require spontaneous motion estimation / stereo matching (e.g., structure from motion, pedestrian detection, lane detection, etc.). The PVA can perform a machine stereo vision function on input from two monocular cameras.
[0091] In some examples, the PVA can be used to perform dense optical flow processing, such as processing raw radar data (e.g., with a 4D Fast Fourier Transform) to produce processed radar data. In other examples, the PVA is used for runtime depth processing, such as processing raw runtime data to produce processed runtime data.
[0092] The DLA can be used to run any type of network to improve control and driving safety, including, for example, a neural network that outputs a measure of confidence for each object detection. Such a confidence score can be interpreted as a probability or as providing a relative "weighting" of each detection compared to other detections. The confidence score allows the system to make further decisions about which detections should be considered true positives and not false positives. For example, the system can set a confidence threshold and only consider detections that exceed the threshold as true positives.In an automatic emergency braking (AEB) system, false positive detections would cause the vehicle to automatically initiate emergency braking, which is obviously undesirable. Therefore, only the most reliable detections should be considered as triggers for AEB. The DLA can run a neural network to regress the confidence value. The neural network can use as its input at least a subset of parameters, such as the dimensions of the boundary frame, the ground plane estimate obtained (e.g., from another subsystem), the output of inertial measurement unit (IMU) sensor 566, which correlates with the vehicle's orientation 500, the distance, the 3D position estimates of the object obtained by the neural network and / or other sensors (e.g., LiDAR sensor(s) 564 or radar sensor(s) 560), etc.
[0093] The SoC(s) 504 may contain data storage 516 (e.g., memory). The data storage 516 may be the on-chip memory of the SoC(s) 504, capable of storing neural networks to be executed on the GPU(s) and / or the DLA. In some examples, the capacity of the data storage 516 may be large enough to store multiple instances of neural networks for redundancy and security. The data storage 512 may include L2 or L3 cache(s) 512. The reference to the data storage 516 may include a reference to the memory allocated to the PVA, DLA, and / or other accelerators 514, as described herein.
[0094] The SoC(s) 504 can contain one or more Processor(s) 510 (e.g., embedded processors). The Processor(s) 510 can include a boot and power management processor, which can be a dedicated processor and subsystem to handle boot power and management functions and associated security enforcement. The boot and power management processor can be part of the SoC(s) 504's boot sequence and provide runtime power management services. The boot and power management processor can provide clock and voltage programming, support for system transitions to a low-power state, management of the SoC(s) 504's thermal and temperature sensors, and / or management of the SoC(s) 504's power state.Each temperature sensor can be implemented as a ring oscillator whose output frequency is proportional to the temperature, and the SoC(s) 504 ring oscillators can be used to detect temperatures from the CPU(s) 506, GPU(s) 508, and / or accelerator(s) 514. If temperatures are determined to exceed a threshold, the booting and power management processor can then enter a temperature fault routine and put the SoC(s) 504 into a lower-power state and / or put the vehicle 500 into a chauffeur-to-safe-stop mode (e.g., bring the vehicle 500 to a safe stop).
[0095] The 510 processor(s) may also include a set of embedded processors that can serve as an audio processing engine. The audio processing engine may be an audio subsystem that provides full hardware support for multi-channel audio across multiple interfaces, as well as a wide and flexible range of audio I / O interfaces. In some examples, the audio processing engine is a dedicated processor core with a digital signal processor and dedicated RAM.
[0096] The 510 processor(s) may also include an always-on processor engine, which can provide the necessary hardware features to support low-power sensor management and wake-up use cases. The always-on processor engine may include a processor core, tightly coupled RAM, supporting peripherals (such as timers and interrupt controllers), various I / O control peripherals, and routing logic.
[0097] The 510 processor(s) can also include a security cluster engine, which comprises a dedicated processor subsystem for handling security management for automotive applications. The security cluster engine can include two or more processor cores, tightly coupled RAM, supporting peripherals (such as timers, interrupt control, etc.), and / or routing logic. In a security mode, the two or more cores can operate in sync and function as a single core with comparison logic to detect any differences between their operations.
[0098] The 510 processor(s) may also include a real-time camera engine, which may include a dedicated processor subsystem for handling real-time camera management.
[0099] The 510 processor(s) may also include a high dynamic range signal processor, which may include an image signal processor, which is a hardware engine that is part of a camera processing pipeline.
[0100] The 510 processor(s) can include a video image compositor, which may be a processing block (e.g., implemented on a microprocessor) that implements video post-processing functions required by a video playback application to generate the final image for the playback program window. The video image compositor can perform lens distortion correction on the 570 wide-view camera(s), the 574 omnidirectional camera(s), and / or the in-cabin surveillance camera sensors. The in-cabin surveillance camera sensor is preferably monitored by a neural network running on a separate instance of the Advanced SoC and configured to detect events in the cabin and respond accordingly.An in-cabin system can perform lip reading to activate mobile service and make a call, dictate emails, change the vehicle's destination, activate or change the vehicle's infotainment system and its settings, or provide voice-activated internet browsing. Certain functions are only available to the driver when the vehicle is operating in autonomous mode and are otherwise disabled.
[0101] The video image compositor can include advanced temporal noise reduction for both spatial and temporal noise reduction. For example, if there is motion in a video, the noise reduction weights the spatial information accordingly by reducing the weight given to information provided by adjacent frames. If an image or a portion of an image does not contain motion, the temporal noise reduction performed by the video image compositor can use information from a previous image to suppress noise in the current image.
[0102] The video image compositor can also be configured to perform stereo equalization on the input stereo lens frames. Furthermore, the video image compositor can be used for user interface composition when the operating system desktop is in use and the GPU(s) 508 are not needed for continuously rendering new surfaces. Even when the GPU(s) 508 are powered on and actively performing 3D rendering, the video image compositor can be used to offload the GPU(s) 508, thus improving performance and responsiveness.
[0103] The SoC(s) 504 may further include a Serial Mobile Industry Processor Interface (MIPI) camera interface for receiving video and camera input, a high-speed interface, and / or a video input block that can be used for camera and associated pixel input functions. The SoC(s) 504 may also include software-controlled input / output control(s) that can be used to receive unassigned I / O signals.
[0104] The SoC(s) 504 can further include a wide range of peripheral interfaces to enable communication with peripheral devices, audio codecs, power management, and / or other devices. The SoC(s) 504 can be used to process data from cameras (e.g., connected via Gigabit Multimedia Serial Link and Ethernet), sensors (e.g., LiDAR sensor(s) 564, RADAR sensor(s) 560, etc., which may be connected via Ethernet), data from the 502 bus (e.g., vehicle speed 500, steering wheel position, etc.), and data from GNSS sensor(s) 558 (e.g., connected via Ethernet or CAN bus). The SoC(s) 504 may also include dedicated high-performance mass storage controllers, which may contain their own DMA engines and which can be used to offload routine data management tasks from the CPU(s) 506.
[0105] The 504 SoC(s) can be an end-to-end platform with a flexible architecture spanning automation levels 3-5, thereby providing a comprehensive functional safety architecture that leverages and efficiently utilizes machine vision and ADAS techniques for diversity and redundancy, and provides a platform for a flexible, reliable driving software stack along with deep learning tools. The 504 SoC(s) can be faster, more reliable, and even more energy- and space-efficient than conventional systems. For example, the 514 accelerator(s), when combined with the 506 CPU(s), 508 GPU(s), and 516 data storage(s), can provide a fast, efficient platform for autonomous vehicles of levels 3-5.
[0106] This technology thus provides capabilities and functions that cannot be achieved with conventional systems. For example, machine vision algorithms can be run on CPUs that can be configured using a high-level programming language, such as C, to execute a wide variety of processing algorithms on a wide variety of visual data. However, these CPUs are often unable to meet the performance requirements of many machine vision applications, such as execution time and power consumption. In particular, many CPUs are unable to execute complex object recognition algorithms in real time, which are required in in-vehicle ADAS applications and in practical Level 3-5 autonomous vehicles.
[0107] Unlike conventional systems, the technology described here, by providing a CPU complex, a GPU complex, and a hardware acceleration cluster, enables the simultaneous and / or sequential execution of multiple neural networks and the combination of their results to enable Level 3-5 autonomous driving functions. For example, a CNN running on the DLA or dGPU (e.g., GPU(s) 520) can include text and word recognition, allowing the supercomputer to read and understand traffic signs, including signs for which the neural network was not specifically trained. The DLA can further include a neural network capable of identifying and interpreting characters, providing a semantic understanding, and passing this semantic understanding to the path planning modules running on the CPU complex.
[0108] As another example, multiple neural networks can be executed simultaneously, as required for driving at levels 3, 4, or 5. For instance, a warning sign reading "Caution: Flashing lights indicate icing," accompanied by an electric light, can be interpreted independently or jointly by several neural networks. The sign itself can be identified as a traffic sign by a first neural network (e.g., a trained neural network), and the text "Flashing lights indicate a diversion" can be interpreted by a second neural network, which informs the vehicle's path planning software (preferably running on the CPU complex) that the presence of icing indicates the presence of flashing lights.The flashing light can be identified by operating a third neural network across multiple frames, which informs the vehicle's path planning software about the presence (or absence) of flashing lights. All three neural networks can run simultaneously, for example, within the DLA and / or on the GPU(s) 508.
[0109] In some examples, a CNN for facial recognition and vehicle owner identification can use data from camera sensors to identify the presence of an authorized driver and / or owner of the vehicle 500. The always-on sensor processing engine can be used to unlock the vehicle when the owner approaches the driver's door and turns on the lights, and, in security mode, to disable the vehicle when the owner leaves. In this way, the SoC(s) 504 provides security against theft and / or carjacking.
[0110] In another example, a CNN for the detection and identification of emergency vehicles can use data from microphones 596 to detect and identify emergency vehicle sirens. Unlike conventional systems that use general classifiers for siren detection and manual feature extraction, the SoC(s) 504 utilize the CNN to classify environmental and urban sounds as well as visual data. In a preferred embodiment, the CNN running on the DLA is trained to identify the relative approach speed of the emergency vehicle (e.g., by using the Doppler effect). The CNN can also be trained to identify emergency vehicles specific to the local area in which the vehicle operates, such as those identified by the GNSS sensor(s) 558.Consequently, for example, when operating in Europe, the CNN attempts to detect European sirens, and when operating in the United States, the CNN attempts to identify only North American sirens. Once an emergency vehicle is detected, a control program can be used to execute an emergency vehicle safety routine to slow the vehicle, pull over to the side of the road, park, and / or let the vehicle idle, using the ultrasonic sensors 562, until the emergency vehicle(s) has passed.
[0111] The vehicle may include one or more CPUs 518 (e.g., discrete CPUs or dCPUs) which may be coupled to the SoC(s) 504 via a high-speed interconnect (e.g., PCIe). The CPU(s) 518 may, for example, include an x86 processor. The CPU(s) 518 may be used to perform any one of a variety of functions, such as mediating potentially inconsistent results between ADAS sensors and the SoC(s) 504 and / or monitoring the status and state of the control unit(s) 536 and / or infotainment SoC 530.
[0112] The Vehicle 500 can include GPU(s) 520 (e.g., discrete GPU(s) or dGPU(s)) that can be coupled to the SoC(s) 504 via a high-speed interconnect (e.g., NVIDIA's NVLINK). The GPU(s) 520 can provide additional functionality for artificial intelligence, such as running redundant and / or different neural networks, and can be used to train and / or update neural networks based on inputs (e.g., sensor data) from the Vehicle 500's sensors.
[0113] The vehicle 500 may also include the network interface 524, which may include one or more wireless antennas 526 (e.g., one or more wireless antennas for different communication protocols, such as a cellular antenna, a Bluetooth antenna, etc.). The network interface 524 can be used to enable a wireless connection via the internet to the cloud (e.g., to a server 578 and / or other network devices), to other vehicles, and / or to computing devices (e.g., passenger client devices). To communicate with other vehicles, a direct link can be established between the two vehicles, and / or an indirect link (e.g., via networks and the internet) can be established. Direct links can be established using a vehicle-to-vehicle communication link.The vehicle-to-vehicle communication link can provide the vehicle 500 with information about vehicles in its vicinity (e.g., vehicles in front of, beside, and / or behind the vehicle 500). The aforementioned functionality can be part of a cooperative adaptive cruise control functionality of the vehicle 500.
[0114] The network interface 524 can include a system-on-a-chip (SoC) that provides modulation and demodulation functionality, enabling the controller(s) 536 to communicate over wireless networks. The network interface 524 can include a high-frequency (RF) front end for upconversion from baseband to RF and downconversion from RF to baseband. The frequency conversions can be performed using well-known processes and / or superposition techniques. In some examples, the RF front-end functionality can be provided by a separate chip. The network interface can include wireless functionality for communication via LTE, WCDMA, UMTS, GSM, CDMA2000, Bluetooth, Bluetooth LE, Wi-Fi, Z-Wave, ZigBee, LoRaWAN, and / or other wireless protocols.
[0115] The vehicle 500 may further include data storage devices 528, which may be located outside the chip (e.g., outside the SoC(s) 504). The data storage device(s) 528 may comprise one or more memory elements, including RAM, SRAM, DRAM, VRAM, flash, hard disks, and / or other components and / or devices capable of storing at least one data bit.
[0116] The Vehicle 500 can also include GNSS sensor(s) 558. The GNSS sensor(s) 558 (e.g., GPS, supported GPS sensors, differential GPS (DGPS) sensors, etc.) assists with mapping, perception, occupancy grid creation, and / or path planning. Any number of GNSS sensor(s) 558 can be used; for example, and without limitation, one GPS using a USB connector with an Ethernet-to-serial (RS-232) bridge.
[0117] The vehicle 500 may also include RADAR sensor(s) 560. The RADAR sensor(s) 560 can be used by the vehicle 500 for long-range vehicle detection, even in darkness and / or adverse weather conditions. The RADAR functional safety level is ASIL B. The RADAR sensor(s) 560 can use the CAN bus and / or the 502 bus (e.g., for transmitting data generated by the RADAR sensor(s) 560) to control and access object tracking data, with some examples providing access to raw data via Ethernet. A wide variety of RADAR sensor types can be used. For example, and without limitation, the RADAR sensor(s) 560 can be suitable for use as front, rear, and side radar. Pulse Doppler RADAR sensors are used in some examples.
[0118] The RADAR sensor(s) 560 can incorporate various configurations, such as long-range with a narrow field of view, short-range with a wide field of view, or with side coverage at short range. In some examples, the long-range RADAR can be used for adaptive cruise control functionality. The RADAR systems can provide a wide field of view at long range, achieved through two or more independent scans, for example, within a range of 250 m. The RADAR sensor(s) 560 can help distinguish between stationary and moving objects and can be used by ADAS systems for emergency braking assistance and forward collision warning. Long-range RADAR sensors can incorporate a monostatic multimodal RADAR with multiple (e.g., six or more) fixed RADAR antennas and a high-speed CAN and FlexRay interface.In an example with six antennas, the four central antennas can generate a focused beam pattern designed to record the area around vehicle 500 at higher speeds with minimal interference from traffic in adjacent lanes. The two other antennas can widen the field of view, making it possible to quickly detect vehicles entering or leaving vehicle 500's lane.
[0119] Medium-range radar systems can, for example, have a range of up to 560 m (front) or 80 m (rear) and a field of view of up to 42 degrees (front) or 550 degrees (rear). Short-range radar systems can, without restriction, include radar sensors designed for installation at both ends of the rear bumper. When the radar sensor system is installed at both ends of the rear bumper, such a system can generate two beams that continuously monitor the blind spots behind and beside the vehicle.
[0120] Short-range radar systems can be used in an ADAS system for blind spot detection and / or lane change assistance.
[0121] The vehicle 500 may also include ultrasonic sensor(s) 562. The ultrasonic sensor(s) 562, which may be positioned at the front, rear, and / or sides of the vehicle 500, may be used as a parking aid and / or for creating and updating an occupancy grid. A wide variety of ultrasonic sensor(s) 562 may be used, and different ultrasonic sensor(s) 562 may be used for different detection ranges (e.g., 2.5 m, 4 m). The ultrasonic sensor(s) 562 may operate at functional safety levels of ASIL B.
[0122] The vehicle 500 can include LiDAR sensor(s) 564. The LiDAR sensor(s) 564 can be used for object and pedestrian detection, emergency braking, collision avoidance, and / or other functions. The LiDAR sensor(s) 564 can meet functional safety level ASIL B. In some examples, the vehicle 500 can include multiple LiDAR sensors 564 (e.g., two, four, six, etc.) that can use Ethernet (e.g., to provide data to a Gigabit Ethernet switch).
[0123] In some examples, the LiDAR sensor(s) 564 may be able to provide a list of objects and their distances for a 360-degree field of view. Commercially available LiDAR sensor(s) 564 may, for example, have an advertised range of approximately 500 m, with an accuracy of 2-3 cm and support for a 500 Mbps Ethernet connection. In some examples, one or more non-protruding LiDAR sensors 564 may be used. In such an example, the LiDAR sensor(s) 564 may be implemented as a small device that can be embedded in the front, rear, sides, and / or corners of the vehicle 500. In such examples, the LiDAR sensor(s) 564 can provide a horizontal field of view of up to 120 degrees and a vertical field of view of up to 35 degrees with a range of 200 m, even with low reflectivity objects.The front-mounted LiDAR sensor(s) 564 can be configured for a horizontal field of view between 45 degrees and 135 degrees.
[0124] In some cases, LiDAR technologies, such as 3D flash LiDAR, can also be used. 3D flash LiDAR uses a laser flash as a source to illuminate the vehicle's surroundings up to a distance of approximately 200 m. A flash LiDAR unit includes a sensor that records the laser pulse travel time and the reflected light at each pixel, which corresponds to the range from the vehicle to the objects. Flash LiDAR can generate highly accurate and distortion-free images of the surroundings with each laser flash. In some cases, four flash LiDAR sensors can be used, one on each side of the vehicle. Available 3D flash LiDAR systems include a solid-state 3D staring array LiDAR camera with no moving parts except for a fan (e.g., a non-scanning LiDAR device).The Blitz LiDAR device can use a 5-nanosecond Class I (eye-safe) laser pulse per image and can capture the reflected laser light in the form of 3D range point clouds and jointly recorded intensity data. Because Blitz LiDAR is a solid-state device with no moving parts, the LiDAR sensor(s) 564 is less susceptible to motion blur, vibration, and / or shock.
[0125] The vehicle may also include IMU sensor(s) 566. The IMU sensor(s) 566 may be located in the center of the rear axle of the vehicle 500 in some examples. The IMU sensor(s) 566 may include, for example, without limitation, an accelerometer, a magnetometer, gyroscope(s), magnetic compass(s), and / or other sensor types. In some examples, such as in six-axis applications, the IMU sensor(s) 566 may include accelerometers and gyroscopes, while in nine-axis applications, the IMU sensor(s) 566 may include accelerometers, gyroscopes, and magnetometers.
[0126] In some embodiments, the IMU sensor(s) 566 can be implemented as a miniaturized, high-performance GPS-aided inertial navigation system (GPS / INS) that combines inertial sensors from micro-electro-mechanical systems (MEMS), a highly sensitive GPS receiver, and advanced Kalman filter algorithms to provide estimates of position, velocity, and orientation. Thus, in some examples, the IMU sensor(s) 566 can enable the vehicle 500 to estimate its course without requiring input from a magnetic sensor by directly observing and correlating changes in velocity from the GPS to the IMU sensor(s) 566. In some examples, the IMU sensor(s) 566 and GNSS sensor(s) 558 can be combined in a single integrated unit.
[0127] The vehicle may include microphone(s) 596, which are mounted in and / or around the vehicle 500. The microphone(s) 596 may be used, among other things, for the detection and identification of emergency vehicles.
[0128] The vehicle may also include any number of camera types, including stereo camera(s) 568, wide-angle camera(s) 570, infrared camera(s) 572, surround-view camera(s) 574, long-range and / or medium-range camera(s) 598, and / or other camera types. The cameras may be used to capture image data around the entire periphery of the vehicle 500. The type of cameras used depends on the embodiment and requirements of the vehicle 500, and any combination of camera types may be used to ensure the necessary coverage around the vehicle 500. Additionally, the number of cameras may vary depending on the embodiment. For example, the vehicle may include six cameras, seven cameras, ten cameras, twelve cameras, and / or another number of cameras. The cameras may, for example, and without limitation, support Gigabit Multimedia Serial Link (GMSL) and / or Gigabit Ethernet.Each of the cameras is discussed here in relation to . Fig. 5A and Fig. 5B.
[0129] The vehicle 500 may also include vibration sensor(s) 542. The vibration sensor(s) 542 can measure vibrations of vehicle components, such as the axle(s). For example, changes in vibration may indicate a change in the road surface. In another example, if two or more vibration sensors 542 are used, the differences between the vibrations can be used to determine the friction or slippage of the road surface (e.g., if the difference in vibration is between a power-driven axle and a freely rotating axle).
[0130] The vehicle 500 may include an ADAS system 538. The ADAS system 538 may include a SoC in some examples. The ADAS system 538 may include autonomous / adaptive / automatic cruise control (ACC), cooperative adaptive cruise control (CACC), forward crash warning (FCW), automatic emergency braking (AEB), lane departure warning (LDW), lane keep assist (LKA), blind spot warning (BSW), rear cross-traffic warning (RCTW), collision warning (CWS), lane centering (LC), and / or other systems, features, and / or functions.
[0131] The ACC systems can use radar sensor(s) 560, LiDAR sensor(s) 564, and / or camera(s). The ACC systems can include longitudinal ACC and / or lateral ACC. Longitudinal ACC controls the distance to the vehicle immediately in front of the vehicle 500 and automatically adjusts the vehicle speed to maintain a safe distance from vehicles ahead. Lateral ACC maintains a safe distance and advises the vehicle 500 to change lanes when necessary. Lateral ACC is integrated with other ADAS applications, such as LCA and CWS.
[0132] A CACC uses information from other vehicles, which can be received via the network interface 524 and / or the wireless antenna(s) 526 from other vehicles either wirelessly or indirectly via a network connection (e.g., the internet). Direct links can be provided through a vehicle-to-vehicle (FF) communication link, while indirect links can be provided through an infrastructure-to-vehicle (IF) communication link. Generally, the FF communication concept provides information about vehicles immediately ahead (e.g., vehicles directly in front of and in the same lane as vehicle 500), while the IF communication concept provides information about traffic further ahead. CACC systems can incorporate either one or both of the IF and FF information sources.Given the information about the vehicles ahead of the vehicle 500, the CACC can be more reliable and has the potential to improve the smoothness of traffic flow and reduce congestion on the road.
[0133] FCW systems are designed to warn a driver of a hazard, allowing the driver to take corrective action. FCW systems utilize a forward-facing camera and / or RADAR sensor(s) coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically connected to driver feedback, such as a display, speaker, and / or vibrating component. The FCW systems can provide a warning, for example, in the form of an audible signal, a visual warning, a vibration, and / or a rapid braking pulse.
[0134] AEB systems detect an impending forward collision with another vehicle or object and can automatically apply the brakes if the driver does not take corrective action within a specified time or distance parameter. AEB systems may use forward-facing camera(s) and / or radar sensor(s) coupled with a dedicated processor, DSP, FPGA, and / or ASIC. When an AEB system detects a hazard, it typically first warns the driver to take a corrective action to avoid a collision. If the driver fails to take corrective action, the AEB system may automatically apply the brakes in an effort to prevent or at least mitigate the impact of the predicted collision. AEB systems may incorporate techniques such as dynamic brake assist and / or anticipatory braking.
[0135] Lane Departure Warning (LDW) systems provide visual, audible, and / or tactile warnings, such as steering wheel or seat vibrations, to alert the driver if the vehicle crosses the lane markings. An LDW system will not activate if the driver indicates an intentional lane departure by using the turn signal. LDW systems can utilize forward-facing cameras coupled with a dedicated processor, DSP, FPGA, and / or ASIC, which are electrically connected to driver feedback, such as a display, speaker, and / or vibrating component.
[0136] LKA systems are a variant of LDW systems. LKA systems provide steering input or braking to correct the vehicle 500 if the vehicle 500 begins to leave its lane.
[0137] Blind Spot Warning (BSW) systems detect and warn the driver of vehicles in a car's blind spot. BSW systems can provide visual, audible, and / or tactile alerts to indicate that merging into or changing lanes is unsafe. The system can issue an additional warning if the driver activates a turn signal. BSW systems can utilize rear-facing camera(s) and / or radar sensor(s) coupled to a dedicated processor, DSP, FPGA, and / or ASIC, electrically coupled to driver feedback such as a display, speaker, and / or vibrating component.
[0138] RCTW systems can provide visual, audible, and / or tactile alerts when an object is detected outside the range of the rear-view camera while the vehicle is reversing. Some RCTW systems include AEB to ensure the vehicle's brakes are applied to avoid a collision. The RCTW systems can use one or more rear-facing radar sensors coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which are electrically coupled to driver feedback, such as a display, speaker, and / or vibrating component.
[0139] Conventional ADAS systems can be prone to false positives, which can be annoying and distracting for the driver, but are typically not catastrophic because the ADAS systems warn the driver and allow the driver to decide whether a safety condition truly exists and to act accordingly. However, in an autonomous vehicle 500, the vehicle 500 itself must decide, in the event of conflicting results, whether to heed the result from a primary computer or a secondary computer (e.g., a first controller 536 or a second controller 536). In some embodiments, for example, the ADAS system 538 can be a backup and / or secondary computer that provides perceptual information to a rationality module of a backup computer.The rationality monitor of the backup computer can run redundant, diverse software on hardware components to detect errors in perception and dynamic driving tasks. The outputs of the ADAS system 538 can be provided to a monitoring MCU. If the outputs of the primary and secondary computers conflict, the higher-level MCU must determine how to resolve the conflict to ensure safe operation.
[0140] In some examples, the primary computer can be configured to provide the monitoring MCU with a confidence score indicating the primary computer's confidence in the chosen outcome. If the confidence score exceeds a certain threshold, the monitoring MCU can follow the primary computer's lead, regardless of whether the secondary computer provides a contradictory or inconsistent result. If one confidence score does not reach the threshold and the primary and secondary computers report different results (e.g., contradiction), the monitoring MCU can mediate between the computers to determine the appropriate outcome.
[0141] The monitoring MCU can be configured to execute neural network(s) trained and configured to determine, at a minimum based on output from the primary and secondary computers, the conditions under which the secondary computer provides false alarms. Consequently, the neural network(s) in the monitoring MCU can learn when the output of a secondary computer can be trusted and when it cannot. For example, if the secondary computer is a radar-based FCW system, neural network(s) in the monitoring MCU can learn when the FCW system identifies metallic objects that are not actually hazards, such as a drain grate or manhole cover, triggering an alarm.If the secondary computer is a camera-based lane departure warning (LDW) system, a neural network in the monitoring MCU can similarly learn to override the LDW when cyclists or pedestrians are present and leaving the lane is indeed the safest maneuver. In embodiments that include neural network(s) running on the monitoring MCU, the monitoring MCU can include at least one DLA or GPU suitable for executing the neural network(s) with associated memory. In preferred embodiments, the monitoring MCU can include and / or be included as a component of one or more of the SoC(s) 504.
[0142] In other examples, the ADAS system 538 may include a secondary computer that performs the ADAS functionality using traditional machine vision rules. Thus, the secondary computer can employ classic machine vision (if-then) rules, and the presence of a neural network(s) in the monitoring MCU can improve reliability, safety, and performance. For example, the diverse implementation and intentional non-identity make the overall system more fault-tolerant, particularly against errors caused by software (or software-hardware interface) functionality.For example, if there is a software bug or error in the software running on the primary computer, and the non-identical software code running on the secondary computer provides the same overall result, the monitoring MCU can have greater confidence that the overall result is correct and that the bug in the software or hardware on the primary computer is not causing a material error.
[0143] In some examples, the output of the ADAS system 538 can be fed into a perception block of the primary computer and / or into a dynamic driving task block of the primary computer. For example, if the ADAS system 538 issues a forward collision warning due to an object immediately ahead, the perception block can use this information in object identification. In other examples, the secondary computer may have its own trained neural network, thus reducing the risk of false positives, as described herein.
[0144] The Vehicle 500 may also include the Infotainment SoC 530 (e.g., an in-vehicle infotainment system - IVI system). Although illustrated and described as a single SoC, the infotainment system may not be a single SoC and may include two or more discrete components. The Infotainment SoC 530 may include a combination of hardware and software that can be used to provide audio (e.g., music, a personal digital assistant, navigation instructions, news, radio, etc.), video (e.g., TV, movies, streaming, etc.), telephone (e.g., hands-free calling), network connectivity (e.g., LTE, Wi-Fi, etc.), and / or information services (e.g., navigation systems, rear parking assistance, a radio data system, vehicle-related information such as fuel level, total distance traveled, brake fluid level, oil level, door open / close status, air filter information, etc.) to the Vehicle 500.The Infotainment SoC 530 can include, for example, radios, record players, navigation systems, video playback devices, USB and Bluetooth connectivity, car computers, in-car entertainment, Wi-Fi, steering wheel audio controls, hands-free voice control, a heads-up display (HUD), an HMI display 534, a telematics device, a control panel (e.g., for controlling and / or interacting with various components, features, and / or systems), and / or other components. The Infotainment SoC 530 can also be used to provide information (e.g., visual and / or audible) to the vehicle's user(s), such as information from the ADAS system 538, autonomous driving information such as planned vehicle maneuvers, trajectories, environmental information (e.g., intersection information, vehicle information, road information, etc.), and / or other information.
[0145] The infotainment SoC 530 may include GPU functions. The infotainment SoC 530 communicates with other devices, systems, and / or components of the vehicle 500 via bus 502 (e.g., CAN bus, Ethernet, etc.). In some examples, the infotainment SoC 530 may be coupled with a monitoring MCU so that the infotainment system's GPU can perform some self-driving functions if the primary controller(s) 536 (e.g., the primary and / or backup computer of the vehicle 500) fails. In such an example, the infotainment SoC 530 can bring the vehicle 500 to a safe stop in a chauffeur-driven mode, as described herein.
[0146] The vehicle 500 may also include an instrument cluster 532 (e.g., a digital instrument panel, an electronic instrument cluster, a digital instrument panel, etc.). The instrument cluster 532 may include a controller and / or a supercomputer (e.g., a discrete controller or a discrete supercomputer). The instrument cluster 532 may include a set of gauges, such as speedometer, fuel level, oil pressure, tachometer, odometer, turn signals, shift lever position indicator, seat belt warning light(s), parking brake warning light(s), engine malfunction light(s), airbag (SRS) system information, lighting controls, safety system controls, navigation information, etc. In some examples, information may be displayed and / or shared by the infotainment SoC 530 and the instrument cluster 532.In other words, the instrument cluster 532 can be included as part of the infotainment SoC 530, or vice versa.
[0147] Fig. 5D is a system diagram for the communication between the cloud-based server(s) and the exemplary autonomous vehicle 500. Fig. 5A, according to some embodiments of the present disclosure. The system 576 may include server(s) 578, network(s) 590, and vehicles, including the vehicle 500. The server(s) 578 may include a plurality of GPUs 584(A)-584(H) (hereinafter collectively referred to as GPUs 584), PCIe switches 582(A)-582(H) (hereinafter collectively referred to as PCIe switches 582), and / or CPUs 580(A)-580(B) (hereinafter collectively referred to as CPUs 580). The GPUs 584, the CPUs 580, and the PCIe switches may be interconnected by high-speed interconnects, such as... B. and without restriction the NVIDIA-developed NVLink interfaces 588 and / or PCIe connections 586. In some examples, the GPUs 584 are connected via an NVLink and / or NVSwitch SoC, and the GPUs 584 and the PCIe switches 582 are connected via PCIe interconnects. Although eight GPUs 584, two CPUs 580, and two PCIe switches are illustrated, this is not intended to be restrictive.Depending on the configuration, each Server 578 can include any number of GPUs 584, CPUs 580, and / or PCIe switches. For example, Server 578(s) can each contain eight, sixteen, thirty-two, and / or more GPUs 584.
[0148] The server(s) 578 can receive image data from the network(s) 590 and the vehicles, representing images that show unexpected or changed road conditions, such as recently started roadworks. The server(s) 578 can transmit neural networks 592, updated neural networks 592, and / or map information 594, including information about traffic and road conditions, to the vehicles via the network(s) 590 and the vehicles. The map information updates 594 can include updates to the HD map 522, such as information about construction sites, potholes, detours, floods, and / or other obstacles.In some examples, the neural networks 592, the updated neural networks 592 and / or the map information 594 may result from new training and / or experience represented in data received from any number of vehicles in the environment, and / or based on training performed in a data center (e.g. using the server(s) 578 and / or other servers).
[0149] Server 578 can be used to train machine learning models (e.g., neural networks) based on training data. The training data can be generated by the vehicles and / or created in a simulation (e.g., using a game engine). In some examples, the training data is tagged (e.g., if the neural network benefits from supervised learning) and / or subjected to other preprocessing, while in other examples, the training data is not tagged and / or preprocessed (e.g., if the neural network does not require supervised learning).Training can be performed using one or more classes of machine learning techniques, including but not limited to classes such as: supervised training, semi-supervised training, unsupervised training, self-learning, reinforcement learning, federated learning, transfer learning, feature learning (including principal component and cluster analysis), multilinear subspace learning, diverse learning, representational learning (including substitute dictionary learning), rule-based machine learning, anomaly detection, and all variants or combinations thereof. Once the machine learning models are trained, they can be used by the vehicles (e.g., transmitted to the vehicles via network(s) 590) and / or the machine learning models can be used by the server(s) 578 to remotely monitor the vehicles.
[0150] In some examples, the Server(s) 578 can receive data from the vehicles and apply that data to current real-time neural networks for intelligent real-time inference. The Server(s) 578 can include deep learning supercomputers and / or dedicated AI computers powered by GPU(s) 584, such as NVIDIA's DGX and DGX Station machines. However, in some examples, the Server(s) 578 can include a deep learning infrastructure that uses only CPU-powered data centers.
[0151] The deep learning infrastructure of the server(s) 578 is capable of fast, real-time inference and can use this capability to assess and verify the state of the processors, software, and / or associated hardware in the vehicle 500. For example, the deep learning infrastructure can receive periodic updates from the vehicle 500, such as a sequence of images and / or objects that the vehicle 500 has located within that sequence of images (e.g., via machine vision and / or other machine learning object classification techniques).The deep learning infrastructure can run its own neural network to identify the objects and compare them with the objects identified by vehicle 500, and if the results do not match and the infrastructure concludes that the AI in vehicle 500 is malfunctioning, then the server(s) 578 can transmit a signal to vehicle 500 instructing a fail-safe computer in vehicle 500 to take control, notify the passengers and perform a safe parking maneuver.
[0152] For derivation, the server(s) 578 can include the GPU(s) 584 and one or more programmable derivation accelerators (e.g., NVIDIA's TensorRT). The combination of GPU-driven servers and derivation acceleration can enable real-time responsiveness. In other examples, where performance is less critical, servers driven by CPUs, FPGAs, and other processors can be used for derivation. EXAMPLE CALCULATION DEVICE
[0153] Fig. Figure 6 is a block diagram of an exemplary computing device(s) 600 suitable for use in implementing some embodiments of the present disclosure. The computing device 600 may include a interconnection system 602 that directly or indirectly couples the following devices: memory 604, one or more central processing units (CPUs) 606, one or more graphics processing units (GPUs) 608, a communication interface 610, input / output (I / O) ports 612, input / output components 614, a power supply 616, one or more presentation components 618 (e.g., display(s)), and one or more logic units 620. In at least one embodiment, the computing device(s) 600 may include one or more virtual machines (VMs) and / or each of its components may consist of virtual components (e.g., virtual hardware components).Non-restrictive examples can include one or more GPUs (608) or one or more vGPUs, one or more CPUs (606) or one or more vCPUs, and / or one or more logic units (620) or one or more virtual logic units. The computing device(s) (600) can include discrete components (e.g., a complete GPU allocated to the computing device (600)), virtual components (e.g., part of a GPU allocated to the computing device (600)), or a combination thereof.
[0154] Even if the various blocks of Fig. Where components 6 are shown connected via the interconnection system 602, this is not intended to be restrictive and serves only for clarity. For example, in some embodiments, a presentation component 618, such as a display device, can be considered an I / O component 614 (e.g., if the display is a touchscreen). As another example, the CPUs 606 and / or GPUs 608 can include memory (e.g., the memory 604 can be representative of a memory device in addition to the memory of the GPUs 608, the CPUs 606, and / or other components). In other words, the computing device consists of Fig. 6 is for illustrative purposes only. No distinction is made between categories such as "workstation", "server", "laptop", "desktop", "tablet", "client device", "mobile device", "handheld device", "game console", "electronic control unit (ECU)," "virtual reality system" and / or other device or system types, as all fall within the scope of the computing device of the Fig. 6 should be considered.
[0155] The interconnection system 602 can represent one or more connections or buses, such as an address bus, a data bus, a control bus, or a combination thereof. The interconnection system 602 can include one or more bus or connection types, such as an Industry Standard Architecture (ISA) bus, an Extended Industry Standard Architecture (EISA) bus, a Video Electronic Standards Association (VESA) bus, a PCI Interconnect (PCI) bus, a PCIe Express Interconnect (PCIe) bus, and / or another type of bus or connection. In some embodiments, there are direct connections between components. For example, the CPU 606 can be directly connected to the memory 604. Furthermore, the CPU 606 can be directly connected to the GPU 608.Where a direct or point-to-point connection exists between components, the interconnection system 602 can include a PCIe connection to implement the connection. In these examples, a PCI bus need not be included in the computing device 600.
[0156] The memory 604 can contain any variety of computer-readable media. The computer-readable media can be any available media to which the computing device 600 can access. The computer-readable media can include both volatile and non-volatile media, as well as removable and non-removable media. By way of example, and without limitation, the computer-readable media can include computer storage media and communication media.
[0157] Computer storage media can include both volatile and non-volatile, and / or removable and non-removable media, implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. For example, memory 604 can store computer-readable instructions (which represent, for example, a program or programs and / or one or more program elements, such as an operating system).Computer storage media may include, but are not limited to, RAM, ROM, EEPROM, flash memory or other storage technology, CD-ROM, Digital Versatile Discs (DVD) or other optical disk storage, magnetic cartridges, magnetic tapes, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and that the Computing Device 600 can access. For the purposes of this text, computer storage media do not include signals per se.
[0158] Computer storage media can embody computer-readable instructions, data structures, program modules, and / or other data types in a modulated data signal, such as a carrier wave or other transport mechanism, and include any information delivery medium. The term "modulated data signal" can refer to a signal that has one or more of its properties modified in such a way that information is encoded in the signal. For example, and not as a limitation, computer storage media can include wired media, such as a wired network or connection, and wireless media, such as acoustic, RF, infrared, and other wireless media. Combinations of any of the above should also be included within the scope of the present computer-readable media.
[0159] The CPU(s) 606 can be configured to execute at least some of the computer-readable instructions, to control one or more components of the computing device 600, and to execute one or more of the procedures and / or processes described herein. The CPU(s) 606 can each include one or more cores (e.g., one, two, four, eight, twenty-eight, seventy-two, etc.) capable of handling a large number of software threads concurrently. The CPU(s) 606 can include any type of processor and may include different types of processors depending on the type of computing device 600 (e.g., processors with fewer cores for mobile devices and processors with more cores for servers).For example, depending on the type of computing device 600, the processor can be an Advanced RISC Machines (ARM) processor implemented using Reduced Instruction Set Computing (RISC), or an x86 processor implemented using Complex Instruction Set Computing (CISC). The computing device 600 can include one or more CPUs 606 in addition to one or more microprocessors or additional coprocessors, such as mathematical coprocessors.
[0160] In addition to or as an alternative to the CPU(s) 606, the GPU(s) 608 may be configured to execute at least some of the computer-readable instructions to control one or more components of the computer device 600 to perform one or more of the methods and / or processes described herein. One or more of the GPU(s) 608 may be an integrated GPU (e.g., with one or more of the CPU(s) 606) and / or one or more of the GPU(s) 608 may be a discrete GPU. In embodiments, one or more of the GPU(s) 608 may be a coprocessor of one or more of the CPU(s) 606. The GPU(s) 608 may be used by the computer device 600 to render graphics (e.g., 3D graphics) or to perform general-purpose computing. For example, the GPU(s) 608 may be used for general-purpose GPU computing (GPGPU).The GPU(s) 608 can contain hundreds or thousands of cores capable of handling hundreds or thousands of software threads simultaneously. The GPU(s) 608 can generate pixel data for output images in response to the rendering of instructions (e.g., rendering instructions from the CPU(s) 606 received via a host interface). The GPU(s) 608 can include graphics memory, such as display memory, to store pixel data or other suitable data, such as GPGPU data. The display memory can be included as part of the memory 604. The GPU(s) 608 can contain two or more GPUs operating in parallel (e.g., via a link). The link can connect the GPUs directly (e.g., using NVLINK) or connect them via a switch (e.g., using NVSwitch).When combined, each GPU can generate 608 pixel data or GPGPU data for different sections of an output or for different outputs (e.g., a first GPU for a first image and a second GPU for a second image). Each GPU can have its own dedicated memory or share memory with other GPUs.
[0161] In addition to or as an alternative to the CPU(s) 606 and / or the GPU(s) 608, the logic unit(s) 620 may be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 600 to execute one or more of the methods and / or processes described herein. In embodiments, the CPU(s) 606, the GPU(s) 608, and / or the logic unit(s) 620 may, individually or jointly, execute any combination of the methods, processes, and / or parts thereof. One or more of the logic units 620 may be part of and / or integrated into one or more of the CPU(s) 606 and / or the GPU(s) 608, and / or one or more of the logic units 620 may be discrete components or otherwise external to the CPU(s) 606 and / or the GPU(s) 608.In embodiments, one or more of the logical units 620 can be a coprocessor of one or more of the CPU(s) 606 and / or one or more of the GPU(s) 608.
[0162] Examples of Logic Unit(s) 620 include one or more processing cores and / or components thereof, such as Tensor Cores (TC), Tensor Processing Units (TPU), Pixel Visual Cores (PVC), Vision Processing Units (VPU), Graphics Processing Clusters (GPC), Texture Processing Clusters (TPC), Streaming Multiprocessors (SM), Tree Traversal Units (TTU), Artificial Intelligence Accelerators (AIA), Deep Learning Accelerators (DLA), Arithmetic Logic Units (ALUs), Application-Specific Integrated Circuits (ASICs), Floating Point Units (FPUs), and Input / Output (I / O) elements.Elements for interconnecting peripheral components (PCI) or express interconnecting of peripheral components (peripheral component interconnect express - PCIe) and / or the like.
[0163] The 610 communication interface can include one or more receivers, transmitters, and / or transceivers that enable the 600 computing device to communicate with other computing devices over an electronic communication network, including wired and / or wireless communication. The 610 communication interface can include components and functionality to enable communication over a number of different networks, such as wireless networks (e.g., Wi-Fi, Z-Wave, Bluetooth, Bluetooth LE, ZigBee, etc.), wired networks (e.g., communication over Ethernet or InfiniBand), low-power wideband networks (e.g., LoRaWAN, SigFox, etc.), and / or the Internet.
[0164] The I / O ports 612 enable the computing device 600 to be logically coupled to other devices, including the I / O components 614, the presentation component(s) 618, and / or other components, some of which may be built into (e.g., integrated into) the computing device 600. Illustrative I / O components 614 include a microphone, mouse, keyboard, joystick, gamepad, game controller, satellite dish, scanner, printer, wireless device, etc. The I / O components 614 can provide a natural user interface (NUI) that processes air gestures, voice, or other physiological input generated by a user. In some cases, input can be transferred to a suitable network element for further processing.A NUI can implement any combination of speech recognition, pen recognition, facial recognition, biometric recognition, gesture recognition (both on-screen and off-screen), air gestures, head and eye tracking, and touch recognition (as further described below) associated with a display of the Computing Device 600. The Computing Device 600 can include depth cameras, such as stereoscopic camera systems, infrared camera systems, RGB camera systems, touchscreen technology, and combinations thereof, for gesture detection and recognition. Additionally, the Computing Device 600 can include accelerometers or gyroscopes (e.g., as part of an inertial measurement unit - IMU) to enable motion detection. In some examples, the output from the accelerometers or gyroscopes can be used by the Computing Device 600 to render immersive augmented reality or virtual reality.
[0165] The power supply 616 can also include a hardwired power supply, a battery power supply, or a combination thereof. The power supply 616 can provide power to the computing device 600 to enable the components of the computing device 600 to operate.
[0166] The presentation component(s) 618 can include a display (e.g., a monitor, a touchscreen, a television screen, a heads-up display (HUD), other display types, or a combination thereof), speakers, and / or other presentation components. The presentation component(s) 618 can receive data from other components (e.g., the GPU(s) 608, the CPU(s) 606, etc.) and output the data (e.g., as an image, video, sound, etc.). EXAMPLE DATA CENTER
[0167] Fig. Figure 7 shows an example of a data center 700 that can be used in at least one embodiment of the present disclosure. The data center 700 can include a data center infrastructure layer 710, a framework layer 720, a software layer 730, and / or an application layer 740.
[0168] As in Fig. As shown in Figure 7, the data center infrastructure layer 710 can include a resource orchestrator 712, clustered compute resources 714, and node compute resources (“node CRs”) 716(1)-716(N), where “N” is any positive integer. In at least one embodiment, the node CRs 716(1)-716(N) can include, but are not limited to, any number of central processing units (“CPUs”) or other processors (including accelerators, field-programmable gate arrays (FPGAs), graphics processing units or graphics processing units (GPUs), etc.), memory devices (e.g., dynamic read-only memory), data storage devices (e.g., solid-state or hard disk drives), network input / output (“NW I / O”) devices, network switches, virtual machines (“VMs”), power modules, and cooling modules, etc. In some embodiments, one or more node CRs can be located among the node CRs.s 716(1)-716(N) correspond to a server that has one or more of the aforementioned computing resources. Furthermore, in some embodiments, the nodes CRs 716(1)-7161(N) may include one or more virtual components, such as vGPUs, vCPUs, and / or the like, and / or one or more of the nodes CRs 716(1)-716(N) may correspond to a virtual machine (VM).
[0169] In at least one embodiment, grouped compute resources 714 can include separate groupings of node CR 716 located in one or more racks (not shown), or in many racks located in data centers at different geographic locations (also not shown). Separate groupings of node CR 716 within grouped compute resources 714 can include grouped compute, network, memory, or data storage resources that can be configured or allocated to support one or more compute loads. In at least one embodiment, multiple node CR 716, including CPUs, GPUs, and / or processors, can be grouped in one or more racks to provide compute resources to support one or more compute loads.One or more racks can also contain any number of power modules, cooling modules, and network switches in any combination.
[0170] The Resource Orchestrator 722 can configure or otherwise control one or more node CR 716(1)-716(N) and / or grouped compute resources 714. In at least one embodiment, the Resource Orchestrator 722 can include a Software Design Infrastructure (“SDI”) management unit for the data center 700. The Resource Orchestrator 722 can consist of hardware, software, or a combination thereof.
[0171] In at least one embodiment, as in Fig. As shown in Figure 7, the framework layer 720 can include a task scheduler 732, a configuration manager 734, a resource manager 736, and a distributed file system 738. The framework layer 720 can include a framework to support software 732 of software layer 730 and / or one or more applications 742 of application layer 740. The software 732 and / or the application 742 can include web-based service software or applications, such as those provided by Amazon Web Services, Google Cloud, and Microsoft Azure. The framework layer 720 can be a type of free and open-source software web application framework, such as Apache Spark™ (hereinafter referred to as "Spark"), which can use a distributed file system 738 for handling large amounts of data (e.g., "Big Data"), but is not limited to it.In at least one embodiment, the task scheduler 732 can include a Spark driver to facilitate the scheduling of workloads supported by different layers of the data center 700. The configuration manager 734 can be able to configure different layers, such as the software layer 730 and the framework layer 720, including Spark and the distributed file system 738 to support the processing of large amounts of data. The resource manager 736 can be able to manage clustered or grouped compute resources allocated or assigned to support the distributed file system 738 and the task scheduler 732. In at least one embodiment, clustered or grouped compute resources can include the grouped compute resource 714 on the data center infrastructure layer 710.The resource manager 1036 and the resource orchestrator 712 can coordinate with each other to manage these allocated or assigned computing resources.
[0172] In at least one embodiment, the software 730 included in software layer 732 may include software used by at least parts of the node CR 716(1)-716(N), grouped computing resources 714, and / or the distributed file system 738 of framework layer 720. One or more types of software may include, but are not limited to, internet web search software, email virus scanning software, database software, and streaming video content software.
[0173] In at least one embodiment, the application(s) 742 included in the application layer 740 may include one or more types of applications used by at least parts of the node CR 716(1)-716(N), grouped compute resources 714, and / or the distributed file system 738 of the framework layer 720. One or more types of applications may include any number of a genomics application, a cognitive computation application, and a machine learning application, including training or derivation software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), and / or other machine learning applications, without limitation, used in conjunction with one or more embodiments.
[0174] In at least one embodiment, any configuration manager 734, resource manager 736, and resource orchestrator 712 can implement any number and type of self-modifying actions based on any amount and type of data acquired in any technically feasible way. Self-modifying actions can relieve a data center operator of the data center 700 of potentially making poor configuration decisions and potentially avoiding underutilized and / or poorly functioning sections of a data center.
[0175] The Data Center 700 may include tools, services, software, or other resources to train one or more machine learning models or to predict or derive information using one or more machine learning models according to one or more embodiments described in this document. For example, a machine learning model(s) may be trained by calculating weighting parameters according to a neural network architecture using software and / or computing resources described above in relation to the Data Center 700.In at least one embodiment, trained or deployed machine learning models corresponding to one or more neural networks can be used to infer or predict information using the resources described above in relation to the Computing Center 700, by using weighting parameters computed by one or more training techniques described herein, but are not limited to this.
[0176] In at least one embodiment, the data center can use 700 CPUs, application-specific integrated circuits (ASICs), GPUs, FPGAs, and / or other hardware (or equivalent virtual computing resources) to perform training and / or inference with the resources described above. Furthermore, one or more of the software and / or hardware resources described above can be configured as a service that allows the user to train or infer information, such as image recognition, speech recognition, or other artificial intelligence services. EXEMPLARY NETWORK ENVIRONMENTS
[0177] Network environments suitable for use in implementing embodiments of the disclosure may include one or more client devices, servers, network-attached storage (NAS), other backend devices, and / or other device types. The client devices, servers, and / or other device types (e.g., each device) may run on one or more instances of the computing device(s) 600 of Fig. 6. For example, each device can include similar components, features, and / or functions of the computing device(s) 600. If backend devices (e.g., servers, NAS, etc.) are implemented, the backend devices can also be part of a data center 700, an example of which is given in Fig. 7 is described in more detail.
[0178] Components in a network environment can communicate with each other over one or more networks, which can be wired, wireless, or both. The network can include multiple networks or a network of networks. For example, the network can include one or more wide area networks (WANs), one or more local area networks (LANs), one or more public networks such as the internet and / or a public switched telephone network (PSTN), and / or one or more private networks. If the network includes a wireless telecommunications network, components such as a base station, a communications tower, or even access points (as well as other components) can provide wireless connectivity.
[0179] Compatible network environments can include one or more peer-to-peer network environments—in which case a server cannot be included in a network environment—and one or more client-server network environments—in which case one or more servers can be included in a network environment. In peer-to-peer network environments, the functionality described herein with respect to one or more servers can be implemented on any number of client devices.
[0180] In at least one embodiment, a network environment can include one or more cloud-based network environments, a distributed computing environment, a combination thereof, etc. A cloud-based network environment can include a framework layer, a job scheduler, a resource manager, and a distributed file system implemented on one or more servers, which may include one or more core network servers and / or edge servers. A framework layer can include a framework to support software of a software layer and / or one or more applications of an application layer. The software or application(s) can each include web-based service software or applications. In embodiments, one or more of the client devices can use the web-based service software or applications (e.g.,by accessing the service software and / or applications via one or more application programming interfaces (APIs). The framework layer can be a type of free and open-source software web application framework that may, for example, use a distributed file system for processing large amounts of data (e.g., "Big Data"), but is not limited to this.
[0181] A cloud-based network environment can provide cloud computing and / or cloud storage, performing any combination (or one or more parts thereof) of the computing and / or data storage functions described herein. Each of these different functions can be distributed across multiple locations of central or core servers (e.g., one or more data centers, which may be distributed across a state, region, country, the globe, etc.). If a connection to a user (e.g., a client device) is relatively close to one or more edge servers, one or more core servers can delegate at least some of the functionality to the edge server(s). A cloud-based network environment can be private (e.g., restricted to a single organization), public (e.g., available to many organizations), and / or a combination of both (e.g., a hybrid cloud environment).
[0182] The client device(s) may include at least one of the components, features, and functionality described herein with respect to Fig.6 described computing device(s) include 600. By way of example, and not as a limitation, a client device may be a personal computer (PC), laptop computer, mobile device, smartphone, tablet computer, smartwatch, portable computer, personal digital assistant (PDA), MP3 player, virtual reality headset, global positioning system (GPS), video player, video camera, surveillance device or system, vehicle, boat, aircraft, virtual machine, drone, robot, portable communication device, device in a hospital, gaming device or system, entertainment system, vehicle computer system, embedded system controller, remote control, household appliance, consumer electronics device, workstation, edge device, any combination of these outlined devices, or any other suitable device.
[0183] The disclosure can be described in the general context of computer code or machine-usable instructions, including computer-executable instructions such as program modules, which are executed by a computer or other machine such as a personal data assistant or other handheld device. In general, program modules, including routines, programs, objects, components, data structures, etc., refer to code that performs specific tasks or implements specific abstract data types. The disclosure can be practiced in a variety of system configurations, including handheld devices, consumer electronics, general-purpose computers, more specialized computing devices, etc. The disclosure can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network.
[0184] As used herein, a recitation of "and / or" in relation to two or more elements should be interpreted as meaning only one element or a combination of elements. For example, "element A, element B and / or element C" may include only element A, only element B, only element C, element A and element B, element A and element C, element B and element C, or elements A, B, and C. Furthermore, "at least one of element A or element B" may include at least one of element A, at least one of element B, or at least one of element A and at least one of element B. Additionally, "at least one of element A and element B" may include at least one of element A, at least one of element B, or at least one of element A and at least one of element B.
[0185] The subject matter of this disclosure is specifically described herein to satisfy legal requirements. However, the description itself is not intended to limit the scope of this disclosure. Rather, the inventors have considered that the claimed subject matter could also be embodied in other ways to include other steps or combinations of steps similar to those described in this document, in conjunction with other present or future technologies. Although the terms "step" and / or "block" may be used herein to denote various elements of the processes employed, these terms should not be interpreted as implying a particular sequence between or between different steps disclosed herein, unless the sequence of each step is explicitly described.
Claims
[1] Procedure, encompassing: Determining an occupant's gaze direction, at least partially, based on initial sensor data (102A) generated by one or more initial sensors of a vehicle; Generating a representation of the viewing direction in a space coordinate system; Determine, in the world coordinate system and at least partially based on second sensor data (102B) generated by one or more second sensors of the vehicle, an object position of at least one object; Comparing the representation of the viewing direction with the object position; and Performing one or more operations that are at least partially based on comparison, Determining a cognitive load score of the occupant, at least partially, based on eye movements, eye measurements, or eye characteristics of the occupant of the vehicle, wherein the performance of one or more operations is based at least partially on the cognitive load score. [2] The method of claim 1, wherein performing the one or more operations comprises performing one of the following operations: Suppress a notification if the rendering overlaps the object position by more than a threshold; or Generate a notification if the representation does not overlap the object position by more than a threshold value. [3] Method according to claim 1 or 2, wherein the one or more first sensors include at least one first sensor with a field of view of the occupant inside the vehicle and the one or more second sensors include at least one second sensor with a field of view outside the vehicle. [4] Method according to any one of the preceding claims, further comprising: Monitoring the eye movements of the vehicle occupant to determine one or more of the following characteristics: gaze patterns, saccadic velocities, fixations, or steady tracking; Determining an inmate's attention score at least partially based on one or more of the gaze patterns, saccade velocities, fixation behaviors, or uniform tracking behaviors, wherein the performance of one or more operations is based at least partially on the attention score. [5] Method according to any one of the preceding claims, further comprising: Generating a heat map that corresponds to the road-searching behavior of the vehicle occupant over a certain period of time, whereby the performance of one or more operations is at least partially based on the heat map. [6] Method according to any of the preceding claims, wherein the determination of the cognitive load value is based at least partially on a cognitive load profile corresponding to the occupant, wherein the cognitive load profile was generated during one or more journeys with the occupant. [7] Method according to any of the preceding claims, wherein determining the object position comprises applying the second sensor data to one or more deep neural networks (DNNs) configured to compute data indicating the object position. [8] Method according to any one of the preceding claims, further comprising: Determining one or more postures, gestures or activities performed by the inmate, the performance of the one or more operations being based at least partially on the posture, gesture or activity. [9] System, comprehensive: one or more sensors; one or more processors; and one or more devices on which instructions are stored which, when executed by the one or more processors, cause the one or more processors to perform operations that include the following: Determining an occupant's gaze direction, at least partially, based on initial sensor data (102A) generated by one or more initial sensors of a vehicle; Generating a representation of the viewing direction in a space coordinate system; Determine, in the world coordinate system and at least partially based on second sensor data (102B) generated by one or more second sensors of the vehicle, an object position of at least one object; Comparing the representation of the viewing direction with the object position; and Performing one or more operations that are at least partially based on comparison, Determining a cognitive load score of the occupant, at least partially, based on eye movements, eye measurements, or eye characteristics of the occupant of the vehicle, wherein the performance of one or more operations is based at least partially on the cognitive load score. [10] System according to claim 9, wherein the system comprises at least one of the following elements: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing simulation operations; a system for performing deep learning operations; a system that is implemented with an edge device; a system that contains one or more virtual machines (VMs); a system that is implemented with a robot; a system that is at least partially implemented in a data center; or a system that was implemented at least partially using cloud computing resources.
Citation Information
Patent Citations
Apparatus and method for detecting a driver's interest in an advertisement by tracking the driver's gaze direction
DE102014109079A1
Methods and systems for processing attention data from a vehicle
DE102015101239A1
Methods for measuring driver attention to objects
DE102015101358A1
EYE-TRACKING OF A VEHICLE OCCUPANT
DE102019122267A1
Method for determining a vehicle occupant's place of interest
DE102020113712A1