Occupant Attention and Cognitive Load Monitoring for Autonomous and Semi-Autonomous Driving Applications

By integrating occupant attention and cognitive load monitoring with external vehicle perception, the system provides a more accurate and reliable assessment of driver state, reducing excessive alerts and enhancing safety in autonomous driving.

JP7742240B2Active Publication Date: 2025-09-19NVIDIA CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2021082051
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-10-21
Filing Date
2021-05-14
Publication Date
2025-09-19
Estimated Expiration
2041-05-14

AI Technical Summary

Technical Problem

Conventional driver monitoring systems inaccurately determine driver attention and cognitive load independently of external environmental conditions, leading to excessive and unreliable warnings or alerts, which can be disabled by drivers, rendering safety systems ineffective.

Method used

Systems and methods that integrate occupant attention and cognitive load monitoring by comparing the driver's estimated field of view with vehicle perception information to determine if the driver has processed external objects or conditions, thereby refining notifications and system actions based on a more holistic and objective assessment.

Benefits of technology

Enhances the accuracy and reliability of driver state determination by considering both internal and external factors, reducing unnecessary alerts and ensuring informed decision-making in autonomous or semi-autonomous driving scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007742240000001
    Figure 0007742240000001
  • Figure 0007742240000002
    Figure 0007742240000002
  • Figure 0007742240000003
    Figure 0007742240000003
Patent Text Reader

Abstract

To provide crew member attention and cognitive load monitoring for autonomous and semi-autonomous driving applications.SOLUTION: In various examples, estimated view or gaze information of a user can be projected onto an outside of a vehicle and compared with vehicle perception information corresponding to an environment outside the vehicle. As a result, internal monitoring of a driver or crew member of the vehicle may be used to determine if the driver or crew member has processed or viewed a certain object type, an environmental condition, or other information outside the vehicle. For a more holistic understanding of a condition of the user, an attention and / or cognitive load of the user may be monitored to determine if one or more actions should be taken. As a result, a notification, an AEB system activation, and / or other actions may be determined based on the user's more complete state as being determined based on a comparison of cognitive load, attention, and / or the comparison of the vehicle's external perception with the user's estimated perception.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This application is a continuation of U.S. patent application Ser. No. 16 / 286,329 filed February 26, 2019, U.S. patent application Ser. No. 16 / 355,328 filed March 15, 2019, U.S. patent application Ser. No. 16 / 356,439 filed March 18, 2019, U.S. patent application Ser. No. 16 / 385,921 filed April 16, 2019, U.S. patent application Ser. No. 16 / 514,230 filed July 17, 2019, and U.S. patent application Ser. No. 16 / 514,230 filed August 17, 2019. U.S. Patent Application No. 535,440 filed on December 8, 2019, U.S. Patent Application No. 16 / 728,595 filed on December 27, 2019, U.S. Patent Application No. 16 / 728,598 filed on December 27, 2019, U.S. Patent Application No. 16 / 813,306 filed on March 9, 2020, U.S. Patent Application No. 16 / 814,351 filed on March 10, 2020, U.S. Patent Application No. 16 / 814,351 filed on April 14, 2020, No. 16 / 848,102, U.S. Patent Application No. 16 / 911,007 filed June 24, 2020, and U.S. Patent Application No. 16 / 915,577 filed June 29, 2020, U.S. Patent Application No. 16 / 363,648 filed October 8, 2018, U.S. Patent Application No. 16 / 544,442 filed August 19, 2019, and U.S. Patent Application No. 17 / 010,2020, filed September 2, 2020. No. 16 / 859,741, filed April 27, 2020; U.S. Patent Application No. 17 / 004,252, filed August 27, 2020; and / or U.S. Patent Application No. 17 / 005,914, filed August 28, 2020; and U.S. Patent Application No. 16 / 907,125, filed June 19, 2020, each of which is incorporated herein by reference in its entirety. [Background technology]

[0002] Cognitive and visual attention play an important role in a driver's ability to detect safety-critical events and, consequently, to safely control the vehicle. For example, if a driver is distracted—e.g., by daydreaming, inattention, using mental faculties for tasks other than the current driving task, observing off-road activity, etc.—the driver may be unable to make appropriate planning and control decisions or may delay making correct decisions. To address this, some vehicle systems—e.g., advanced driver assistance systems (ADAS)—are used to generate audible, visual, and / or tactile warnings or alerts for the driver to inform them of environmental and / or road conditions (e.g., vulnerable road users (VRUs), traffic lights, traffic congestion, potential collisions, etc.). However, as the number of warning systems increases (e.g., automatic emergency braking (AEB), blind spot detection (BSD), forward collision warning (FCW), etc.), the number of alerts or warnings generated can become overwhelming to the driver. As a result, the driver may disable the safety system functionality of one or more of these systems, thereby rendering these systems ineffective in mitigating the occurrence of a dangerous event.

[0003] In some conventional systems, a driver monitoring system and / or a driver drowsiness detection system may be used to determine a driver's current state. However, these conventional systems measure cognitive load or driver attention separately from each other. For example, once a driver's state is determined, the driver may receive a warning or alert regarding their determined inattention or increased cognitive load. This not only results in a warning or alert in addition to the already existing warning alerts of other ADAS systems, but the determination of attention or high cognitive load may be inaccurate or imprecise. For example, with regard to attention, a user's gaze may be measured within a vehicle—e.g., where the user is looking within the vehicle cabin. However, because the state of the external environment—e.g., the position of static or dynamic objects, road conditions, waiting states, etc.—is not taken into account, a driver who is determined to be attentive because their gaze is directed toward the windshield may actually be inattentive. With regard to cognitive load, conventional systems use deep neural networks (DNNs) trained with simulated data that link pupil size or other eye characteristics, eye movements, blink rate or other eye measurements, and / or other information to current cognitive load. However, these measurements are subjective and can be inaccurate or imprecise for certain users—for example, cognitive load may appear differently for different drivers and / or some drivers may perform better than others when under high cognitive load. As such, using driver inattention or cognitive load independently of each other and ignoring environmental conditions outside the vehicle can result in excessive warnings and alerts being generated based on inaccurate, unreliable, or fragmented information about the driver's current state. Summary of the Invention [Means for solving the problem]

[0004] Embodiments of the present disclosure relate to occupant attention and cognitive load monitoring for semi-autonomous or autonomous driving applications. In addition to monitoring attention and / or cognitive load, systems and methods are disclosed that compare a user's estimated field of view or gaze information with vehicle perception information corresponding to the environment outside the vehicle. As a result, the vehicle driver's or occupant's internal monitoring can be extended to the outside of the vehicle to determine whether the driver or occupant has processed or seen certain object types, environmental conditions, or other information outside the vehicle—e.g., dynamic actors, static objects, vulnerable road users (VRUs), queue information, signs, potholes, bumps, debris, etc. If a projected representation—e.g., in a world space coordinate system—of the user's field of view or gaze is determined to overlap with a detected object, state, and / or the like, the system can infer that the user has seen the object, state, etc. and can refrain from performing an action (e.g., generating a notification, activating an AEB system, etc.).

[0005] In some examples, for a more holistic understanding of a user's state—e.g., corresponding to the user's current ability to process seen or visualized information—the user's attention and / or cognitive load may be monitored to determine whether one or more actions should be taken (e.g., to generate a visual, audible, tactile, or other type of notification, to assume control of the vehicle, to activate one or more AEB systems, etc.). As such, even if an object, road condition, etc. may be determined to be in the user's field of view, if the user's determined attention and / or cognitive load indicates that the user may not have fully processed the information, causing an informed driving decision to be no action, action may be taken in response to the object, road condition, and / or the like. As a result, and in contrast to conventional systems, notification, AEB system activation, and / or other action may be determined based on cognitive load, attention, and / or a more objective and complete state of the user as determined based on a comparison of the vehicle's external perceptions and the user's estimated perception as projected from outside the vehicle. Additionally, by using more objective measures of attention and / or cognitive load, simulations of real-world scenarios—e.g., in virtual simulation environments—may be more accurate and reliable, and therefore more suitable for designing, testing, and ultimately deploying in real-world systems.

[0006] The present systems and methods for occupant attention and cognitive load monitoring for semi-autonomous or autonomous driving applications are described in detail below with reference to the accompanying drawings. [Brief explanation of the drawings]

[0007] [Figure 1] 1 is a data flow diagram of a process of attention and / or cognitive load monitoring, according to some embodiments of the present disclosure. [Figure 2A] 10A-10C illustrate example plots generated using eye movement information, according to some embodiments of the present disclosure. [Figure 2B] 1A-1C illustrate example plots including vehicle region visualizations generated using eye movement information, according to some embodiments of the present disclosure. [Figure 2C] 1 is an exemplary illustration of eye positions at time steps or frames used to determine eye movement information, according to some embodiments of the present disclosure. [Figure 2D] 1 is an exemplary illustration of eye positions at time steps or frames used to determine eye movement information, according to some embodiments of the present disclosure. [Figure 2E] 1 is an exemplary heat map corresponding to eye movement information collected over a period of time, according to some embodiments of the present disclosure. [Figure 3] 10 is an example visualization of an extended gaze or field of view representation outside the vehicle for comparing vehicle perception with estimated occupant perception, according to some embodiments of the present disclosure. [Figure 4] 1 is a flow diagram illustrating a method for determining an action based on driver attention and / or cognitive load, according to some embodiments of the present disclosure. [Figure 5A] 1 is an illustration of an exemplary autonomous vehicle, according to some embodiments of the present disclosure. [Figure 5B] 5B is an illustration of camera positions and fields of view for the example autonomous vehicle of FIG. 5A, according to some embodiments of the present disclosure. [Figure 5C] FIG. 5B is a block diagram of an example system architecture of the example autonomous vehicle of FIG. 5A, in accordance with some embodiments of the present disclosure. [Figure 5D] FIG. 5B is a system diagram of communication between a cloud-based server and the example autonomous vehicle of FIG. 5A, according to some embodiments of the present disclosure. [Figure 6] FIG. 1 is a block diagram of an exemplary computing device suitable for use in implementing some embodiments of the present disclosure. [Figure 7] FIG. 1 is a block diagram of an exemplary data center suitable for use in implementing some embodiments of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0008] Systems and methods are disclosed for occupant attention and cognitive load monitoring for semi-autonomous or autonomous driving applications. The present disclosure may be described with reference to an exemplary autonomous vehicle 500 (an example of which is described herein with reference to FIGS. 5A-5D and alternatively referred to herein as “vehicle 500” or “ego vehicle 500”), but this is not intended to be limiting. For example, the systems and methods described herein may be used by non-autonomous vehicles, semi-autonomous vehicles (e.g., in one or more Advanced Driver Assistance System(s) (ADAS)), robots, warehouse vehicles, off-road vehicles, airships, boats, and / or other types of vehicles. Additionally, the present disclosure may be described with reference to autonomous driving, but this is not intended to be limiting. For example, the systems and methods described herein may be used in robotics, aviation systems (e.g., for determining attention and / or cognitive load), marine systems, simulation environments (e.g., for simulating actions based on the attention and / or cognitive load of a human operator of a virtual vehicle within a virtual simulation environment), and / or other technical fields.

[0009] Referring to FIG. 1, FIG. 1 is a data flow diagram of a process 100 for attention and / or cognitive load monitoring according to some embodiments of the present disclosure. It should be understood that this and other configurations described herein are provided by way of example only. Other configurations and elements (e.g., machines, interfaces, functions, sequences, groupings of functions, etc.) may be used in addition to or instead of those illustrated, and some elements may be omitted entirely. Furthermore, many of the elements described herein are functional entities that may be implemented as separate or distributed components or in conjunction with other components, and in any suitable combination and location. Various functions described herein as being performed by entities may be implemented by hardware, firmware, software, and / or any combination thereof. For example, various functions may be performed by a processor executing instructions stored in a memory.

[0010] Process 100 may include generating and / or receiving sensor data 102A and / or 102B (collectively referred to herein as “sensor data 102”) from one or more sensors of vehicle 500 (which may be similar to vehicle 500 or may include a non-autonomous or semi-autonomous vehicle). Sensor data 102 may be used within process 100 to track body movements or postures of one or more occupants of vehicle 500, track eye movements of occupants of vehicle 500, determine occupant attention and / or cognitive load, project a representation of the occupant's gaze or field of view outside of vehicle 500, generate output 114 using one or more deep neural networks (DNNs) 112, compare the field of view or gaze representation to output 114, determine occupant state, determine one or more actions to take based on the state, and / or other tasks or operations. The sensor data 102 may, in some instances, include, without limitation, sensor data 102 from any type of sensor, such as, but not limited to, those described herein with respect to the vehicle 500 and / or other vehicles or objects, e.g., robotic devices, VR systems, AR systems, etc.By way of non-limiting example, and with reference to FIGS. 5A-5C , the sensor data 102 may include, without limitation, global navigation satellite system (GNSS) sensors 558 (e.g., global positioning system (GPS) sensors, differential GPS (DGPS) sensors, etc.), RADAR sensors 560, ultrasonic sensors 562, LIDAR sensors 564, inertial measurement units (IMUs), and the like. The vehicle 500 may include data generated by sensors 566 (e.g., accelerometers, gyroscopes, magnetic compasses, magnetometers, etc.), microphones 596, stereo cameras 568, wide-view cameras 570 (e.g., fisheye cameras), infrared cameras 572, surround cameras 574 (e.g., 360-degree cameras), long-range and / or medium-range cameras 598, in-cabin cameras, in-cabin heat, pressure, or touch sensors, in-cabin motion sensors, in-cabin microphones, speed sensors 544 (e.g., for measuring the speed and / or distance traveled of the vehicle 500), and / or other sensor types.

[0011] In some examples, sensor data 102A may correspond to sensor data generated using one or more in-cabin sensors, e.g., one or more in-cabin cameras, in-cabin near-infrared (NIR) sensors, in-cabin microphones, and / or the like, and sensor data 102B may correspond to sensor data generated using one or more external sensors of vehicle 500, e.g., one or more cameras, RADAR sensors 560, ultrasonic sensors 562, LIDAR sensors 564, and / or the like. As such, sensor data 102A may correspond to sensors having a perceptual field or field of view interior to vehicle 500 (e.g., cameras having an occupant, e.g., the driver, within their field of view), and sensor data 102B may correspond to sensors having a perceptual field or field of view exterior to vehicle 500 (e.g., cameras, LiDAR sensors, etc. having a perceptual field that includes the environment outside vehicle 500). However, in some embodiments, sensor data 102A and sensor data 102B may include sensor data from any sensor having a perceptual field inside and / or outside vehicle 500.

[0012] The sensor data 102A may be used by the body tracker 104 and / or the eye tracker 106 to determine gestures, posture, activity, eye movements (e.g., saccade velocity, smooth pursuit, gaze position, direction, or vector, pupil size, blink rate, road scan range and distribution, etc.), and / or other information about an occupant, such as the driver, of the vehicle 500. This information may then be used by an attention determiner 108 to determine the occupant's attentiveness, a cognitive load determiner 110 to determine the occupant's cognitive load, and / or a field of view (FOV) projector 116 for comparison—via a comparator 120—with external sensory output 114 from one or more deep neural networks (DNNs) 112. Information representing attention, cognitive load, and / or the output of the comparator 120 may be analyzed by the state machine 122 to determine the occupant's state, which may then be used by the action determiner 124 to determine one or more actions or behaviors to perform (e.g., to issue a visual, audible, and / or tactile notification, suppress a notification, engage an ADAS system, assume autonomous control of the vehicle 500, etc.).

[0013] The body tracker 104 can use sensor data 102A—e.g., sensor data from one or more in-cabin cameras, microphones, pressure sensors, temperature sensors, etc.—to determine posture, pose, activity, or state (e.g., both hands on the steering wheel, one hand on the steering wheel, texting, reading, slouched, suddenly ill, incapacitated, distracted, etc.) and / or other information about the occupant. For example, the body tracker 104 can execute one or more machine learning algorithms, deep neural networks, computer vision algorithms, image processing algorithms, mathematical algorithms, and / or the like to determine the body tracking information. In some non-limiting examples, the body tracker 104 may include similar features, functionality, and / or components as described in U.S. patent application Ser. No. 16 / 915,577, filed June 29, 2020, and / or U.S. patent application Ser. No. 16 / 907,125, filed June 19, 2020, each of which is incorporated by reference in its entirety herein.

[0014] The eye tracker 106 can use sensor data 102A—e.g., sensor data from one or more in-cabin cameras, NIR cameras or sensors, and / or other eye-tracking sensor types—to determine gaze direction and movement, fixations, road scanning behavior (e.g., road scanning pattern, distribution, and range), saccade information (e.g., velocity, direction, etc.), blink rate, smooth pursuit information (e.g., velocity, direction, etc.), and / or other information. The eye tracker 106 can determine the duration corresponding to a particular state, e.g., how long the fixation lasted, and / or track how many times a particular state is determined—e.g., how many fixations, how many saccades, how many smooth pursuits, etc. The eye tracker 106 can monitor or analyze each eye individually and / or both eyes together. For example, both eyes can be monitored to use triangulation to measure the depth of the occupant's gaze. In some embodiments, the eye tracker 106 may execute one or more machine learning algorithms, deep neural networks, computer vision algorithms, image processing algorithms, mathematical algorithms, and / or the like to determine the eye-tracking information. In some non-limiting examples, the eye tracker 106 may include similar features, functionality, and / or components as described in U.S. patent application Ser. No. 16 / 363,648, filed October 8, 2018, U.S. patent application Ser. No. 16 / 544,442, filed August 19, 2019, U.S. patent application Ser. No. 17 / 010,205, filed September 2, 2020, U.S. patent application Ser. No. 16 / 859,741, filed April 27, 2020, U.S. patent application Ser. No. 17 / 004,252, filed August 27, 2020, and / or U.S. patent application Ser. No. 17 / 005,914, filed August 28, 2020, each of which is incorporated by reference in its entirety.

[0015] The attentiveness determiner 108 may be used to determine the occupant's attentiveness. For example, outputs from the body tracker 104 and / or the eye tracker 106 may be processed or analyzed by the attentiveness determiner 108 to generate an attentiveness value, score, or level. The attentiveness determiner 108 may execute one or more machine learning algorithms, deep neural networks, computer vision algorithms, image processing algorithms, mathematical algorithms, and / or the like to determine attentiveness. For example, with reference to FIGS. 2A-2E, FIG. 2A includes a graph 202 corresponding to a current (e.g., corresponding to a current time or period—e.g., 1 second, 3 seconds, 5 seconds, etc.) gaze direction and gaze information. For example, the gaze direction may be represented by point 212, where an (x, y) location in the graph 202 may have a corresponding location relative to the vehicle 500. Graph 202 may be used to determine eye movement types such as saccades—e.g., recent eye movement types—or smooth pursuit, gaze, etc., as referenced in FIG. 2A. As a further example, with reference to FIG. 2B, points 212 on graph 202 may be reflected in graph 204, which may reflect the occupant's current and / or recent (e.g., within the last second, three seconds, etc.) gaze areas. For example, any number of gaze areas 214 (e.g., gaze areas 214A-214F) may be used to determine road scanning behavior, gaze, and / or other information that may be used by attentiveness determiner 108.By way of non-limiting example, gaze areas 214 may include left gaze area 214A (e.g., corresponding to the driver's side window, driver's side mirror, etc.), left front gaze area 214B (e.g., corresponding to the left half or portion of the windshield), right front gaze area 214C (e.g., corresponding to the right half or portion of the windshield), right gaze area 214D (e.g., corresponding to the passenger's side window, passenger's side mirror, etc.), instrument cluster gaze area 214E (e.g., corresponding to the instrument cluster 532 or instrument panel behind, below, and / or above the steering wheel), and / or center console gaze area 214F (e.g., corresponding to a control-bearing surface, a display, a touch screen interface, radio controls, climate control controls, hazard light controls, an in-vehicle infotainment system (IVI), in-car entertainment (ICE), and / or other center console features).

[0016] 2C-2D, diagrams 206 and 208 include visualizations of the occupant—e.g., more focused on the occupant's eyes in diagram 206 and a broader focus on the occupant in diagram 208—that may be used to generate diagrams 202, 204, and / or 210. For example, a determination that the occupant is looking to the right front of vehicle 500—e.g., toward right front gaze area 214C—and that the occupant's last movement was a saccade may be determined using one or more instances of diagrams 206 and 208. The (x,y) position of the occupant's head and / or eyes may have a known correlation to gaze area 214 and / or to the (x,y) position of the gaze area diagrams—e.g., diagrams 202 and / or 204. As such, the orientation of the occupant's head and / or eyes may be determined and used to determine the gaze direction and / or position of the current frame. Additionally, results over any number of frames (e.g., 2 seconds of frame capture at 30 frames per second, or 60 frames) can be used to track movement types—e.g., saccades, blink rate, smooth pursuit, gaze, road-scanning behavior, and / or the like. In some embodiments, as shown in FIG. 2E , occupant gaze information can be tracked over a period of time and used to generate a heat map 210 (e.g., darker regions correspond to more frequent gaze locations or directions than regions with lighter or lower point-density patterns) corresponding to the occupant's gaze location and direction over time (e.g., over 30 seconds, 1 minute, 3 minutes, 5 minutes, etc.). For example, for heat map 210—whose coordinate system is similar to that of charts 202 and / or 204—heat map 210 may indicate that the occupant gazed more frequently toward right-front gaze region 214C and / or right-side gaze region 214D than toward other gaze regions 214. In some embodiments, heat map 210 may be updated at each frame to include information for the current frame. For example, weighting may be applied to give more weight to more recent gaze information when generating heat map 210. Heat map 210 may then represent the occupant's road scanning behavior, patterns, and / or frequency, which may be used to determine the occupant's attentiveness.

[0017] The attention determiner 108 can then use some or all of the eye-tracking information to determine the occupant's attentiveness. For example, if the heat map indicates that the occupant was frequently and extensively scanning the road and that the current gaze direction was toward the driving surface or its immediate vicinity, the current attention value or score may be determined to be high. As another example, if the heat map indicates that the occupant was more focused on a location off the road—e.g., toward a sidewalk, building, scenery, etc.—and the current gaze was also toward the side of the vehicle 500 (e.g., when the actual driving surface was only within the occupant's peripheral field of view), the attention score, at least for the eye-tracking-based information, may be low—e.g., the occupant may be determined to be in a daydreaming or gazing state.

[0018] In some instances, as described in more detail herein, the attentiveness determiner 108 can determine attentiveness using the output of the comparator 120. For example, if several objects—e.g., a VRU, a traffic sign, a waiting state, etc.—are identified using the vehicle 500's external perception, the number of those objects seen and / or processed by the occupant can be determined over some period of time. For example, if 30 objects are detected and the occupant has seen and / or processed 28 of them—e.g., as determined by the comparator 120 using the output of the FOV projector 116 and the output 114 of the DNN 112—the occupant's attentiveness can be determined to be high. However, if the occupant sees and / or processes only 15 of the 30, the occupant's attentiveness can be determined to be low (at least with respect to the comparator 120 calculation). In some examples, the determination of whether the occupant has processed the objects can be based on other attention factors and / or cognitive load determinations. For example, if a gaze projection (e.g., as determined from a calculated gaze vector) or a vehicle occupant's field of view overlaps—at least partially—with a detected object, the vehicle occupant may be determined to have seen the object. However, if other attention information indicates that the vehicle occupant was actually fixated on a particular gaze or momentarily looked down at their phone, or if the vehicle occupant is determined to have a high cognitive load, the vehicle occupant may be determined to have not processed, registered, or paid attention to the object. In such instances, the vehicle occupant's attention may be determined to be lower than if the vehicle occupant processed or registered the object.

[0019] In addition to or instead of using eye-tracking information, the attentiveness determiner 108 can use body tracking information from the body tracker 104 to determine attentiveness. For example, pose, posture, activity, and / or other information about the occupant can be used to determine attentiveness. As such, if the driver has both hands on the steering wheel and an upright posture, the driver's body-based contribution to attentiveness may correspond to a high attentiveness value or score. In contrast, if the driver has one hand on the steering wheel and the other holding their phone in front of their face, a low attentiveness value or score may be determined for at least the body-based contribution.

[0020] The cognitive load determiner 110 may be used to determine the occupant's cognitive load. For example, outputs from the body tracker 104 and / or the eye tracker 106 may be processed or analyzed by the cognitive load determiner 110 to generate a cognitive load value, score, or level. The cognitive load determiner 110 may execute one or more machine learning algorithms, deep neural networks, computer vision algorithms, image processing algorithms, mathematical algorithms, and / or the like to determine the cognitive load. For example, eye-based metrics such as pupil size (e.g., diameter) or other eye characteristics, eye movements, blink rate, or other eye measurements, and / or other information may be used to calculate a cognitive load value, score, or level (e.g., low, medium, high)—e.g., by computer vision algorithms and / or DNNs. For example, dilated pupils may indicate high cognitive load, and the cognitive load determiner 110 may use pupil size, in addition to other information, to determine the cognitive load score.

[0021] Road scanning behavior may also indicate cognitive load—e.g., whether the occupant is attentive or daydreaming. For example, if a heat map corresponding to road scanning pattern or behavior indicates that the occupant was scanning the entire scene from left to right, this information may indicate a lower cognitive load. However, if the occupant is not scanning the road or fixates (e.g., gazing blankly, daydreaming, etc.) on less important parts of the environment—e.g., the sidewalk, scenery, etc.—for a period of time, this may indicate a higher cognitive load. In some examples, environmental (e.g., weather, time of day, etc.) and / or driving conditions (e.g., traffic, road conditions, etc.) may be factored into the cognitive load determination (and / or attentiveness determination), such that snow, sleet, hail, rain, direct sunlight, darkness, heavy traffic, and / or other conditions may affect the cognitive load determination. For example, for the same set of visual target and / or body tracking information, cognitive load may be determined to be low when weather conditions are clear and there is little traffic, and high when it is snowing and there is heavy traffic.

[0022] In some embodiments, the occupant's cognitive load and / or attentiveness may be based on a profile corresponding to the occupant. For example, the occupant's eye-tracking information, body-tracking information, and / or other information over the occupant's current drive and / or one or more previous drives may be monitored and used to determine a customized profile for the occupant that indicates when a particular occupant is attentive, inattentive, has a higher cognitive load, has a lower cognitive load, etc. For example, a first occupant may scan the road less but process objects and make correct decisions faster or more frequently, while a second occupant may scan the road less but process objects less quickly or less frequently and make correct decisions. As such, by customizing the profiles for the first and second occupants, when both the first and second occupants are performing similar road-scanning behaviors, the first occupant may not be overly alerted, while the second occupant may receive more frequent alerts to ensure safe driving. An occupant's profile may also include increased granularity, such that occupant behavior may be tracked across a particular street, highway, route, weather conditions, etc., and this information may be used to determine the occupant's attention and / or cognitive load during future instances of traveling the same street, highway, in the same weather, etc. In addition to cognitive load and / or attention determination, occupants may be able to customize their profile to include certain notifications or other action types, objects for which they would like to be notified (e.g., a first occupant may not want a notification for a crosswalk, while a second occupant may want one), etc.

[0023] In some examples, the sensor data 102B may be applied to one or more deep neural networks (DNNs) 112 trained to calculate a variety of different outputs 114. Prior to application or input to the DNNs 112, the sensor data 102 may undergo preprocessing, for example, to transform, crop, expand, shrink, zoom in, rotate, and / or otherwise modify the sensor data 102. For example, if the sensor data 102B corresponds to camera image data, the image data may be cropped, reduced, expanded, flipped, rotated, and / or otherwise adjusted for an appropriate input format for the respective DNNs 112. In some examples, the sensor data 102B may include image data representing an image, image data representing a video (e.g., a snapshot of a video), and / or sensor data representing a representation of a sensor's field of perception (e.g., a depth map for a LIDAR sensor, a value graph for an ultrasonic sensor, etc.). In some instances, the sensor data 102B may be used without preprocessing (e.g., in its raw or captured format), while in other instances, the sensor data 102B may undergo preprocessing (e.g., noise balancing, demosaicing, scaling, cropping, enhancement, white balancing, tone curve adjustment, etc., using a sensor data preprocessor (not shown), etc.).

[0024] Although examples are described herein with respect to the use of DNNs 112 (and / or the use of DNNs, computer vision algorithms, image processing algorithms, machine learning models, etc. in connection with body tracker 104, eye tracker 106, attention determiner 108, and / or cognitive load determiner 110), this is not intended to be limiting. For example, and without limitation, the DNNs 112 and / or computer vision algorithms, image processing algorithms, machine learning models, etc. described herein with respect to the body tracker 104, the eye tracker 106, the attention determiner 108, and / or the cognitive load determiner 110 may include any type of machine learning model or algorithm, such as linear regression, logistic regression, decision trees, support vector machines (SVMs), naive Bayes, k-nearest neighbor (Knn), K-means clustering, random forests, dimensionality reduction algorithms, gradient boosting algorithms, neural networks (e.g., autoencoders, convolutional, recurrent, perceptrons, long / short-term memory / LSTM, Hopfield, Boltzmann, deep belief, deconvolution, generative adversarial, liquid state machines, etc.), area-of-interest detection algorithms, machine learning models using computer vision algorithms, and / or other types of algorithms or machine learning models.

[0025] As an example, the DNN 112 can process the sensor data 102 to produce detections of lane markings, road boundaries, signs, poles, trees, static objects, vehicles and / or other dynamic objects, wait states, intersections, distance, depth, object dimensions, etc. For example, the detections may correspond to location (e.g., in 2D image space, in 3D space, etc.), geometry, pose, semantic information, and / or other information related to the detection. As such, for lane lines, the location of the lane line and / or the type of lane line (e.g., dashed, solid, yellow, white, crosswalk, bike lane, etc.) may be detected by the DNN 112 processing the sensor data 102. For signs, the location of the sign or other wait state information and / or its type (e.g., yield ahead, stop, crosswalk, traffic light, yield ahead warning light, construction, speed limit, exit, etc.) may be detected using the DNN 112. For detected vehicles, motorcyclists, and / or other dynamic actors or road users, the position and / or type of the dynamic actor may be identified and / or tracked and / or used to determine waiting states within the scene (e.g., if a vehicle moves in a certain way with respect to an intersection, e.g., by coming to a stop, its corresponding intersection or waiting state may be detected as an intersection with a stop sign or traffic lights).

[0026] The output 114 of the DNN 112, in embodiments, may undergo post-processing, for example, by converting the raw output into a useful output—for example, if the raw output corresponds to a confidence level for each point (e.g., in LiDAR, RADAR, etc.) or pixel (e.g., in a camera image) that the point or pixel corresponds to a particular object type, post-processing may be performed to determine each point or pixel that corresponds to a single instance of the object type. This post-processing may include temporal filtering, weighting, outlier removal (e.g., removing pixels or points determined to be outliers), upscaling (e.g., the output may be predicted at a lower resolution than the input sensor data instances, and the output may then be scaled back to the input resolution), downscaling, curve fitting, and / or other post-processing techniques. The output 114—after post-processing, in embodiments—may be in a 2D coordinate space (e.g., image space, LiDAR range image space, etc.) and / or in a 3D coordinate system. In an embodiment, if output 114 is in a 2D coordinate space and / or in a 3D coordinate space other than 3D world space, output 114 may be transformed into the same coordinate system as the projected FOV output by FOV projection 116 (e.g., into a 3D world space coordinate system with an origin at a location on the vehicle).

[0027] In some non-limiting examples, the DNN 112 and / or output 114 may be implemented using a method such as those described in U.S. patent application Ser. No. 16 / 286,329 filed February 26, 2019, U.S. patent application Ser. No. 16 / 355,328 filed March 15, 2019, U.S. patent application Ser. No. 16 / 356,439 filed March 18, 2019, U.S. patent application Ser. No. 16 / 385,921 filed April 16, 2019, U.S. patent application Ser. No. 535,440 filed August 8, 2019, U.S. patent application Ser. No. 535,440 filed December 27, 2019, each of which is incorporated herein by reference in its entirety. No. 16 / 728,595, filed Dec. 27, 2019; U.S. Patent Application No. 16 / 728,598, filed Mar. 9, 2020; U.S. Patent Application No. 16 / 813,306, filed Mar. 9, 2020; U.S. Patent Application No. 16 / 848,102, filed Apr. 14, 2020; U.S. Patent Application No. 16 / 814,351, filed Mar. 10, 2020; U.S. Patent Application No. 16 / 911,007, filed Jun. 24, 2020; and / or U.S. Patent Application No. 16 / 514,230, filed Jul. 17, 2019.

[0028] The FOV projector 116 can use eye-tracking information from the eye tracker 106 to determine the gaze direction and projected occupant gaze or field of view. For example, with reference to FIG. 3 , if the occupant is looking at the right-front gaze area 214C ( FIG. 2B ), specifically through the lower-right portion of the windshield 302, a two-dimensional (2D) or three-dimensional (3D) (as shown) projection 306 of the occupant's field of view or gaze can be generated and extended into the environment outside the vehicle 500. Additionally, as described herein, one or more DNNs 112 can be used to process the sensor data 102B to determine the locations (e.g., bounding shapes, bounding contours, bounding boxes, points, pixels, etc.) of people 304 (e.g., VRUs) and signs 310 within the environment, and can generate output 114—e.g., directly and / or after post-processing. While a sign 310 and a person 304 are illustrated, this is not intended to be limiting, and the objects or information of interest in the environment may include any road, object, environmental, or other information, depending on the embodiment. For example, static objects, dynamic actors, intersections (or information corresponding thereto), queues, lane lines, road boundaries, sidewalks, obstacles, signs, poles, trees, scenery, information determined from a map (e.g., from an HD map), environmental conditions, and / or other information may be subjected to overlap comparison using comparator 120.

[0029] As a non-limiting example, a first DNN 112 may be used to detect the sign 310, and a second DNN 112 may be used to detect the person 304. In other examples, the same DNN 112 may be used to detect both the sign 310 and the person 304. The projection 306 (and / or one or more additional projections corresponding to previous time steps or frames) may then be compared by the comparator 120 to the detected object—e.g., the sign 310 and the person 304—to determine whether the occupant saw the detected object. If overlap occurs—at least partially—it may be determined that the occupant saw the object. In some examples, the overlap determination may include an overlap threshold, for example, 50% overlap (e.g., 50% of the bounding shape overlaps with a portion of the projection 306), 70% overlap, 90% overlap, etc. In other examples, any amount of overlap may satisfy the overlap determination, or complete overlap may satisfy the overlap determination.

[0030] To determine overlap, projection 306 and output 114 may be calculated in—or transformed—the same coordinate system. For example, projection 306 and output 114 may be determined with respect to a shared (e.g., world space) coordinate system—e.g., having an origin at vehicle 500, e.g., an axle location (e.g., the center of the rear axle of the vehicle), the front bumper, the windshield, etc. As such, if output 114 is otherwise calculated in 2D image space, a 3D space with a different origin, and / or in a 2D or 3D coordinate space other than the shared coordinate space used by comparator 120, output 114 may be transformed to the shared coordinate space.

[0031] Although shown as a narrow projection 306 (e.g., including the focal region of the occupant's field of view), this is not intended to be limiting. For example, in some embodiments, the projection 306 may include the focus of gaze within the occupant's field of view in addition to some or all of the periphery of the occupant's estimated field of view. In some embodiments, the overlap determination may be weighted based on the portion of the projection 306 that overlaps with the object (e.g., the focal or center of the field of view may be weighted more highly toward determining that the occupant saw the object compared to the peripheral portion of the field of view). For example, if a portion of the projection 306 corresponding only to the periphery of the occupant's field of view overlaps, the determination may be that the occupant did not see the object, or some other criteria may also have to be met to determine that the occupant saw the object (e.g., a threshold for time in the periphery may have to be met, a recency threshold may have to be met—e.g., the object was in the periphery within the last second—etc.).

[0032] In some embodiments, if it is determined that the occupant has seen the object, one or more additional criteria may be used by the comparator 120 to make a final determination. For example, duration (e.g., continuous, cumulative, etc. over a period of time) and / or recency may be factored in. In such instances, the occupant may need to see the object (e.g., determined to have seen the object using an overlap threshold) for a duration threshold—e.g., several frames, e.g., 5 or 10 frames, a period, e.g., 0.5 seconds, 1 second, 2 seconds, etc.—to constitute a positive determination that the occupant has seen the object. In some embodiments, the duration may be a cumulative duration over a period or a (sliding) time window. As a non-limiting example, if the cumulative duration is 1 second and the period is 5 seconds, and an overlap occurs for 1 / 4 second, and then 3 seconds later, another overlap occurs for 3 / 4 second, the cumulative duration may be met within the period, and the occupant may be determined to have seen the object. In some instances, a recency determination may additionally or alternatively be used, such that the occupant may have to be determined to have seen the object within a time window. For example, the occupant may have to see the object within 3 seconds, 5 seconds, 10 seconds, etc. of the current frame or time step for the final determination to be that the occupant saw the object.

[0033] Comparator 120 can output a confidence level (e.g., based on the amount of overlap, the duration of the overlap, the recency of the overlap, etc.) that the occupant saw the object, an overlap value (e.g., a percentage), the duration of the overlap, the recency of the overlap, and / or another value, score, or probability that the occupant saw the object at each time step. This determination can be made for each object type and / or for each instance thereof. For example, if action decision 124 includes a notification decider and the system is configured to issue notifications for unseen (and / or unprocessed) lane markers and for unseen (and / or unprocessed) pedestrians, comparator 120 can output a prediction corresponding to each lane marker and a prediction corresponding to each pedestrian (e.g., at each time step or frame). As such, state machine 122 and / or action decider 124 can generate outputs and / or make decisions corresponding to lane markers, pedestrians, or a combination thereof. As another example, if the action determiner 124 corresponds to one or more ADAS systems, state information corresponding to a lane marker may correspond to a different action than state information corresponding to a pedestrian. As such, separate outputs from the comparator 120 corresponding to different object types and / or instances thereof may be useful not only to the action determiner 124 in determining the correct action to take (or refrain from taking), but also to the system in determining the state of the occupant with respect to a particular object type or instance.

[0034] The state machine 122 can receive as inputs the outputs from the comparator 120, the attentiveness determiner 108, and / or the cognitive load determiner 110 and can determine the state of the occupant (e.g., the driver). In some instances, the state machine 122 can determine a single state (e.g., aware, distracted, focused, etc.) that can be used by the action determiner 124 when deciding which action to take or not take (e.g., suppress notification). In other instances, the state machine 122 can determine a state for each object type and / or its instance that can be used by the action determiner 124 to make a decision regarding one or more different types of actions. In some embodiments, the state machine 122 can determine the state based on various levels or processes (e.g., a hierarchical state determination process). For example, an unaware or inattentive state can be determined if an overlap threshold (e.g., overlap amount, overlap duration, overlap recency, etc.) is not met—indicating, for example, that the occupant may not have seen and / or processed the object. In one such instance, the next level of processing—e.g., for processing the output from the attentiveness determiner 108 and / or the cognitive load determiner 110—may not be performed. As such, if the occupant is not determined by the comparator 120 to have seen the object—e.g., the VRU—actions that would be performed as a result of not seeing the object may be performed (e.g., to generate and / or output a notification or warning, to activate an ADAS system such as an automatic emergency braking (AEB) system, etc.).

[0035] In some examples, even if the overlap threshold is not met, the state machine 122 may further process the output from the attentiveness determiner 108 and / or the cognitive load determiner 110 to determine a combined state—e.g., attentive (e.g., road scanning behavior indicates attentiveness) but unaware (e.g., the overlap threshold was not met and the occupant may not have seen the object). In one such example, different tiers or types of actions may be performed. In the example described above with an AEB system, if the occupant is determined to be attentive (and / or have low cognitive load) but unaware, the AEB system may not be executed, but a notification or warning may be generated and output to the occupant—e.g., audibly, tactilely, visually, and / or otherwise. In one example, if the overlap threshold is not met and attentiveness is low and / or cognitive load is high, more preventative action may be taken. For example, a notification may be output, an ADAS system may be executed, and / or - if the vehicle 500 is an autonomous-enabled vehicle - autonomous control may be assumed (e.g., to perform a safety maneuver such as pulling over to the side of the road).

[0036] If the overlap threshold is met—e.g., indicating that the driver saw the object—state machine 122 can determine a final state using output from attentiveness determiner 108 and / or cognitive load determiner 110. For example, if the occupant is aware—e.g., as indicated by overlap information from comparator 120—attention score, value, level, etc. and / or cognitive load score, value, level, etc. can be analyzed by state machine 122 to determine a final state. As such, even if it is determined that the occupant saw an object (e.g., a pedestrian, a vehicle, a queue, an intersection, a pole, a sign, etc.), state machine 122 can be used to determine the occupant's ability (e.g., based on cognitive load) or likelihood (e.g., based on attentiveness) to process the perception of the object. For example, if the driver sees a vehicle some distance ahead of ego vehicle 500 (e.g., determined to see based on a comparison of the driver's estimated field of view to the vehicle's position) but is determined to have high cognitive load and / or low attentiveness (e.g., based on heat maps, gaze, etc.), the state may include aware (e.g., to the vehicle) but inattentive (e.g., potentially did not process the vehicle's presence). In such an instance, action determiner 124 may perform an action as if the occupant had not noticed the object and / or may issue a lower level action—e.g., issuing a notification or warning and not executing an ADAS system (e.g., AEB).

[0037] In some embodiments, the state determination of state machine 122 can incorporate environmental conditions (e.g., weather, time of day, etc.) and / or driving conditions (e.g., traffic, road conditions, etc.) when determining the state. For example, in some embodiments, the attention, cognitive load, and / or redundancy determination can incorporate environmental and / or driving conditions, and state machine 122 can determine the state based on this information. In other embodiments, in addition to or instead of attentiveness determiner 108, cognitive load determiner 110, and / or comparator 120 incorporating environmental and / or driving conditions, state machine 122 can incorporate environmental and / or driving conditions when determining the state. For example, a certain value, score, and / or level of attentiveness may indicate an aware and / or attentive driver when it is sunny and light outside (e.g., daytime), but may indicate an unaware and / or inattentive driver when it is snowing and / or dark outside (e.g., nighttime).

[0038] The action determiner 124 may include or be part of a warning and notification system, an ADAS system, a control system or control layer of an autonomous driving software stack, a planning system or layer of an autonomous driving stack, a collision avoidance system or layer or autonomous driving stack, an actuation system or layer of an autonomous driving stack, etc. The action determiner 124 may perform different actions depending on the entity (or object) to which the state corresponds. By way of non-limiting example, for a dynamic object on the driving surface (e.g., a vehicle, a pedestrian, a bicyclist, etc.), the action determiner 124 may determine a notification type, an ADAS system to execute or engage, an autonomous system to execute or engage, etc. For an off-road object, such as a pole, a pedestrian, and / or the like, the action determiner 124 may determine a notification type but may not engage an ADAS or autonomous function unless the vehicle 500 is also currently off-road or heading off-road. For wait conditions (e.g., street lights, stop signs, etc.), the action determiner 124 can determine the notification type and / or ADAS system to perform or engage—e.g., stopping the vehicle 500 at a stop light if it is determined that the light is red and the occupant did not see the red light.

[0039] Referring now to FIG. 4 , each block of method 400 described herein includes a computational process that may be implemented using any combination of hardware, firmware, and / or software. For example, various functions may be implemented by a processor executing instructions stored in a memory. Method 400 may also be implemented as computer-usable instructions stored on a computer storage medium. Method 400 may be provided by a standalone application, a service, or a hosted service (standalone or in combination with another hosted service), or a plug-in to another product, to name a few. Additionally, method 400 is described with respect to process 100 of FIG. 1 , by way of example. However, method 400 may additionally or alternatively be performed by any one process and / or system, or any combination of processes and / or systems, including, but not limited to, those described herein.

[0040] 4 is a flow diagram illustrating a method 400 for determining an action based on driver attention and / or cognitive load, according to some embodiments of the present disclosure. The method 400 includes, at block B402, determining gaze information of a vehicle occupant. For example, the eye tracker 106 can determine gaze information (e.g., a gaze vector) and / or other information that can be used to determine the current gaze or field of view of an occupant of the vehicle 500, e.g., the driver.

[0041] The method 400 includes projecting a representation of the gaze information outside the vehicle at block B404. For example, a projection (e.g., projection 306 of FIG. 3) may be generated based on the gaze information, and the projection may correspond to at least a portion of the occupant's field of view.

[0042] The method 400 includes, at block B406, calculating, using one or more DNNs, object position information corresponding to one or more objects external to the vehicle. For example, the one or more DNNs 112 may be used to calculate position and / or semantic information corresponding to one or more objects, subjects, and / or other information in the environment external to the vehicle 500. By way of non-limiting example, position and / or semantic information regarding dynamic actors, static objects, queues, lane lines, road boundaries, poles, and / or signs may be calculated using the DNNs 112.

[0043] The method 400 includes, at block B408, determining whether a threshold overlap between the representation and one or more objects is met. For example, the comparator 120 may compare the representation of the gaze (e.g., the projection 306) with objects in the environment—e.g., in the same coordinate system—to determine whether a threshold amount of overlap has occurred. The threshold amount of overlap may also include determining whether the overlap has occurred within a time window from the current time, over a cumulative time within the time window, and / or for a sufficiently long continuous duration.

[0044] If the threshold overlap is not met in block B408, method 400 may proceed to block B410. Block B410 includes performing a first action. For example, if a threshold amount of overlap does not exist (and / or other thresholds are not met, e.g., recency, duration, etc.), which may indicate that the occupant did not see the object, the first action may be performed. For example, a warning or notification may be generated and output, one or more ADAS systems may be activated, and / or the like.

[0045] If the threshold overlap is met in block B408, the method 400 may proceed to block B412. Block B412 may include determining at least one of the occupant's cognitive load or the occupant's attentiveness. For example, meeting the overlap threshold—e.g., indicating that the occupant saw the object—may enable the cognitive load determiner 110 to determine the occupant's cognitive load and / or the attentiveness determiner 108 to determine the occupant's attentiveness. This information may be used in block B414 to determine whether the occupant was able to process the perception of the object.

[0046] Method 400 includes determining whether cognitive load exceeds a first threshold and / or attention is less than a second threshold at block B414. For example, if cognitive load exceeds the first threshold—e.g., indicating that the occupant does not currently have the capacity to process the perception of the object—method 400 may proceed to block B416. Similarly, if attention is less than the second threshold—indicating that the occupant does not currently have the capacity to process the perception of the object—method 400 may proceed to block B416. In some embodiments, some combination of attention and cognitive load may be used to determine to proceed to block B416. Block B416 includes performing a second action. For example, the second action may be the same as the first action in some embodiments. As such, the action performed may be the same if it is determined in block B408 that the occupant did not see the object and / or if it is determined that the occupant saw the object but did not fully process it—e.g., due to high cognitive load and / or low attention. However, in some embodiments, the second action may differ from the first action. For example, if it is determined that the occupant saw the object, but the determination in block B414 is that the occupant may not have fully processed the object, a notification or warning may be issued, but one or more ADAS systems may not be executed (at least initially). As another example, the first action may include an audible and visual warning, and the second action may include only a visual warning—e.g., a less severe warning because it was determined that the occupant saw the object.

[0047] As another example, if the cognitive load is below a first threshold—e.g., indicating that the occupant currently has a higher ability to process the object's perception—method 400 may proceed to block B418. Similarly, if the attention exceeds a second threshold—e.g., indicating that the occupant currently has a higher ability to process the object's perception—method 400 may proceed to block B418. Block B418 includes performing a third action. For example, if cognitive load is low and / or attention is high and it is determined in block B408 that the occupant saw the object, the third action may be performed. In some examples, the third action may include not performing an action or suppressing an action—e.g., suppressing a notification. As such, if it is determined that the occupant saw the object and has low cognitive load and / or high attention, notifications, ADAS systems, and / or other actions may be suppressed to reduce the number of vehicle warnings, notifications, and / or autonomous or semi-autonomous activations. As such, the process 100 and / or method 400 may implement a hierarchical decision tree to determine the action to take based on the state of the occupant.

[0048] Exemplary Autonomous Vehicle 5A is a diagram of an example autonomous vehicle 500 according to some embodiments of the present disclosure. Autonomous vehicle 500 (alternatively referred to herein as “vehicle 500”) may include, but is not limited to, a passenger vehicle, such as a car, a truck, a bus, a first responder vehicle, a shuttle, an electric or moped, a motorcycle, a fire engine, a police vehicle, an ambulance, a boat, a construction vehicle, a submarine, a drone, and / or another type of vehicle (e.g., unmanned and / or carrying one or more passengers). Autonomous vehicles are generally described in terms of levels of automation as defined by the National Highway Traffic Safety Administration (NHTSA), a division of the U.S. Department of Transportation, and the Society of Automotive Engineers (SAE) "Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles" (Standard No. J3016-201806, published June 15, 2018; Standard No. J3016-201609, published September 30, 2016; and previous and future versions of this standard). Mobile vehicle 500 may be capable of functioning according to one or more of levels 3 through 5 of autonomous driving. For example, mobile vehicle 500 may be capable of conditional automation (Level 3), high automation (Level 4), and / or full automation (Level 5), depending on the embodiment.

[0049] The mobile vehicle 500 may include components such as a chassis, body, wheels (e.g., 2, 4, 6, 8, 18, etc.), tires, axles, and other components of the mobile vehicle. The mobile vehicle 500 may include a propulsion system 550, such as an internal combustion engine, a hybrid power plant, a fully electric engine, and / or another propulsion system type. The propulsion system 550 may be connected to a drive train of the mobile vehicle 500, which may include a transmission, to enable propulsion of the mobile vehicle 500. The propulsion system 550 may be controlled in response to receiving a signal from a throttle / accelerator 552.

[0050] A steering system 554, which may include a steering wheel, may be used to steer the vehicle 500 (e.g., along a desired course or route) when the propulsion system 550 is operating (e.g., when the vehicle is moving). The steering system 554 may receive signals from a steering actuator 556. A steering wheel may be optional for fully automated (Level 5) functionality.

[0051] Brake sensor system 546 may be used to operate vehicle brakes in response to receiving signals from brake actuators 548 and / or brake sensors.

[0052] A controller 536, which may include one or more system on chip (SoC) 504 (FIG. 5C) and / or a GPU, can provide signals (e.g., representations of commands) to one or more components and / or systems of the vehicle 500. For example, the controller can send signals to operate vehicle brakes via one or more brake actuators 548, to operate a steering system 554 via one or more steering actuators 556, and to operate a propulsion system 550 via one or more throttle / accelerators 552. The controller 536 may include one or more on-board (e.g., integrated) computing devices (e.g., supercomputers) that process sensor signals and output operational commands (e.g., signals representing commands) to enable rhythmic driving and / or assist a driver in operating the vehicle 500. The controllers 536 may include a first controller 536 for autonomous driving functions, a second controller 536 for functional safety functions, a third controller 536 for artificial intelligence functions (e.g., computer vision), a fourth controller 536 for infotainment functions, a fifth controller 536 for redundancy in emergency situations, and / or other controllers. In some instances, a single controller 536 may handle two or more of the foregoing functions, and two or more controllers 536 may handle a single function and / or any combination thereof.

[0053] The controller 536 may provide signals to control one or more components and / or systems of the vehicle 500 in response to sensor data (e.g., sensor inputs) received from one or more sensors. Sensor data may be received from, for example, and without limitation, global navigation satellite system sensors 558 (e.g., global positioning system sensors), RADAR sensors 560, ultrasonic sensors 562, LIDAR sensors 564, inertial measurement unit (IMU) sensors 566 (e.g., accelerometers, gyroscopes, magnetic compasses, magnetometers, etc.), microphones 596, stereo cameras 568, wide-view cameras 570 (e.g., fisheye cameras), infrared cameras 572, surround cameras 574 (e.g., 360-degree cameras), long-range and / or medium-range cameras 598, speed sensors 544 (e.g., for measuring the speed of the moving vehicle 500), vibration sensors 542, steering sensors 540, brake sensors (e.g., as part of a brake sensor system 546), and / or other sensor types.

[0054] One or more of the controllers 536 may receive input (e.g., represented by input data) from the instrument cluster 532 of the vehicle 500 and provide output (e.g., represented by output data, display data, etc.) via a human-machine interface (HMI) display 534, an audible annunciator, a loudspeaker, and / or other components of the vehicle 500. The output may include information such as vehicle velocity, speed, time, map data (e.g., HD map 522 of FIG. 5C ), position data (e.g., the location of the vehicle 500 on a map, etc.), direction, the locations of other vehicles (e.g., occupancy grid), information about objects and object situations as known by the controller 536, etc. For example, the HMI display 534 may display information regarding the presence of one or more objects (e.g., road signs, warning signs, traffic light changes, etc.) and / or a driving maneuver that the moving vehicle has performed, is performing, or will perform (e.g., changing lanes now, taking exit 34B in 3.22 km (2 miles), etc.).

[0055] The mobile vehicle 500 further includes a network interface 524 capable of communicating over one or more networks using one or more wireless antennas 526 and / or a modem. For example, the network interface 524 may be capable of communicating over LTE, WCDMA, UMTS, GSM, CDMA2000, etc. The wireless antenna 526 may also enable communication between objects in the environment (e.g., mobile vehicles, mobile devices, etc.) using local area networks such as Bluetooth, Bluetooth LE, Z-Wave, Zigbee, etc., and / or low power wide-area networks (LPWANs) such as LoRaWAN, SigFox, etc.

[0056] 5B is an illustration of camera positions and fields of view of the exemplary autonomous vehicle 500 of FIG. 5A, according to some embodiments of the present disclosure. The cameras and their respective fields of view are one illustrative example and are not intended to be limiting. For example, additional and / or alternative cameras may be included and / or the cameras may be located at different positions on the vehicle 500.

[0057] The camera type may include, but is not limited to, a digital camera adapted for use with components and / or systems of the mobile vehicle 500. The camera may be capable of operating at Automotive Safety Integrity Level (ASIL) B and / or another ASIL. The camera type may be capable of any image capture rate, such as 60 frames per second (fps), 120 fps, 240 fps, etc., depending on the embodiment. The camera may be capable of using a rolling shutter, a global shutter, another type of shutter, or a combination thereof. In some instances, the color filter array may include a red clear clear clear (RCCC) color filter array, a red clear clear blue (RCCB) color filter array, a red blue green clear (RBGC) color filter array, a Foveon X3 color filter array, a Bayer sensor (RGGB) color filter array, a monochrome sensor color filter array, and / or another type of color filter array. In some embodiments, clear pixel cameras, such as cameras with RCCC, RCCB, and / or RBGC color filter arrays, may be used in an effort to increase light sensitivity.

[0058] In some instances, one or more of the cameras may be used to perform advanced driver assistance system (ADAS) functions (e.g., as part of a redundant or fail-safe design). For example, a multi-function mono camera may be installed to provide functions including lane departure warning, traffic sign assist, and intelligent headlamp control. One or more of the cameras (e.g., all cameras) may simultaneously record and provide image data (e.g., video).

[0059] One or more of the cameras may be mounted in a mounting part, such as a custom-designed (e.g., 3D printed) part, to filter out stray light and reflections from within the vehicle (e.g., reflections from the dashboard reflected in the windshield mirror) that may interfere with the camera's image data capture ability. Referring to a side mirror mounting part, the side mirror part may be custom 3D printed so that the camera mounting plate fits the shape of the side mirror. In some instances, the camera may be integrated into the side mirror. For side view cameras, the camera may also be integrated into four posts at each corner of the cabin.

[0060] A camera (e.g., a forward-facing camera) with a field of view that includes a portion of the environment in front of the vehicle 500 may be used for surround view to aid in identifying a forward path and obstacles and, with the assistance of one or more controllers 536 and / or control SoCs, to provide information essential for generating an occupancy grid and / or determining a preferred vehicle path. Forward-facing cameras may be used to perform many of the same ADAS functions as LIDAR, including emergency braking, pedestrian detection, and collision avoidance. Forward-facing cameras may also be used for ADAS functions and systems, including other functions such as lane departure warning (LDW), autonomous cruise control (ACC), and / or traffic sign recognition.

[0061] Various cameras can be used in a forward-facing configuration, including, for example, a monocular camera platform including a complementary metal oxide semiconductor (CMOS) color imager. Another example can be a wide-view camera 570 that can be used to understand objects that come into view from the periphery (e.g., pedestrians, crossing traffic, or bicycles). While only one wide-view camera is shown in FIG. 5B, any number of wide-view cameras 570 can be present in the vehicle 500. Additionally, a long-range camera 598 (e.g., a long-view stereo camera pair) can be used for depth-based object detection, particularly for objects for which a neural network has not yet been trained. The long-range camera 598 can also be used for object detection and classification, as well as basic object tracking.

[0062] One or more stereo cameras 568 may also be included in the forward-facing configuration. The stereo camera 568 may include an integrated control unit with an extensible processing unit, which may provide programmable logic (FPGA) and a multi-core microprocessor with a CAN or Ethernet interface integrated on a single chip. Such a unit may be used to generate a 3D map of the vehicle's environment, including distance estimates for all points in the image. An alternative stereo camera 568 may include a compact stereo vision sensor, which may include two camera lenses (one on the left and one on the right) and an image processing chip that can measure the distance from the vehicle to objects of interest and use the generated information (e.g., metadata) to activate autonomous emergency braking and lane departure warning functions. Other types of stereo cameras 568 may be used in addition to or instead of those described herein.

[0063] Cameras having a field of view that includes portions of the environment to the sides of the mobile vehicle 500 (e.g., side-view cameras) may be used for surround view, providing information used to create and update the occupancy grid and generate side-impact collision warnings. For example, surround cameras 574 (e.g., four surround cameras 574 as shown in FIG. 5B ) may be positioned on the mobile vehicle 500. The surround cameras 574 may include wide-view cameras 570, fisheye cameras, 360-degree cameras, and / or the like. For example, four fisheye cameras may be positioned on the front, rear, and sides of the mobile vehicle. In an alternative arrangement, the mobile vehicle may use three surround cameras 574 (e.g., left, right, and rear) and utilize one or more other cameras (e.g., a forward-facing camera) as a fourth surround-view camera.

[0064] A camera having a field of view that includes the portion of the environment behind the mobile vehicle 500 (e.g., a rearview camera) may be used for parking assistance, surround view, rear collision warning, and creating and updating an occupancy grid. As described herein, a wide variety of cameras may be used, including, but not limited to, cameras that are also suitable as forward-facing cameras (e.g., long-range and / or medium-range cameras 598, stereo cameras 568, infrared cameras 572, etc.).

[0065] FIG. 5C is a block diagram of an example system architecture for the example autonomous vehicle 500 of FIG. 5A , in accordance with some embodiments of the present disclosure. It should be understood that this and other arrangements described herein are merely illustrative. Other arrangements and elements (e.g., machines, interfaces, functions, sequences, groupings of functions, etc.) may be used in addition to or instead of those shown, and some elements may be omitted altogether. Furthermore, many of the elements described herein are functional entities that may be implemented as separate or distributed components or in conjunction with other components, and in any suitable combination and location. Various functions described herein as being performed by entities may be implemented by hardware, firmware, and / or software. For example, various functions may be implemented by a processor executing instructions stored in a memory.

[0066] Each of the components, features, and systems of the mobile vehicle 500 in FIG. 5C is shown connected via a bus 502. The bus 502 may include a controller area network (CAN) data interface (alternatively referred to as a "CAN bus"). The CAN may be a network within the mobile vehicle 500 used to help control various features and functions of the mobile vehicle 500, such as braking, acceleration, braking, steering, windshield wiper operation, etc. The CAN bus may be configured to have dozens or even hundreds of nodes, each with its own unique identifier (e.g., CAN ID). The CAN bus may be read to determine steering angle, ground speed, engine revolutions per minute (RPM), button position, and / or other vehicle status indicators. The CAN bus may be ASIL B compliant.

[0067] Although the bus 502 is described herein as being a CAN bus, this is not intended to be limiting. For example, FlexRay and / or Ethernet may be used in addition to or as an alternative to a CAN bus. Additionally, although a single line is used to represent the bus 502, this is not intended to be limiting. There may be any number of buses 502, which may include, for example, one or more CAN buses, one or more FlexRay buses, one or more Ethernet buses, and / or one or more other types of buses using different protocols. In some instances, two or more buses 502 may be used to perform different functions and / or for redundancy. For example, a first bus 502 may be used for collision avoidance functions, and a second bus 502 may be used for actuation control. In any instance, each bus 502 may communicate with any of the components of the vehicle 500, and two or more buses 502 may communicate with the same component. In some instances, each SoC 504, each controller 536, and / or each computer in the vehicle may have access to the same input data (e.g., input from sensors in the vehicle 500) and may be connected to a common bus, such as a CAN bus.

[0068] Mobile vehicle 500 may include one or more controllers 536, such as those described herein with respect to FIG. 5A. Controller 536 may be used for a variety of functions. Controller 536 may be coupled to any of a variety of other components and systems of mobile vehicle 500 and may be used for control of mobile vehicle 500, artificial intelligence of mobile vehicle 500, infotainment for mobile vehicle 500, and / or the like.

[0069] The vehicle 500 may include a system-on-chip (SoC) 504. The SoC 504 may include a CPU 506, a GPU 508, a processor 510, a cache 512, an accelerator 514, a data store 516, and / or other components and features not shown. The SoC 504 may be used to control the vehicle 500 in a variety of platforms and systems. For example, the SoC 504 may be coupled in a system (e.g., the system of the vehicle 500) with an HD map 522 that can obtain map refreshes and / or updates via a network interface 524 from one or more servers (e.g., server 578 of FIG. 5D ).

[0070] The CPU 506 may include a CPU cluster or CPU complex (alternatively referred to as a "CCPLEX"). The CPU 506 may include multiple cores and / or L2 caches. For example, in some embodiments, the CPU 506 may include eight cores in a coherent multiprocessor configuration. In some embodiments, the CPU 506 may include four dual-core clusters, each with its own dedicated L2 cache (e.g., a 2M L2 cache). The CPU 506 (e.g., a CCPLEX) may be configured to support simultaneous cluster operation, allowing any combination of clusters of CPUs 506 to be active at any given time.

[0071] The CPU 506 may implement power management capabilities including one or more of the following features: individual hardware blocks may be automatically clock gated when idle to conserve dynamic power; each core clock may be gated when the core is not actively executing instructions by executing a WFI / WFE instruction; each core may be independently power gated; each core cluster may be independently clock gated when all cores are clock gated or power gated; and / or each core cluster may be independently power gated when all cores are power gated. The CPU 506 may further implement an enhanced algorithm for managing power states, where allowable power states and expected wake-up times are specified and hardware / microcode determines the best power state for entering the cores, clusters, and CCPLEX. The processing cores may support simplified power state entry sequences in software with work offloaded to microcode.

[0072] The GPU 508 may include an integrated GPU (alternatively referred to herein as an "iGPU"). The GPU 508 may be programmable and efficient for parallel workloads. In some instances, the GPU 508 may use an enhanced tensor instruction set. The GPU 508 may include one or more streaming microprocessors, where each streaming microprocessor may include an L1 cache (e.g., an L1 cache having at least 96 KB of storage capacity) and two or more of the streaming microprocessors may share a cache (e.g., an L2 cache having 512 KB of storage capacity). In some embodiments, the GPU 508 may include at least eight streaming microprocessors. The GPU 508 may use a compute application programming interface (API). Additionally, the GPU 508 may use one or more parallel computing platforms and / or programming models (e.g., NVIDIA's CUDA).

[0073] The GPU 508 may be power-optimized for best performance in automotive and embedded use cases. For example, the GPU 508 may be fabricated on FinFET (Fin field-effect transistor) chips. However, this is not intended to be limiting, and the GPU 508 may be fabricated using other semiconductor manufacturing processes. Each streaming microprocessor may incorporate several mixed-precision processing cores partitioned into multiple blocks. For example, and without limitation, 64 PF32 cores and 32 PF64 cores may be partitioned into four processing blocks. In such an example, each processing block may be assigned 16 FP32 cores, 8 FP64 cores, 16 INT32 cores, two mixed-precision NVIDIA tensor cores for deep learning matrix operations, an L0 instruction cache, a warp scheduler, a dispatch unit, and / or a 64KB register file. Additionally, the streaming microprocessor may include independent parallel integer and floating-point data paths to provide efficient execution of workloads with a mix of computational and addressing operations. Streaming microprocessors may include independent thread scheduling capabilities to allow finer-grained synchronization and coordination among concurrent threads. Streaming microprocessors may include a combined L1 data cache and shared memory unit to improve performance while simplifying programming.

[0074] The GPU 508 may, in some instances, include a high bandwidth memory (HBM) and / or 16GB HBM2 memory subsystem to provide up to 900GB / s of peak memory bandwidth. In some instances, synchronous graphics random-access memory (SGRAM), such as graphics double data rate type five synchronous random-access memory (GDDR5), may be used in addition to or in place of the HBM memory.

[0075] The GPU 508 may include unified memory technology that includes access counters to enable more accurate movement of memory pages to the processors that access them most frequently, thereby improving the efficiency of storage areas shared between processors. In some instances, address translation service (ATS) support may be used to enable the GPU 508 to directly access the CPU 506 page tables. In such instances, when the GPU 508 memory management unit (MMU) experiences a miss, an address translation request may be sent to the CPU 506. In response, the CPU 506 may consult its page table for a virtual-to-real mapping of addresses and send the translation back to the GPU 508. As such, unified memory technology may enable a single unified virtual address space for both the CPU 506 and GPU 508 memories, thereby simplifying GPU 508 programming and porting of applications to the GPU 508.

[0076] Additionally, GPU 508 may include access counters that can record the frequency of GPU 508's accesses to the memory of other processors. The access counters can help ensure that memory pages are moved to the physical memory of the processors that are accessing the pages most frequently.

[0077] The SoC 504 may include any number of caches 512, including those described herein. For example, the cache 512 may include an L3 cache available to both the CPU 506 and the GPU 508 (e.g., connected to both the CPU 506 and the GPU 508). The cache 512 may include a write-back cache that can record line state, such as by using a cache coherence protocol (e.g., MEI, MESI, MSI, etc.). The L3 cache may include 4 MB or more, depending on the implementation, although smaller cache sizes may also be used.

[0078] The SoC 504 may include an arithmetic logic unit (ALU) that may be utilized in performing processing for any of various tasks or operations (e.g., processing DNNs) of the vehicle 500. Additionally, the SoC 504 may include a floating point unit (FPU) (or other math co-processor or math co-processor type) for performing mathematical operations within the system. For example, the SoC 104 may include one or more FPUs integrated as execution units within the CPU 506 and / or GPU 508.

[0079] The SoC 504 may include one or more accelerators 514 (e.g., hardware accelerators, software accelerators, or a combination thereof). For example, the SoC 504 may include a hardware acceleration cluster, which may include optimized hardware accelerators and / or large on-chip memory. The large on-chip memory (e.g., 4 MB of SRAM) may enable the hardware acceleration cluster to accelerate neural networks and other operations. The hardware acceleration cluster may be used to complement the GPU 508 and to offload some of the GPU 508's tasks (e.g., to free up more cycles for the GPU 508 to perform other tasks). As an example, the accelerator 514 may be used for target workloads that are sufficiently stable to be suitable for acceleration (e.g., perception, convolutional neural networks (CNNs), etc.). As used herein, the term "CNN" may include all types of CNNs, including region-based or regional convolutional neural networks (RCNNs) and Faster RCNNs (e.g., as used for object detection).

[0080] The accelerator 514 (e.g., a hardware acceleration cluster) may include a deep learning accelerator (DLA). The DLA may include one or more tensor processing units (TPUs) that can be configured to provide an additional 10 trillion operations per second for deep learning applications and inference. The TPU may be an accelerator configured and optimized to perform image processing functions (e.g., CNN, RCNN, etc.). The DLA may also be optimized for a specific set of neural network types and floating-point operations, as well as inference. The DLA design can provide more performance per millimeter than a general-purpose GPU, significantly exceeding the performance of a CPU. The TPU can perform several functions, including, for example, single-instance convolution functions, supporting INT8, INT16, and FP16 data types for both features and weights, and post-processor functions.

[0081] The DLA can quickly and efficiently run neural networks, particularly CNNs, on processed or unprocessed data for any of a variety of functions, including, but not limited to: CNNs for object identification and detection using data from camera sensors, CNNs for distance estimation using data from camera sensors, CNNs for emergency vehicle detection and identification using data from microphones, CNNs for face recognition and moving vehicle owner identification using data from camera sensors, and / or CNNs for security and / or safety related events.

[0082] The DLA can perform any function of the GPU 508, and by using an inference accelerator, for example, a designer can target either the DLA or the GPU 508 for any function. For example, a designer can focus on processing CNNs and floating-point operations on the DLA and offload other functions to the GPU 508 and / or other accelerators 514.

[0083] The accelerator 514 (e.g., a hardware acceleration cluster) may include a programmable vision accelerator (PVA), which may alternatively be referred to herein as a computer vision accelerator. The PVA may be designed and configured to accelerate computer vision algorithms for advanced driver assistance systems (ADAS), autonomous driving, and / or augmented reality (AR) and / or virtual reality (VR) applications. The PVA may provide a balance between performance and flexibility. For example, each PVA may include, but is not limited to, any number of reduced instruction set computer (RISC) cores, direct memory access (DMA), and / or any number of vector processors.

[0084] The RISC cores may interact with an image sensor (e.g., an image sensor in any of the cameras described herein), an image signal processor, and / or the like. Each RISC core may include any amount of memory. The RISC cores may use any of several protocols, depending on the embodiment. In some instances, the RISC cores may execute a real-time operating system (RTOS). The RISC cores may be implemented using one or more integrated circuit devices, application specific integrated circuits (ASICs), and / or memory devices. For example, the RISC cores may include an instruction cache and / or tightly coupled RAM.

[0085] The DMA may enable components of the PVA to access system memory independent of the CPU 506. The DMA may support any number of features used to provide optimizations to the PVA, including, but not limited to, supporting multi-dimensional addressing and / or circular addressing. In some instances, the DMA may support up to six or more dimensions of addressing, which may include block width, block height, block depth, horizontal block stepping, vertical block stepping, and / or depth stepping.

[0086] A vector processor may be a programmable processor that can be designed to efficiently and flexibly execute computer vision algorithm programming and provide signal processing capabilities. In some instances, a PVA may include a PVA core and two vector processing subsystem partitions. The PVA core may include a processor subsystem, a DMA engine (e.g., two DMA engines), and / or other peripherals. The vector processing subsystem may act as the PVA's primary processing engine and may include a vector processing unit (VPU), an instruction cache, and / or a vector memory (e.g., VMEM). The VPU core may include a digital signal processor, such as a single instruction, multiple data (SIMD), or very long instruction word (VLIW) digital signal processor. The combination of SIMD and VLIW can increase throughput and speed.

[0087] Each vector processor may include an instruction cache and may be coupled to dedicated memory. As a result, in some instances, each vector processor may be configured to execute independently of other vector processors. In other instances, the vector processors included in a particular PVA may be configured to employ data parallelism. For example, in some embodiments, multiple vector processors included in a single PVA may execute the same computer vision algorithm, but on different regions of an image. In other instances, the vector processors included in a particular PVA may execute different computer vision algorithms simultaneously on the same image, or even execute different algorithms on sequential images or portions of an image. In particular, any number of PVAs may be included in a hardware-accelerated cluster, and any number of vector processors may be included in each PVA. Additionally, the PVA may include additional error correcting code (ECC) memory to enhance overall system security.

[0088] The accelerator 514 (e.g., a hardware acceleration cluster) may include a computer vision network-on-chip and SRAM to provide high-bandwidth, low-latency SRAM for the accelerator 514. In some instances, the on-chip memory may include, for example, and without limitation, at least 4 MB of SRAM consisting of eight field-configurable memory blocks that may be accessible by both the PVA and DLA. Each pair of memory blocks may include an advanced peripheral bus (APB) interface, configuration circuitry, a controller, and a multiplexer. Any type of memory may be used. The PVA and DLA can access the memory through a backbone that provides the PVA and DLA with high-speed access to the memory. The backbone may include a computer vision network-on-chip that interconnects the PVA and DLA to the memory (e.g., using the APB).

[0089] The computer vision network-on-chip may include an interface that determines, prior to the transmission of any control signals, addresses, or data, that both the PVA and DLA provide ready and valid signals. Such an interface may provide separate phases and separate channels for transmitting control signals, addresses, and data, as well as burst-type communication for continuous data transfer. This type of interface may conform to the ISO 26262 or IEC 61508 standards, although other standards and protocols may also be used.

[0090] In some instances, SoC 504 may include a real-time ray tracing hardware accelerator, such as that described in U.S. patent application Ser. No. 16 / 101,232, filed Aug. 10, 2018. The real-time ray tracing hardware accelerator may be used to quickly and efficiently determine the location and scale of objects (e.g., within a world model) to generate real-time visualization simulations for RADAR signal interpretation, for acoustic propagation synthesis and / or analysis, for SONAR system simulation, for general wave propagation simulation, for comparison to LIDAR data for localization and / or other functions, and / or other uses. In some embodiments, one or more tree traversal units (TTUs) may be used to perform one or more ray tracing-related operations.

[0091] The accelerator 514 (e.g., a hardware accelerator cluster) has diverse applications for autonomous driving. The PVA may be a programmable vision accelerator that can be used for critical processing stages in ADAS and autonomous vehicles. The capabilities of the PVA make it well suited to algorithmic domains that require predictable processing at low power and low latency. In other words, the PVA works well for semi-dense or dense regular computations on small data sets that require predictable execution times with low latency and low power. Therefore, because the PVA is efficient at object detection and integer computation, in the context of a platform for autonomous vehicles, the PVA is designed to run classic computer vision algorithms.

[0092] For example, according to one embodiment of the present technology, PVA is used to perform computer stereo vision. A semi-global matching-based algorithm may be used in some instances, but this is not intended to be limiting. Many applications for Level 3-5 autonomous driving require motion estimation / stereo matching on the fly (e.g., structure from motion, pedestrian recognition, lane detection, etc.). PVA can perform computer stereo vision functions with input from two monocular cameras.

[0093] In some instances, PVA may be used to perform dense optical flow by processing raw RADAR data (e.g., using a 4D Fast Fourier Transform) to provide processed RADAR. In other instances, PVA is used in time of flight depth processing, for example, by processing raw time of flight data to provide processed time of flight data.

[0094] DLA can be used to implement any type of network to enhance control and driving safety, including, for example, a neural network that outputs a confidence measure for each object detection. Such a confidence value can be interpreted as a probability or as providing the relative "weight" of each detection compared to other detections. This confidence value allows the system to make further decisions regarding which detections should be considered true positives rather than false positives. For example, the system can set a confidence threshold and consider only detections above the threshold as true positives. In an automatic emergency braking (AEB) system, a false positive detection would cause a moving vehicle to automatically apply emergency braking, which is clearly undesirable. Therefore, only the most confident detections should be considered to trigger AEB. DLA can implement a neural network that regresses the confidence value. The neural network may receive as its inputs at least some subset of parameters, such as bounding box dimensions, ground plane estimates obtained (e.g., from another subsystem), inertial measurement unit (IMU) sensor 566 outputs that correlate with vehicle 500 orientation, range, and 3D position estimates of objects obtained from the neural network and / or other sensors (e.g., LIDAR sensor 564 or RADAR sensor 560), and others.

[0095] The SoC 504 may include a data store 516 (e.g., memory). The data store 516 may be on-chip memory of the SoC 504 and may store neural networks to be executed by the GPU and / or DLA. In some instances, the data store 516 may have a capacity large enough to store multiple instances of the neural network for redundancy and safety. The data store 516 may comprise an L2 or L3 cache 512. References to the data store 516 may include references to memory associated with the GPU, DLA, and / or other accelerators 514, as described herein.

[0096] The SoC 504 may include one or more processors 510 (e.g., embedded processors). The processors 510 may include a boot and power management processor, which may be a dedicated processor and subsystem for handling boot power and management capabilities and related security enforcement. The boot and power management processor may be part of the SoC 504 boot sequence and may provide run-time power management services. The boot power and management processor may provide clock and voltage programming, assist with system low-power state transitions, manage the SoC 504 thermal and temperature sensors, and / or manage the SoC 504 power state. Each temperature sensor may be implemented as a ring oscillator whose output frequency is proportional to temperature, and the SoC 504 may use the ring oscillator to detect the temperature of the CPU 506, GPU 508, and / or accelerator 514. If the temperature is determined to exceed a threshold, the boot and power management processor may enter a temperature fault routine, place the SoC 504 in a lower power state and / or place the vehicle 500 in a Chauffeur safe shutdown mode (e.g., bring the vehicle 500 to a safe shutdown).

[0097] The processor 510 may further include a set of embedded processors that can perform the functions of an audio processing engine. The audio processing engine may be an audio subsystem that allows full hardware support for multi-channel audio through multiple interfaces and a wide and flexible range of audio I / O interfaces. In some instances, the audio processing engine is a dedicated processor core that includes a digital signal processor with dedicated RAM.

[0098] The processor 510 may further include an always-on processor engine that can provide the necessary hardware features to support low-power sensor management and wake use cases. The always-on processor engine may include a processor core, tightly coupled RAM, support peripherals (e.g., timers and interrupt controllers), various I / O controller peripherals, and routing logic.

[0099] The processor 510 may further include a safety cluster engine that includes a processor subsystem dedicated to handling safety management for automotive applications. The safety cluster engine may include two or more processor cores, tightly coupled RAM, supporting peripherals (e.g., timers, interrupt controllers, etc.), and / or routing logic. In safety mode, the two or more cores may operate in lockstep mode and function as a single core with comparison logic to detect any differences between their operations.

[0100] The processor 510 may further include a real-time camera engine, which may include a dedicated processor subsystem for handling real-time camera management.

[0101] The processor 510 may further include a high dynamic range signal processor, which may include an image signal processor, which is a hardware engine that is part of the camera processing pipeline.

[0102] The processor 510 may include a video image compositor, which may be a processing block (e.g., implemented in a microprocessor) that implements video post-processing functions required by the video playback application to produce the final image for the player window. The video image compositor may perform lens distortion correction on the wide-view camera 570, the surround camera 574, and / or the in-cabin surveillance camera sensor. The in-cabin surveillance camera sensor is preferably monitored by a neural network running on a separate instance of the advanced SoC, configured to identify in-cabin events and respond appropriately. The in-cabin system may perform lip reading to activate cellular service and make phone calls, dictate emails, change the vehicle's destination, activate or change the vehicle's infotainment system and settings, or provide voice-activated web surfing. Certain features are available to the driver only when operating in autonomous mode and are disabled otherwise.

[0103] The video image combiner may include enhanced temporal noise reduction for both spatial and temporal noise reduction. For example, when motion occurs in the video, the noise reduction reduces the weight of information provided by adjacent frames and appropriately weights spatial information. When an image or portion of an image does not contain motion, the temporal noise reduction performed by the video image combiner can use information from previous images to reduce noise in the current image.

[0104] The video image compositor may also be configured to perform stereo rectification on the input stereo lens frames. The video image compositor may further be used for user interface compositing when the operating system desktop is in use, and the GPU 508 is not required to continuously render new surfaces. Even when the GPU 508 is powered on and actively performing 3D rendering, the video image compositor may be used to offload the GPU 508 to improve performance and responsiveness.

[0105] The SoC 504 may further include a mobile industry processor interface (MIPI) camera serial interface, a high-speed interface for receiving video and input from a camera, and / or a video input block that may be used for camera and related pixel input functions. The SoC 504 may further include an input / output controller that may be controlled by software and that may be used to receive I / O signals that are not committed to a specific role.

[0106] The SoC 504 may further include a wide range of peripheral interfaces to enable communication with peripherals, audio codecs, power management, and / or other devices. The SoC 504 may be used to process data from cameras (e.g., connected via gigabit multimedia serial links and Ethernet), sensors (e.g., LIDAR sensors 564, RADAR sensors 560, etc., which may be connected via Ethernet), data from the bus 502 (e.g., vehicle 500 speed, steering wheel position, etc.), and GNSS sensors 558 (e.g., connected via Ethernet or CAN bus). The SoC 504 may further include a dedicated high-performance mass storage controller, which may include its own DMA engine and may be used to offload routine data management tasks from the CPU 506.

[0107] The SoC 504 may be an end-to-end platform with a flexible architecture that spans levels 3-5 of automation, thereby providing a comprehensive functional safety architecture that leverages and efficiently uses computer vision and ADAS techniques for diversity and redundancy, and provides a platform for a flexible, reliable driving software stack along with deep learning tools. The SoC 504 may be faster, more reliable, and more energy- and space-efficient than conventional systems. For example, when the accelerator 514 is combined with the CPU 506, the GPU 508, and the data store 516, it can provide a fast and efficient platform for levels 3-5 of autonomous vehicles.

[0108] This technology therefore offers capabilities and functionality not achievable by conventional systems. For example, computer vision algorithms can be implemented on a central processing unit (CPU), which can be configured using a high-level programming language, such as the C programming language, to execute a wide variety of processing algorithms across a wide variety of visual data. However, CPUs often cannot meet the performance requirements of many computer vision applications, including those related to execution time and power consumption. Specifically, many CPUs cannot execute complex object detection algorithms in real time, a requirement for in-vehicle ADAS applications and practical Level 3-5 autonomous vehicles.

[0109] In contrast to conventional systems, by providing a CPU complex, a GPU complex, and a hardware acceleration cluster, the technology described herein allows multiple neural networks to run simultaneously and / or serially and the results to be combined to enable Level 3-5 autonomous driving capabilities. For example, a CNN running on the DLA or dGPU (e.g., GPU520) can include text and word recognition, enabling the supercomputer to read and understand traffic signs, including signs for which the neural network was not specifically trained. The DLA can further include a neural network that can identify, interpret, and provide a semantic understanding of the signs and pass the semantic understanding to a route planning module running on the CPU complex.

[0110] As another example, multiple neural networks may be run simultaneously, as required for Level 3, 4, or 5 operation. For example, a warning sign consisting of "Caution: Flashing lights indicate icy conditions" along with a lightning flash may be interpreted independently or collectively by several neural networks. The sign itself may be identified as a traffic sign by a first deployed neural network (e.g., a trained neural network), and the text "Flashing lights indicate icy conditions" may be interpreted by a second deployed neural network that notifies the vehicle's route planning software (preferably running on a CPU complex) that icy conditions exist when the flashing light is detected. The flashing light may be identified by running a third deployed neural network over multiple frames, informing the vehicle's route planning software of the presence (or absence) of the flashing light. All three neural networks may run simultaneously, such as within the DLA and / or on the GPU 508.

[0111] In some instances, a CNN for facial recognition and vehicle owner identification can use data from the camera sensor to identify the presence of a legitimate driver and / or owner of the vehicle 500. An always-on sensor processing engine can be used to unlock the vehicle and turn on the lights when the owner approaches the driver's side door, and in security mode, to disable operation of the vehicle when the owner leaves the vehicle. In this manner, the SoC 504 provides security against theft and / or vehicle hijacking.

[0112] In another example, a CNN for emergency vehicle detection and identification can detect and identify emergency vehicle sirens using data from microphone 596. In contrast to conventional systems that use general classifiers to detect sirens and manually extract features, SoC 504 uses CNNs for environmental and urban sound classification, as well as visual data classification. In a preferred embodiment, the CNN running on the DLA is trained to identify the relative terminal velocity of emergency vehicles (e.g., by using the Doppler effect). The CNN can also be trained to identify emergency vehicles specific to the local area in which the mobile vehicle is operating, as identified by GNSS sensor 558. Thus, for example, when operating in Europe, the CNN would attempt to detect European sirens, and when in the United States, the CNN would attempt to identify only North American sirens. After an emergency vehicle is detected, a control program can be used to perform emergency vehicle safety routines, such as slowing down the mobile vehicle, stopping it at the side of the road, parking it, and / or idling it, with the assistance of ultrasonic sensor 562, until the emergency vehicle has passed.

[0113] The vehicle may include a CPU 518 (e.g., a discrete CPU or dCPU) that may be coupled to the SoC 504 via a high-speed interconnect (e.g., PCIe). The CPU 518 may include, for example, an X86 processor. The CPU 518 may be used to perform any of a variety of functions, including, for example, reconciling potentially inconsistent results between the ADAS sensors and the SoC 504 and / or monitoring the status and health of the controller 536 and / or infotainment SoC 530.

[0114] Vehicle 500 may include a GPU 520 (e.g., a discrete GPU or dGPU) that may be coupled to SoC 504 via a high-speed interconnect (e.g., NVIDIA's NVLINK). GPU 520 may provide additional artificial intelligence functionality, such as by running redundant and / or different neural networks, and may be used to train and / or update neural networks based on input (e.g., sensor data) from sensors in vehicle 500.

[0115] The mobile vehicle 500 may further include a network interface 524, which may include one or more wireless antennas 526 (e.g., one or more wireless antennas for different communication protocols, such as a cellular antenna, a Bluetooth antenna, etc.). The network interface 524 may be used to enable wireless connections with the cloud via the Internet (e.g., with the server 578 and / or other network devices), with other mobile vehicles, and / or with computing devices (e.g., passenger client devices). To communicate with other mobile vehicles, a direct link may be established between the two mobile vehicles, and / or an indirect link may be established (e.g., through a network and via the Internet). A direct link may be provided using a vehicle-to-vehicle communication link. The vehicle-to-vehicle communication link may provide the mobile vehicle 500 with information about mobile vehicles in its vicinity (e.g., vehicles in front of, beside, and / or behind the mobile vehicle 500). This functionality may be part of a cooperative adaptive cruise control function of the mobile vehicle 500.

[0116] The network interface 524 may include an SoC that provides modulation and demodulation functions and enables the controller 536 to communicate over a wireless network. The network interface 524 may include a radio frequency front end for upconversion from baseband to radio frequency and downconversion from radio frequency to baseband. The frequency conversion may be performed through well-known processes and / or may be performed using a superheterodyne process. In some instances, the radio frequency front end functionality may be provided by a separate chip. The network interface may include wireless functionality for communicating via LTE, WCDMA, UMTS, GSM, CDMA2000, Bluetooth, Bluetooth LE, Wi-Fi, Z-Wave, Zigbee, LoRaWAN, and / or other wireless protocols.

[0117] Mobile vehicle 500 may further include a data store 528, which may include off-chip (e.g., off-SoC 504) storage. Data store 528 may include one or more memory elements, including RAM, SRAM, DRAM, VRAM, flash, hard disk, and / or other components and / or devices capable of storing at least one bit of data.

[0118] The vehicle 500 may further include a GNSS sensor 558. The GNSS sensor 558 (e.g., a GPS, an aided GPS sensor, a differential GPS (DGPS) sensor, etc.) aids in mapping, perception, occupancy grid generation, and / or route planning functions. Any number of GNSS sensors 558 may be used, including, for example, but not limited to, a GPS using a USB connector with an Ethernet to serial (RS-232) bridge.

[0119] The mobile vehicle 500 may further include a RADAR sensor 560. The RADAR sensor 560 may be used by the mobile vehicle 500 for long-range mobile vehicle detection, even in darkness and / or severe weather conditions. The RADAR functional safety level may be ASIL B. In some instances, the RADAR sensor 560 may use the CAN and / or bus 502 for control and to access object tracking data (e.g., to transmit data generated by the RADAR sensor 560), with access to Ethernet for accessing raw data. A wide variety of RADAR sensor types may be used. For example, and without limitation, the RADAR sensor 560 may be suitable for front, rear, and side RADAR use. In some instances, a pulse-Doppler RADAR sensor is used.

[0120] The RADAR sensor 560 may include different configurations, such as long-range with a narrow field of view, short-range with a wide field of view, and short-range side coverage. In some instances, long-range RADAR may be used for adaptive cruise control functions. Long-range RADAR systems may provide a wide field of view achieved by two or more independent scans, such as within a 250-meter range. The RADAR sensor 560 may help distinguish between static and moving objects and may be used by ADAS systems for emergency brake assist and forward collision warning. Long-range RADAR sensors may include monostatic multimodal RADAR with multiple (e.g., six or more) fixed RADAR antennas and high-speed CAN and FlexRay interfaces. In one example with six antennas, the center four antennas may create a focused beam pattern designed to record the surroundings of the moving vehicle 500 at high speeds with minimal interference from traffic in adjacent lanes. The other two antennas may widen the field of view, allowing for rapid detection of moving vehicles entering or leaving the lane of the moving vehicle 500.

[0121] As an example, a medium-range RADAR system may include a range of up to 560 meters (front) or 80 meters (rear) and a field of view of up to 42 degrees (front) or 550 degrees (rear). A short-range RADAR system may include, but is not limited to, a RADAR sensor designed to be mounted on either end of a rear bumper. When mounted on either end of a rear bumper, such a RADAR sensor system can create two beams that constantly monitor the blind spots behind and adjacent to a moving vehicle.

[0122] Short-range RADAR systems may be used in ADAS systems for blind spot detection and / or lane change assist.

[0123] The mobile vehicle 500 may further include ultrasonic sensors 562. The ultrasonic sensors 562, which may be positioned on the front, rear, and / or sides of the mobile vehicle 500, may be used for parking assistance and / or for creating and updating an occupancy grid. A variety of ultrasonic sensors 562 may be used, with different ultrasonic sensors 562 being used for different ranges of detection (e.g., 2.5 m, 4 m). The ultrasonic sensors 562 may operate at an ASIL B functional safety level.

[0124] The mobile vehicle 500 may include a LIDAR sensor 564. The LIDAR sensor 564 may be used for object and pedestrian detection, emergency braking, collision avoidance, and / or other functions. The LIDAR sensor 564 may be functional safety level ASIL B. In some instances, the mobile vehicle 500 may include multiple (e.g., two, four, six, etc.) LIDAR sensors 564 that can use Ethernet (e.g., to provide data to a Gigabit Ethernet switch).

[0125] In some instances, the LIDAR sensor 564 may be capable of providing a list of objects and their distances in a 360-degree field of view. Commercially available LIDAR sensors 564 may have an advertised range of approximately 500 m, with an accuracy of 2 cm to 3 cm, and support for a 500 Mbps Ethernet connection, for example. In some instances, one or more non-protruding LIDAR sensors 564 may be used. In such instances, the LIDAR sensor 564 may be implemented as a small device that may be integrated into the front, rear, sides, and / or corners of the vehicle 500. In such instances, the LIDAR sensor 564 may have a range of 200 m, even for low-reflecting objects, and provide up to a 120-degree horizontal and 35-degree vertical field of view. A front-mounted LIDAR sensor 564 may be configured for a horizontal field of view between 45 and 135 degrees.

[0126] In some implementations, LIDAR technology such as 3D flash LIDAR may also be used. 3D flash LIDAR uses a laser flash as a transmitter to illuminate the surroundings of the vehicle up to approximately 200 meters. The flash LIDAR unit includes a receptor that records the laser pulse transit time and the reflected light at each pixel, which in turn corresponds to the range from the vehicle to the object. Flash LIDAR may enable a highly accurate and distortion-free image of the surroundings to be generated with every laser flash. In some implementations, four flash LIDAR sensors may be deployed, one on each side of the vehicle 500. Available 3D flash LIDAR systems include solid-state 3D steering array LIDAR cameras (e.g., non-scanning LIDAR devices) with no moving parts other than the blower. Flash LIDAR devices may use 5 nanosecond Class I (eye-safe) laser pulses per frame and may capture reflected laser light in the form of a 3D range point cloud and coregistered intensity data. By using flash LIDAR, and because flash LIDAR is a solid-state device with no moving parts, the LIDAR sensor 564 may be less susceptible to motion blur, vibration, and / or shock.

[0127] The mobile vehicle may further include an IMU sensor 566. In some instances, the IMU sensor 566 may be positioned at the center of the rear axle of the mobile vehicle 500. The IMU sensor 566 may include, for example, but not limited to, an accelerometer, a magnetometer, a gyroscope, a magnetic compass, and / or other sensor types. In some instances, such as in a six-axis application, the IMU sensor 566 may include an accelerometer and a gyroscope, while in a nine-axis application, the IMU sensor 566 may include an accelerometer, a gyroscope, and a magnetometer.

[0128] In some embodiments, the IMU sensor 566 may be implemented as a miniature, high-performance GPS-Aided Inertial Navigation System (GPS / INS) that combines micro-electro-mechanical system (MEMS) inertial sensors, a highly sensitive GPS receiver, and advanced Kalman filtering algorithms to provide estimates of position, velocity, and attitude. As such, in some instances, the IMU sensor 566 may enable the vehicle 500 to estimate heading without requiring input from a magnetic sensor by directly observing and correlating changes in velocity from the GPS to the IMU sensor 566. In some instances, the IMU sensor 566 and the GNSS sensor 558 may be combined in a single integrated unit.

[0129] The mobile vehicle may include microphones 596 placed within and / or around the mobile vehicle 500. The microphones 596 may be used for emergency vehicle detection and identification, among other things.

[0130] The vehicle may further include any number of camera types, including stereo cameras 568, wide-view cameras 570, infrared cameras 572, surround cameras 574, long-range and / or mid-range cameras 598, and / or other camera types. The cameras may be used to capture image data around the entire exterior of the vehicle 500. The types of cameras used depend on the implementation and requirements of the vehicle 500, and any combination of camera types may be used to achieve the desired coverage around the vehicle 500. Additionally, the number of cameras may vary depending on the implementation. For example, the vehicle may include six cameras, seven cameras, ten cameras, twelve cameras, and / or another number of cameras. The cameras may support, by way of example only, Gigabit Multimedia Serial Link (GMSL) and / or Gigabit Ethernet. Each camera is described in further detail herein with reference to FIGS. 5A and 5B.

[0131] The vehicle 500 may further include a vibration sensor 542. The vibration sensor 542 may measure vibrations of vehicle components, such as an axle. For example, a change in vibration may indicate a change in the road surface. In another example, when two or more vibration sensors 542 are used, the difference in vibration may be used to determine friction or slippage of the road surface (e.g., when the difference in vibration is between a powered axle and a free-spinning axle).

[0132] The mobile vehicle 500 may include an ADAS system 538. In some instances, the ADAS system 538 may include an SoC. The ADAS system 538 may include autonomous / adaptive / automatic cruise control (ACC), cooperative adaptive cruise control (CACC), forward crash warning (FCW), automatic emergency braking (AEB), lane departure warning (LDW), lane keep assist (LKA), blind spot warning (BSW), rear cross-traffic warning (RCTW), collision warning system (CWS), lane centering (LC), and / or other features and functions.

[0133] The ACC system may use a RADAR sensor 560, a LIDAR sensor 564, and / or a camera. The ACC system may include longitudinal ACC and / or lateral ACC. The longitudinal ACC monitors and controls the distance to the vehicle directly ahead of the vehicle 500 and automatically adjusts the vehicle speed to maintain a safe distance from the vehicle ahead. The lateral ACC performs distance keeping and advises the vehicle 500 to change lanes when necessary. The lateral ACC is related to other ADAS applications such as LCA and CWS.

[0134] CACC uses information from other moving vehicles, which may be received from other moving vehicles via a wireless link via the network interface 524 and / or wireless antenna 526, or indirectly via a network connection (e.g., via the Internet). A direct link may be provided by a vehicle-to-vehicle (V2V) communication link, while an indirect link may be an infrastructure-to-vehicle (I2V) communication link. Generally, V2V communication concepts provide information about the immediately preceding moving vehicle (e.g., the moving vehicle directly ahead of the moving vehicle 500 that is in the same lane as the moving vehicle 500), while I2V communication concepts provide information about traffic further ahead. A CACC system may include either or both I2V and V2V information sources. Given information about moving vehicles ahead of the moving vehicle 500, CACC may be more reliable, potentially allowing for smoother traffic flow and reducing road congestion.

[0135] FCW systems are designed to warn the driver of hazards so that the driver can take corrective action. FCW systems use forward-facing cameras and / or RADAR sensors 560 coupled to dedicated processors, DSPs, FPGAs, and / or ASICs, electrically coupled to driver feedback such as displays, speakers, and / or vibration components. FCW systems can provide warnings in the form of audio, visual alerts, vibrations, and / or quick brake pulses.

[0136] An AEB system can detect an imminent forward collision with another moving vehicle or other object and automatically apply the brakes if the driver does not take corrective action within specified time or distance parameters. The AEB system can use a forward-facing camera and / or RADAR sensor 560 coupled to a dedicated processor, DSP, FPGA, and / or ASIC. When the AEB system detects a hazard, it typically first alerts the driver to take corrective action to avoid the collision; if the driver does not take corrective action, the AEB system can automatically apply the brakes as part of an effort to prevent, or at least mitigate, the effects of the predicted collision. The AEB system can include techniques such as dynamic brake support and / or collision imminent braking.

[0137] The LDW system provides visual, audible, and / or tactile warnings, such as vibration of the steering wheel or seat, to alert the driver when the mobile vehicle 500 crosses a lane marking. The LDW system does not activate when the driver indicates an intentional lane departure by activating a turn signal. The LDW system may use a forward-facing camera coupled to a dedicated processor, DSP, FPGA, and / or ASIC electrically coupled to driver feedback, such as a display, speaker, and / or vibration components.

[0138] The LKA system is a modification of the LDW system, which provides steering input or braking to correct the vehicle 500 if it begins to drift out of its lane.

[0139] The BSW system detects and warns the driver of a moving vehicle in the vehicle's blind spot. The BSW system can provide visual, audible, and / or tactile warnings to indicate that merging or changing lanes is unsafe. The system can provide additional warnings when the driver uses a turn signal. The BSW system can use a rear-facing camera and / or RADAR sensor 560 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, electrically coupled to driver feedback, e.g., a display, speaker, and / or vibration component.

[0140] The RCTW system can provide visual, audible, and / or tactile notifications when an object is detected outside the range of the rear camera when the vehicle 500 is backing up. Some RCTW systems include AEB to ensure vehicle brakes are applied to avoid a collision. The RCTW system can use one or more rear-facing RADAR sensors 560 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, electrically coupled to driver feedback, e.g., a display, speaker, and / or vibration components.

[0141] Because conventional ADAS systems alert the driver and allow the driver to determine whether a safety condition truly exists and act accordingly, conventional ADAS systems can be prone to producing false positives that, while not usually catastrophic, can be annoying and distracting to the driver. However, in an autonomous vehicle 500, when results conflict, the vehicle 500 itself must decide whether to listen to results from a primary computer or a secondary computer (e.g., the first controller 536 or the second controller 536). For example, in some embodiments, the ADAS system 538 may be a backup and / or secondary computer that provides perception information to a backup computer rationality module. The backup computer rationality monitor can run redundant software on hardware components to detect failures in perception and dynamic driving tasks. Output from the ADAS system 538 may be provided to a supervisory MCU. When the outputs from the primary and secondary computers conflict, the supervisory MCU must decide how to reconcile the conflict to ensure safe operation.

[0142] In some instances, the primary computer may be configured to provide a reliability score to the supervising MCU indicating the reliability of the primary computer in a selected outcome. If the reliability score exceeds a threshold, the supervising MCU may follow the primary computer's instructions regardless of whether the secondary computers provide conflicting or inconsistent results. If the reliability score does not meet the threshold, and the primary and secondary computers provide different (e.g., conflicting) results, the supervising MCU may arbitrate between the computers to determine the appropriate outcome.

[0143] The supervisory MCU may be configured to execute a neural network trained and configured to determine, based on outputs from the primary and secondary computers, conditions under which the secondary computer will provide a false alarm. Thus, the neural network in the supervisory MCU can learn when the output of the secondary computer can be trusted and when it cannot be trusted. For example, when the secondary computer is a RADAR-based FCW system, the neural network in the supervisory MCU can learn when the FCW identifies a metal object that is not actually dangerous, such as a sewer grate or manhole cover, which triggers an alarm. Similarly, when the secondary computer is a camera-based LDW system, the neural network in the supervisory MCU can learn to ignore the LDW when a bicyclist or pedestrian is present and lane departure is, in fact, the safest maneuver. In embodiments including a neural network running on the supervisory MCU, the supervisory MCU may include at least one of a DLA or a GPU suitable for executing the neural network with associated memory. In a preferred embodiment, the supervising MCU may comprise and / or be included as a component of the SoC 504 .

[0144] In other instances, the ADAS system 538 may include a secondary computer that performs ADAS functions using traditional rules of computer vision. As such, the secondary computer may use classical computer vision rules (if-then), and the presence of a neural network in the supervisory MCU may improve reliability, safety, and performance. For example, diverse implementations and intentional non-identity may make the overall system more fault-tolerant, particularly to failures caused by software (or software-hardware interface) functions. For example, if a software bug or error exists in software running on the primary computer and non-identical software code running on the secondary computer provides the same overall result, the supervisory MCU may have greater confidence that the overall result is correct and that a bug in the software or hardware on the primary computer has not caused a critical error.

[0145] In some instances, the output of the ADAS system 538 can be fed to the perception block of the primary computer and / or the dynamic driving task block of the primary computer. For example, if the ADAS system 538 indicates a forward collision warning due to an object directly ahead, the perception block can use this information when identifying the object. In other instances, the secondary computer can have its own neural network that is trained as described herein, thus reducing the risk of false positives.

[0146] The mobile vehicle 500 may further include an infotainment SoC 530 (e.g., an in-vehicle infotainment system (IVI)). Although shown and described as an SoC, the infotainment system need not be an SoC and may include two or more separate components. The infotainment SoC 530 may include a combination of hardware and software that may be used to provide audio (e.g., music, personal digital assistants, navigation instructions, news, radio, etc.), video (e.g., TV, movies, streaming, etc.), telephony (e.g., hands-free calling), network connectivity (e.g., LTE, Wi-Fi, etc.), and / or information services (e.g., navigation system, reverse parking assist, wireless data system, vehicle-related information such as fuel level, total distance traveled, brake fuel level, oil level, door opening / closing, air filter information, etc.) to the mobile vehicle 500. For example, the infotainment SoC 530 may be a radio, a disc player, a navigation system, a video player, USB and Bluetooth® connectivity, a car computer, in-car entertainment, Wi-Fi®, steering wheel audio controls, hands-free voice control, a heads-up display (HUD), an HMI display 534, a telematics device, a control panel (e.g., for controlling and / or interacting with various components, features, and / or systems), and / or other components. The infotainment SoC 530 may further be used to provide information (e.g., visual and / or audible) to a user of the vehicle, such as information from an ADAS system 538, autonomous driving information such as planned vehicle maneuvers, trajectory, surrounding environment information (e.g., intersection information, vehicle information, road information, etc.), and / or other information.

[0147] The infotainment SoC 530 may include GPU functionality. The infotainment SoC 530 may communicate with other devices, systems, and / or components of the mobile vehicle 500 via the bus 502 (e.g., CAN bus, Ethernet, etc.). In some instances, the infotainment SoC 530 may be coupled to the supervisory MCU so that the infotainment system's GPU can perform some self-driving functions in the event of a failure of the primary controller 536 (e.g., the primary and / or backup computer of the mobile vehicle 500). In such instances, the infotainment SoC 530 may place the mobile vehicle 500 in a Chauffeur safe stop mode, as described herein.

[0148] The mobile vehicle 500 may further include an instrument cluster 532 (e.g., a digital dash, an electronic instrument cluster, a digital instrument panel, etc.). The instrument cluster 532 may include a controller and / or a supercomputer (e.g., a separate controller or supercomputer). The instrument cluster 532 may include a set of instruments such as a speedometer, fuel level, oil pressure, a tachometer, an odometer, turn signals, a gear shift position indicator, a seat belt warning light, a parking brake warning light, an engine malfunction light, an airbag (SRS) system information, lighting controls, safety system controls, navigation information, etc. In some instances, information may be displayed and / or shared between the infotainment SoC 530 and the instrument cluster 532. In other words, the instrument cluster 532 may be included as part of the infotainment SoC 530, or vice versa.

[0149] 5D is a system diagram of communication between the cloud-based server and the example autonomous vehicle 500 of FIG. 5A in accordance with some embodiments of the present disclosure. System 576 may include a server 578, a network 590, and a mobile vehicle including the mobile vehicle 500. Server 578 may include multiple GPUs 584(A)-584(H) (collectively referred to herein as GPUs 584), PCIe switches 582(A)-582(H) (collectively referred to herein as PCIe switches 582), and / or CPUs 580(A)-580(B) (collectively referred to herein as CPUs 580). GPUs 584, CPUs 580, and PCIe switches may be interconnected with a high-speed interconnect, such as, but not limited to, an NVLink interface 588 developed by NVIDIA and / or a PCIe connection 586. In some instances, the GPUs 584 are connected via NVLink and / or NVSwitch SoCs, and the GPUs 584 and PCIe switches 582 are connected via PCIe interconnects. While eight GPUs 584, two CPUs 580, and two PCIe switches are illustrated, this is not intended to be limiting. Depending on the embodiment, each server 578 may include any number of GPUs 584, CPUs 580, and / or PCIe switches. For example, the servers 578 may each include 8, 16, 32, and / or more GPUs 584.

[0150] Server 578 can receive image data from the mobile vehicles over network 590, representing images showing unexpected or changed road conditions, such as recently begun road construction. Server 578 can transmit neural network 592, updated neural network 592, and / or map information 594, including information about traffic and road conditions, to the mobile vehicles over network 590. Updates to map information 594 can include updates to HD map 522, such as information about construction sites, potholes, detours, flooding, and / or other obstacles. In some instances, neural network 592, updated neural network 592, and / or map information 594 can result from new training and / or experience represented in data received from any number of mobile vehicles in the environment and / or based on training performed at a data center (e.g., using server 578 and / or other servers).

[0151] The server 578 may be used to train a machine learning model (e.g., a neural network) based on training data. The training data may be generated by a mobile vehicle and / or generated in a simulation (e.g., using a game engine). In some instances, the training data is tagged (e.g., if the neural network benefits from supervised learning) and / or undergoes other preprocessing, while in other instances, the training data is not tagged and / or preprocessed (e.g., if the neural network does not require supervised learning). The training may be performed according to any one or more classes of machine learning techniques, including, but not limited to, the following classes: supervised training, semi-supervised training, unsupervised training, self-learning, reinforcement learning, federated learning, transfer learning, feature learning (including principal component and cluster analysis), multilinear subspace learning, manifold learning, representation learning (including preliminary dictionary learning), rule-based machine learning, anomaly detection, and variations or combinations thereof. After the machine-learned model is traced, it may be used by the vehicle (e.g., transmitted to the vehicle via network 590) and / or it may be used by server 578 to remotely monitor the vehicle.

[0152] In some instances, server 578 can receive data from mobile vehicles and apply the data to state-of-the-art real-time neural networks for real-time intelligent inference. Server 578 can include deep learning supercomputers and / or dedicated AI computers powered by GPUs 584, such as the DGX and DGX Station machines developed by NVIDIA. However, in some instances, server 578 can include deep learning infrastructure that uses only CPU-powered data centers.

[0153] The deep learning infrastructure of server 578 may be capable of rapid real-time inference and may use that capability to evaluate and verify the health of the processor, software, and / or associated hardware within mobile vehicle 500. For example, the deep learning infrastructure may receive periodic updates from mobile vehicle 500 (e.g., via computer vision and / or other machine learning object classification techniques), such as a sequence of images and / or objects where mobile vehicle 500 was located within the sequence of images. The deep learning infrastructure may run its own neural network to identify objects and compare them to objects identified by mobile vehicle 500; if the results are inconsistent and the infrastructure concludes that the AI ​​within mobile vehicle 500 is not functioning properly, server 578 may send a signal to mobile vehicle 500 instructing its failsafe computer to take control, notify passengers, and complete a safe parking maneuver.

[0154] For inference, the server 578 may include a GPU 584 and one or more programmable inference accelerators (e.g., NVIDIA's TensorRT). The combination of a GPU-powered server and inference acceleration can enable real-time responsiveness. In other instances, such as when less performance is required, servers powered by CPUs, FPGAs, and other processors may be used for inference.

[0155] Exemplary Computing Device 6 is a block diagram of an example computing device 600 suitable for use in implementing some embodiments of the present disclosure. Computing device 600 may include an interconnection system 602 that indirectly or directly couples the following devices: memory 604, one or more central processing units (CPUs) 606, one or more graphics processing units (GPUs) 608, a communication interface 610, input / output (I / O) ports 612, input / output components 614, a power supply 616, one or more presentation components 618 (e.g., displays), and one or more logic units 620. In at least one embodiment, computing device 600 may include one or more virtual machines (VMs), and / or any of its components may include virtual components (e.g., virtual hardware components). As non-limiting examples, one or more of GPUs 608 may include one or more vGPUs, one or more of CPUs 606 may include one or more vCPUs, and / or one or more of logical units 620 may include one or more virtual logical units. As such, computing device 600 may include discrete components (e.g., an entire GPU dedicated to computing device 600), virtual components (e.g., a portion of a GPU dedicated to computing device 600), or a combination thereof.

[0156] While the various blocks in FIG. 6 are depicted as connected by lines via interconnection system 602, this is not intended to be limiting and is merely for clarity. For example, in some embodiments, a presentation component 618, such as a display device, may be considered an I / O component 614 (e.g., if the display is a touch screen). As another example, CPU 606 and / or GPU 608 may include memory (e.g., memory 604 may represent a storage device in addition to the memory of GPU 608, CPU 606, and / or other components). In other words, the computing devices in FIG. 6 are merely exemplary. Categories such as “workstation,” “server,” “laptop,” “desktop,” “tablet,” “client device,” “mobile device,” “handheld device,” “gaming console,” “electronic control unit (ECU),” “virtual reality system,” and / or other device or system types are all intended to be within the scope of the computing devices in FIG. 6 and therefore will not be distinguished from one another.

[0157] Interconnect system 602 may represent one or more links or buses, such as an address bus, a data bus, a control bus, or a combination thereof. Interconnect system 602 may include one or more bus or link types, such as an industry standard architecture (ISA) bus, an extended industry standard architecture (EISA) bus, a video electronics standards association (VESA) bus, a peripheral component interconnect (PCI) bus, a peripheral component interconnect express (PCIe) bus, and / or another type of bus or link. In some embodiments, direct connections exist between components. As an example, CPU 606 may be directly connected to memory 604. Further, CPU 606 may be directly connected to GPU 608. When direct or point-to-point connections exist between components, interconnect system 602 may include a PCIe link to implement the connections. In these examples, a PCI bus need not be included in computing device 600.

[0158] Memory 604 may include any of a variety of computer-readable media. Computer-readable media may be any available media that can be accessed by computing device 600. Computer-readable media may include both volatile and nonvolatile media, and removable and non-removable media. By way of example, and not limitation, computer-readable media may comprise computer storage media and communication media.

[0159] Computer storage media may include both volatile and nonvolatile media, and / or removable and non-removable media implemented in any method or technology for storage of information, such as computer-readable instructions, data structures, program modules, and / or other data types. For example, memory 604 may store computer-readable instructions (e.g., representing programs and / or program elements), such as an operating system. Computer storage media may include, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and that can be accessed by computing device 600. As used herein, computer storage media does not include the signals themselves.

[0160] Computer storage media may embody computer-readable instructions, data structures, program modules, and / or other data types in a modulated data signal such as a carrier wave or other transport mechanism and include any information delivery media. The term "modulated data signal" may refer to a signal that has one or more of its characteristics set or changed in such a manner as to encode information in the signal. By way of example, and not limitation, computer storage media may include wired media such as a wired network or direct-wired connection, and wireless media such as acoustic, RF, infrared and other wireless media. Combinations of any of the above should also be included within the scope of computer-readable media.

[0161] The CPU 606 may be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 600 to perform one or more of the methods and / or processes described herein. The CPU 606 may include one or more (e.g., 1, 2, 4, 8, 28, 72, etc.) cores, each capable of simultaneously processing multiple software threads. The CPU 606 may include any type of processor, and may include different types of processors depending on the type of computing device 600 implemented (e.g., a processor with fewer cores for a mobile device and a processor with more cores for a server). For example, depending on the type of computing device 600, the processor may be an Advanced RISC Machines (ARM) processor implemented using Reduced Instruction Set Computing (RISC) or an x86 processor implemented using Complex Instruction Set Computing (CISC). Computing device 600 may include one or more CPUs 606 within one or more microprocessors or auxiliary coprocessors, such as computational coprocessors.

[0162] In addition to or instead of CPU 606, GPU 608 may be configured to execute at least some of the computer-readable instructions to control one or more components of computing device 600 to perform one or more of the methods and / or processes described herein. One or more of GPUs 608 may be integrated GPUs (e.g., with one or more of CPUs 606 and / or one or more of GPUs 608 may be discrete GPUs. In an embodiment, one or more of GPUs 608 may be coprocessors of one or more of CPUs 606. GPU 608 may be used by computing device 600 to render graphics (e.g., 3D graphics) or perform general-purpose computing. For example, GPU 608 may be used with GPGPU (General-Purpose Computing on a GPU) The GPU 608 may be used for graphics processing (GPU). The GPU 608 may include hundreds or thousands of cores capable of processing hundreds or thousands of software threads simultaneously. The GPU 608 may generate pixel data for an output image in response to rendering commands (e.g., rendering commands from the CPU 606 received via a host interface). The GPU 608 may include graphics memory, e.g., display memory, for storing pixel data or any other suitable data, e.g., GPGPU data. The display memory may be included as part of the memory 604. GPU 608 may include two or more GPUs operating in parallel (e.g., via links). The links may connect the GPUs directly (e.g., using NVLINK) or may connect the GPUs via a switch (e.g., using NVSwitch). When coupled together, each GPU 608 may generate pixel data or GPGPU data for a different portion of the output or for a different output (e.g., a first GPU for a first image and a second GPU for a second image). Each GPU may include its own memory or may share memory with other GPUs.

[0163] In addition to or instead of CPU 606 and / or GPU 608, logic unit 620 may be configured to execute at least some of the computer-readable instructions to control one or more of computing devices 600 to perform one or more of the methods and / or processes described herein. In an embodiment, CPU 606, GPU 608, and / or logic unit 620 may discretely or jointly execute any combination of methods, processes, and / or portions thereof. One or more of logic units 620 may be part of and / or integrated with one or more of CPU 606 and / or GPU 608, and / or one or more of logic units 620 may be discrete components to or otherwise external to CPU 606 and / or GPU 608. In an embodiment, one or more of logic units 620 may be a coprocessor of one or more of CPU 606 and / or GPU 608.

[0164] Examples of logic unit 620 include one or more processing cores and / or components thereof, such as a tensor core (TC), a tensor processing unit (TPU), a pixel visual core (PVC), a vision processing unit (VPU), a graphics processing cluster (GPC), a texture processing cluster (TPC), a streaming multiprocessor (SM), a tree traversal unit (TTU), an artificial intelligence accelerator (AIA), a deep learning accelerator (DLA), an arithmetic logic unit (ALU), an application specific integrated circuit (ASIC), a floating point unit (FPU), an input / output (I / O) element, a peripheral component interconnect (PCI) or peripheral component interconnect express (PCIe) element, and / or the like.

[0165] The communications interface 610 may include one or more receivers, transmitters, and / or transceivers that enable the computing device 600 to communicate with other computing devices over electronic communications networks, including wired and / or wireless communications. The communications interface 610 may include components and functionality to enable communication over any of several different networks, such as a wireless network (e.g., Wi-Fi, Z-Wave, Bluetooth, Bluetooth LE, Zigbee, etc.), a wired network (e.g., communicating over Ethernet or InfiniBand), a low-power wide area network (e.g., LoRaWAN, SigFox, etc.), and / or the Internet.

[0166] The I / O ports 612 may enable the computing device 600 to be logically coupled to other devices, including I / O components 614, presentation components 618, and / or other components, some of which may be built into (e.g., integrated with) the computing device 600. Exemplary I / O components 614 include a microphone, mouse, keyboard, joystick, game pad, game controller, satellite dish, scanner, printer, wireless device, etc. The I / O components 614 may provide a natural user interface (NUI) that processes air gestures, voice, or other physiological input generated by the user. In some cases, the input may be sent to an appropriate network element for further processing. The NUI may implement any combination of voice recognition, stylus recognition, facial recognition, biometric recognition, on-screen and adjacent-screen gesture recognition, air gestures, head and eye tracking, and touch recognition in connection with the display of the computing device 600 (as described in more detail below). Computing device 600 may include a depth camera, such as a stereoscopic camera system, an infrared camera system, an RGB camera system, touch screen technology, and combinations thereof, for gesture detection and recognition. Additionally, computing device 600 may include an accelerometer or gyroscope (e.g., as part of an inertia measurement unit (IMU)) to enable detection of movement. In some instances, the output of the accelerometer or gyroscope may be used by computing device 600 to render immersive augmented or virtual reality.

[0167] The power supply 616 may include a hardwired power supply, a battery power supply, or a combination thereof. The power supply 616 may provide power to the computing device 600 to enable the components of the computing device 600 to operate.

[0168] The presentation component 618 may include a display (e.g., a monitor, a touch screen, a television screen, a heads-up display (HUD), other display types, or a combination thereof), speakers, and / or other presentation components. The presentation component 618 can receive data from other components (e.g., GPU 608, CPU 606, etc.) and output data (e.g., as images, video, sound, etc.).

[0169] Exemplary Data Center 7 illustrates an example data center 700 that may be used in at least one embodiment of the present disclosure. The data center 700 may include a data center infrastructure layer 710, a framework layer 720, a software layer 730, and / or an application layer 740.

[0170] 7, the data center infrastructure layer 710 may include a resource orchestrator 712, grouped computational resources 714, and node computational resources (“node CRs”) 716(1) through 716(N), where “N” represents any integer, natural number. In at least one embodiment, the node CRs 716(1) through 716(N) may include, but are not limited to, any number of central processing units (“CPUs”) or other processors (including accelerators, field programmable gate arrays (FPGAs), graphics processors or graphics processing units (GPUs), etc.), memory devices (e.g., dynamic read-only memory), storage devices (e.g., solid-state or disk drives), network input / output (“NW I / O”) devices, network switches, virtual machines (“VMs”), power modules, and / or cooling modules. In some embodiments, one or more of the nodes CR716(1)-716(N) may correspond to a server having one or more of the aforementioned computing resources. Additionally, in some embodiments, the nodes CR716(1)-716(N) may include one or more virtual components, such as a vGPU, a vCPU, and / or the like, and / or one or more of the nodes CR716(1)-716(N) may correspond to a virtual machine (VM).

[0171] In at least one embodiment, the grouped computing resources 714 may include separate groups of nodes CR716 housed within one or more racks (not shown), or multiple racks housed in data centers in various geographic locations (also not shown). The separate groups of nodes CR716 within the grouped computing resources 714 may include grouped computing, network, memory, or storage resources that can be configured or assigned to support one or more workloads. In at least one embodiment, several nodes CR716 including CPUs, GPUs, and / or other processors may be grouped within one or more racks to provide computing resources to support one or more workloads. The one or more racks may also include any number of power modules, cooling modules, and / or network switches, in any combination.

[0172] The resource orchestrator 722 can configure or otherwise control one or more nodes CR 716(1)-716(N) and / or grouped computational resources 714. In at least one embodiment, the resource orchestrator 722 can include a software design infrastructure (“SDI”) management entity of the data center 700. The resource orchestrator 722 can include hardware, software, or some combination thereof.

[0173] In at least one embodiment, as shown in FIG. 7 , framework layer 720 may include a job scheduler 732, a configuration manager 734, a resource manager 736, and / or a distributed file system 738. Framework layer 720 may include a framework to support software 732 in software layer 730 and / or one or more applications 742 in application layer 740. Software 732 or applications 742 may include web-based service software or applications, such as those offered by Amazon Web Services, Google Cloud, and Microsoft Azure, respectively. Framework layer 720 may be a type of free and open source software web application framework, such as, but not limited to, Apache Spark™ (hereinafter “Spark”), which may use distributed file system 738 for large-scale data processing (e.g., “big data”). In at least one embodiment, job scheduler 732 may include a Spark driver to facilitate scheduling of workloads supported by various tiers of data center 700. The configuration manager 734 may be capable of configuring different layers, for example, the software layer 730 and the framework layer 720, which includes Spark and a distributed file system 738 to support large-scale data processing. The resource manager 736 may be capable of managing clustered or grouped computing resources that are mapped or allocated to support the distributed file system 738 and the job scheduler 732. In at least one embodiment, the clustered or grouped computing resources may include the computing resources 714 grouped in the data center infrastructure layer 710. The resource manager 1036 may coordinate with the resource orchestrator 712 to manage these mapped or allocated computing resources.

[0174] In at least one embodiment, software 732 included in software layer 730 may include software used by at least a portion of nodes CR 716(1)-716(N), grouped computational resources 714, and / or distributed file system 738 of framework layer 720. The one or more types of software may include, but are not limited to, internet web page searching software, email virus scanning software, database software, and streaming video content software.

[0175] In at least one embodiment, the applications 742 included in the application layer 740 may include one or more types of applications used by at least a portion of the nodes CR 716(1)-716(N), the grouped computational resources 714, and / or the distributed file system 738 of the framework layer 720. The one or more types of applications may include, but are not limited to, any number of genomics applications, cognitive computing, and machine learning applications, including training or inference software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), and / or other machine learning applications used in conjunction with one or more embodiments.

[0176] In at least one embodiment, any of configuration manager 734, resource manager 736, and resource orchestrator 712 can implement any number and type of self-modifying actions based on any amount and type of data obtained in any technically possible manner. The self-modifying actions can free data center operators of data center 700 from making potentially poor configuration decisions and possibly avoiding underutilized and / or underperforming portions of the data center.

[0177] Data center 700 may include tools, services, software, or other resources for training one or more machine learning models or for predicting or inferring information using one or more machine learning models according to one or more embodiments described herein. For example, a machine learning model may be trained by calculating weight parameters according to a neural network architecture using the software and / or computing resources described above with respect to data center 700. In at least one embodiment, a trained or deployed machine learning model corresponding to one or more neural networks may be used to infer or predict information using the resources described above with respect to data center 700, for example, by using weight parameters calculated via one or more training techniques, including but not limited to those described herein.

[0178] In at least one embodiment, data center 700 may use CPUs, application specific integrated circuits (ASICs), GPUs, FPGAs, and / or other hardware (or corresponding virtual computing resources) for training and / or performing inference using such resources. Additionally, one or more of such software and / or hardware resources may be configured as services, such as image recognition, speech recognition, or other artificial intelligence services, to enable users to train or perform inference on information.

[0179] Example Network Environment A network environment suitable for use in implementing embodiments of the present disclosure may include one or more client devices, servers, network attached storage (NAS), other back-end devices, and / or other device types. The client devices, servers, and / or other device types (e.g., each device) may be implemented with one or more instances of computing device 600 of FIG. 6, e.g., each device may include similar components, features, and / or functionality of computing device 600. Additionally, if a back-end device (e.g., server, NAS, etc.) is implemented, the back-end device may be included as part of data center 700, examples of which are further detailed herein with respect to FIG. 7.

[0180] Components of a network environment may communicate with each other via a network, which may be wired, wireless, or both. A network may include multiple networks or a network of networks. Illustratively, a network may include one or more wide area networks (WANs), one or more local area networks (LANs), one or more public networks, such as the Internet and / or the Public Switched Telephone Network (PSTN), and / or one or more private networks. When a network includes a wireless telecommunications network, components such as base stations, communication towers, or access points (as well as other components) may provide wireless connectivity.

[0181] Compatible network environments may include one or more peer-to-peer network environments (wherein a server may not be included in the network environment) and one or more client-server network environments (wherein a server or servers may be included in the network environment). In a peer-to-peer network environment, functionality described herein with respect to a server may be implemented in any number of client devices.

[0182] In at least one embodiment, the network environment may include one or more cloud-based network environments, distributed computing environments, combinations thereof, etc. The cloud-based network environment may include a framework layer, a job scheduler, a resource manager, and a distributed file system implemented on one or more of the servers, which may include one or more core network servers and / or edge servers. The framework layer may include a framework to support software in the software layer and / or one or more applications in the application layer. The software or applications may each include web-based service software or applications. In an embodiment, one or more of the client devices may use the web-based service software or applications (e.g., by accessing the service software and / or applications via one or more application programming interfaces (APIs)). The framework layer may be, but is not limited to, a type of free and open source software web application framework that may use a distributed file system for large-scale data processing (e.g., “big data”).

[0183] A cloud-based network environment may provide cloud computing and / or cloud storage that implements any combination of the computing and / or data storage functions (or one or more portions thereof) described herein. Any of these various functions may be distributed across multiple locations from a central or core server (e.g., one or more data centers that may be distributed across a state, region, country, or the world). When a user (e.g., a client device) is connected relatively close to an edge server, the core server may delegate at least a portion of its functionality to the edge server. A cloud-based network environment may be private (e.g., limited to a single organization), public (e.g., available to multiple organizations), and / or a combination thereof (e.g., a hybrid cloud environment).

[0184] A client device may include at least some of the components, features, and functionality of the exemplary computing device 600 described herein with respect to Figure 6. By way of illustration, and not limitation, a client device may be embodied as a personal computer (PC), a laptop computer, a mobile device, a smartphone, a tablet computer, a smart watch, a wearable computer, a personal digital assistant (PDA), an MP3 player, a virtual reality headset, a global positioning system (GPS) or device, a video player, a video camera, a surveillance device or system, a vehicle, a boat, an airship, a virtual machine, a drone, a robot, a handheld communications device, a hospital device, a gaming device or system, an entertainment system, a vehicle computing system, an embedded system controller, a remote control, an appliance, a consumer electronic device, a workstation, an edge device, any combination of these depicted devices, or any other suitable device.

[0185] The present disclosure may be described in the general context of computer code or machine-usable instructions, including computer-executable instructions, such as program modules, being executed by a computer or other machine, such as a personal digital assistant or other handheld device. Generally, program modules, including routines, programs, objects, components, data structures, etc., refer to code that performs particular tasks or implements particular abstract data types. The present disclosure may be implemented in a variety of configurations, including handheld devices, consumer electronics, general-purpose computers, more specialized computing devices, etc. The present disclosure may also be implemented in distributed computing environments where tasks are performed by remote processing devices linked through a communications network.

[0186] As used herein, the term "and / or" in reference to two or more elements should be interpreted to mean one element only or a combination of elements. For example, "element A, element B, and / or element C" may include element A only, element B only, element C only, elements A and B, elements A and C, elements B and C, or elements A, B, and C. Additionally, "at least one of element A or element B" may include at least one of element A, at least one of element B, or at least one of element A and at least one of element B. Furthermore, "at least one of element A and element B" may include at least one of element A, at least one of element B, or at least one of element A and at least one of element B.

[0187] The subject matter of the present disclosure has been described with specificity to meet statutory requirements. However, that description itself is not intended to limit the scope of the disclosure. Rather, the inventors contemplate that the claimed subject matter may be implemented in other ways, including different steps or combinations of steps similar to those described herein, in conjunction with other current or future technologies. Furthermore, although the terms "step" and / or "block" may be used herein to connote different elements of the method used, these terms should not be construed as implying any particular order among the various steps disclosed herein unless and when the order of individual steps is explicitly described.

Claims

1. determining a gaze direction of an occupant based at least in part on first sensor data generated using one or more first sensors of the vehicle; generating a three-dimensional representation of the occupant's field of view in a world space coordinate system based on the gaze direction; determining an object position of at least one object in the world coordinate system and based at least in part on second sensor data generated using one or more second sensors of the vehicle; comparing the representation to the object location and weighting the representation based on the portion of the representation that overlaps with the object location to determine whether a threshold amount of overlap between the representation and the object location is met; performing one or more actions based at least in part on the comparison; A method comprising:

2. performing the one or more operations suppressing notifications when the representation overlaps the object location by more than the threshold amount; or generating said notification when said representation does not overlap said object location by more than said threshold amount; The method of claim 1 , comprising performing one of the following steps:

3. 2. The method of claim 1, wherein the one or more first sensors include at least one first sensor having the view of the occupant inside the vehicle, and the one or more second sensors include at least one second sensor having a view outside the vehicle.

4. monitoring eye movements of the occupant of the vehicle to determine one or more of gaze patterns, saccade velocity, fixations, or smooth pursuit; determining an attentiveness score for the occupant based at least in part on one or more of the gaze pattern, the saccade velocity, the gaze behavior, or the smooth pursuit behavior; further comprising The method of claim 1 , wherein the performing the one or more actions is further based at least in part on the attentiveness score.

5. generating a heat map corresponding to the road scanning behavior of the occupant of the vehicle over a period of time; further comprising performing the one or more actions is further based at least in part on the heat map. The method of claim 1.

6. determining a cognitive load score for the vehicle occupant based at least in part on at least one of eye movements, eye measurements, or eye characteristics of the occupant; further comprising performing the one or more actions further based at least in part on the cognitive load score. The method of claim 1.

7. 7. The method of claim 6, wherein the determining the cognitive load score is based at least in part on a cognitive load profile corresponding to the occupant, the cognitive load profile being generated during one or more drives involving the occupant.

8. 2. The method of claim 1 , wherein the determining the object location comprises applying the second sensor data to one or more deep neural networks (DNNs) configured to calculate data indicative of the object location.

9. determining one or more of a posture, a gesture, or an activity being performed by the occupant; further comprising the performing the one or more actions is further based at least in part on one or more of the posture, the gesture, or the activity. The method of claim 1.

10. A step of generating a three-dimensional representation projecting at least a portion of the user's field of view based on the user's direction of gaze, in a coordinate system and based at least in part on first sensor data generated using one or more first sensors of the vehicle; determining an object position of at least one object in the coordinate system and based at least in part on second sensor data generated using one or more second sensors of the vehicle; weighting the representation based on the portion of the representation that overlaps with the object location to determine whether a threshold amount of overlap between the representation and the object location is met; the representation overlaps the object location by less than a first threshold amount; the cognitive load value is greater than a second threshold amount; or Attention value is less than a third threshold amount determining to generate a notification based at least in part on at least one of A method comprising:

11. generating the representation, determining the gaze direction of the user; determining at least the portion of the field of view based at least in part on the direction of gaze of the user; The method of claim 10, comprising:

12. The method of claim 10 , wherein the coordinate system is a three-dimensional (3D) world space coordinate system.

13. monitoring eye movements of the user of the vehicle to determine one or more of gaze patterns, saccade velocity, gaze behavior, or smooth pursuit behavior; determining the attention value of the user based at least in part on one or more of the gaze pattern, the saccade velocity, the fixation, or the smooth pursuit; The method of claim 10 further comprising:

14. generating a heat map corresponding to road scanning behavior of the user of the vehicle over a period of time; determining the attentiveness value of the user based at least in part on the heat map; The method of claim 10 further comprising:

15. monitoring one or more of eye movements, eye characteristics, or eye measurements of the user of the vehicle; determining the cognitive load value of the user based at least in part on one or more of the eye movements, the eye characteristics, or the eye measurements of the user; The method of claim 10 further comprising:

16. 16. The method of claim 15, wherein the determining the cognitive load value comprises comparing at least one of the eye movements, the eye characteristics, or the eye measurements to a profile of the user, the profile being generated during one or more operations of the vehicle by the user.

17. one or more sensors; one or more processors; one or more memory devices storing instructions that, when executed by the one or more processors, provide to the one or more processors: determining a gaze direction of an occupant of the vehicle based at least in part on first sensor data generated using the one or more sensors of a first subset; generating a three-dimensional representation in a coordinate system of the occupant's field of view based on the gaze direction; determining an object position of at least one object in the coordinate system and based at least in part on second sensor data generated using the one or more sensors of a second subset; comparing the representation of the gaze direction to the object position; and determining whether a threshold amount of overlap between the representation and the object location is met, weighted based on the portion of the representation that overlaps with the object location; determining to generate a notification or suppress the notification based at least in part on whether the overlap is satisfied; one or more memory devices that perform operations including A system comprising:

18. 18. The system of claim 17, wherein the one or more sensors of the first subset include at least one first sensor having the field of view of the occupant inside the vehicle, and the one or more sensors of the second subset include at least one second sensor having a field of view outside the vehicle.

19. The operation further comprises: monitoring eye movements of the occupants of the vehicle to determine one or more of gaze patterns, saccade velocities, gaze behavior, or smooth pursuit behavior; determining one or more of an attentiveness score or a cognitive load score for the occupant based at least in part on one or more of the gaze pattern, the saccade velocity, the gaze behavior, or the smooth pursuit behavior; 20. The system of claim 17, wherein the determining to generate or suppress the notification is further based at least in part on one or more of the attentiveness scores of the cognitive load scores.

20. The system comprises: Control systems for autonomous or semi-autonomous machines, Perception systems for autonomous or semi-autonomous machines, a system for performing a simulation operation; a system for performing deep learning operations; a system implemented using edge devices; a system incorporating one or more virtual machines (VMs); Systems implemented using robots, a system at least partially implemented in a data center; or Systems implemented at least in part using cloud computing resources 20. The system of claim 17, included in at least one of:

Citation Information

Patent Citations

  • Information providing device for vehicle

    JP2001357498A

  • Device for providing information on viewing action

    JP2011000278A

  • Predictive human-machine interface using eye-tracking technology, blind spot indicators, and driver experience.

    JP2013514592A

  • Driving assist system

    JP2020149431A