Path prediction for autonomous or semi-autonomous systems and applications
By using a probabilistic technique to calculate the machine occupant intention path score, the problem of inaccurate path prediction in existing technologies is solved, achieving more stable and accurate path prediction and improving the machine's navigation capabilities in complex environments.
Patent Information
- Application Number
- CN202511150949.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-08-21
- Filing Date
- 2025-08-15
- Publication Date
- 2026-03-03
AI Technical Summary
Existing technologies struggle to accurately predict a machine's intended path, especially when multiple paths are available, leading to errors in downstream systems or components.
By employing a probabilistic technique that combines input data from a perception system, a positioning system, and a mapping system, the algorithm calculates the score of the machine occupant's intended path. Through multiple iterations and time-related summaries, the algorithm selects the highest-scoring path, thereby improving the stability and accuracy of the prediction.
It improves the accuracy and stability of path prediction, reduces abrupt changes in path prediction, and enhances the machine's navigation decision-making ability in complex environments.
Smart Images

Figure CN121594907A_ABST
Abstract
Description
Background Technology
[0001] Accurately predicting a machine's path through its environment before it traverses the path (e.g., a path that follows the occupants' intentions) can be a key component of autonomous or semi-autonomous navigation. For example, predicting the path in advance allows various systems or components of the machine to anticipate upcoming road conditions and monitor (e.g., assess, confirm, adjust, etc.) the machine's speed and / or trajectory accordingly. Furthermore, accurate path prediction allows the machine to make informed navigation decisions in advance, such as changing lanes or making smoother turns, which can reduce the potential for sudden maneuvers that could disrupt traffic flow or lead to adverse events.
[0002] However, accurately predicting a machine's intended path in advance can be challenging. For example, in some cases, multiple valid paths may exist, but the data relied upon for predicting the intended path may be insufficient to determine which path the machine driver intends to use. Furthermore, traditional systems may rely on rules and / or heuristics for path prediction. Therefore, there are many instances where these traditional systems fail to correctly predict the machine's path in advance. Moreover, since other systems or components of the machine may take the predicted path as input and depend on accurate path prediction, problems related to the predicted path can lead to errors in these downstream systems or components. Summary of the Invention
[0003] Embodiments of this disclosure relate to path prediction for autonomous or semi-autonomous systems and applications. For example, the systems and methods described herein can use probabilistic techniques to predict the intended path of a machine through its environment. Various input data from sensing systems, positioning systems, mapping systems, and / or any other data sources can be used to determine the intent of a machine occupant (e.g., driver, passenger, operator, etc.) and to calculate scores associated with road segments in the environment. These scores can indicate the probability that certain road segments are part of the occupant's intended machine path, and scores can be aggregated for each road segment across multiple instances of receiving and analyzing input data. In some cases, one or more road segments with the highest scores can be selected as part of the machine's predicted path. For example, at one or more intersections where multiple road segments intersect, the road segment with the highest score can be selected as the predicted path.
[0004] Unlike conventional systems, the systems disclosed herein, in some embodiments, are capable of more accurately predicting a machine's intended path using probabilistic techniques. These techniques are more robust, use one or more simplified architectures, and address inherent problems typically associated with conventional systems. For example, unlike conventional systems, the systems disclosed herein can compute intent scores to apply occupant intent to multiple intersections ahead of the machine. These intersections may include multiple options for different road segments available to the machine for the path, and the system can apply intent scores to the road segments to determine which option the machine's occupant intends to use. Furthermore, unlike conventional systems, the systems disclosed herein can perform time-related aggregation, tracking, and weighting of intent scores for each road segment corresponding to an upcoming intersection across multiple iterations or frames, selecting the option with the highest score. This improves the stability of the predicted path and reduces flicker, while also achieving accurate path prediction related to occupant intent. Attached Figure Description
[0005] The system and method for path prediction in autonomous or semi-autonomous systems and applications will be described in detail below with reference to the accompanying drawings, wherein:
[0006] Figure 1 An example data flow diagram of a path prediction process according to some embodiments of the present disclosure is shown;
[0007] Figure 2 An example map showing an environment and a machine positioned relative to a map according to some embodiments of this disclosure;
[0008] Figure 3A An example of the default path of a machine according to some embodiments of this disclosure is shown;
[0009] Figure 3B Examples of path options for a machine according to some embodiments of this disclosure are shown;
[0010] Figure 3C Examples of predicted paths for machines according to some embodiments of this disclosure are shown;
[0011] Figure 4 This is a visualization of example probability distributions associated with road segments corresponding to intersections in front of a machine, according to some embodiments of this disclosure;
[0012] Figure 5 This is a data flow diagram illustrating an example of a process for training one or more machine learning models to predict machine paths of intent, according to some embodiments of this disclosure.
[0013] Figure 6 Examples of systems that can perform one or more of the processes described herein are shown according to some embodiments of this disclosure;
[0014] Figure 7 According to some embodiments of this disclosure, an example flowchart of a method for predicting the path of a machine is shown;
[0015] Figure 8 According to some embodiments of this disclosure, an example flowchart is shown for a method of scoring a road segment intersection in front of a machine using occupant intent;
[0016] Figure 9A These are illustrations of example autonomous vehicles according to some embodiments of the present disclosure;
[0017] Figure 9B According to some embodiments of this disclosure Figure 9A Examples of camera positions and field of view for autonomous vehicles;
[0018] Figure 9C According to some embodiments of this disclosure Figure 9A A block diagram of an example system architecture for an example autonomous vehicle;
[0019] Figure 9D Cloud-based servers and according to some embodiments of this disclosure Figure 9A A system diagram illustrating communication between autonomous vehicles;
[0020] Figure 10 This is a block diagram of an example computing device suitable for implementing some embodiments of the present disclosure; and
[0021] Figure 11 This is a block diagram of an example data center suitable for implementing some embodiments of the present disclosure. Detailed Implementation
[0022] Systems and methods related to path prediction for autonomous or semi-autonomous systems and applications are disclosed. Although this disclosure may be combined with exemplary autonomous or semi-autonomous vehicles or machines 900 (also referred to herein as "vehicle 900", "self-vehicle 900", "self-machine 900" or "machine 900"), examples are combined with... Figures 9A to 9DThe invention is described herein (as described herein), but this is not intended to limit the invention. For example, the systems and methods described herein can be used by (but are not limited to) non-autonomous vehicles or machines, semi-autonomous vehicles or machines (e.g., in one or more adaptive driver assistance systems (ADAS)), manned and unmanned robots or robotic platforms, warehouse vehicles, off-road vehicles, vehicles coupled to one or more trailers, aircraft, ships, space shuttles, emergency response vehicles, motorcycles, electric or motorized bicycles, aircraft, engineering vehicles, underwater vehicles, drones, and / or other types of vehicles. Furthermore, although this disclosure may be described with regard to path prediction for vehicles or machines, this is not intended to limit the invention; the systems and methods described herein can be used in augmented reality, virtual reality, mixed reality, robotics, security and supervision, autonomous or semi-autonomous machine applications, and / or any other technical field where path prediction is applicable.
[0023] In some examples, the system of this disclosure can logically separate path prediction operations based on the distance between the machine and intersections and / or road segments. For example, to predict the path of the machine through intersections beyond a threshold distance or range (e.g., 100 meters, 200 meters, 5 seconds, etc.), the system may use default path prediction logic that may not consider occupant intent. Instead, the default path prediction logic may use map data and / or other road information (e.g., distance of road segments to defined routes, road classification (e.g., highway, roadway, arterial road, etc.), turning angles between road segments, number of lanes, etc.) to determine the predicted path of the machine through these intersections. In some examples, the system may use the default path prediction logic to make an initial prediction related to the predicted path of the machine. For example, when the machine begins to move, occupant intent may not be known, so the system may initially use the default path prediction logic to make an initial prediction of the machine's path.
[0024] For intersections and / or road segments located within or at a machine threshold distance, the system may use probabilistic algorithms (or any other techniques disclosed herein) that can leverage occupant intent to predict the machine's path. In some examples, this may include a system that uses probabilistic algorithms to evaluate an initial path prediction (based on default logic) and determine whether the machine is likely to deviate from the initially predicted path at one or more intersections within the machine threshold distance.
[0025] For example, one or more systems may receive input data from various components or systems of the machine and use the input data to determine one or more occupant intentions corresponding to one or more upcoming intersections within a threshold distance of the machine. For each road segment associated with the upcoming one or more intersections, one or more systems may compute one or more scores (also referred to herein as “intent scores”) based on the occupant’s one or more intentions. Intent scores may indicate the probability that certain road segments are part of the machine’s intended path, and one or more systems may use these scores to select road segments for the machine’s predicted path. One or more systems may update the initial predicted path to include road segments determined using a probability-based algorithm and road segments outside the range determined using a default algorithm.
[0026] In some examples, the input data used to determine one or more occupant intentions can include various forms of data from various systems or components of the machine. For example, input data can include input / output (I / O) data associated with the machine. Such I / O data can indicate whether the machine's steering signal is activated, whether the machine's driver is using the accelerator or brake pedal, the steering wheel angle, etc. I / O data can also include voice or text input, such as input representing the occupant's voice. For example, if the I / O data represents the occupant saying "turn left at road x," then that I / O data can be a strong indication of the occupant's intention (e.g., the machine's path should include turning left at road x).
[0027] In some examples, the input data may also include perception data generated using the machine's perception systems or components. The perception data may be used by one or more systems to determine occupant intent based on information associated with the machine's surrounding environment. For example, the perception data may indicate that the machine is operating in a specific type or class of lane, such as a turning lane or conversely, a no-turn lane, which can be a strong indication of occupant intent. Additionally or alternatively, the perception data may indicate surface markings associated with the machine's lane or road segment, such as right-turn markings, left-turn markings, straight-ahead markings, or other markings (e.g., arrow markings indicating the machine's expected path / direction of travel within the lane).
[0028] Input data may additionally or alternatively include positioning data. Positioning data may indicate the machine's current location, current attitude (e.g., orientation, heading, yaw, etc.). In some cases, positioning data may be used in conjunction with map data. For example, positioning data and map data can be used to indicate the machine's position, attitude, etc., relative to an environmental map. In some examples, positioning data and / or map data can be used to detect lane changes by the machine, which can also serve as a strong indicator of occupant intent (e.g., changing lanes to exit a highway, turning onto another road, etc.).
[0029] While machine I / O data, perception data, positioning data, and / or map data are just a few examples of input data that a system can use to determine occupant intentions and / or predict machine paths, in additional or alternative examples, the system may use any other type of data, as described in more detail herein, to achieve these purposes. For example, input data may include, but is not limited to, machine I / O data, perception data, positioning data, map data, sensor data (e.g., raw or processed image data, LiDAR data, RADAR data, ultrasonic data, audio data, etc.), route data, machine status data, road marking data, and / or lane marking data.
[0030] As described herein, in various examples, one or more systems may use map data, positioning data, and / or perception data to determine whether a machine is within a threshold distance of an intersection (e.g., one or more intersections). One or more intersections may contain one or more options that cause the machine to deviate from its initial path or another previously predicted path. For example, the machine's predicted path may include a first road segment connecting to at least one of the one or more intersections, but one or more second road segments may also connect to that first intersection. Accordingly, one or more systems of this disclosure may use input data to compute scores (also referred to herein as "intent scores") for the first road segment and one or more second road segments at the first intersection. These scores may represent the probability, at least based on occupant intent and / or other factors, that the first road segment and one or more second road segments are part of the machine's intended path (e.g., the machine path of the occupant's intent).
[0031] In some examples, one or more systems may apply scores differently to road segments at intersections based on whether the occupant's intention corresponds to a specific road segment. For example, one or more systems may calculate a confidence score for the occupant's intention. In at least some examples, the confidence score may indicate the degree of confidence that the occupant's intention applies to a specific road segment. For example, consider a scenario where a machine is approaching two intersections along a straight path. At the first intersection, the machine may continue straight to the second intersection, or it may veer off course and avoid / miss the second intersection. If the machine's occupant activates the machine's turn signal before reaching the first intersection, it may be difficult to determine whether the occupant's intention (e.g., the turn signal) should apply to the first or second intersection. In other words, based solely on the turn signal, it may be difficult to determine whether the occupant intends to turn at the first or second intersection. In this scenario, the system of this disclosure can calculate a confidence score for whether the occupant's intention applies to a first road segment associated with the machine turning at a first intersection, or whether the occupant's intention applies to a second road segment associated with the machine turning at a second intersection. One or more systems can then apply the confidence score to scores for different road segments to apply the intention to each segment, but possibly at different proportions. For example, if the machine speed is high and slowing down to turn at the first intersection may be difficult, the system can assign a higher confidence score to the occupant's intention applied to the second road segment associated with the machine turning at the second intersection. As another example, past occupant behavior can be taken into account. For example, if approaching two intersections and the turn signal is on, usage of the second intersection during one or more previous trips can be incorporated into the scores for the first and second intersections. In some examples, the weights of different intersections can be tracked over time, such that repeated use at a particular intersection may have a higher weight than at an intersection that has never been used before.
[0032] Furthermore, in some examples, the system may apply multiple scores associated with occupant intent to multiple road segments evaluated as options for path prediction, and aggregate scores across multiple frames / instances. For example, consider an intersection associated with a first road segment for continuing straight and a second road segment associated with turning right at that intersection. If one or more systems detect that a machine's right-turn signal is activated, the system may apply a first intent score to the second road segment associated with the right turn. If one or more systems also detect that the machine is in a dedicated right-turn lane, one or more systems may apply a second intent score to the second road segment. One or more systems may add the first and second intent scores together, and if the aggregated intent score for the second road segment is greater than the aggregated intent score for the first road segment, the system may predict that the machine's path includes the second road segment. Additionally, when one or more systems acquire additional new data, one or more systems may recalculate the intent score for each road segment using the new data and aggregate the new intent scores with the previous intent scores, thereby creating a dynamic cumulative total. In some examples, one or more systems may apply different weights or priorities to certain occupant intents.
[0033] As described herein, in some examples, the system may use the scores of road segments associated with intersections located within a threshold distance of the machine to determine the machine's predicted path (e.g., predicting the occupant's intended path). For example, one or more systems may select one or more road segments with the highest intent scores to be included as part of the predicted path. In some examples, one or more systems may select one or more road segments in descending order based on the sequence of (one or more) intersections. For example, one or more systems may first select a first road segment corresponding to a first intersection, which may be the nearest intersection or the next intersection. Then, the system may select a second road segment corresponding to a second intersection, which may be the first intersection and the next intersection after the first road segment, and so on, until a predicted path is determined. Furthermore, in some cases, once no intersections and / or road segments remain within the range, one or more systems may use default logic to determine any remaining portion of the predicted path after the intersections and / or road segments within the range.
[0034] In some examples, one or more systems may generate one or more outputs based at least on the predicted path of the determined machine. For example, in addition to outputting the predicted path of the machine (or as an alternative), one or more systems may additionally or alternatively output one or more road transition types (e.g., an array describing how each segment in the predicted path is linked, indicating turns or straight connections), one or more route input states (e.g., indicating whether a navigation route has been established and / or whether the predicted path follows that route), one or more map identifiers (e.g., to ensure that downstream systems or components are using the same map as the path prediction system), one or more map timestamps (e.g., timestamps for when the map was generated), one or more lane map timestamps, and / or one or more predicted lane trajectories.
[0035] In some examples, one or more systems may use the predicted path and / or other outputs as input to perform one or more additional operations associated with the machine. As an example, a system may determine road curvature and fuse the curvature of the machine's trajectory at intersections, and in some cases, recalculate curvature already provided in map data based on the predicted path. As another example, a system may use the predicted path to perform localization to determine which road segment the machine is operating on. As yet another example, one or more systems may use the predicted path to access map data attributes along the predicted path to resolve speed limit issues that conflict with perception results. In further examples, one or more systems may acquire events along the predicted path, such as T-junctions and toll plazas.
[0036] In various examples, one or more systems may enable a machine to perform one or more actions based at least on the machine's predicted path and / or any other outputs or actions described herein. For example, a system may use the predicted path, curvature, speed limits, etc., to determine whether the machine requires additional braking or acceleration. As an example, the predicted path might indicate that the machine is merging into a road or highway, and one or more systems may apply additional acceleration to bring the machine to the speed required for merging. As another example, the predicted path might indicate that the machine is about to make a sudden turn, and additional braking may be applied to ensure that the machine turns at a comfortable and / or safe speed. In some examples, one or more systems may communicate the machine's predicted path to one or more other machines, such as machines operating nearby or machines with trajectories intersecting with or adjacent to the machine's predicted path.
[0037] In some embodiments, the systems and methods described herein can be performed in a simulated environment (e.g., NVIDIA's DriveSIM) using simulated data (e.g., simulated sensor data from virtual machines or simulated sensors of simulated machines). For example, simulated input data (e.g., map data, perception data, machine I / O data, or any other input data described herein) can be used to determine road segments / intersections and score them based on occupant intentions in the simulated environment, and this information can be used to perform operations associated with virtual machines in the environment (e.g., acceleration, deceleration, curvature fusion, etc.). These simulated operations can be used to test the performance of underlying algorithms, systems, and / or processes before deploying them to the real world. In some cases, simulations can be used to generate synthetic training data, such as training data from simulations indicating occupant intentions, map data, etc. The synthetic training data (in addition to or as a substitute for real-world data) can then be processed to determine a predicted path for the machine in the environment, such as a path for a machine to pass through a warehouse. In any example, such as when using a simulated environment for testing, validation, training, etc., one or more optical transport algorithms (such as ray tracing and / or path tracing algorithms) can be used to render or otherwise generate the simulated environment and / or related training data.
[0038] In some embodiments, a simulated environment and / or one or more of its objects, features, or components can be generated or managed within a 3D content collaboration platform (such as NVIDIA's OMNIVERSE) for use in industrial digitization, generative physics AI, and / or other use cases, applications, or services. For example, the content collaboration platform or system may include systems for using or developing generic scene descriptors (USD) (such as OpenUSD) data to manage objects, features, scenes, etc., within simulated environments, digital environments, etc. To simulate real physics and physical interactions with simulations hosted on the platform, the platform may include realistic physics simulations, such as using NVIDIA's PhysX SDK. The platform can integrate OpenUSD with ray tracing / path tracing / light transport simulations (such as NVIDIA's RTX rendering technology) into software tools and simulation workflows for building, training, deploying, or testing AI systems, such as systems for testing, validating, training (e.g., machine learning models, neural networks, etc.) and / or other tasks related to automobiles, robots, machines, or other applications. In some examples, the simulated environment can include a digital twin of a real environment, such as a specific section of road, warehouse, data center, airport, geographic area, ocean area, or any other real environment in which autonomous or semi-autonomous machines can operate.
[0039] The systems and methods described herein can be used by, but are not limited to, non-autonomous vehicles or machines, semi-autonomous vehicles or machines (e.g., in one or more adaptive driver assistance systems (ADAS)), autonomous vehicles or machines, manned and unmanned robots or robotic platforms, warehouse vehicles, off-road vehicles, vehicles coupled to one or more trailers, aircraft, boats, shuttles, emergency response vehicles, motorcycles, electric or motorized bicycles, aircraft, construction vehicles, underwater vehicles, drones, and / or other vehicle types. Furthermore, the systems and methods described herein can be used for a wide range of purposes, such as (but not limited to) machine control, machine mobility, machine driving, synthetic data generation, model training, perception, augmented reality, virtual reality, mixed reality, robotics, security and surveillance, simulation and digital twins, autonomous or semi-autonomous machine applications, deep learning, environmental simulation, object or participant simulation and / or digital twins, data center processing, conversational AI, optical transport simulation (e.g., ray tracing, path tracing, etc.), collaborative content creation for 3D assets, cloud computing, and / or any other suitable applications. For example, the systems and methods disclosed herein can be used in the aviation or aerospace field to communicate the predicted path of an aircraft to other aircraft or air traffic controllers.
[0040] The disclosed embodiments can be included in a variety of different systems, such as automotive systems (e.g., control systems for autonomous or semi-autonomous machines, perception systems for autonomous or semi-autonomous machines), systems implemented using robots, aviation systems, medical systems, marine systems, smart area monitoring systems, systems for performing deep learning operations, systems for performing simulation operations, systems for performing digital twin operations, systems implemented using edge devices, systems implementing language models (e.g., large language models (LLM), visual language models (VLM), and / or multimodal language models), systems containing one or more virtual machines (VMs), systems for performing synthetic data generation operations, systems implemented at least partially in a data center, systems for performing conversational AI operations, systems for performing optical transmission simulations, systems for performing collaborative content creation for 3D assets, systems for performing generative AI operations, systems implemented at least partially using cloud computing resources, and / or other types of systems.
[0041] refer to Figure 1 , Figure 1Example data flow diagrams of path prediction processes according to some embodiments of this disclosure are shown. It should be understood that such and other arrangements described herein are merely illustrative. Other arrangements and elements (e.g., machines, interfaces, functions, sequences, functional groups, etc.) may be used in addition to the arrangements and elements shown, and certain elements may be omitted together. Furthermore, many of the elements described herein are functional entities that can be implemented as discrete or distributed components, or in combination with other components, and can be implemented in any suitable combination and location. The various functions performed by the entities described herein can be implemented by hardware, firmware, and / or software. For example, various functions can be implemented by a processor executing instructions stored in memory. In some embodiments, the systems, methods, and processes described herein can be used with… Figures 9A-9D The exemplary autonomous vehicle 900 Figure 10 Exemplary computing device 1000 and / or Figure 11 The components, features, and / or functions of the exemplary data center 1100 are similar to those of other components, features, and / or functions used in the implementation.
[0042] Process 100 can be implemented using one or more route prediction systems 102 and driving stack 104 (as well as additional or alternative components). Route prediction system 102 may include a default route component 106, an intent component 108, a scoring component 110, and a route prediction component 112, as well as additional or alternative components. One or more route prediction systems 102 may run on one or more computing devices in a machine (e.g., an autonomous or semi-autonomous vehicle).
[0043] In summary, process 100 may include one or more path prediction systems 102 that receive: one or more map data 114 representing one or more maps of the environment; route data 116 representing one or more routes of the machine (e.g., predefined routes); state data 118 representing one or more states associated with the machine; sensor data 120 (e.g., image data, LiDAR data, RADAR data, ultrasonic data, audio data, etc.); I / O data 122 indicating one or more inputs and / or outputs of the machine; positioning data 124 indicating one or more of the machine's position, attitude, current route, previous behavior data, etc.; and additional or alternative input data. Using at least some of the input data, a default path component 106 of one or more path prediction systems 102 may determine one or more default paths of the machine, which may be represented using default path data 126. Furthermore, an intent component 108 may use at least some of the input data to determine one or more occupant intents, which may be represented using intent data 128. Scoring component 110 can use intent data 128 to calculate one or more scores 130 for road segments in the environment that may potentially be used as machine paths, including one or more road segments of one or more default paths. Route prediction component 112 can use one or more scores 130 to determine one or more predicted paths for the machine, which can be represented using path prediction data 132. One or more path prediction systems 102 can provide the path prediction data 132 to driving stack 104, which can use the predicted paths as input to one or more other algorithms or processes, and / or also to perform one or more machine-related operations.
[0044] In some examples, map data 114 may represent one or more maps of the environment in which the machine operates. Map data 114 may correspond to a standard-resolution (SD) map of the environment or a high-resolution (HD) map of the environment. In various examples, the system of this disclosure may use an SD map of the environment to predict the machine path. For example, Figure 2 An example of a map 202 of an environment according to some embodiments of the present disclosure is shown, along with a machine 204 positioned relative to the map 202. The map 202 may contain multiple road segments 206. Each road segment 206(1)-206(6) may represent a portion of a driving surface between a pair of intersections 208(1)-208(5). For example, as... Figure 2As shown in the example, map 202 may include a first road segment 206(1) between a first intersection 208(1) and a second intersection 208(2) (e.g., the current road segment of machine 204), a second road segment 206(2) between a first intersection 208(1) and a third intersection 208(3), a third road segment 206(3) connected to the first intersection 208(1), and so on.
[0045] In some examples, machine 204 may correspond to machine 900 as described herein. Machine 204 may include any component of the components described herein. For example, machine 204 may include a positioning component for determining the position, orientation, etc., of the machine relative to map 202. In some examples, machine 204 may include... Figure 1 The example describes one or more components and systems, and these systems and components can be used to predict the path of machine 204.
[0046] review Figure 1 For example, one or more path prediction systems 102 may use a default path component 106 to predict the path a machine takes through intersections beyond a threshold distance or range (e.g., 100 meters, 200 meters, 5 seconds, etc.) and to make an initial prediction about the machine's path (e.g., before occupant intent can be determined). In some examples, the default path component 106 may use map data 114, route data 116, status data 118, location data 124, and / or any other input data to determine default path data 126. The default path component 106 may determine the initial predicted path for the machine based on several factors and / or rules (e.g., when occupant intent is minimal and / or not considered).
[0047] In some cases, to predict out-of-range portions of a path and / or to predict an initial path (also referred to herein as the "default path"), the default path component 106 may sort or prioritize potential road segments based on the distance of the segment to a defined route, the road type associated with the segment (e.g., highway, road, etc.), the turning angle between consecutive segments, and / or the number of lanes per segment. For example, if a route is defined between two locations, the default path may include the segment closest to that route. If no route is defined, the default path component 106 may select the segment with the highest road priority form as the default path. In some examples, if all subsequent segments (segments following another segment) have the same road priority form, the default path component 106 may select the segment with the smallest turning angle for the default path. As an example, the default path component 106 may select a path that proceeds straight through an intersection rather than a path that includes a turn at an angle at the intersection. In some cases, if subsequent road segments have the same or similar turning angles, the default path component 106 can select the road segment with the most lanes (for example, selecting a road segment with two lanes instead of a road segment with one lane).
[0048] For example, Figure 3A Examples of default paths 302 for machine 204 according to some embodiments of the present disclosure are shown. As shown, default path 302 may be a straight path based on the default logic described above and herein. For example, default path 302 includes a first segment 206(1), a second segment 206(2), and a fifth segment 206(5). In some examples, default path component 106 may use the rules described above to determine the segments 206 of default path 302. For example, first segment 206(1), second segment 206(2), and fifth segment 206(5) may have a road form corresponding to a highway or expressway, while other perpendicular segments may be exits or side roads. Furthermore, or alternatively, first segment 206(1) and second segment 206(2) may have a 0-degree or other low-degree turning angle at a first intersection 208(1), and second segment 206(2) and fifth segment 206(5) may also have a 0-degree or low-degree turning angle at a third intersection 208(3). In various examples, the default path data 126 may contain a list of segment identifiers corresponding to the first segment 206(1), the second segment 206(2), and the fifth segment 206(5). In some examples, the default path component 106 may score each segment using the default algorithm techniques described above and herein when determining the default path 302, and use the segment with the highest score as the default path 302.
[0049] review Figure 1For example, process 100 may also include intent components 108 of one or more path prediction systems 102 that determine intent data 128 corresponding to intersections and / or road segments located within a machine threshold distance or range. Intent data 128 may indicate occupant intent associated with the machine path. That is, intent data 128 may indicate the likelihood or probability that a machine occupant (e.g., driver, passenger, remote operator, etc.) intends for the machine path to include a particular road segment. For example, while a default path may indicate the most likely path for the machine based on road geometry, route, road type or class (e.g., prioritizing highways or roads over ordinary streets) and / or other factors, intent data 128 may indicate which road segments in the environment the machine occupant might actually intend to use as a path.
[0050] In some examples, occupant intent may include, but is not limited to, the machine's yaw rate, whether the machine's steering signal is activated, lane changes, lane assignment (e.g., which lane the machine is positioned in), arrow markings detected in the lane, the machine's perceived path, the machine's steering angle, and whether the machine is operating in a dedicated lane. For example, if the machine is operating on a multi-vehicle road segment and the perception system or other components of the machine determine that a lane change is in progress, the intent component 108 can detect the lane change intent. As another example, if a steering signal is activated, the intent component can detect the occupant's intent to turn at an intersection or exit a highway.
[0051] In some examples, the intent component 108 may determine the intent data 128 based on a variety of inputs. For example, the intent component 108 may use one or more of map data 114, route data 116, status data 118, sensor data 120, I / O data 122, previous behavior data, and / or location data 124 to determine the intent data 128. For example, the intent component 108 may use I / O data 122 associated with the machine to determine whether the machine's steering signal is activated, whether the machine's driver is using the accelerator or brake pedal, the machine's steering wheel angle, etc. The I / O data 122 may also include voice or text input, such as input representing one or more occupant voices. For example, if the I / O data 122 represents the occupant saying "turn left at X road," then the I / O data may be a strong indication of the occupant's intent (e.g., the machine's path should include a left turn at X road).
[0052] In some examples, the input data may also include perception data (not shown) generated using the machine's perception system or components. The intent component 108 may use the perception data to determine the occupant's intent based on information associated with the machine's surrounding environment. For example, the perception data may indicate that the machine is operating in a lane of a specific type or class, such as a turning lane, or conversely, a lane where turning is prohibited, which could be a strong indication of the occupant's intent. Alternatively, the perception data may indicate surface markings associated with the lane or road segment in which the machine is located, such as right-turn markings, left-turn markings, straight-ahead markings, or other markings (e.g., arrow markings). For example, when the perception data indicates that the machine is traveling in a lane with double white lines on each side (e.g., indicating no entry or exit from that lane), the intent to remain in that lane may increase, while the intent to enter other lanes that may lead to an intersection may decrease.
[0053] In some examples, location data 124 may indicate the machine's current position, current attitude (e.g., orientation, heading, yaw, etc.). In some cases, location data 124 may be used in association with map data 114. For example, location data 124 and map data 114 may be used to indicate the machine's position, attitude, etc., relative to an environmental map (e.g., map 202). In some examples, intent component 108 may use location data 124 and / or map data 114 to detect lane changes by the machine, which can also serve as a strong indication of occupant intent (e.g., changing lanes to exit the highway, turning onto another road, etc.). Alternatively, intent component 108 may determine which road to associate occupant intent with, at least based on the detection of a lane change. For example, if intent component 108 detects a lane change by the machine, intent component 108 (and / or scoring component 110) may determine that the machine's turn signal is being used for the lane change, and that the turn signal may not indicate an occupant intent to turn onto a different road, etc.
[0054] In some examples, the intent data 128 and / or specific occupant intent detectable by the intent component 108 can be applied to one or more road segments and / or intersections in the environment. As an example, a machine's turn signal may indicate that the machine intends to turn; however, the turn signal may not indicate the specific road the machine intends to turn onto. For example, Figure 3B Examples of different path options for a machine according to some embodiments of this disclosure are shown. Figure 3BThe path options shown in the example include first path option 304A and second path option 304B, both of which can be valid options as the intended path for machine 204 in a scenario where intent data 128 indicates a future intention to turn left. For example, if machine 204 (and / or the occupants of machine 204) activates a left-turn signal, that left-turn signal can be used to determine whether the intent is to turn left at the first intersection 208(1) or the third intersection 208(3). In some examples, as described in more detail herein, additional input data can be used to determine whether the intent is first path option 304A or second path option 304B. As an example, if in Figure 3B In the example, machine 204 is moving too fast in its current position, such that the rate of deceleration required for machine 204 to use the first path option 304A would exceed the comfort limit (e.g., 3.5 m / s). 2 If the machine 204 has slowed down at the first intersection 208(1) and / or is in a lane dedicated to left turns, then the intent data 128 may indicate that the left turn signal is applied with a higher confidence level to the sixth segment 206(6) associated with the second route option 304B. As a second example, if the machine 204 has slowed down at the first intersection 208(1) and / or is in a lane dedicated to left turns, then the intent data 128 may indicate that the left turn signal is applied with a higher confidence level to the third segment 206(3) associated with the first route option 304A.
[0055] review Figure 1 For example, process 100 may include a scoring component 110 that calculates one or more scores 130 (e.g., intent scores) for one or more road segments in the environment that can be used as machine paths. In some examples, scoring component 110 may use intent data 128 and / or default path data 126 to calculate one or more scores 130. One or more scores 130 may indicate the probability that certain road segments are part of a machine path intended by the occupants, and path prediction component 112 may use the scores 130 to select road segments for machine-predicted paths.
[0056] In some examples, the scoring component 110 may determine a score 130 and / or apply the score 130 to different road segments based on whether the intent data 128 corresponds to a specific road segment. For example, the scoring component 110 may calculate a confidence score for the occupant's intent. In at least some examples, the confidence score may refer to the confidence level of whether the schematic diagram data 128 is applied to a specific road segment. For example, with Figure 3BFor example, machine 204 is approaching a first intersection 208(1) and a third intersection 208(3). At the first intersection 208(1), machine 204 may (i) turn right; (ii) continue straight along the second road segment 206(2) to the third intersection 208(3); or (iii) turn left into the third road segment 206(3) associated with the first route option 304A. If the occupant of machine 204 activates the machine's left-turn signal before reaching the first intersection 208(1), it may be difficult to determine whether the occupant's intention (e.g., the left-turn signal) should apply to the first intersection 208(1) or the third intersection 208(3). In other words, based solely on the turn signal, it may be difficult to determine whether the occupant intends to turn at the first intersection 208(1) or at the third intersection 208(3), as illustrated by the two potential route options 304A and 304B. In this scenario, the scoring component 110 can calculate a confidence score for whether the occupant's intention applies to the third segment 206(3) associated with the machine 204 turning left at the first intersection 208(1), or to the sixth segment 206(6) associated with the machine 204 turning left at the third intersection 208(3). The scoring component 110 can then apply the confidence score to one or more scores 130 for different segments to apply the intention to each segment, but possibly in different proportions. For example, if the machine 204 is traveling at a high speed and slowing down at the first intersection 208(1) to turn may be difficult, the scoring component 110 could assign a higher confidence score to the intention score for the sixth segment 206(6) associated with the machine 204 turning at the third intersection 208(3).
[0057] Furthermore, in some examples, the scoring component 110 may apply multiple scores 130 associated with occupant intent to multiple road segments that are evaluated as options for the predicted path, and aggregate the scores across multiple frames / instances. For example, consider an intersection associated with a first road segment for continuing straight and a second road segment for turning right at that intersection. If the intent component 108 detects that the machine's right-turn signal light is activated, the scoring component 110 may apply a first intent score to the second road segment associated with the right turn. If the intent component 108 also detects that the machine is in a dedicated right-turn lane, the scoring component 110 may apply a second intent score to the second road segment. The scoring component 110 may add the first intent score and the second intent score together, and if the aggregated intent score for the second road segment is greater than the aggregated intent score for the first road segment, the path prediction component 112 may predict that the machine's path includes the second road segment.
[0058] Furthermore, when the intent component 108 receives additional and / or new input and generates additional and / or new intent data 128, the scoring component 110 can use the new intent data 128 to recalculate the intent score for each segment and sum the new intent scores with the previous intent scores to create a dynamic cumulative total. For example, the individual intent scores for each segment can be combined to obtain a total score. In some cases, the scoring component 110 can implement a maximum score that can be added from a single intent to influence the total score for that segment. For example, since intents may differ in duration and scores are summarized across frames, intents with longer durations may become dominant when path prediction is involved. Therefore, to address this issue, the scoring component 110 can set a threshold for the highest possible score that a particular intent can achieve.
[0059] In some examples, the scoring component 110 may apply different weights or priorities to certain occupant intentions. For example, the scoring component 110 may assign higher values to certain intention scores in the score 130 based on the underlying context. As an example, the machine's yaw rate or turn signal may have higher priority in indicating occupant intentions than the machine being within a dedicated turning lane. For instance, a driver may sometimes mistakenly enter a designated turning lane and activate their turn signal to change lanes, so the scoring component 110 may assign higher priority to the turn signal in indicating occupant intentions. In some examples, to give one occupant intention a higher priority than another, the scoring component 110 may assign a higher value to the score of the higher-priority intention.
[0060] In some examples, the scoring component 110 may calculate one or more scores 130 that correspond to a probability distribution of road segments, where the probability distribution indicates which road segments have the highest probability of being part of the intended path set by the occupant for the machine. For example, Figure 4 This is a visualization of an example probability distribution 402 associated with road segments corresponding to intersections ahead of a machine, according to some embodiments of this disclosure. Probability distribution 402 may include a plurality of road segments 404A-404N (where “N” represents any number) and a plurality of scores 406A(1)-406N(N) corresponding to a plurality of intentions 408(1)-408(N) applied to each of the road segments 404A-404N. Probability distribution 402 may also include a plurality of summary scores 410A-410N for each of the road segments 404A-404N.
[0061] In some examples, probability distribution 402 can be used as follows: Figure 4The table shown in the example is used to represent or otherwise maintain the data. Alternatively, other data structures may be used. In some examples, the field corresponding to road segments 404A-404N may include an identifier corresponding to the actual physical road segment, with the row of probability distribution 402 corresponding to that actual physical road segment. In some cases, the field corresponding to intentions 408(1)-408(N) may include the occupant intention type corresponding to each of scores 406A(1)-406N(N). For example, the first column of probability distribution 402 for the first intention 408(1) may include the intention score for the machine yaw rate, the second column of probability distribution 402 for the second intention 408(2) may include the intention score for the machine's turn signal, and so on. That is, for the first segment 404A, fraction 406A(1) may include a component value of the machine yaw rate applicable to the first segment 404A (e.g., 0-10), while fraction 406A(2) may include a component value of the machine turn signal applicable to the first segment 404A (e.g., 0-10). In some examples, the intent fraction may be applicable to multiple segments 404A-404N. For example, if the machine turn signal is activated, the corresponding intent fraction may be calculated and / or applied to multiple segments. In this case, and as an example, if the turn signal is applicable to segments 404A, 404C, and 404F, fractions 406A(2), 406C(2), and 406F(2) may each include a component value (e.g., a non-zero value).
[0062] In some examples, the summary scores 410A-410N may include values corresponding to the sum and / or time summary of scores 406A(1)-406N(N) across one or more frames (e.g., across a period of time). Summary score 410A may correspond to the sum or summary of scores 406A(1)-406A(N), summary score 410B may correspond to the sum or summary of scores 406B(1)-406B(N), and so on. As an example, suppose the scores 406(1)-406A(N) of road segment 404A contain the following values: 1, 3, 0, 0, 5, ..., and 2, respectively. In this example, the summary score 410A of road segment 404A may contain the value 11. Furthermore, suppose that after additional intent data is acquired during the next frame, scores 406(1)-406A(N) are updated to 2, 5, 0, 0, 10, ..., and 3. In this example, the summary score 410A may contain the value 20.
[0063] In various examples, refer to Figure 1 and Figure 4Secondly, the path prediction component 112 can use the summary scores 410A-410N to predict the machine's path, which can be represented using path prediction data. For example, the path prediction component 112 can select the road segment with the highest summary scores 410A-410N to be included as the machine's predicted path.
[0064] In some examples, the path prediction component 112 can select road segments in descending order based on the sequence in which the machine is approaching the intersection. For example, the path prediction component 112 can first select a first road segment corresponding to a first intersection, which could be the nearest intersection or the next intersection. Then, the path prediction component 112 can select a second road segment corresponding to a second intersection, which could be the first intersection and the next intersection after the first road segment, and so on, until a predicted path is determined. Furthermore, in some cases, once there are no remaining intersections and / or road segments within the range, the path prediction component 112 can use default path data 126 to determine any remaining portion of the predicted path following the intersections and / or road segments within the range.
[0065] For example, Figure 3C An example of a predicted path 306 for machine 204 according to some embodiments of the present disclosure is shown. Predicted path 306 may correspond to... Figure 3B The example shows the second path option 304B, and in some cases, the path prediction component 112 can use one or more scores 130 to determine the predicted path 306. Figure 3C The example shown illustrates a predicted path 306 comprising a first segment 206(1), a second segment 206(2), and a sixth segment 206(6). In determining the segments of the predicted path 306, the path prediction component 112 may first determine which segment has the highest intent score for the first intersection 208(1). For example, the path prediction component 112 may compare the intent scores between the second segment 206(2), the third segment 206(3), and the fourth segment 206(4). Figure 3C In the example, the path prediction component 112 can select the second road segment 206(2) as the predicted path 306 based on the highest aggregate intent score. Next, the path prediction component 112 can determine which road segment has the highest intent score for the third intersection 208(3). For example, the path prediction component 112 can compare the intent scores between the fifth road segment 206(5) and the sixth road segment 206(6). Figure 3CIn the example, the path prediction component 112 may select the sixth road segment 206(6) as the predicted path 306 based on the highest aggregate intent score of the sixth road segment 206(6). If the next intersection (not shown) is out of range, the path prediction component 112 may use the default path data 126 to determine any remaining portion of the predicted path 306 following the sixth road segment 206(6).
[0066] review Figure 1 For example, process 100 may include driving stack 104 receiving path prediction data 132. In some examples, path prediction data 132 may include or otherwise indicate one or more road transition types (e.g., an array showing how each segment in the predicted path is linked, indicating turn or straight connections), one or more route input states (e.g., indicating whether a navigation route has been established and / or whether the predicted path follows that route), one or more map identifiers (e.g., to ensure that downstream systems or components are using the same map as the path prediction system), one or more map timestamps (e.g., timestamps when the map was generated), one or more lane map timestamps, and / or one or more predicted lane map trajectories.
[0067] In some examples, driving stack 104 may include Figure 1 Various systems, components, layers, and / or modules are not shown in the examples. For example, driving stack 104 may include perception components, mapping components, planning components, control components, actuation components, and / or other components corresponding to additional and / or alternative layers of driving stack 104. Driving stack 104 may correspond to a machine, such as machine 204 and / or machine 900 as described herein.
[0068] In some examples, the driving stack 104 can use path prediction data 132 as input to perform one or more additional operations associated with the machine. As an example, the driving stack 104 can use path prediction data 132 to determine road curvature and fuse the curvature of the machine's trajectory at intersections, and in some cases, recalculate the curvature already provided in the map data based on the predicted path. As another example, the driving stack 104 can use path prediction data 132 to perform localization using the predicted path to determine which road segment the machine is operating on. As yet another example, the driving stack 104 can use path prediction data 132 to access map data attributes along the predicted path to resolve conflicts between speed limits and perception results. In a further example, the driving stack 104 can use path prediction data 132 to acquire events along the predicted path, such as T-junctions and toll plazas.
[0069] In various examples, the driving stack 104 enables the machine to perform one or more actions based at least on the path prediction data 132. For example, the driving stack 104 can use the predicted path, curvature, speed limit, etc. (all of which can be included in and / or determined using the path prediction data 132) to determine whether the machine requires additional braking or acceleration. As an example, the predicted path may indicate that the machine is merging into a road or highway, and the driving stack 104 (e.g., a planning component or control component of the driving stack 104) can apply additional acceleration to bring the machine to the speed required for merging. As another example, the predicted path may indicate that the machine is about to make a sudden turn, and the driving stack 104 can apply additional braking to ensure that the machine makes the turn at a comfortable and / or safe speed.
[0070] Now for reference Figure 5 , Figure 5 This is a data flow diagram illustrating an example of a process 500 for training one or more machine learning models 502 to predict intent machine paths according to some embodiments of the present disclosure. For example, one or more machine learning models 502 may correspond to one or more path prediction systems 102 and / or one or more components of one or more path prediction systems 102, such as default path component 106, intent component 108, scoring component 110, and / or path prediction component 112.
[0071] As illustrated, one or more machine learning models 502 can be trained using various input data 504 (e.g., training input data), which may include one or more of map data 114, route data 116, state data 118, sensor data 120, I / O data 122, location data 124, and / or any other data described herein, such as intent data 128 in some cases. In some examples, input data 504 may include one or more actual (e.g., previously generated and / or stored) versions of map data 114, route data 116, state data 118, sensor data 120, I / O data 122, location data 124, intent data 128, etc. Alternatively, input data 504 may be based on actual versions of map data 114, route data 116, state data 118, sensor data 120, I / O data 122, location data 124, and / or intent data. For example, input data 504 may include one or more modified versions of map data 114, route data 116, status data 118, sensor data 120, I / O data 122, location data 124, and / or intent data 128.
[0072] One or more machine learning models 502 can be trained using input data 504 and corresponding ground truth data 506. Ground truth data 506 may include annotations, labels, masks, and / or the like. For example, in some embodiments, ground truth data 506 may indicate the actual values of parameters associated with machine-predicted paths, occupant intentions, intention scores, and / or input data 504. For example, parameters in ground truth data 506 may include, but are not limited to, predicted path geometry, predicted lanes, predicted occupant intention values, predicted segment intention scores, predicted map data, and / or any other parameters. Ground truth data 506 may be generated within a drawing program (e.g., annotating program), a computer-aided design (CAD) program, a marker program, or other programs suitable for generating ground truth data 506, and / or in some examples may be hand-drawn. In any example, the ground truth data 506 can be synthetically generated (e.g., generated from a computer model or rendering), realistically generated (e.g., designed and generated from real-world data), machine-automated generated (e.g., using feature analysis and learning to extract features from the data and then generate labels), human-annotated (e.g., a tagger or annotation expert defines the location of the labels), and / or a combination thereof (e.g., human identification of the vertices of a polyline, and a machine generating polygons using a polygon rasterizer).
[0073] Training engine 508 can use one or more loss functions to measure the loss (e.g., error) of the output data 510 generated by machine learning model 502 compared to ground truth data 506. Output data 510 may include intent data 128 indicating predicted occupant intent, scores 130 indicating the probability that a road segment corresponds to a machine intent path, path prediction data 132 indicating a predicted machine path, or any other output. In some examples, any type of loss function can be used, such as cross-entropy loss, mean squared error, mean absolute error, mean bias error, and / or other loss function types. In some examples, different outputs may have different loss functions. For example, a first predicted occupant intent may include a first loss, a second predicted occupant intent may include a second loss, a third predicted occupant intent may include a third loss, and so on. In these examples, loss functions can be combined to form a total loss, and training engine 508 can use this total loss to, in some cases, update one or more parameters 512 (e.g., weights, biases, etc.) of one or more machine learning models 502 to train one or more machine learning models 502. In any example, backpropagation can be performed to recursively compute the gradient of the loss function with respect to the training parameters. In some examples, these gradients can be computed using the weights and biases of one or more machine learning models.
[0074] One or more machine learning models 502 may use any type of machine learning technique and / or algorithm. For example, but not limited to, any of the various machine learning models discussed herein may include one or more machine learning models of any type, such as those using linear regression, logistic regression, decision trees, support vector machines (SVM), Naive Bayes, k-nearest neighbors (Knn), K-means clustering, random forests, dimensionality reduction algorithms, gradient boosting algorithms, neural networks (e.g., autoencoders, convolutions, recurrent neural networks, perceptrons, long / short-term memory (LSTM), Hopfield, Boltzmann, deep belief, deconvolution, generative adversarial, liquid machines, large language models (LLM), visual language models (VLM), multimodal language models, diffusion, transformers, encoder-only, decoder-only, encoder-decoder, etc.), and / or other types of machine learning models.
[0075] In some examples, the machine learning model 502 can be packaged as a microservice (e.g., an inference microservice (e.g., NVIDIA NIM)), which can contain containers (e.g., operating system (OS) level virtualization packages) that may contain an application programming interface (API) layer, a server layer, a runtime layer, and / or a model "engine". For example, the inference microservice may include the container itself and the model 502 (e.g., weights and biases). In some cases, such as when the machine learning model 502 is small enough (e.g., has a sufficiently small number of parameters), the model may be included within the container itself. In some embodiments, the machine learning model 502 described herein can be deployed as an inference microservice to accelerate model deployment on any cloud, data center, or edge computing system while ensuring data security. For example, an inference microservice may include one or more APIs, pre-configured containers for simplified deployment, an optimized inference engine (e.g., built using standardized AI model deployment and execution software, such as NVIDIA's Triton Inference Server, and / or one or more APIs for high-performance deep learning inference, which may include inference runtime and model optimizations to provide low latency and high throughput for production applications, such as NVIDIA's TensorRT), and / or enterprise management data for telemetry (e.g., including identity, metrics, health checks, and / or monitoring). One or more machine learning models 502 described herein may be included as part of a microservice, along with an accelerated infrastructure capable of deployment with a single command, and / or orchestrated and automatically scaled (e.g., reaching data center scale on a single device) using a container orchestration system on the accelerated infrastructure. Therefore, an inference microservice may include one or more machine learning models 502 (e.g., models optimized for high-performance inference), inference runtime software for executing one or more machine learning models 502 and providing output / response to inputs (e.g., user queries, prompts, etc.), and enterprise management software for providing health checks, identity verification, and other monitoring. In some embodiments, the inference microservice may include software for performing in-situ replacements and / or updates to one or more machine learning models 502. During replacement or update, the software performing the replacement / update may maintain user configurations for the inference runtime software and enterprise management software.
[0076] Now for reference Figure 6 , Figure 6An example of a system 602 according to some embodiments of the present disclosure is shown, which can perform one or more processes described herein. As shown, system 602 (which may represent and / or include one or more example computing devices 1000 and / or example data centers 1100) may include one or more processors 604 (which may resemble and / or include one or more CPUs 1006 and / or one or more GPUs 1008) and memory 606 (which may resemble memory 1004 and / or include memory 1004). For example, memory 606 may store one or more components of one or more path prediction systems 102, such as default path component 106, intent component 108, scoring component 110, and path prediction component 112, as well as one or more machine learning models 502 and training engines 508. In addition, one or more processors 604 may execute the default path component 106, the intent component 108, the scoring component 110, and the path prediction component 112, as well as one or more machine learning models 502 and / or training engines 508, to perform one or more processes described herein.
[0077] For example, system 602 may receive input data 608 generated by one or more components 610 of one or more machines 612, which may correspond to machine 204 or machine 900. Input data 608 may include map data 114, route data 116, status data 118, sensor data 120, I / O data 122, positioning data 124, perception data, or any other data. System 602 may then process and evaluate the input data 608 to determine a predicted path for one or more machines 612. System 602 may send output data 614, which may include path prediction data 132 and / or any other outputs of one or more path prediction systems 102 described herein. The driving stack 104 of one or more machines 612 may use the output data 614 to control one or more operations of one or more machines 612. Although depicted as separate systems, in some examples, system 602 and one or more machines 612 may be the same or different systems. For example, processor 604 and memory 606 may be part of one or more machines 612 (e.g., included within a computing device of one or more machines 612).
[0078] Now for reference Figure 7 and Figure 8Each block of methods 700 and 800 described herein contains a computational process that can be executed using any combination of hardware, firmware, and / or software. For example, various functions can be implemented by a processor executing instructions stored in memory. These methods can also be embodied as computer-usable instructions stored on a computer storage medium. These methods can be provided by standalone applications, services, or managed services (standalone or in combination with other managed services) or plug-ins to another product, to name a few. Furthermore, methods 700 and 800 are implemented by combining... Figure 1 These methods are described by way of example. However, these methods may be additionally or alternatively implemented by any system or combination of systems, including but not limited to those described herein.
[0079] Figure 7 This is a flowchart illustrating an example of a method 700 for predicting a machine path according to some embodiments of the present disclosure. Method 700 may include at block B702: calculating a first plurality of scores indicating a first probability that a first plurality of road segments correspond to an intended path of the machine, the first plurality of road segments being associated with a first intersection. For example, scoring component 110 may calculate one or more scores 130 that may indicate the first probability that the first plurality of road segments correspond to an intended path of the machine. In some examples, the first plurality of scores may be calculated at least based on intent data indicating one or more occupant intentions for specific road segments within the first plurality of road segments.
[0080] Method 700 at box B704 may include: calculating a second plurality of scores indicating a second probability that a second plurality of road segments correspond to a machine's intended path, the second plurality of road segments being associated with a second intersection. For example, scoring component 110 may calculate one or more scores 130 that may indicate a second probability that the second plurality of road segments correspond to a machine's intended path. In some examples, the second plurality of scores may be calculated at least based on intent data indicating one or more occupant intentions for specific road segments within the second plurality of road segments.
[0081] Method 700 at box B706 may include: determining a predicted path for the machine based at least on a first plurality of scores and a second plurality of scores, the predicted path including at least a first road segment of the first plurality of road segments and a second road segment of the second plurality of road segments. For example, path prediction component 112 may generate path prediction data 132 based at least on one or more scores 130. Path prediction data 132 may indicate a predicted path for the machine, which may include a first road segment of a first plurality of road segments associated with a first intersection and a second road segment of a second plurality of road segments associated with a second intersection. In some examples, the first road segment may be the highest-scoring road segment of the first plurality of road segments, and the second road segment may be the highest-scoring road segment of the second plurality of road segments.
[0082] Method 700 at box B708 may include performing one or more machine-related operations based at least on the predicted path. For example, driving stack 104 may use path prediction data 132 to perform one or more machine-related operations. In some examples, these operations may include, but are not limited to: determining road curvature, incorporating road curvature into intersections in map data, using the predicted path to perform machine localization to determine which road segment the machine is operating on, accessing map data attributes along the predicted path to resolve conflicts between speed limits and perception results, acquiring events along the predicted path (e.g., T-junctions and toll booths), etc. Furthermore, in some examples, driving stack 104 may use the predicted path, curvature, speed limits, etc., to adjust the machine's braking, acceleration, steering angle, etc.—all of which may be included in and / or determined using path prediction data 132.
[0083] Figure 8 A flowchart illustrating an example of a method 800 for scoring an intersection ahead of a machine using occupant intent, according to some embodiments of the present disclosure. Method 800 may include, at block B802, determining, based at least one of map data or perception data, that the machine's location is within a threshold distance of one or more intersections containing one or more options for a predicted path. For example, intent component 108 may use map data 114 and / or location data 124 to determine that the machine's location is within a threshold distance of one or more intersections containing one or more options for a predicted path. In some examples, the threshold distance may be the actual distance between the machine's location and one or more locations of the intersection, and / or may be a time period (e.g., 5 seconds, 8 seconds, etc.) associated with the machine's arrival at one or more locations of the intersection. That is, if the machine will arrive at the intersection within the threshold time period (e.g., the next 5 seconds), the machine is likely within the threshold distance of one or more intersections. Furthermore, in some examples, intent component 108 may determine one or more occupant intents applicable to one or more options. In some examples, one or more options may include one or more road segments that the machine may traverse.
[0084] Method 800 at box B804 may include: calculating multiple scores, at least based on one or more occupant intentions corresponding to a plurality of options, indicating whether a respective option among the plurality of options corresponds to a machine intention path. For example, scoring component 110 may calculate score 130, which may indicate whether a respective option corresponds to a machine intention path. In some examples, one or more scores may include intention scores determined at least based on intention data determined by intention component 108. In some examples, one or more scores may be aggregated across multiple frames, and one or more options with the highest scores may be selected for a predicted path of the machine, which preferably corresponds to a machine intention path.
[0085] Method 800 at box B806 may include using multiple scores to determine a predicted path for the machine. For example, path prediction component 112 may use one or more scores 130 to determine a predicted path. In some examples, path prediction component 112 may iteratively select the road segment with the highest score at each intersection. For example, path prediction component 112 may select the first highest-scoring road segment at the first intersection closest to the machine, then "follow" the first road segment to the next intersection, and select a second road segment as the highest-scoring road segment for the next intersection, and so on, until a predicted path is determined.
[0086] Method 800 at box B808 may include performing one or more machine-related operations based at least on the machine's predicted path. For example, driving stack 104 may use path prediction data 132 to perform one or more machine-related operations. In some examples, one or more operations may include, but are not limited to: determining road curvature, incorporating road curvature into intersections in map data, using the predicted path to perform machine localization to determine which road segment the machine is operating on, accessing map data attributes along the predicted path to resolve conflicts between speed limits and perception results, acquiring events along the predicted path (e.g., T-junctions and toll booths), etc. Furthermore, in some examples, driving stack 104 may use the predicted path, curvature, speed limits, etc., to adjust the machine's braking, acceleration, steering angle, etc.—all of which may be included in and / or determined using path prediction data 132.
[0087] Example autonomous vehicles
[0088] Figure 9AThis is an illustration of an example autonomous vehicle 900 according to some embodiments of the present disclosure. The autonomous vehicle 900 (also referred to herein as “vehicle 900”) may include, but is not limited to, passenger vehicles such as automobiles, trucks, buses, ambulances, shuttles, electric or motorized bicycles, motorcycles, fire trucks, police cars, ambulances, boats, engineering vehicles, submarines, robotic vehicles, aircraft, drones, trailer-mounted vehicles (e.g., semi-trailer trucks for transporting goods), and / or other types of vehicles (e.g., driverless and / or capable of accommodating one or more passengers). Autonomous vehicles are generally described according to the level of automation defined by the National Highway Traffic Safety Administration (NHTSA) of the U.S. Department of Transportation and the Society of Automotive Engineers (SAE) in their standard “Classification and Definition of Terms Related to Driving Automation Systems for Road Motor Vehicles” (Standard No. J3016-201806, published June 15, 2018; Standard No. J3016-201609, published September 30, 2016; and previous and future versions of this standard). Vehicle 900 may be capable of having one or more functions according to Level 3-5 of the autonomous driving level. Vehicle 900 may be capable of functioning according to one or more Level 1-5 of the autonomous driving level. For example, vehicle 900 may be able to provide driver assistance (Level 1), partial automation (Level 2), conditional automation (Level 3), high automation (Level 4), and / or full automation (Level 5), depending on the embodiment. As used herein, the term "autonomy" may include any and / or all types of autonomy for vehicle 900 or other machines, such as full autonomy, high autonomy, conditional autonomy, partial autonomy, providing assisted autonomy, semi-autonomy, primary autonomy, or other names.
[0089] Vehicle 900 may include components such as chassis, body, wheels (e.g., 2, 4, 6, 8, 18, etc.), tires, axles, and other vehicle components. Vehicle 900 may include a propulsion system 950, such as an internal combustion engine, a hybrid power plant, an all-electric motor, and / or another type of propulsion system. Propulsion system 950 may be connected to the drivetrain of vehicle 900, which may include a transmission, to enable propulsion of vehicle 900. Propulsion system 950 may be controlled in response to receiving a signal from throttle / accelerator 952.
[0090] A steering system 954, which may include a steering wheel, can be used to steer the vehicle 900 (e.g., along a desired path or route) when the propulsion system 950 is operating (e.g., when the vehicle is in motion). The steering system 954 may receive signals from the steering actuator 956. For fully automatic (level 5) functions, the steering wheel may be optional.
[0091] The brake sensor system 946 can be used to operate the vehicle brakes in response to receiving signals from the brake actuator 948 and / or the brake sensor.
[0092] It can include one or more System-on-a-Chip (SoC) 904 ( Figure 9C One or more controllers 936, including one or more GPUs, can provide signals (e.g., signals representing commands) to one or more components and / or systems of vehicle 900. For example, one or more controllers can send signals to operate vehicle brakes via one or more brake actuators 948, to operate steering system 954 via one or more steering actuators 956, and to operate propulsion system 950 via one or more throttles / accelerators 952. One or more controllers 936 may include one or more onboard (e.g., integrated) computing devices (e.g., supercomputers) that process sensor signals and output operating commands (e.g., signals representing commands) to enable autonomous driving and / or assist a human driver in driving vehicle 900. One or more controllers 936 may include a first controller 936 for autonomous driving functions, a second controller 936 for functional safety functions, a third controller 936 for artificial intelligence functions (e.g., computer vision), a fourth controller 936 for infotainment functions, a fifth controller 936 for redundancy in emergency situations, and / or other controllers. In some examples, a single controller 936 can handle two or more of the functions described above, two or more controllers 936 can handle a single function, and / or any combination thereof.
[0093] One or more controllers 936 may provide signals for controlling one or more components and / or systems of vehicle 900 in response to sensor data (e.g., sensor inputs) received from one or more sensors. Sensor data may be received from, for example, but not limited to, a global navigation satellite system sensor 958 (e.g., a Global Positioning System sensor), a RADAR sensor 960, an ultrasonic sensor 962, a LIDAR sensor 964, an inertial measurement unit (IMU) sensor 966 (e.g., an accelerometer, gyroscope, magnetic compass, magnetometer, etc.), a microphone 996, a stereo camera 968, a wide-angle camera 970 (e.g., a fisheye camera), an infrared camera 972, a surround camera 974 (e.g., a 360-degree camera), a long-range and / or medium-range camera 998, a speed sensor 944 (e.g., for measuring the rate of vehicle 900), a vibration sensor 942, a steering sensor 940, a braking sensor (e.g., as part of a braking sensor system 946), and / or other sensor types.
[0094] One or more controllers 936 may receive inputs (e.g., represented by input data) from the instrument cluster 932 of the vehicle 900 and provide outputs (e.g., represented by output data, display data, etc.) via a human-machine interface (HMI) display 934, an auditory signaling device, a speaker, and / or via other components of the vehicle 900. These outputs may include information such as vehicle speed, rate, time, map data (e.g., [missing information]). Figure 9C Information such as a high-definition (“HD”) map 922, location data (e.g., the location of vehicle 900 on the map), direction, the location of other vehicles (e.g., occupying a grid), and information about objects and their states perceived by the controller 936, etc. For example, the HMI display 934 may display information about the existence of one or more objects (e.g., street signs, warning signs, traffic light changes, etc.) and / or information about driving maneuvers that the vehicle has made, is making, or will make (e.g., changing lanes now, leaving 34B in two miles, etc.).
[0095] The vehicle 900 further includes a network interface 924, which can communicate via one or more networks using one or more wireless antennas 926 and / or a modem. For example, the network interface 924 may be able to communicate via Long Term Evolution (“LTE”), Wideband Code Division Multiple Access (“WCDMA”), Universal Mobile Telecommunications System (“UMTS”), Global System for Mobile Communications (“GSM”), IMT-CDMA Multicarrier (“CDMA2000”), etc. One or more wireless antennas 926 may also enable communication between objects in the environment (e.g., vehicles, mobile devices, etc.) using one or more local area networks such as Bluetooth, Bluetooth Low Energy (“LE”), Z-Wave, ZigBee, etc., and / or one or more low-power wide area networks (LPWANs such as LoRaWAN, SigFox, etc.).
[0096] Figure 9B For use in accordance with some embodiments of this disclosure Figure 9A This is an example of the camera position and field of view of an example autonomous vehicle 900. The camera and its respective field of view are an example embodiment and are not intended to be limiting. For example, additional and / or replaceable cameras may be included, and / or these cameras may be located at different positions on the vehicle 900.
[0097] The camera type used for the camera may include, but is not limited to, a digital camera suitable for use with components and / or systems of vehicle 900. The camera may operate at Automotive Safety Integrity Level (ASIL) B and / or another ASIL. The camera type may have any image capture rate, such as 60 frames per second (fps), 120 fps, 240 fps, etc., depending on the embodiment. The camera may be able to use a rolling shutter, a global shutter, another type of shutter, or a combination thereof. In some examples, the color filter array may include a red-transparent (RCCC) color filter array, a red-transparent-blue (RCCB) color filter array, a red-blue-green (RBGC) color filter array, a Foveon X3 color filter array, a Bayer sensor (RGGB) color filter array, a monochrome sensor color filter array, and / or another type of color filter array. In some embodiments, a high-resolution camera, such as a camera with RCCC, RCCB, and / or RBGC color filter arrays, may be used in efforts to improve light sensitivity.
[0098] In some examples, one or more of the cameras can be used to perform advanced driver assistance system (ADAS) functions (e.g., as part of a redundant or fail-safe design). For example, a multi-function monocular camera can be installed to provide functions including lane departure warning, traffic sign assistance, and intelligent headlight control. One or more of the cameras (e.g., all cameras) can simultaneously record and provide image data (e.g., video).
[0099] One or more of the cameras can be mounted in mounting components such as custom-designed (3D-printed) components to cut off stray light and reflections from inside the vehicle (e.g., reflections from the dashboard reflected in the windshield mirror) that may interfere with the camera's image data capture capabilities. Regarding the wing mirror mounting components, the wing mirror components can be custom-3D printed so that the camera mounting plate matches the shape of the wing mirror. In some examples, one or more cameras can be integrated into the wing mirror. For side-view cameras, one or more cameras can also be integrated into the four pillars at each corner of the cab.
[0100] A camera with a field of view that includes the environment in front of the vehicle 900 (e.g., a front-facing camera) can be used for surround view to help identify forward paths and obstacles, and, with the assistance of one or more controllers 936 and / or control SoCs, to provide information crucial for generating an occupancy grid and / or determining a preferred vehicle path. The front-facing camera can be used to perform many of the same ADAS functions as LiDAR, including emergency braking, pedestrian detection, and collision avoidance. The front-facing camera can also be used in ADAS functions and systems, including Lane Departure Warning (“LDW”), Autonomous Cruise Control (“ACC”), and / or other functions such as traffic sign recognition.
[0101] A variety of cameras can be used in front-facing configurations, including, for example, monocular camera platforms including complementary metal-oxide-semiconductor (“CMOS”) color imagers. Another example could be a wide-angle camera 970, which can be used to perceive objects entering the field of view from the periphery (such as pedestrians, traffic at intersections, or bicycles). Although Figure 9B The middle image shows only one wide-angle camera, but any number (including zero) of wide-angle cameras 970 can exist on vehicle 900. Furthermore, any number of remote cameras 998 (e.g., long-view stereo camera pairs) can be used for depth-based object detection, especially for objects for which neural networks have not yet been trained. Remote cameras 998 can also be used for object detection and classification, as well as basic object tracking.
[0102] Any number of stereo cameras 968 can also be included in the front-mounted configuration. In at least one embodiment, one or more of the stereo cameras 968 may include an integrated control unit that includes a scalable processing unit that can provide a multi-core microprocessor and programmable logic (FPGA) with an integrated controller area network (“CAN”) or Ethernet interface on a single chip. Such a unit can be used to generate a 3D map of the vehicle environment, including distance estimates for all points in the image. Alternative stereo cameras 968 may include a compact stereo vision sensor that may include two camera lenses (one on each side) and an image processing chip that can measure the distance from the vehicle to a target object and use the generated information (e.g., metadata) to activate autonomous emergency braking and lane departure warning functions. Other types of stereo cameras 968 may be used in addition to those described herein or alternatively.
[0103] Cameras with a field of view including the side portion of the vehicle 900 (e.g., side-view cameras) can be used for surround view, providing information for creating and updating occupancy grids and generating side-impact collision warnings. For example, surround camera 974 (e.g., ... Figure 9B The four surround cameras 974 shown can be mounted on the vehicle 900. The surround cameras 974 can include wide-angle cameras 970, fisheye cameras, 360-degree cameras, and / or the like. Four examples are provided; the four fisheye cameras can be positioned at the front, rear, and sides of the vehicle. In an alternative arrangement, the vehicle can use three surround cameras 974 (e.g., left, right, and rear) and can utilize one or more other cameras (e.g., forward-facing cameras) as a fourth surround-view camera.
[0104] A camera with a field of view that includes the environment behind the vehicle 900 (e.g., a rear-view camera) can be used for parking assistance, surround view, rear collision warning, and creating and updating occupancy grids. A wide variety of cameras can be used, including but not limited to those also suitable as front-facing cameras as described herein (e.g., long-range and / or mid-range camera 998, stereo camera 968, infrared camera 972, etc.).
[0105] Figure 9C For use in accordance with some embodiments of this disclosure Figure 9A The example autonomous vehicle 900 is illustrated in the block diagram of an example system architecture. It should be understood that this arrangement, and other arrangements described herein, are merely illustrative. Other arrangements and elements (e.g., machines, interfaces, functions, sequences, functional groupings, etc.) may be used in addition to or in place of those shown, and some elements may be omitted entirely. Furthermore, many of the elements described herein are functional entities, which may be implemented as discrete or distributed components or in combination with other components, and in any suitable combination and location. The various functions described herein as being performed by these entities can be implemented via hardware, firmware, and / or software. For example, the various functions can be implemented by a processor executing instructions stored in memory.
[0106] Figure 9C Each component, feature, and system in vehicle 900 is illustrated as being connected via bus 902. Bus 902 may include a Controller Area Network (CAN) data interface (or, alternatively, referred to herein as the "CAN bus"). CAN may be a network within vehicle 900 used to assist in the control of various features and functions of vehicle 900, such as the actuation of brakes, acceleration, braking, steering, windshield wipers, etc. The CAN bus can be configured to have dozens or even hundreds of nodes, each with its own unique identifier (e.g., CAN ID). The CAN bus can be read to find steering wheel angle, ground speed, engine speed per minute (RPM), button positions, and / or other vehicle status indicators. The CAN bus may be ASIL B compliant.
[0107] Although bus 902 is described herein as a CAN bus, this is not intended to be limiting. For example, FlexRay and / or Ethernet may be used in addition to or alternatively to a CAN bus. Furthermore, although bus 902 is represented by a single line, this is not intended to be limiting. For example, any number of buses 902 may exist, which may include one or more CAN buses, one or more FlexRay buses, one or more Ethernet buses, and / or one or more other types of buses using different protocols. In some examples, two or more buses 902 may be used to perform different functions and / or may be used for redundancy. For example, a first bus 902 may be used for collision avoidance functions, and a second bus 902 may be used for drive control. In any example, each bus 902 may communicate with any component of vehicle 900, and two or more buses 902 may communicate with the same component. In some examples, each SoC 904, each controller 936, and / or each computer within the vehicle may have access to the same input data (e.g., input from sensors of vehicle 900) and may be connected to a common bus such as a CAN bus.
[0108] Vehicle 900 may include one or more controllers 936, such as those described herein. Figure 9A The controllers described. Controller 936 can be used for a wide variety of functions. Controller 936 can be coupled to any other different components and systems of vehicle 900 and can be used for the control of vehicle 900, artificial intelligence of vehicle 900, infotainment and / or the like for vehicle 900.
[0109] Vehicle 900 may include one or more System-on-Chip (SoC) 904s. SoC 904 may include a CPU 906, GPU 908, processor 910, cache 912, accelerator 914, data storage 916, and / or other components and features not shown. SoC 904 can be used to control vehicle 900 across a wide variety of platforms and systems. For example, one or more SoCs 904s may be combined with an HD map 922 in a system (e.g., the system of vehicle 900), the HD map being transmitted via a network interface 924 from one or more servers (e.g., [server name missing]). Figure 9D One or more servers (978) receive map refresh and / or updates.
[0110] The CPU 906 may include a CPU cluster or a CPU complex (or, alternatively, referred to herein as "CCPLEX"). The CPU 906 may include multiple cores and / or L2 cache. For example, in some embodiments, the CPU 906 may include eight cores in a coherent multiprocessor configuration. In some embodiments, the CPU 906 may include four dual-core clusters, each with a dedicated L2 cache (e.g., 2MB L2 cache). The CPU 906 (e.g., CCPLEX) may be configured to support simultaneous cluster operation, such that any combination of CPU 906 clusters can be active at any given time.
[0111] The CPU 906 can implement power management capabilities including one or more of the following features: automatic clock gating of hardware blocks when idle to conserve dynamic power; clock gating of each core when the core is not actively executing instructions due to the execution of WFI / WFE instructions; independent power gating of each core; independent clock gating of each core cluster when all cores are clock-gated or power-gated; and / or independent power gating of each core cluster when all cores are power-gated. The CPU 906 can further implement enhanced algorithms for managing power states, where allowed power states and desired wake-up times are specified, and the hardware / microcode determines the optimal power state to enter for the core, cluster, and CCPLEX. The processing core can support simplified power state entry sequences in software, with this work offloaded to the microcode.
[0112] The GPU 908 may include an integrated GPU (or, alternatively, referred to herein as an "iGPU"). The GPU 908 may be programmable and efficient for parallel workloads. In some examples, the GPU 908 may use an enhanced tensor instruction set. The GPU 908 may include one or more streaming microprocessors, each of which may include an L1 cache (e.g., an L1 cache with at least 96KB of storage capacity), and two or more of these streaming microprocessors may share an L2 cache (e.g., an L2 cache with 512KB of storage capacity). In some embodiments, the GPU 908 may include at least eight streaming microprocessors. The GPU 908 may use a computation application programming interface (API). Furthermore, the GPU 908 may use one or more parallel computing platforms and / or programming models (e.g., NVIDIA's CUDA).
[0113] In automotive and embedded applications, the GPU 908 can be power-optimized for optimal performance. For example, the GPU 908 can be fabricated on FinFETs. However, this is not intended to be limiting, and the GPU 908 can be fabricated using other semiconductor manufacturing processes. Each streaming microprocessor can combine several mixed-precision processing cores divided into multiple blocks. For example, and without limitation, 64 PF32 cores and 32 PF64 cores can be divided into four processing blocks. In such an example, each processing block can be allocated 16 FP32 cores, 8 FP64 cores, 16 INT32 cores, two mixed-precision NVIDIA Tensor cores for deep learning matrix arithmetic, an L0 instruction cache, a warp scheduler, dispatch units, and / or a 64KB register file. Furthermore, the streaming microprocessor can include independent parallel integer and floating-point data paths to leverage the mixture of computation and addressing computations for efficient execution of workloads. The streaming microprocessor can include independent thread scheduling capabilities to allow for finer-grained synchronization and cooperation between parallel threads. Streaming microprocessors can include a combination of L1 data cache and shared memory units to improve performance while simplifying programming.
[0114] The GPU 908 may include, in some examples, a high-bandwidth memory (HBM) and / or a 16GB HBM2 memory subsystem providing peak memory bandwidth of approximately 900GB / s. In some examples, in addition to HBM memory or alternatively, synchronous graphics random access memory (SGRAM), such as fifth-generation graphics double data rate synchronous random access memory (GDDR5), may be used.
[0115] The GPU 908 may include unified memory technology, which includes access counters to allow memory pages to be migrated more precisely to the processors that access them most frequently, thereby improving the efficiency of shared memory ranges between processors. In some examples, Address Translation Service (ATS) support can be used to allow the GPU 908 to directly access the CPU 906 page tables. In such examples, when the GPU 908 Memory Management Unit (MMU) experiences a miss, the address translation request can be transferred to the CPU 906. In response, the CPU 906 can look up the virtual-physical mapping for the address in its page tables and transfer the translation back to the GPU 908. Thus, unified memory technology can allow a single unified virtual address space for the memory of both the CPU 906 and GPU 908, simplifying GPU 908 programming and porting applications to the GPU 908.
[0116] In addition, the GPU 908 may include access counters that track how frequently the GPU 908 accesses the memory of other processors. These access counters help ensure that memory pages are moved to the physical memory of the processor that accesses those pages most frequently.
[0117] SoC 904 may include any number of caches 912, including those described herein. For example, cache 912 may include an L3 cache available to both CPU 906 and GPU 908 (e.g., it is connected to both CPU 906 and GPU 908). Cache 912 may include a write-back cache, which can track the state of rows, for example, using a cache coherence protocol (e.g., MEI, MESI, MSI, etc.). Depending on the embodiment, the L3 cache may include 4MB or more, but a smaller cache size may also be used.
[0118] The SoC 904 may include an arithmetic logic unit (ALU), which can be used to perform processing of any of a variety of tasks or operations related to the vehicle 900—such as processing a DNN. Additionally, the SoC 904 may include a floating-point unit (FPU)—or other mathematical coprocessor or digital coprocessor type—for performing mathematical operations within the system. For example, the SoC 904 may include one or more FPUs integrated as execution units within the CPU 906 and / or GPU 908.
[0119] SoC 904 may include one or more accelerators 914 (e.g., hardware accelerators, software accelerators, or a combination thereof). For example, SoC 904 may include a hardware acceleration cluster, which may include optimized hardware accelerators and / or large on-chip memory. This large on-chip memory (e.g., 4MB SRAM) can enable the hardware acceleration cluster to accelerate neural networks and other computations. The hardware acceleration cluster can be used to supplement GPU 908 and offload some tasks from GPU 908 (e.g., freeing up more cycles of GPU 908 to perform other tasks). As an example, accelerator 914 can be used for targeted workloads (e.g., perception, convolutional neural networks (CNNs), etc.) that are stable enough to be easily controlled for acceleration. When used herein, the term "CNN" can include all types of CNNs, including region-based or region convolutional neural networks (RCNNs) and fast RCNNs (e.g., for object detection).
[0120] Accelerator 914 (e.g., hardware acceleration clusters) may include a Deep Learning Accelerator (DLA). A DLA may include one or more Tensor Processing Units (TPUs) that can be configured to provide an additional 10 trillion operations per second for deep learning applications and inference. TPUs may be accelerators configured to perform image processing functions (e.g., for CNNs, RCNNs, etc.) and optimized for performing image processing functions. DLAs may be further optimized for a specific set of neural network types and floating-point operations as well as inference. DLAs are designed to provide higher performance per millimeter than general-purpose GPUs and significantly outperform CPUs. TPUs can perform several functions, including single-instance convolution functions, support for INT8, INT16, and FP16 data types for both features and weights, and post-processor functions.
[0121] DLA can execute neural networks, especially CNNs, quickly and efficiently on processed or unprocessed data for any function across a wide variety of applications, such as, but not limited to: CNNs for object recognition and detection using data from camera sensors; CNNs for distance estimation using data from camera sensors; CNNs for emergency vehicle detection and recognition using data from microphones; CNNs for face recognition and vehicle owner recognition using data from camera sensors; and / or CNNs for safety and / or safety-related events.
[0122] The DLA can perform any function of the GPU 908, and by using inference accelerators, for example, designers can target either the DLA or the GPU 908 for any function. For instance, a designer can focus the CNN processing and floating-point operations on the DLA and leave other functions to the GPU 908 and / or other accelerators 914.
[0123] Accelerator 914 (e.g., a hardware acceleration cluster) may include a programmable vision accelerator (PVA), which may alternatively be referred to herein as a computer vision accelerator. A PVA can be designed and configured to accelerate computer vision algorithms for advanced driver assistance systems (ADAS), autonomous driving, and / or augmented reality (AR) and / or virtual reality (VR) applications. A PVA can provide a balance between performance and flexibility. For example, each PVA may include, for example, but not limited to, any number of reduced instruction set computer (RISC) cores, direct memory access (DMA), and / or any number of vector processors.
[0124] RISC cores can interact with image sensors (such as the image sensor of any camera described herein), image signal processors, and / or the like. Each of these RISC cores may include any amount of memory. Depending on the embodiment, the RISC core may use any of several protocols. In some examples, the RISC core may execute a real-time operating system (RTOS). RISC cores may be implemented using one or more integrated circuit devices, application-specific integrated circuits (ASICs), and / or memory devices. For example, a RISC core may include an instruction cache and / or tightly coupled RAM.
[0125] DMA enables PVA components to access system memory independently of the CPU 906. DMA can support any number of features to provide optimizations to the PVA, including but not limited to support for multidimensional addressing and / or circular addressing. In some examples, DMA can support addressing in up to six or more dimensions, which can include block width, block height, block depth, horizontal block step, vertical block step, and / or depth step.
[0126] A vector processor can be a programmable processor designed to efficiently and flexibly execute programming for computer vision algorithms and provide signal processing capabilities. In some examples, a PVA may include a PVA core and two vector processing subsystem partitions. The PVA core may include a processor subsystem, one or more DMA engines (e.g., two DMA engines), and / or other peripherals. The vector processing subsystem may operate as the main processing engine of the PVA and may include a vector processing unit (VPU), an instruction cache, and / or a vector memory (e.g., a VMEM). The VPU core may include a digital signal processor, such as, for example, a Single Instruction Multiple Data (SIMD) or Very Long Instruction Word (VLIW) digital signal processor. The combination of SIMD and VLIW can enhance throughput and speed.
[0127] Each of the vector processors may include an instruction cache and may be coupled to dedicated memory. Consequently, in some examples, each of the vector processors may be configured to execute independently of other vector processors. In other examples, the vector processors included in a particular PVA may be configured to employ data parallelism. For example, in some embodiments, multiple vector processors included in a single PVA may execute the same computer vision algorithm, but on different regions of an image. In other examples, vector processors included in a particular PVA may execute different computer vision algorithms simultaneously on the same image, or even different algorithms on a sequence of images or portions of an image. Among other things, any number of PVAs may be included in a hardware-accelerated cluster, and any number of vector processors may be included in each of these PVAs. Furthermore, the PVA may include additional error-correcting code (ECC) memory to enhance overall system security.
[0128] Accelerator 914 (e.g., a hardware acceleration cluster) may include an on-chip computer vision network and SRAM to provide high-bandwidth, low-latency SRAM for accelerator 914. In some examples, on-chip memory may include at least 4MB of SRAM consisting of, for example, but not limited to, eight field-configurable memory blocks accessible by both PVA and DLA. Each pair of memory blocks may include an Advanced Peripheral Bus (APB) interface, configuration circuitry, a controller, and a multiplexer. Any type of memory can be used. PVA and DLA can access memory via a backbone that provides high-speed memory access to PVA and DLA. The backbone may include (e.g., using an APB) an on-chip computer vision network that interconnects PVA and DLA to memory.
[0129] On-chip computer vision networks can include interfaces that ensure both the PVA and DLA provide ready and valid signals before transmitting any control signals / addresses / data. Such interfaces can provide separate phases and channels for transmitting control signals / addresses / data, as well as burst communication for continuous data transmission. This type of interface can conform to ISO 26262 or IEC 61508 standards, but other standards and protocols can also be used.
[0130] In some examples, the SoC 904 may include, for example, a real-time ray tracing hardware accelerator as described in U.S. Patent Application No. 16 / 101,232, filed August 10, 2018. This real-time ray tracing hardware accelerator can be used to quickly and efficiently determine the location and extent of objects (e.g., within a world model) to generate real-time visualization simulations for RADAR signal interpretation, sound propagation synthesis and / or analysis, SONAR system simulation, general wave propagation simulation, comparison with LiDAR data for localization and / or other functional purposes, and / or for other uses. In some embodiments, one or more Tree Traversal Units (TTUs) may be used to perform one or more ray tracing-related operations.
[0131] Accelerator 914 (e.g., hardware accelerator clusters) has broad applications in autonomous driving. PVAs can be programmable vision accelerators used in critical processing stages of ADAS and autonomous vehicles. The capabilities of PVAs are a good match for algorithmic domains requiring predictable processing, low power, and low latency. In other words, PVAs perform well in semi-dense or dense rule computation, even on small datasets requiring predictable runtimes with low latency and low power. Therefore, in the context of platforms for autonomous vehicles, PVAs are designed to run classical computer vision algorithms because they are efficient in object detection and integer mathematical operations.
[0132] For example, according to one embodiment of this technology, PVA is used to perform computer stereo vision. In some examples, semi-global matching-based algorithms may be used, but this is not intended to be limiting. Many applications for Level 3-5 autonomous driving require instantaneous motion estimation / stereo matching (e.g., from moving structures, pedestrian recognition, lane detection, etc.). PVA can perform computer stereo vision functions on input from two monocular cameras.
[0133] In some examples, PVA can be used to perform intensive optical flow, providing processed RADAR data from the raw RADAR data (e.g., using 4D Fast Fourier Transform). In other examples, PVA is used for time-of-flight depth processing, which, for example, involves processing raw time-of-flight data to provide processed time-of-flight data.
[0134] DLA can be used to run any type of network to enhance control and driving safety, including, for example, neural networks that output a confidence metric for each object detection. Such a confidence value can be interpreted as a probability or as providing a relative “weight” for each detection compared to other detections. This confidence value allows the system to make further decisions about which detections should be considered true positives rather than false positives. For example, the system can set a threshold for the confidence and only consider detections exceeding the threshold as true positives. In an Automatic Emergency Braking (AEB) system, false positives can cause the vehicle to automatically perform emergency braking, which is clearly undesirable. Therefore, only the most confident detections should be considered as triggers for AEB. DLA can run a neural network to regress the confidence value. This neural network can take at least a subset of parameters as input, such as bounding box dimensions, ground plane estimates (e.g., from another subsystem), inertial measurement unit (IMU) sensor 966 outputs related to vehicle orientation and distance, 3D position estimates of objects obtained from the neural network and / or other sensors (e.g., LiDAR sensor 964 or RADAR sensor 960), etc.
[0135] The SoC 904 may include one or more data storage units 916 (e.g., memory). The data storage unit 916 may be on-chip memory of the SoC 904, which may store neural networks to be executed on the GPU and / or DLA. In some examples, for redundancy and security, the data storage unit 916 may be large enough to store multiple instances of the neural network. The data storage unit 916 may include L2 or L3 cache 912. References to the data storage unit 916 may include references to memory associated with PVA, DLA, and / or other accelerators 914 as described herein.
[0136] The SoC 904 may include one or more processors 910 (e.g., embedded processors). Processor 910 may include a startup and power management processor, which may be a dedicated processor and subsystem for handling startup power and management functions, as well as safety implementation. The startup and power management processor may be part of the SoC 904 startup sequence and may provide runtime power management services. The startup power and management processor may provide clock and voltage programming, auxiliary system low-power state transitions, SoC 904 thermal and temperature sensor management, and / or SoC 904 power state management. Each temperature sensor may be implemented as a ring oscillator whose output frequency is proportional to the temperature, and the SoC 904 may use the ring oscillator to detect the temperature of the CPU 906, GPU 908, and / or accelerator 914. If it is determined that the temperature exceeds a threshold, the startup and power management processor may enter a temperature fault routine and place the SoC 904 into a lower power state and / or place the vehicle 900 into a driver-safe parking mode (e.g., safely stopping the vehicle 900).
[0137] The processor 910 may further include a set of embedded processors that can be used as an audio processing engine. The audio processing engine can be an audio subsystem that allows for full hardware support for multi-channel audio via multiple interfaces, as well as a wide and flexible range of audio I / O interfaces. In some examples, the audio processing engine is a dedicated processor core with a digital signal processor and dedicated RAM.
[0138] The processor 910 may further include an always-on-processor engine that can provide the necessary hardware features to support low-power sensor management and wake-up use cases. This always-on-processor engine may include a processor core, tightly coupled RAM, support for peripherals (such as timers and interrupt controllers), various I / O controller peripherals, and routing logic.
[0139] The processor 910 may further include a secure cluster engine, which includes a dedicated processor subsystem for handling security management for automotive applications. The secure cluster engine may include two or more processor cores, tightly coupled RAM, support for peripheral devices (e.g., timers, interrupt controllers, etc.), and / or routing logic. In secure mode, the two or more cores may operate in lockstep mode and function as a single core with comparison logic that detects any differences between their operations.
[0140] The processor 910 may further include a real-time camera engine, which may include a dedicated processor subsystem for handling real-time camera management.
[0141] The processor 910 may further include a high dynamic range signal processor, which may include an image signal processor, which is a hardware engine that is part of the camera processing pipeline.
[0142] Processor 910 may include a video image compositer, which may be (e.g., implemented on a microprocessor) a processing block, implementing video post-processing functions required by the video playback application to generate the final image for the player window. The video image compositer may perform lens distortion correction on the wide-angle camera 970, the surround camera 974, and / or the in-cabin monitoring camera sensor. The in-cabin monitoring camera sensor is preferably monitored by a neural network running on another instance of an advanced SoC, configured to recognize in-cabin events and respond accordingly. The in-cabin system may perform lip reading to activate mobile phone services and make calls, dictate emails, change vehicle destinations, activate or change the vehicle's infotainment system and settings, or provide voice-activated web browsing. Some functions are only available to the driver when the vehicle is operating in autonomous mode and are disabled in other situations.
[0143] Video image compositers can include enhanced temporal denoising for both spatial and temporal noise reduction. For example, in the case of motion in the video, denoising appropriately weights spatial information, reducing the weight of information provided by neighboring frames. In cases where the image or part of the image does not contain motion, the temporal denoising performed by the video image compositer can use information from previous images to reduce noise in the current image.
[0144] The video image compositer can also be configured to perform stereo correction on input stereo lens frames. When the operating system desktop is in use and the GPU 908 does not need to continuously render new surfaces, the video image compositer can be further used for user interface components. Even when the GPU 908 is powered on and active, performing 3D rendering, the video image compositer can be used to offload the GPU 908 to improve performance and responsiveness.
[0145] The SoC 904 may further include a Mobile Industry Processor Interface (MIPI) camera serial interface, a high-speed interface, and / or a video input block that can be used for camera and related pixel input functions for receiving video and input from a camera. The SoC 904 may further include an input / output controller that can be software-controlled and can be used to receive I / O signals not assigned to a specific role.
[0146] SoC 904 may further include a wide range of peripheral interfaces to enable communication with peripherals, audio codecs, power management and / or other devices. SoC 904 can be used to process data from cameras and sensors (e.g., LIDAR sensor 964, RADAR sensor 960, etc., which can be connected via Gigabit Multimedia Serial Link and Ethernet), data from bus 902 (e.g., vehicle 900 speed, steering wheel position, etc.), and data from GNSS sensor 958 (connected via Ethernet or CAN bus). SoC 904 may further include a dedicated high-performance, high-capacity memory controller, which may include its own DMA engine, and which can be used to free up CPU 906 from routine data management tasks.
[0147] The SoC 904 can be an end-to-end platform with a flexible architecture spanning Automation Levels 3-5, providing a comprehensive functional safety architecture that leverages and efficiently utilizes computer vision and ADAS technologies for diversity and redundancy, along with deep learning tools to deliver a flexible and reliable driving software stack. The SoC 904 can be faster, more reliable, and even more energy- and space-efficient than conventional systems. For example, when combined with the CPU 906, GPU 908, and data storage 916, the accelerator 914 can provide a fast and efficient platform for Level 3-5 autonomous vehicles.
[0148] Therefore, this technology offers capabilities and functionalities that cannot be achieved through conventional systems. For example, computer vision algorithms can be executed on CPUs, which can be configured using high-level programming languages such as C to execute a wide variety of processing algorithms across a diverse range of visual data. However, CPUs often cannot meet the performance requirements of many computer vision applications, such as those related to execution time and power consumption. In particular, many CPUs cannot execute complex object detection algorithms in real time, which is a requirement for automotive ADAS applications and practical Level 3-5 autonomous vehicles.
[0149] In contrast to conventional systems, the techniques described in this paper, by providing CPU complexes, GPU complexes, and hardware acceleration clusters, allow multiple neural networks to be executed simultaneously and / or sequentially, and the results to be combined to achieve Level 3–5 autonomous driving capabilities. For example, a CNN executed on a DLA or dGPU (e.g., GPU 920) could include text and word recognition, allowing a supercomputer to read and understand traffic signs, including those for which neural networks have not yet been specifically trained. The DLA could further include a neural network capable of recognizing, interpreting, and providing semantic understanding of the signs, and passing that semantic understanding to a path planning module running on a CPU complex.
[0150] As another example, multiple neural networks can operate simultaneously, as required for Level 3, 4, or 5 driving. For instance, a warning sign consisting of "Caution: Flashing lights indicate icy conditions" along with a light can be interpreted independently or jointly by several neural networks. The sign itself can be recognized as a traffic sign by a deployed first neural network (e.g., a trained neural network), and the text "Flashing lights indicate icy conditions" can be interpreted by a deployed second neural network that informs the vehicle's path planning software (preferably executing on a CPU complex) that icy conditions exist when the flashing lights are detected. The flashing lights can be identified by a deployed third neural network operating across multiple frames, informing the vehicle's path planning software of the presence (or absence) of the flashing lights. All three neural networks can operate simultaneously, for example, within a DLA and / or on a GPU 908.
[0151] In some examples, the CNN used for facial recognition and owner identification can use data from camera sensors to identify the presence of an authorized driver and / or owner of vehicle 900. A processing engine always on the sensors can be used to unlock the vehicle and turn on the lights when the owner approaches the driver's door, and in safe mode, to disable the vehicle when the owner leaves. In this way, SoC 904 provides security against theft and / or carjacking.
[0152] In another example, the CNN used for emergency vehicle detection and identification can use data from microphone 996 to detect and identify emergency vehicle siren. In contrast to conventional systems that use a general classifier to detect siren and manually extract features, SoC 904 uses a CNN to classify environmental and urban sounds as well as visual data. In a preferred embodiment, the CNN running on the DLA is trained to identify the relative shut-off rate of emergency vehicles (e.g., by using the Doppler effect). The CNN can also be trained to identify emergency vehicles specific to the localized area in which the vehicle operates, as identified by GNSS sensor 958. Thus, for example, when operating in Europe, the CNN will seek to detect European siren, and when operating in the United States, the CNN will seek to identify siren only in North America. Once an emergency vehicle is detected, with the assistance of ultrasonic sensor 962, the control program can be used to execute emergency vehicle safety routines, causing the vehicle to slow down, pull over to the side of the road, stop, and / or idle until the emergency vehicle passes.
[0153] The vehicle may include a CPU 918 (e.g., a discrete CPU or dCPU) that can be coupled to the SoC 904 via a high-speed interconnect (e.g., PCIe). The CPU 918 may include, for example, an x86 processor. The CPU 918 can be used to perform any of a wide variety of functions, including, for example, arbitrating the results of potential inconsistencies between ADAS sensors and the SoC 904, and / or monitoring the status and health of the controller 936 and / or the infotainment SoC 930.
[0154] Vehicle 900 may include a GPU 920 (e.g., a discrete GPU or dGPU) that can be coupled to SoC 904 via a high-speed interconnect (e.g., NVIDIA's NVLINK). GPU 920 may provide additional artificial intelligence capabilities, for example by executing redundant and / or different neural networks, and can be used to train and / or update neural networks based on inputs from sensors of vehicle 900 (e.g., sensor data).
[0155] Vehicle 900 may further include a network interface 924, which may include one or more wireless antennas 926 (e.g., one or more wireless antennas for different communication protocols, such as cellular antennas, Bluetooth antennas, etc.). Network interface 924 can be used to enable wireless connectivity via the Internet to the cloud (e.g., with server 978 and / or other network devices), with other vehicles, and / or with computing devices (e.g., a passenger's client device). For communication with other vehicles, a direct link can be established between the two vehicles, and / or an indirect link can be established (e.g., across a network and via the Internet). A direct link can be provided using a vehicle-to-vehicle communication link. The vehicle-to-vehicle communication link can provide vehicle 900 with information about vehicles approaching vehicle 900 (e.g., vehicles in front, to the side, and / or behind vehicle 900). This functionality can be part of vehicle 900's cooperative adaptive cruise control function.
[0156] Network interface 924 may include a SoC that provides modulation and demodulation functions and enables controller 936 to communicate over a wireless network. Network interface 924 may include an RF front-end for up-conversion from baseband to RF and down-conversion from RF to baseband. Frequency conversion can be performed using known processes and / or using a superheterodyne process. In some examples, the RF front-end functionality may be provided by a separate chip. The network interface may include wireless functions for communication via LTE, WCDMA, UMTS, GSM, CDMA2000, Bluetooth, Bluetooth LE, Wi-Fi, Z-Wave, ZigBee, LoRaWAN, and / or other wireless protocols.
[0157] Vehicle 900 may further include data storage 928, which may include off-chip (e.g., outside of SoC 904) storage devices. Data storage 928 may include one or more storage elements, including RAM, SRAM, DRAM, VRAM, flash memory, hard disk, and / or other components and / or devices capable of storing at least one bit of data.
[0158] Vehicle 900 may further include a GNSS sensor 958. The GNSS sensor 958 (e.g., GPS, assisted GPS sensor, differential GPS (DGPS) sensor, etc.) is used for auxiliary mapping, sensing, occupancy grid generation, and / or path planning functions. Any number of GNSS sensors 958 can be used, including, for example, but not limited to, GPS using a USB connector with an Ethernet-to-serial (RS-232) bridge.
[0159] Vehicle 900 may further include a RADAR sensor 960. The RADAR sensor 960 can be used by vehicle 900 for remote vehicle detection even in dark and / or inclement weather conditions. The RADAR functional safety level can be ASIL B. The RADAR sensor 960 can use CAN and / or bus 902 (e.g., to transmit data generated by the RADAR sensor 960) for control and access to object tracking data, and in some examples, Ethernet access for accessing raw data. A wide variety of RADAR sensor types can be used. For example, and without limitation, the RADAR sensor 960 can be adapted for front, rear, and side RADAR use. In some examples, a pulse Doppler RADAR sensor is used.
[0160] The RADAR sensor 960 can include different configurations, such as long-range with a narrow field of view, short-range with a wide field of view, short-range side coverage, etc. In some examples, the long-range RADAR can be used for adaptive cruise control functions. The long-range RADAR system can provide a wide field of view (e.g., within 250m) achieved through two or more independent scans. The RADAR sensor 960 can help distinguish between stationary and moving objects and can be used by ADAS systems for emergency braking assist and forward collision warning. The long-range RADAR sensor can include a single-site multi-mode RADAR with multiple (e.g., six or more) fixed RADAR antennas and high-speed CAN and FlexRay interfaces. In an example with six antennas, the four central antennas can create a focused beam pattern designed to record the vehicle 900's surroundings at a higher rate with minimal traffic interference from adjacent lanes. The other two antennas can extend the field of view, enabling rapid detection of vehicles entering or leaving the vehicle 900's lane.
[0161] As an example, a mid-range RADAR system can include a range of up to 960m (front) or 80m (rear) and a field of view of up to 42 degrees (front) or 950 degrees (rear). Short-range RADAR systems can include, but are not limited to, RADAR sensors designed to be mounted at both ends of the rear bumper. When mounted at both ends of the rear bumper, such a RADAR sensor system can create two beams that continuously monitor blind spots behind and beside the vehicle.
[0162] Short-range RADAR systems can be used in ADAS systems for blind spot detection and / or lane change assistance.
[0163] Vehicle 900 may further include ultrasonic sensors 962. Ultrasonic sensors 962, which may be positioned at the front, rear, and / or sides of vehicle 900, can be used for parking assistance and / or creating and updating occupancy grids. A wide variety of ultrasonic sensors 962 can be used, and different ultrasonic sensors 962 can be used for different detection ranges (e.g., 2.5m, 4m). Ultrasonic sensors 962 can operate at functional safety level ASIL B.
[0164] Vehicle 900 may include a LIDAR sensor 964. The LIDAR sensor 964 may be used for object and pedestrian detection, emergency braking, collision avoidance, and / or other functions. The LIDAR sensor 964 may be of functional safety level ASIL B. In some examples, vehicle 900 may include multiple LIDAR sensors 964 (e.g., two, four, six, etc.) that can use Ethernet (e.g., to provide data to a Gigabit Ethernet switch).
[0165] In some examples, the LiDAR sensor 964 may be able to provide a list of objects and their distances within a 360-degree field of view. Commercially available LiDAR sensors 964 may have an advertising range of, for example, approximately 900m, with an accuracy of 2cm-3cm, and support 900Mbps Ethernet connectivity. In some examples, one or more non-protruding LiDAR sensors 964 may be used. In such examples, the LiDAR sensor 964 may be implemented as a small device that can be embedded in the front, rear, sides, and / or corners of a vehicle 900. In such examples, the LiDAR sensor 964 may provide a horizontal field of view of up to 120 degrees and a vertical field of view of 35 degrees, even for low-reflectivity objects, with a range of 200m. Front-mounted LiDAR sensors 964 may be configured for a horizontal field of view between 45 and 135 degrees.
[0166] In some examples, LiDAR technologies such as 3D flash LiDAR can also be used. 3D flash LiDAR uses a flash of laser light as the emission source to illuminate the vehicle's surroundings up to approximately 200 meters. A flash LiDAR unit includes a receiver that records the laser pulse propagation time and reflected light on each pixel, which in turn corresponds to the range from the vehicle to the object. Flash LiDAR allows for the generation of highly accurate and distortion-free images of the surrounding environment using each laser flash. In some examples, four flash LiDAR sensors can be deployed, one on each side of the vehicle. Available 3D flash LiDAR systems include solid-state 3D staring array LiDAR cameras (e.g., non-scanning LiDAR devices) without moving parts other than a fan. Flash LiDAR devices can use 5 nanosecond Class I (eye-safe) laser pulses per frame and can capture reflected laser light in the form of a 3D range point cloud and co-registered intensity data. By using a flash LiDAR, and because a flash LiDAR is a solid-state device with no moving parts, the LiDAR sensor 964 is less susceptible to motion blur, vibration, and / or shock.
[0167] The vehicle may further include an IMU sensor 966. In some examples, the IMU sensor 966 may be located at the center of the rear axle of the vehicle 900. The IMU sensor 966 may include, for example, but not limited to, an accelerometer, a magnetometer, a gyroscope, a magnetic compass, and / or other sensor types. In some examples, such as in a six-axis application, the IMU sensor 966 may include an accelerometer and a gyroscope, while in a nine-axis application, the IMU sensor 966 may include an accelerometer, a gyroscope, and a magnetometer.
[0168] In some embodiments, the IMU sensor 966 can be implemented as a miniature, high-performance GPS-assisted inertial navigation system (GPS / INS) that combines a microelectromechanical system (MEMS) inertial sensor, a high-sensitivity GPS receiver, and an advanced Kalman filter algorithm to provide estimates of position, velocity, and attitude. Thus, in some examples, the IMU sensor 966 can enable the vehicle 900 to estimate heading by directly observing and correlating velocity changes from GPS to the IMU sensor 966 without input from a magnetic sensor. In some examples, the IMU sensor 966 and the GNSS sensor 958 can be combined into a single integrated unit.
[0169] The vehicle may include a microphone 996 placed in and / or around the vehicle 900. Among other things, the microphone 996 may be used for emergency vehicle detection and identification.
[0170] The vehicle may further include any number of camera types, including stereo camera 968, wide-angle camera 970, infrared camera 972, surround camera 974, long-range and / or mid-range camera 998, and / or other camera types. These cameras can be used to capture image data around the entire perimeter of the vehicle 900. The camera types used depend on the embodiment and the requirements of the vehicle 900, and any combination of camera types can be used to provide the necessary coverage around the vehicle 900. Furthermore, the number of cameras may vary depending on the embodiment. For example, the vehicle may include six cameras, seven cameras, ten cameras, twelve cameras, and / or another number of cameras. As an example and without limitation, these cameras may support Gigabit Multimedia Serial Link (GMSL) and / or Gigabit Ethernet. Each of the cameras is described herein with respect to... Figure 9A and Figure 9B It was described in more detail.
[0171] Vehicle 900 may further include vibration sensor 942. Vibration sensor 942 can measure vibrations of vehicle components such as axles. For example, changes in vibration can indicate changes in the road surface. In another example, when two or more vibration sensors 942 are used, differences between vibrations can be used to determine friction or slippage of the road surface (e.g., when there is a vibration difference between a power drive shaft and a freely rotating shaft).
[0172] Vehicle 900 may include ADAS system 938. In some examples, ADAS system 938 may include SoC. ADAS system 938 may include autonomous / adaptive / automatic cruise control (ACC), cooperative adaptive cruise control (CACC), forward collision warning (FCW), automatic emergency braking (AEB), lane departure warning (LDW), lane keeping assist (LKA), blind spot warning (BSW), rear cross traffic warning (RCTW), collision warning system (CWS), lane centering (LC) and / or other features and functions.
[0173] The ACC system can use a RADAR sensor 960, a LIDAR sensor 964, and / or a camera. The ACC system can include longitudinal ACC and / or lateral ACC. Longitudinal ACC monitors and controls the distance to vehicles immediately in front of vehicle 900 and automatically adjusts the vehicle speed to maintain a safe distance. Lateral ACC performs distance holding and, if necessary, advises vehicle 900 to change lanes. Lateral ACC is associated with other ADAS applications such as LCA and CWS.
[0174] CACC uses information from other vehicles, which can be received indirectly from other vehicles via a wireless link or network connection (e.g., via the Internet) through network interface 924 and / or wireless antenna 926. Direct links can be provided by vehicle-to-vehicle (V2V) communication links, while indirect links can be infrastructure-to-vehicle (I2V) communication links. Typically, the V2V communication concept provides information about vehicles immediately ahead (e.g., vehicles immediately in front of vehicle 900 and in the same lane), while the I2V communication concept provides information about traffic further ahead. A CACC system can include either or both of these I2V and V2V information sources. Given information about vehicles ahead of vehicle 900, CACC can be more reliable, and it has the potential to improve traffic flow and reduce road congestion.
[0175] The Forward-Looking Warning (FCW) system is designed to alert the driver to hazards, enabling the driver to take corrective action. The FCW system uses a front-facing camera and / or RADAR sensor 960 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback such as a display, speaker, and / or vibrating components. The FCW system can provide warnings in the form of, for example, audible, visual, haptic, and / or rapid braking pulses.
[0176] An AEB (Autonomous Emergency Braking) system detects an impending forward collision with another vehicle or other object and can automatically apply the brakes if the driver does not take corrective action within a specified time or distance parameter. The AEB system can use a front-facing camera and / or RADAR sensor 960 coupled to a dedicated processor, DSP, FPGA, and / or ASIC. When the AEB system detects a hazard, it typically first alerts the driver to take corrective action to avoid a collision, and if the driver does not take corrective action, the AEB system can automatically apply the brakes to attempt to prevent or at least mitigate the effects of the predicted collision. The AEB system may include technologies such as dynamic brake support and / or collision proximity braking.
[0177] The Lane Departure Warning (LDW) system provides visual, auditory, and / or tactile warnings, such as steering wheel or seat vibrations, to alert the driver when the vehicle crosses lane markings. When the driver indicates intentional lane departure, the LDW system is deactivated by activating a turn signal. The LDW system can utilize a front-facing camera coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback such as a display, speaker, and / or vibrating components.
[0178] The LKA system is a variation of the LDW system. If vehicle 900 begins to leave the lane, the LKA system provides steering input or braking to correct vehicle 900.
[0179] The BSW system detects and warns the driver of vehicles in the blind spot. The BSW system can provide visual, auditory, and / or tactile alerts to indicate that merging or changing lanes is unsafe. The system can provide additional warnings when the driver uses turn signals. The BSW system can utilize a rear-facing camera and / or RADAR sensor 960 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback such as a display, speaker, and / or vibrating components.
[0180] RCTW systems can provide visual, auditory, and / or tactile notifications when an object is detected outside the range of a rear-view camera while the vehicle is reversing. Some RCTW systems include AEB (Autonomous Emergency Braking) to ensure the application of the vehicle's brakes to avoid a collision. RCTW systems may use one or more rear-view RADAR sensors 960 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback such as a display, speaker, and / or vibrating components.
[0181] Conventional ADAS systems can be prone to false positives, which can be annoying and distracting for the driver, but typically not catastrophic, as they alert the driver and allow them to determine whether a safe condition truly exists and take appropriate action. However, in an autonomous vehicle 900, in the event of conflicting results, the vehicle 900 itself must decide whether to heed the results from the main computer or auxiliary computer (e.g., the first controller 936 or the second controller 936). For example, in some embodiments, the ADAS system 938 may be a backup and / or auxiliary computer used to provide perception information to a backup computer rationality module. The backup computer rationality monitor may run redundant and varied software on hardware components to detect faults in perception and dynamic driving tasks. Outputs from the ADAS system 938 may be provided to a supervisory MCU. If the outputs from the main computer and the auxiliary computer conflict, the supervisory MCU must determine how to reconcile the conflict to ensure safe operation.
[0182] In some examples, the master computer can be configured to provide a confidence score to the supervisory MCU, indicating the master computer's confidence level in the selected result. If the confidence score exceeds a threshold, the supervisory MCU can follow the master computer's direction regardless of whether the auxiliary computer provides conflicting or inconsistent results. If the confidence score does not meet the threshold and the master and auxiliary computers indicate different results (e.g., conflict), the supervisory MCU can arbitrate between these computers to determine the appropriate result.
[0183] The supervisory MCU can be configured to run a neural network trained and configured to determine the conditions under which the auxiliary computer provides a false alarm based on outputs from both the host and auxiliary computers. Thus, the neural network in the supervisory MCU can learn when the output of the auxiliary computer can be trusted and when it cannot. For example, when the auxiliary computer is a RADAR-based FCW system, the neural network in the supervisory MCU can learn when the FCW system is identifying a metallic object that is not actually dangerous, such as a drain grid or manhole cover that triggers an alarm. Similarly, when the auxiliary computer is a camera-based LDW system, the neural network in the supervisory MCU can learn to ignore the LDW when a cyclist or pedestrian is present and lane departure is actually the safest strategy. In embodiments that include a neural network running on the supervisory MCU, the supervisory MCU may include at least one of a DLA or GPU suitable for running the neural network using associated memory. In a preferred embodiment, the supervisory MCU may include components of and / or be included as components of the SoC 904.
[0184] In other examples, ADAS system 938 may include an auxiliary computer that performs ADAS functions using conventional computer vision rules. This allows the auxiliary computer to use classic computer vision rules (if-then), and the presence of neural networks in the supervising MCU can improve reliability, safety, and performance. For example, diverse implementations and intentional non-identity make the entire system more fault-tolerant, especially for failures caused by software (or software-hardware interface) functionality. For instance, if a software vulnerability or bug exists in the software running on the host computer and non-identical software code running on the auxiliary computer provides the same overall result, the supervising MCU can be more confident that the overall result is correct and that the vulnerability in the software or hardware on the host computer does not cause a substantial error.
[0185] In some examples, the output of the ADAS system 938 can be fed to the perception block and / or the dynamic driving task block of the main computer. For example, if the ADAS system 938 issues a forward collision warning because an object is immediately in front, the perception block can use this information when identifying the object. In other examples, the assistance computer can have its own neural network, which is trained and thus reduces the risk of false positives as described herein.
[0186] Vehicle 900 may further include an infotainment SoC 930 (e.g., an in-vehicle infotainment system (IVI)). Although illustrated and described as an SoC, the infotainment system may not be an SoC and may include two or more discrete components. The infotainment SoC 930 may include a combination of hardware and software that can be used to provide vehicle 900 with audio (e.g., music, personal digital assistant, navigation instructions, news, radio, etc.), video (e.g., TV, movies, streaming media, etc.), telephone (e.g., hands-free calling), network connectivity (e.g., LTE, Wi-Fi, etc.) and / or information services (e.g., navigation system, rear parking assistance, radio data system, vehicle-related information such as fuel level, total coverage distance, brake fuel level, fuel level, door opening / closing, air filter information, etc.). For example, the infotainment SoC 930 may include a radio, disc player, navigation system, video player, USB and Bluetooth connectivity, in-vehicle computer, in-vehicle entertainment, Wi-Fi, steering wheel audio controls, hands-free voice controls, head-up display (HUD), HMI display 934, telematics device, control panel (e.g., for controlling and / or interacting with various components, features, and / or systems) and / or other components. The infotainment SoC 930 may further be used to provide information (e.g., visual and / or auditory) to the vehicle's users, such as information from the ADAS system 938, autonomous driving information such as planned vehicle maneuvers, trajectories, surrounding environment information (e.g., intersection information, vehicle information, road information, etc.), and / or other information.
[0187] The infotainment SoC 930 may include GPU functionality. The infotainment SoC 930 can communicate with other devices, systems, and / or components of the vehicle 900 via bus 902 (e.g., CAN bus, Ethernet, etc.). In some examples, the infotainment SoC 930 may be coupled to a supervisory MCU, allowing the GPU of the infotainment system to perform some autonomous driving functions in the event of a failure of the main controller 936 (e.g., the primary and / or backup computer of the vehicle 900). In such an example, the infotainment SoC 930 may place the vehicle 900 into a driver-safe parking mode as described herein.
[0188] Vehicle 900 may further include instrument cluster 932 (e.g., digital instrument panel, electronic instrument cluster, digital instrument panel, etc.). Instrument cluster 932 may include a controller and / or supercomputer (e.g., a discrete controller or supercomputer). Instrument cluster 932 may include a set of instruments such as speedometer, fuel level, oil pressure, tachometer, odometer, turn indicator, shift position indicator, seatbelt warning light, parking brake warning light, engine malfunction indicator, airbag (SRS) system information, lighting controls, safety system controls, navigation information, etc. In some examples, information may be displayed and / or shared between infotainment SoC 930 and instrument cluster 932. In other words, instrument cluster 932 may be included as part of infotainment SoC 930, or vice versa.
[0189] Figure 9D For cloud-based servers and according to some embodiments of this disclosure Figure 9A The diagram illustrates a system for communication between example autonomous vehicles 900. System 976 may include server 978, network 990, and vehicles including vehicle 900. Server 978 may include multiple GPUs 984(A)-984(H) (collectively referred to herein as GPU 984), PCIe switches 982(A)-982(H) (collectively referred to herein as PCIe switch 982), and / or CPUs 980(A)-980(B) (collectively referred to herein as CPU 980). GPUs 984, CPUs 980, and PCIe switches may be interconnected with high-speed interconnects and / or PCIe connections 986, such as, but not limited to, NVLink interface 988 developed by NVIDIA. In some examples, GPUs 984 are connected via NVLink and / or NVSwitch SoCs, and GPUs 984 and PCIe switches 982 are connected via PCIe interconnects. Although eight GPUs 984, two CPUs 980, and two PCIe switches are shown in the diagram, this is not intended to be limiting. Depending on the embodiment, each of the servers 978 may include any number of GPUs 984, CPUs 980, and / or PCIe switches. For example, each of the servers 978 may include eight, sixteen, thirty-two, and / or more GPUs 984.
[0190] Server 978 can receive image data from vehicles via network 990, representing images of unexpected or changed road conditions such as recently commenced roadworks. Server 978 can also transmit neural network 992, updated neural network 992, and / or map information 994, including information about traffic and road conditions, to vehicles via network 990. Updates to map information 994 may include updates to HD map 922, such as information about construction sites, potholes, bends, floods, or other obstacles. In some examples, neural network 992, updated neural network 992, and / or map information 994 may have been generated from new training and / or data received from any number of vehicles in the environment, and / or based on experience gained from training performed at a data center (e.g., using server 978 and / or other servers).
[0191] Server 978 can be used to train machine learning models (e.g., neural networks) based on training data. Training data can be generated by the vehicle and / or generated in a simulation (e.g., using a game engine). In some examples, the training data is labeled (e.g., where the neural network benefits from supervised learning) and / or undergoes other preprocessing, while in other examples, the training data is not labeled and / or preprocessed (e.g., where the neural network does not require supervised learning). Training can be performed according to any one or more categories of machine learning techniques, including but not limited to: categories such as supervised training, semi-supervised training, unsupervised training, self-learning, reinforcement learning, joint learning, transfer learning, feature learning (including principal component and cluster analysis), multilinear subspace learning, manifold learning, representation learning (including alternative dictionary learning), rule-based machine learning, anomaly detection, and any variations or combinations thereof. Once the machine learning model is trained, it can be used by the vehicle (e.g., transmitted to the vehicle via network 990), and / or the machine learning model can be used by Server 978 to remotely monitor the vehicle.
[0192] In some examples, server 978 can receive data from vehicles and apply that data to state-of-the-art real-time neural networks for real-time intelligent inference. Server 978 may include a deep learning supercomputer powered by GPU 984 and / or a dedicated AI computer, such as the DGX and DGX Station machines developed by NVIDIA. However, in some examples, server 978 may include a deep learning infrastructure in a data center that uses only CPU power.
[0193] The deep learning infrastructure of server 978 is capable of rapid, real-time inference and can be used to assess and verify the health of the processor, software, and / or associated hardware in vehicle 900. For example, the deep learning infrastructure can receive periodic updates from vehicle 900, such as image sequences and / or objects located within those image sequences by vehicle 900 (e.g., via computer vision and / or other machine learning object classification techniques). The deep learning infrastructure can run its own neural network to identify objects and compare them to those identified by vehicle 900. If the results do not match and the infrastructure concludes that the AI in vehicle 900 has malfunctioned, server 978 can transmit a signal to vehicle 900 instructing its fail-safe computer to take control, notify passengers, and complete a safe stopping operation.
[0194] For inference, the server 978 can include a GPU 984 and one or more programmable inference accelerators (such as NVIDIA's TensorRT). The combination of a GPU-powered server and inference acceleration enables real-time response. In other examples, such as where performance is less critical, CPU, FPGA, and other processor-powered servers can be used for inference.
[0195] Example computing device
[0196] Figure 10 This is a block diagram suitable for implementing some embodiments of the present disclosure of an example computing device 1000. The computing device 1000 may include an interconnect system 1002 directly or indirectly coupled to the following devices: a memory 1004, one or more central processing units (CPUs) 1006, one or more graphics processing units (GPUs) 1008, a communication interface 1010, input / output (I / O) ports 1012, input / output components 1014, a power supply 1016, one or more presentation components 1018 (e.g., displays), and one or more logic units 1020. In at least one embodiment, one or more computing devices 1000 may include one or more virtual machines (VMs), and / or any component thereof may include virtual components (e.g., virtual hardware components). For a non-limiting example, one or more GPUs 1008 may include one or more vGPUs, one or more CPUs 1006 may include one or more vCPUs, and / or one or more logic units 1020 may include one or more virtual logic units. Thus, one or more computing devices 1000 may include discrete components (e.g., a full GPU dedicated to computing device 1000), virtual components (e.g., a portion of the GPU dedicated to computing device 1000), or combinations thereof.
[0197] although Figure 10 The various blocks are shown as connected via an interconnect system 1002 with wiring, but this is not intended to be limiting and is merely for clarity. For example, in some embodiments, a presentation component 1018, such as a display device, may be considered an I / O component 1014 (e.g., if the display is a touchscreen). As another example, the CPU 1006 and / or GPU 1008 may include memory (e.g., memory 1004 may represent a storage device other than the memory of the GPU 1008, CPU 1006, and / or other components). In other words, Figure 10 The computing devices mentioned are merely illustrative. No distinction is made between categories such as "workstation," "server," "laptop," "desktop," "tablet," "client device," "mobile device," "handheld device," "game console," "electronic control unit (ECU)," "virtual reality system," and / or other device or system types, as all of these are considered within the same category. Figure 10 Within the scope of computing devices.
[0198] Interconnect system 1002 may represent one or more links or buses, such as address buses, data buses, control buses, or combinations thereof. Interconnect system 1002 may include one or more bus or link types, such as Industry Standard Architecture (ISA) bus, Extended Industry Standard Architecture (EISA) bus, Video Electronics Standards Association (VESA) bus, Peripheral Component Interconnect (PCI) bus, Peripheral Component Interconnect Fast (PCIe) bus, and / or another type of bus or link. In some embodiments, there is a direct connection between components. For example, CPU 1006 may be directly connected to memory 1004. Furthermore, CPU 1006 may be directly connected to GPU 1008. In cases where there is a direct or point-to-point connection between components, interconnect system 1002 may include a PCIe link to perform the connection. In these examples, a PCI bus is not required in computing device 1000.
[0199] The memory 1004 may include any of a wide variety of computer-readable media. Computer-readable media can be any available medium that can be accessed by the computing device 1000. Computer-readable media may include volatile and non-volatile media, as well as removable and non-removable media. For example and without limitation, computer-readable media may include computer storage media and communication media.
[0200] Computer storage media may include volatile and non-volatile media and / or removable and non-removable media, implemented in any method or technique for storing information such as computer-readable instructions, data structures, program modules, and / or other data types. For example, memory 1004 may store computer-readable instructions (e.g., representing programs and / or program elements, such as an operating system). Computer storage media may include, but is not limited to, RAM, ROM, EEPROM, flash memory or other storage technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage devices, magnetic tape cassettes, magnetic tape, disk storage devices or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by computing device 1000. As used herein, computer storage media does not include the signal itself.
[0201] Computer storage media may include computer-readable instructions, data structures, program modules, and / or other data types in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium. The term "modulated data signal" may refer to a signal whose characteristics are set or altered in a manner that encodes information into that signal. For example and without limitation, computer storage media may include wired media such as wired networks or direct wired connections, and wireless media such as sound, RF, infrared, and other wireless media. Any combination of the above should also be included within the scope of computer-readable media.
[0202] CPU 1006 may be configured to execute at least some of computer-readable instructions to control one or more components of computing device 1000 to perform one or more of the methods and / or processes described herein. Each of CPU 1006 may include one or more cores (e.g., one, two, four, eight, twenty-eight, seventy-two, etc.) capable of processing a large number of software threads simultaneously. CPU 1006 may include any type of processor and may include different types of processors depending on the type of computing device 1000 implemented (e.g., processors with fewer cores for mobile devices and processors with more cores for servers). For example, depending on the type of computing device 1000, the processor may be an advanced RISC machine (ARM) processor implemented using Reduced Instruction Set Computing (RISC) or an x86 processor implemented using Complex Instruction Set Computing (CISC). In addition to one or more microprocessors or supplementary coprocessors such as math coprocessors, computing device 1000 may also include one or more CPUs 1006.
[0203] In addition to or as a replacement for CPU 1006, one or more GPUs 1008 may be configured to execute at least some of computer-readable instructions to control one or more components of computing device 1000 to perform one or more of the methods and / or processes described herein. One or more GPUs 1008 may be integrated GPUs (e.g., having one or more CPUs 1006) and / or one or more GPUs 1008 may be discrete GPUs. In embodiments, one or more GPUs 1008 may be coprocessors of one or more CPUs 1006. Computing device 1000 may use GPUs 1008 to render graphics (e.g., 3D graphics) or perform general-purpose computing. For example, one or more GPUs 1008 may be used for general-purpose computing on a GPU (GPGPU). One or more GPUs 1008 may include hundreds or thousands of cores capable of processing hundreds or thousands of software threads simultaneously. GPUs 1008 may generate pixel data for outputting an image in response to rendering commands (e.g., rendering commands received from CPU 1006 via a host interface). GPU 1008 may include graphics memory, such as display memory, for storing pixel data or any other suitable data, such as GPGPU data. Display memory may be included as part of memory 1004. One or more GPUs 1008 may include two or more GPUs operating in parallel (e.g., via a link). The link may be directly connected to the GPUs (e.g., using NVLINK) or connected via a switch (e.g., using NVSwitch). When combined, each GPU 1008 may generate pixel data or GPGPU data for different portions of the output or different outputs (e.g., the first GPU for the first image, the second GPU for the second image). Each GPU may include its own memory or may share memory with other GPUs.
[0204] In addition to or as an alternative to CPU 1006 and / or GPU 1008, logic unit 1020 may be configured to execute at least some of computer-readable instructions to control one or more components of computing device 1000 to perform one or more of the methods and / or processes described herein. In embodiments, CPU 1006, GPU 1008, and / or logic unit 1020 may execute any combination of methods, processes, and / or portions thereof, discretely or jointly. One or more logic units 1020 may be part of and / or integrated into one or more of CPU 1006 and / or GPU 1008, and / or one or more logic units 1020 may be discrete components or otherwise separate from CPU 1006 and / or GPU 1008. In embodiments, one or more logic units 1020 may be coprocessors of one or more CPUs 1006 and / or one or more GPUs 1008.
[0205] Examples of logic unit 1020 include one or more processing cores and / or components thereof, such as data processing unit (DPU), tensor core (TC), tensor processing unit (TPU), pixel vision core (PVC), vision processing unit (VPU), graphics processing cluster (GPC), texture processing cluster (TPC), streaming multiprocessor (SM), tree traversal unit (TTU), artificial intelligence accelerator (AIA), deep learning accelerator (DLA), arithmetic logic unit (ALU), application-specific integrated circuit (ASIC), floating-point unit (FPU), input / output (I / O) element, peripheral component interconnect (PCI) or peripheral component interconnect fast (PCIe) element, etc.
[0206] The communication interface 1010 may include one or more receivers, transmitters, and / or transceivers that enable the computing device 1000 to communicate with other computing devices via electronic communication networks, including wired and / or wireless communications. The communication interface 1010 may include components and functions that enable communication via any of several different networks, such as wireless networks (e.g., Wi-Fi, Z-Wave, Bluetooth, Bluetooth LE, ZigBee, etc.), wired networks (e.g., communication via Ethernet or InfiniBand), low-power wide area networks (e.g., LoRaWAN, SigFox, etc.), and / or the Internet. In one or more embodiments, one or more logic units 1020 and / or the communication interface 1010 may include one or more data processing units (DPUs) to directly transmit data received via a network and / or via interconnect system 1002 to (e.g., memory) one or more GPUs 1008.
[0207] I / O port 1012 enables computing device 1000 to be logically coupled to other devices, including I / O component 1014, presentation component 1018, and / or other components, some of which may be built into (e.g., integrated into) computing device 1000. Illustrative I / O component 1014 includes microphones, mice, keyboards, joysticks, game pads, game controllers, satellite dish antennas, scanners, printers, wireless devices, and so on. I / O component 1014 can provide a Natural User Interface (NUI) for processing user-generated air gestures, voice, or other physiological input. In some instances, the input may be transmitted to appropriate network elements for further processing. The NUI can implement any combination of voice recognition, stylus recognition, facial recognition, biometric recognition, on-screen and adjacent-screen gesture recognition, air gestures, head and eye tracking, and touch recognition associated with the display of computing device 1000 (described in more detail below). Computing device 1000 may include depth cameras such as stereo camera systems, infrared camera systems, RGB camera systems, touchscreen technology, and combinations thereof for gesture detection and recognition. In addition, the computing device 1000 may include an accelerometer or gyroscope that enables motion detection (e.g., as part of an inertial measurement unit (IMU)). In some examples, the output of the accelerometer or gyroscope may be used by the computing device 1000 to render immersive augmented reality or virtual reality.
[0208] The power supply 1016 may include a hard-wired power supply, a battery power supply, or a combination thereof. The power supply 1016 may supply power to the computing device 1000 so that the components of the computing device 1000 can operate.
[0209] The presentation component 1018 may include a display (such as a monitor, touch screen, television screen, head-up display (HUD), other display types, or combinations thereof), speakers, and / or other presentation components. The presentation component 1018 may receive data from other components (such as GPU 1008, CPU 1006, DPU, etc.) and output that data (such as as images, videos, sounds, etc.).
[0210] Example Data Center
[0211] Figure 11 An example data center 1100 that may be used in at least one embodiment of this disclosure is shown. The data center 1100 may include a data center infrastructure layer 1110, a framework layer 1120, a software layer 1130, and / or an application layer 1140.
[0212] like Figure 11As shown, the data center infrastructure layer 1110 may include a resource coordinator 1112, grouped computing resources 1114, and node computing resources (“nodes CR”) 1116(1)-1116(N), where “N” represents any complete positive integer. In at least one embodiment, nodes CR 1116(1)-1116(N) may include, but are not limited to, any number of central processing units (“CPUs”) or other processors (including accelerators, field-programmable gate arrays (FPGAs), graphics processing units or graphics processing units (GPUs), etc.), memory devices (e.g., dynamic read-only memory), storage devices (e.g., solid-state or disk drives), network input / output (“NW I / O”) devices, network switches, virtual machines (“VMs”), power modules and / or cooling modules, etc. In some embodiments, one or more node CRs from nodes CR 1116(1)-1116(N) may correspond to servers having one or more of the aforementioned computing resources. In addition, in some embodiments, nodes CR1116(1)-11161(N) may include one or more virtual components, such as vGPU, vCPU, etc., and / or one or more of nodes CR1116(1)-1116(N) may correspond to virtual machines (VMs).
[0213] In at least one embodiment, the grouped computing resources 1114 may include individual groups of nodes CR1116 housed within one or more racks (not shown), or multiple racks housed within a data center at different geographical locations (also not shown). Individual groups of nodes CR1116 within the grouped computing resources 1114 may include grouped computing, networking, memory, or storage resources that can be configured or allocated to support one or more workloads. In at least one embodiment, several nodes CR1116, including CPUs, GPUs, DPUs, and / or other processors, may be grouped within one or more racks to provide computing resources to support one or more workloads. One or more racks may also include any number of power modules, cooling modules, and / or network switches in any combination.
[0214] Resource coordinator 1122 may be configured or otherwise control one or more nodes CR1116(1)-1116(N) and / or grouped computing resources 1114. In at least one embodiment, resource coordinator 1122 may include a Software Design Infrastructure (“SDI”) management entity for data center 1100. Resource coordinator 1122 may include hardware, software, or some combination thereof.
[0215] In at least one embodiment, such as Figure 11As shown, framework layer 1120 may include job scheduler 1133, configuration manager 1134, resource manager 1136, and / or distributed file system 1138. Framework layer 1120 may include a framework of software 1132 supporting software layer 1130 and / or one or more applications 1142 supporting application layer 1140. Software 1132 or application 1142 may respectively contain web-based service software or applications, such as those provided by Amazon Web Services, Google Cloud, and Microsoft Azure. Framework layer 1120 may be, but is not limited to, a free and open-source software web application framework (such as Apache Spark™ (hereinafter “Spark”)) that can leverage distributed file system 1138 for large-scale data processing (e.g., “big data”). In at least one embodiment, job scheduler 1133 may include Spark drivers to facilitate the scheduling of workloads supported by different layers of data center 1100. Configuration manager 1134 may be able to configure different layers, such as software layer 1130 and framework layer 1120 (which includes Spark and distributed file system 1138 for supporting large-scale data processing). Resource manager 1136 may be able to manage computing resources mapped to or allocated to clusters of distributed file system 1138 and job scheduler 1133 to support distributed file system 1138 and job scheduler 1133. In at least one embodiment, clustered or grouped computing resources may include grouped computing resources 1114 in data center infrastructure layer 1110. Resource manager 1136 may coordinate with resource coordinator 1112 to manage these mapped or allocated computing resources.
[0216] In at least one embodiment, the software 1132 included in software layer 1130 may include software used in at least a portion of the nodes CRs 1116(1)-1116(N), the grouped computing resources 1114, and / or the distributed file system 1138 of framework layer 1120. One or more types of software may include, but are not limited to, internet web search software, email virus scanning software, database software, and streaming video content software.
[0217] In at least one embodiment, the application 1142 included in the application layer 1140 may include one or more types of applications used at least in part by nodes CR1116(1)-1116(N), grouped computing resources 1114, and / or the distributed file system 1138 of the framework layer 1120. One or more types of applications may include, but are not limited to, any number of genomics applications, cognitive computing, and machine learning applications, including training or inference software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), and / or other machine learning applications used in combination with one or more embodiments.
[0218] In at least one embodiment, any of the configuration manager 1134, resource manager 1136, and resource coordinator 1112 can implement any number and type of self-modification actions based on any amount and type of data acquired in any technically feasible manner. Self-modification actions can free data center operators of data center 1100 from making potentially poor configuration decisions and may prevent underutilization and / or poor performance of the data center.
[0219] According to one or more embodiments described herein, data center 1100 may include tools, services, software, or other resources to train one or more machine learning models or to use one or more machine learning models to predict or infer information. For example, one or more machine learning models may be trained by using the software and / or computing resources described above with respect to data center 1100 to compute weight parameters according to a neural network architecture. In at least one embodiment, a trained or deployed machine learning model corresponding to one or more neural networks may be used to infer or predict information using the resources described above with respect to data center 1100 by using weight parameters computed through one or more training techniques, such as, but not limited to, those described herein.
[0220] In at least one embodiment, the data center 1100 may use a CPU, application-specific integrated circuit (ASIC), GPU, FPGA, and / or other hardware (or corresponding virtual computing resources) to perform training and / or inference using the aforementioned resources. Furthermore, one or more of the software and / or hardware resources described above may be configured to allow a user to train or perform services that infer information, such as image recognition, speech recognition, or other artificial intelligence services.
[0221] Example network environment
[0222] A network environment suitable for implementing embodiments of this disclosure may include one or more client devices, servers, network-attached storage (NAS), other backend devices, and / or other device types. Client devices, servers, and / or other device types (e.g., each device) may... Figure 10 This can be implemented on one or more instances of computing device 1000—for example, each device may include similar components, features, and / or functions of computing device 1000. Furthermore, in the case of implementing backend devices (e.g., servers, NAS, etc.), the backend devices may be included as part of data center 1100, examples of which are described in this document. Figure 11 To describe in more detail.
[0223] Components of a network environment can communicate with each other via a network, which can be wired, wireless, or both. A network can include multiple networks, or one of multiple networks. For example, a network can include one or more wide area networks (WANs), one or more local area networks (LANs), one or more public networks such as the Internet and / or the Public Switched Telephone Network (PSTN), and / or one or more private networks. In cases where the network includes a wireless telecommunications network, components such as base stations, communication towers, or even access points (and other components) can provide wireless connectivity.
[0224] A compatible network environment may include one or more peer-to-peer network environments (in which case the server may not be included in the network environment) and one or more client-server network environments (in which case one or more servers may be included in the network environment). In a peer-to-peer network environment, the server functionality described herein can be implemented on any number of client devices.
[0225] In at least one embodiment, the network environment may include one or more cloud-based network environments, distributed computing environments, combinations thereof, etc. A cloud-based network environment may include a framework layer, a job scheduler, a resource manager, and a distributed file system implemented on one or more servers, which may include one or more core network servers and / or edge servers. The framework layer may include a framework for supporting software at the software layer and / or one or more applications at the application layer. The software or applications may respectively include network-based service software or applications. In embodiments, one or more client devices may use network-based service software or applications (e.g., by accessing the service software and / or applications via one or more application programming interfaces (APIs)). The framework layer may be, but is not limited to, a type of free and open-source software network application framework, such as one that can use a distributed file system for large-scale data processing (e.g., "big data").
[0226] A cloud-based network environment can provide cloud computing and / or cloud storage for any combination of the computing and / or data storage functions (or one or more portions thereof) described herein. Any of these various functions can be distributed across multiple locations, such as a central or core server (e.g., distributed across one or more data centers at the state, region, country, global, etc.). If the connection to a user (e.g., a client device) is relatively close to an edge server, the core server can assign at least a portion of the functionality to the edge server. A cloud-based network environment can be private (e.g., limited to a single organization), public (e.g., available to many organizations), and / or a combination thereof (e.g., a hybrid cloud environment).
[0227] Client devices may include those described in this article. Figure 10 The example computing device 1000 described includes at least some components, features, and functions. By way of example and not limitation, the client device may be a personal computer (PC), laptop computer, mobile device, smartphone, tablet computer, smartwatch, wearable computer, personal digital assistant (PDA), MP3 player, virtual reality headset, global positioning system (GPS) or device, video player, camera, surveillance equipment or system, vehicle, boat, aircraft, virtual machine, drone, robot, handheld communication device, hospital equipment, gaming device or system, entertainment system, in-vehicle computer system, embedded system controller, remote control, electrical appliance, consumer electronics device, workstation, edge device, any combination of these devices described, or any other suitable device.
[0228] This disclosure can be described in the general context of machine-usable instructions or computer code, including computer-executable instructions such as program modules, which are executed by a computer or other machine such as a personal digital assistant or other handheld device. Typically, a program module, including routines, programs, objects, components, data structures, etc., refers to code that performs a specific task or implements a specific abstract data type. This disclosure can be practiced in a wide variety of system configurations, including handheld devices, consumer electronics, general-purpose computers, more specialized computing devices, etc. This disclosure can also be practiced in distributed computing environments where tasks are performed by remote processing devices linked via a communication network.
[0229] As used herein, the phrase "and / or" relating to two or more elements should be interpreted as referring to only one element or a combination of elements. For example, "element A, element B, and / or element C" could include only element A, only element B, only element C, element A and element B, element A and element C, element B and element C, or elements A, B, and C. Furthermore, "at least one of element A or element B" could include at least one of element A, at least one of element B, or at least one of element A and at least one of element B. Further, "at least one of element A and element B" could include at least one of element A, at least one of element B, or at least one of element A and at least one of element B.
[0230] This document describes in detail the subject matter of this disclosure to satisfy legal requirements. However, the description itself is not intended to limit the scope of this disclosure. Rather, the inventors have envisioned that the claimed subject matter may also be embodied in other ways to include steps different from or similar combinations of steps described herein in conjunction with other current or future techniques. Moreover, although the terms "step" and / or "block" may be used herein to imply different elements of the method employed, these terms should not be construed as implying any particular order among or between the various steps disclosed herein, unless the order of the steps is explicitly described.
[0231] Example paragraph
[0232] A. A method comprising: calculating a probability distribution based at least on input data indicating occupant intent associated with an intent path of a machine, the probability distribution comprising at least: a first plurality of scores corresponding to a first plurality of road segments associated with a first intersection, the first plurality of scores representing a first probability that the first plurality of road segments correspond to the intent path of the machine; and a second plurality of scores corresponding to a second plurality of road segments associated with a second intersection, the second plurality of scores representing a second probability that the second plurality of road segments correspond to the intent path of the machine; determining a predicted path of the machine based at least on the first plurality of scores and the second plurality of scores, the predicted path comprising at least one of a first road segment of the first plurality of road segments or a second road segment of the second plurality of road segments; and performing one or more operations associated with the machine based at least on the predicted path.
[0233] B. The method according to paragraph A further includes: determining, based on at least one of map data or sensor data, that the machine is located within a threshold distance between the first intersection and the second intersection, wherein the determination of the predicted path of the machine, including at least one of the first road segment or the second road segment, is also based at least on the machine being located within the threshold distance between the first intersection and the second intersection.
[0234] C. The method according to any one of paragraphs A and B, further comprising: determining, at least based on the first plurality of scores, that the first road segment is the highest-scoring road segment among the first plurality of road segments; and determining, at least based on the second plurality of scores, that the second road segment is the highest-scoring road segment among the second plurality of road segments, wherein the predicted path of the machine including at least one of the first road segment or the second road segment is at least based on the first road segment being the highest-scoring road segment among the first plurality of road segments, and at least based on the second road segment being the highest-scoring road segment among the second plurality of road segments.
[0235] D. The method according to any one of paragraphs A, C, and D, further comprising: acquiring second input data over a period of time indicating the occupant's intention associated with the intended path of the machine; updating the first plurality of scores and the second plurality of scores over the period of time, and based at least on the second input data; and determining the predicted path of the machine based at least on the updates to the first plurality of scores and the second plurality of scores.
[0236] E. The method according to any one of paragraphs A and D, further comprising: calculating one or more first intent scores for the first plurality of road segments, the first intent scores indicating at least whether one or more detected occupant intents correspond to one or more first road segments of the first plurality of road segments; and calculating one or more second intent scores for the second plurality of road segments, the second intent scores indicating at least whether the one or more detected occupant intents correspond to one or more second road segments of the second plurality of road segments, wherein the first plurality of scores and the second plurality of scores are calculated at least based on the one or more first intent scores and the one or more second intent scores.
[0237] F. The method according to any one of paragraphs AE, wherein the input data includes at least one of the following: map data representing a map of an environment; positioning data indicating the position of the machine relative to the map of the environment; route data indicating a predefined route of the machine through the environment; preference or behavior data corresponding to one or more previous trips; or status data indicating one or more states associated with one or more components or systems of the machine.
[0238] G. A system comprising: one or more processors configured to: determine, based at least one of map data or perception data, that the location of a machine is within a threshold distance of one or more intersections, the intersections including multiple options for the machine's path; calculate, based at least on one or more occupant intentions corresponding to the multiple options, multiple scores indicating whether each of the multiple options corresponds to an intended path of the machine; determine a predicted path of the machine using the multiple scores; and perform one or more operations associated with the machine, based at least on the predicted path of the machine.
[0239] H. According to the system described in paragraph G, wherein at least a subset of the plurality of options corresponds to one or more road segments associated with the one or more intersections, and the one or more processors are further configured to select at least a first road segment from the one or more road segments for the machine's predicted path.
[0240] I. The system according to any one of paragraphs GH, wherein the one or more processors are further configured to determine that the time period associated with the arrival of the machine at the one or more intersections is less than a threshold time period, wherein the determination that the location of the machine is within the threshold distance of the one or more intersections is based at least on the time period being less than the threshold time period.
[0241] J. The system according to any one of paragraphs GI, wherein the one or more processors are further configured to calculate one or more confidence scores associated with the one or more occupant intentions, wherein the calculation of the plurality of scores is further based at least on the one or more confidence scores.
[0242] K. A system according to any one of paragraphs GJ, wherein the intention of one or more occupants includes at least one of the following: yaw rate associated with the machine; state of a steering signal associated with the machine; lane change associated with the machine; lane assignment associated with the machine; presence of an arrow marker in the lane used by the machine; or steering angle associated with the machine.
[0243] L is a system according to any one of paragraphs GK, wherein the predicted path of the machine corresponds at least partially to the intended path of the machine, and the determination of the predicted path includes: using the plurality of scores, determining a subset of the respective options having the highest scores among the plurality of options for the one or more intersections.
[0244] M. The system according to any one of paragraphs GL, wherein the calculation of the plurality of scores includes calculating the plurality of scores over a period of time based at least on a time series of the one or more occupant intentions corresponding to the plurality of options.
[0245] N. The system according to any one of paragraphs GM, wherein the one or more processors are further configured to: calculate a first plurality of intent scores for a first option among a plurality of options, based at least on a first plurality of occupant intents corresponding to the first option; calculate a second plurality of intent scores for one or more second options among the plurality of options, based at least on a second plurality of occupant intents corresponding to the one or more second options; calculate a first score among the plurality of scores using the first plurality of intent scores, the first score indicating whether the first option corresponds to the intent path of the machine; and calculate one or more second scores among the plurality of scores using the second plurality of intent scores, the second score indicating whether the one or more second options correspond to the intent path of the machine.
[0246] O. The system according to any one of paragraphs GN, wherein one or more of the first plurality of occupant intentions are included in the second plurality of occupant intentions.
[0247] P. A system according to any one of paragraphs GO, wherein the system includes at least one of: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing one or more simulated operations; a system for performing one or more digital twin operations; a system for performing optical transmission simulation; a system for performing collaborative content creation for 3D assets; a system for performing one or more deep learning operations; a system implemented using an edge device; a system implemented using a robot; a system for performing one or more generative AI operations; a system for performing operations using a large language model; a system for performing operations using one or more visual language models (VLMs); a system for performing operations using one or more multimodal language models; a system for performing one or more conversational AI operations; a system for generating synthetic data; a system for presenting at least one of virtual reality content, augmented reality content, or mixed reality content; a system containing one or more virtual machines (VMs); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources.
[0248] Q. One or more processors, comprising: processing circuitry for evaluating one or more path prediction algorithms in a simulation rendered using one or more optical transport simulation algorithms, the algorithms being used to predict the machine's intended path using a probability distribution associated with a plurality of options for an intended path of the machine in an environment, the probability distribution including a plurality of scores indicating whether each of the plurality of options corresponds to the machine's intended path.
[0249] R. According to one or more processors described in paragraph Q, wherein the probability distribution is generated based on at least the following operations: detecting one or more occupant intentions corresponding to each of the plurality of options; calculating a respective intention score for each of the options based at least on the one or more occupant intentions; and summing the respective intention scores to calculate a respective score among the plurality of scores for each of the options.
[0250] S. One or more processors according to any one of paragraphs QR, wherein the simulation is generated at least in part using a 3D content collaboration platform for 3D assets.
[0251] T. One or more processors according to any one of paragraphs QS, wherein the 3D content collaboration platform for 3D assets uses generic scene descriptor (USD) data to manage one or more attributes of the simulation environment associated with the simulation.
Claims
1. A method comprising: Based at least on input data indicating occupant intent associated with the machine's intent path, a probability distribution is calculated, which includes at least: Corresponding to a first plurality of road segments associated with a first intersection, the first plurality of scores represent a first probability that the first plurality of road segments correspond to the intended path of the machine; and A second plurality of scores corresponding to a second plurality of road segments associated with a second intersection, the second plurality of scores representing a second probability that the second plurality of road segments correspond to the machine’s intended path; Based at least on the first plurality of scores and the second plurality of scores, a predicted path of the machine is determined, the predicted path of the machine including at least one of a first segment of the first plurality of road segments or a second segment of the second plurality of road segments; and At least one or more operations associated with the machine are performed based on the predicted path.
2. The method according to claim 1, further comprising: Based on at least one of map data or sensor data, it is determined that the machine is within a threshold distance between the first intersection and the second intersection. The determination of the predicted path of the machine, which includes at least one of the first road segment or the second road segment, is also based at least on the fact that the machine is located within the threshold distance of the first intersection and the second intersection.
3. The method according to claim 1, further comprising: Based at least on the first plurality of scores, the first road segment is determined to be the highest-scoring road segment among the first plurality of road segments; as well as Based at least on the second plurality of scores, the second road segment is determined to be the highest-scoring road segment among the second plurality of road segments. The predicted path of the machine, which includes at least one of the first road segment or the second road segment, is based at least on the first road segment being the highest-scoring road segment among the first plurality of road segments, and at least on the second road segment being the highest-scoring road segment among the second plurality of road segments.
4. The method according to claim 1, further comprising: Acquire second input data over a period of time that indicates the occupant's intent as associated with the intent path of the machine; Update the first plurality of scores and the second plurality of scores over the time period and based at least on the second input data; as well as The machine's predicted path is determined at least based on the updates to the first plurality of scores and the second plurality of scores.
5. The method according to claim 1, further comprising: For the first plurality of road segments, one or more first intent scores are calculated, wherein the first intent scores indicate at least whether one or more detected occupant intents correspond to one or more first road segments from the first plurality of road segments; as well as For the second plurality of road segments, one or more second intent scores are calculated, wherein the second intent score indicates at least whether the one or more detected occupant intentions correspond to one or more of the second plurality of road segments. The first plurality of scores and the second plurality of scores are calculated based at least on the one or more first intent scores and the one or more second intent scores.
6. The method according to claim 1, wherein, The input data includes at least one of the following: Map data representing the environment; Location data indicating the position of the machine relative to the map of the environment; Route data instructing the machine to traverse a predefined route through the environment; Preference or behavioral data corresponding to one or more previous trips; or Status data indicating one or more states associated with one or more components or systems of the machine.
7. A system comprising: One or more processors are used for: Based on at least one of map data or perception data, determine that the machine's location is within a threshold distance of one or more intersections, said intersections including multiple options for the machine's path; Based at least on one or more occupant intentions corresponding to the plurality of options, calculate multiple scores indicating whether each of the plurality of options corresponds to the intention path of the machine; The predicted path of the machine is determined using the multiple scores; as well as At least based on the predicted path of the machine, perform one or more operations associated with the machine.
8. The system according to claim 7, wherein, At least a subset of the plurality of options corresponds to one or more road segments associated with the one or more intersections, and the one or more processors are further configured to select at least a first road segment from the one or more road segments for the machine's predicted path.
9. The system of claim 7, wherein the one or more processors are further configured to determine that the time period associated with the arrival of the machine at the one or more intersections is less than a threshold time period, wherein, The determination that the machine's location is within the threshold distance of the one or more intersections is based at least on the time period being less than the threshold time period.
10. The system of claim 7, wherein the one or more processors are further configured to calculate one or more confidence scores associated with the one or more occupant intentions, wherein the calculation of the plurality of scores is further based at least on the one or more confidence scores.
11. The system according to claim 7, wherein, The intent of one or more occupants includes at least one of the following: The yaw rate associated with the machine; The status of the steering signal associated with the machine; Lane changes associated with the machine; Lane assignment associated with the machine; The presence of arrow markings in the lane used by the machine; or The steering angle associated with the machine.
12. The system according to claim 7, wherein, The predicted path of the machine corresponds at least partially to the machine's intended path, and the determination of the predicted path includes: using the plurality of scores, determining a subset of the respective options with the highest scores among the plurality of options for the one or more intersections.
13. The system according to claim 7, wherein, The calculation of the plurality of scores includes calculating the plurality of scores over a period of time based at least on a time series of one or more occupant intentions corresponding to the plurality of options.
14. The system of claim 7, wherein the one or more processors are further configured to: For the first option among the plurality of options, a first plurality of intent scores are calculated based at least on a first plurality of occupant intents corresponding to the first option; For one or more of the multiple options, a second plurality of intent scores are calculated based at least on a second plurality of occupant intents corresponding to the one or more of the second options; Using the first plurality of intent scores, a first score is calculated among the plurality of scores, the first score indicating whether the first option corresponds to the intent path of the machine; as well as Using the second plurality of intent scores, one or more second scores are calculated, the one or more second scores indicating whether the one or more second options correspond to the intent path of the machine.
15. The system according to claim 14, wherein, One or more of the first plurality of occupant intentions are included in the second plurality of occupant intentions.
16. The system according to claim 7, wherein, The system is included in at least one of the following: Control systems for autonomous or semi-autonomous machines; Sensing systems for autonomous or semi-autonomous machines; A system for performing one or more simulation operations; A system for performing one or more digital twin operations; A system for performing optical transmission simulation; A system for performing collaborative content creation for 3D assets; A system for performing one or more deep learning operations; Systems implemented using edge devices; Systems implemented using robots; A system for performing one or more generative AI operations; A system for performing operations using large language models; A system for performing operations using one or more visual language models (VLMs); A system for performing operations using one or more multimodal language models; A system for performing one or more conversational AI operations; A system for generating synthetic data; A system for presenting at least one of virtual reality content, augmented reality content, or mixed reality content; A system containing one or more virtual machines (VMs); A system that is at least partially implemented in a data center; or A system that utilizes cloud computing resources at least in part.
17. One or more processors, comprising: Processing circuitry is configured to evaluate one or more path prediction algorithms in a simulation rendered using one or more optical transmission simulation algorithms to predict the machine's intended path using a probability distribution associated with multiple options of the machine's intended path in the environment, the probability distribution including multiple scores indicating whether each of the multiple options corresponds to the machine's intended path.
18. One or more processors according to claim 17, wherein, The probability distribution is generated based on at least the following operations: Detect one or more occupant intentions corresponding to each of the multiple options; Based on at least one or more occupant intentions, a respective intention score is calculated for each of the respective options; as well as The respective intent scores are aggregated to calculate a score from the plurality of scores for each of the respective options.
19. One or more processors according to claim 18, wherein, The simulation was generated at least in part using a 3D content collaboration platform for 3D assets.
20. One or more processors according to claim 19, wherein, The 3D content collaboration platform for 3D assets uses Universal Scene Descriptor (USD) data to manage one or more attributes of the simulation environment associated with the simulation.
Citation Information
Patent Citations
Method for programmable timeouts of tree traversal mechanisms in hardware
US10885698B2