Positioning using specific path trackers for autonomous or semi-autonomous systems and applications
By initializing and evaluating multiple path trackers in autonomous or semi-autonomous machines, the positioning difficulties caused by the lack of HD maps were resolved, enabling high-precision autonomous navigation in SD maps, reducing processing burden and latency, and improving the real-time navigation performance of the system.
Patent Information
- Application Number
- CN202511392164.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-09-27
- Filing Date
- 2025-09-26
- Publication Date
- 2026-03-27
AI Technical Summary
In existing technologies, the lack or inaccuracy of HD maps during autonomous or semi-autonomous machine localization leads to increased processing bandwidth and latency, making it difficult to accurately locate in SD maps, especially at road intersections, thus affecting the system's real-time navigation performance.
By initializing multiple path-specific trackers, evaluating candidate locations based on sensor and map data, calculating scores to determine the machine's accurate location after road intersections, using Kalman filters to track machine movement, and dynamically adjusting trackers to improve positioning accuracy and reduce processing load.
Without relying on HD maps, high-precision positioning of autonomous or semi-autonomous machines in SD maps was achieved, reducing processing bandwidth and latency, and improving the navigation safety and real-time performance of the system.
Smart Images

Figure CN121740073A_ABST
Abstract
Description
Background Technology
[0001] To enable autonomous or semi-autonomous machines (e.g., vehicles) to navigate safely in their environment, the machines may rely on maps corresponding to the environment in which they intend to operate, such as navigation, standard definition (SD), and / or high definition (HD) maps. For example, many traditional systems can locate autonomous or semi-autonomous machines by matching features in an HD map with corresponding features perceived in the real-world environment. However, HD maps are not always available, or where they are, it may be necessary to process sensor data from various modalities to align the perceived data with the map data. This can burden processing bandwidth and / or increase system latency, making real-time or near-real-time deployments unsuitable. Thus, in some scenarios, it may be desirable (or in some cases, necessary) to locate autonomous or semi-autonomous machines to features in navigation maps and / or SD maps. However, because SD maps may contain a lower level of detail than HD maps (e.g., they may lack information such as lane numbers, categories, and geometry, and only have GNSS-level accuracy or precision (e.g., within 3 meters on a road, not lane-level positioning), rather than the centimeter-level accuracy or precision of HD maps), it can be challenging in some instances to accurately locate autonomous or semi-autonomous machines to features on SD maps. Summary of the Invention
[0002] Various embodiments of this disclosure relate to localization and applications using path-specific trackers for autonomous and semi-autonomous systems. For example, the systems and methods described herein can initialize and use multiple path-specific trackers to localize a machine (e.g., an autonomous or semi-autonomous machine or vehicle) relative to a specific path in an environment. As an example, when a machine passes through the junction of multiple road segments, a corresponding tracker can be initialized for each segment, and the tracker can be placed at a corresponding candidate location along each of the multiple road segments. The candidate location can represent the possible position of the machine along the road segment based on the machine's tracking motion. Using various inputs from sensors, perception systems, and / or other systems of the machine, it can be determined which tracker (or candidate location) most closely corresponds to the machine's actual position after the junction. In some examples, a tracker can be selected to track the machine's position along the current road segment, while other trackers can be terminated or removed.
[0003] Compared to conventional systems, the systems of this disclosure, in some embodiments, are capable of accurately locating autonomous or semi-autonomous machines with respect to features included in navigation or SD maps. For example, by initializing trackers along potential road segments the machine might traverse after passing an intersection, the systems of this disclosure are able to more accurately determine which road segment the machine is traversing after an intersection. Additionally, by calculating a score for each tracker and / or candidate location among the trackers and / or candidate locations, the systems of this disclosure are able to determine and indicate when a location result is unavailable or has low confidence. Thus, and as described in more detail herein, by performing these processes, the systems of this disclosure are able to accurately determine the machine's position in an environment without using HD map data, thereby facilitating safer traversal of autonomous or semi-autonomous machines while minimizing or at least reducing processing bandwidth and / or latency. Attached Figure Description
[0004] The following describes in detail, with reference to the accompanying drawings, existing systems and methods for localization using specific path trackers for autonomous and semi-autonomous systems and applications, wherein:
[0005] Figure 1 An example data flow diagram is shown illustrating the process of locating a machine using a specific path tracker according to some embodiments of the present disclosure;
[0006] Figure 2 An example architecture of a tracker that can be used to track the location of a machine according to some embodiments of this disclosure is shown;
[0007] Figure 3 Example visualizations of candidate locations determined by a machine are shown according to some embodiments of this disclosure;
[0008] Figures 4A-4C Some embodiments of the present disclosure are shown in relation to the present disclosure. Figure 3 The example of scoring the candidate positions of the machines is associated with the example;
[0009] Figures 5A-5C Examples of cached maps associated with determining whether to update the environment are shown according to some embodiments of this disclosure;
[0010] Figure 6 Examples of systems that can perform one or more processes described herein are shown according to some embodiments of the present disclosure;
[0011] Figure 7 A flowchart illustrating an example of a method for locating a machine after an intersection using a specific path tracker, according to some embodiments of the present disclosure;
[0012] Figure 8A flowchart illustrating an example of a method for determining the location of a machine based on scoring multiple candidate locations, according to some embodiments of the present disclosure;
[0013] Figure 9 A flowchart illustrating an example of a method for determining a road segment that a machine is using after passing an intersection, according to some embodiments of the present disclosure;
[0014] Figure 10A These are illustrations of example autonomous vehicles according to some embodiments of the present disclosure;
[0015] Figure 10B According to some embodiments of this disclosure Figure 10A Examples of camera positions and fields of view for autonomous vehicles;
[0016] Figure 10C According to some embodiments of this disclosure Figure 10A A block diagram of an example system architecture for an example autonomous vehicle;
[0017] Figure 10D This is based on some embodiments of the present disclosure for use in cloud-based servers and Figure 10A A system diagram illustrating communication between autonomous vehicles;
[0018] Figure 11 This is a block diagram of an example computing device suitable for implementing some embodiments of the present disclosure; and
[0019] Figure 12 This is a block diagram of an example data center suitable for implementing some embodiments of the present disclosure. Detailed Implementation
[0020] A system and method are disclosed relating to localization using a specific path tracker for autonomous and semi-autonomous systems and applications. Although this disclosure may relate to an example autonomous or semi-autonomous vehicle or machine 1000 (which may alternatively be referred to herein as "vehicle 1000," "self-vehicle 1000," or "self-machine 1000," examples of which are referred to herein), Figure 10A-10DThe description herein is intended to be limiting. For example, the systems and methods described herein can be used by, but are not limited to, non-autonomous vehicles or machines, semi-autonomous vehicles or machines (e.g., in one or more adaptive driver assistance systems (ADAS)), manned and unmanned robots or robotic platforms, warehouse vehicles, off-road vehicles, vehicles coupled to one or more trailers, spacecraft, ships, space shuttles, emergency response vehicles, motorcycles, electric or motorized bicycles, aircraft, engineering vehicles, underwater vehicles, drones, and / or other vehicle types. Furthermore, while this disclosure may describe road segment localization, this is not intended to be limiting, and the systems and methods described herein can be used in augmented reality, virtual reality, mixed reality, robotics, security and supervision, autonomous or semi-autonomous machine applications, and / or any other technological space where path localization can be used.
[0021] For example, one or more systems can determine that a machine (e.g., an autonomous or semi-autonomous machine or vehicle) is about to pass (e.g., within a threshold distance and / or time from arrival at the intersection), is passing (within a threshold distance or time at the intersection), and / or has already passed an intersection associated with a road network in the environment (e.g., within a threshold distance or time after passing the intersection). In some examples, an intersection may correspond to an area in the environment where two or more road segments meet or intersect. The complexity of the intersection can vary across different instances, from simple T-junctions or Y-junctions to more complex designs (such as roundabouts or multi-lane grade-separated intersections). Thus, an intersection can include multiple options for the machine's path. For example, the intersection may correspond to the intersection of a first road segment, a second road segment, a third road segment, etc. Therefore, at the intersection, the machine can choose one of the first road segment, the second road segment, the third road segment, etc., as part of its path to traverse.
[0022] In some instances, one or more systems may determine the presence of an intersection and / or the machine's position relative to the intersection (e.g., whether the machine is about to pass, is passing, or has already passed the intersection) based at least on the position of the tracking machine with respect to a map of the environment (e.g., a navigation map or an SD map). For example, one or more systems may initialize a first tracker (also referred to herein as the "initial tracker") at the approximate location of the machine with respect to the map based at least on determining the machine's initial position / location with respect to the map. In some examples, the first tracker may include a Kalman filter for tracking the machine's position using state variables updated per frame, every other frame, etc., using a Kalman filter. In some instances, the state variables may include at least a first state variable corresponding to a segment identifier, a second state variable corresponding to the machine's offset with respect to the segment (e.g., the machine's position relative to the start or end point of the segment), and a third state variable corresponding to a confidence level associated with the offset. In some examples, if, after applying a new measurement to the Kalman filter of the first tracker, the offset state has moved past the end point of the current segment and onto a subsequent segment, it can be determined whether the machine has passed the intersection.
[0023] As described herein, in some examples, based on the determination that the machine has passed through the intersection, one or more systems may initialize multiple trackers to track multiple possible locations (also referred to herein as “candidate locations” and / or “candidate poses”) of the machine along one or more subsequent road segments away from the intersection. For example, if the intersection is a four-way intersection associated with four road segments and a first tracker moves from a first road segment to a second road segment via the intersection, one or more systems may initialize two additional trackers to track two possible locations of the machine along two additional subsequent road segments away from the intersection (e.g., road segments away from the end of the first road segment and / or the start of the second road segment).
[0024] In some examples, a first tracker can move to a subsequent road segment at the intersection corresponding to the machine's predicted path (e.g., the most likely path), while additional trackers can be initialized to track possible locations along other subsequent road segments that are not part of the predicted path. In other words, suppose, for example, at a four-way intersection, the machine's predicted path is to travel straight through the intersection from the first road segment to the second road segment. In this example, by default, the first tracker can move from the first road segment to one or more first possible locations of the machine along the second road segment. However, additional trackers can be initialized to track possible locations of the machine along a third road segment (e.g., turning right at the intersection) and a fourth road segment (e.g., turning left at the intersection). For example, a second tracker can be initialized to track one or more second possible locations of the machine along the third road segment, and a third tracker can be initialized to track one or more third possible locations of the machine along the fourth road segment.
[0025] In some instances, one or more systems can calculate a score for the possible location of the machine after each tracker and / or intersection. The score can indicate which tracker or possible location corresponds to the machine's actual location. For example, in the example above where the intersection is a four-way intersection, one or more systems can calculate one or more first scores for the first tracker, one or more second scores for the second tracker, and one or more third scores for the third tracker. In some instances, one or more systems can calculate new scores at each round and aggregate the scores over a period of time to determine the machine's location. In some examples, the score values can be calculated using the same or similar measurements used to update the state of the Kalman filter of the first tracker, such as machine motion (relative and global motion), perceived curvature, yaw rate, whether a candidate location corresponds to the machine's most likely path (e.g., located on the machine's most likely path), the number of lanes, the color / style of lane markings, etc.
[0026] As described herein, one or more systems can evaluate scores to determine which of the possible locations / trackers corresponds to the actual location of the machine. In some examples, the highest score (which is higher than one or more other scores by a threshold) can be identified as corresponding to the actual location of the machine and can be provided as a positioning result. For example, and continuing with the four-way intersection example above, one or more systems can determine that one or more first scores associated with one or more first possible locations are greater than one or more second scores associated with one or more second possible locations and one or more third scores associated with one or more third possible locations. In this example, one or more systems can determine that the machine's location corresponds to one or more first possible locations based at least on the fact that one or more first scores are greater than one or more second scores and one or more third scores. Additionally, in some examples, one or more systems can determine whether one or more differences between one or more first scores and one or more second scores and one or more third scores reach or exceed a threshold, and determining that the machine's location corresponds to one or more first possible locations can also be based on one or more differences reaching or exceeding that threshold.
[0027] In some examples, once it is determined that one score associated with one of the possible locations is significantly larger than the other scores by more than a threshold, one or more systems can set the tracker as the prevailing location and remove / terminate other trackers for other locations. In some instances, if the prevailing location / tracker is not the initially predicted path or the most likely path, one or more systems can reinitialize the tracker at the segment corresponding to the prevailing location / tracker (e.g., reinitialize the Kalman filter). In contrast, if the prevailing location / tracker is the initially predicted path or the most likely path, in some examples, one or more systems may not need to make any changes because the initial tracker can simply continue its normal operation. In some examples, if the machine travels beyond a threshold distance and / or beyond a threshold time period across multiple possible locations and trackers, and none of the scores reach or exceed the threshold, one or more systems may mark the localization result as unavailable.
[0028] In some examples, one or more systems may perform one or more operations associated with the machine based at least on determining the machine's position relative to the road segment following the intersection. That is, one or more systems may perform one or more operations based on determining which road segment the machine is traversing after the intersection. For example, one or more operations may be based on one or more features, attributes, etc., associated with the road segment. As an example, if the road segment the machine is traversing is an exit ramp, one or more systems might slow the machine down. As another example, if the road segment associated with the machine's predicted path has a low curvature level, but the machine turns onto a different road segment with higher curvature at the intersection, one or more systems may update one or more operational parameters associated with the machine based on the curvature of the different road segment, such as adjusting the maximum speed the machine can operate under. Additionally or alternatively, one or more systems may plan the path the machine should follow based on positioning the machine on a specific road segment after the intersection.
[0029] In at least one embodiment, one or more systems may use the determined location of the machine to update a cached map of the environment. For example, as the machine traverses the environment, it may continuously receive portions of the map (e.g., navigation, SD map, etc.) for use in traversing the environment. Thus, the machine may receive and temporarily store relevant portions of the map, such as map portions corresponding to the machine's current and / or future location. In some examples, once a map portion is deemed valid (e.g., for the machine's path), the machine may cache the map portion until a new map portion is received, and the newly received portion is also considered valid. Thus, one or more systems may use the determined location of the machine regarding the road segment it is using after the intersection to determine whether a newly received map portion is valid and / or whether to update the cached map.
[0030] In these instances, if the newly received map portion does not include road segments currently being used by the machine, one or more systems may determine that the newly received map portion is invalid and avoid updating the cached map. That is, because the new map portion is invalid, one or more systems may allow the machine to continue relying on the cached map for navigation. As another example, if the new map portion does not include road segments corresponding to the machine's predicted path or most likely path, one or more systems may avoid updating the cached map. However, if the new map portion includes all relevant road segments, one or more systems may determine that the new map portion is valid and allow the cached map to be updated to the newly received map.
[0031] In some embodiments, the systems and methods described herein can be performed in a simulated environment (e.g., NVIDIA's DriveSIM) using simulated data (e.g., simulated sensor data of a virtual machine or simulated sensors of a simulated machine). For example, simulated input data (e.g., map data, perception data, or any other data described herein) can be used to initialize a tracker to track the possible locations of a machine along different road segments away from the intersection, and the tracker and / or possible locations can be scored to determine the actual location of the virtual machine in the simulated environment with respect to the different road segments. These simulation operations can be used to test the performance of the underlying algorithms, systems, and / or processes before deploying them to the real world. In some instances, simulations can be used to generate synthetic training data, e.g., perception and / or map training data indicating road segments and intersections in the simulation. The synthetic training data (in addition to real-world data, or alternatively, data from the real world) can then be processed to locate the machine with respect to a specific path after the intersection, such as, for example, locating a machine (e.g., a robot) in a warehouse to a specific path in the warehouse after the machine has passed the intersection. In any example (such as one where a simulated environment is used for testing, validation, training, etc.), one or more optical transport algorithms (such as ray tracing and / or path tracing algorithms) may be used to render or generate the simulated environment and / or associated training data.
[0032] In some embodiments, simulated environments and / or one or more of their objects, features, or components can be generated or managed in a 3D content collaboration platform (e.g., NVIDIA's OMNIVERSE) for industrial digitization, generative physics AI, and / or other use cases, applications, or services. For example, the content collaboration platform or system may include a system for using or developing generic scene descriptors (USD) (e.g., OpenUSD) data to manage objects, features, scenes, etc., in simulated environments, digital environments, etc. The platform may include realistic physics simulations, such as using NVIDIA's PhysX SDK, to simulate real-world physical phenomena and physical interactions with simulations hosted on the platform. The platform may integrate OpenUSD with ray tracing / path tracing / light transport simulations (e.g., NVIDIA's RTX rendering technology) into software tools and simulation workflows for building, training, deploying, or testing AI systems, such as systems for testing, validating, training (e.g., machine learning models, neural networks, etc.) and / or other tasks related to automobiles, robots, machines, or other applications. In some examples, the simulated environment can include a digital twin of a real environment, such as a stretch of roadway, warehouse, data center, airport, geographic region, ocean region, or any other real environment in which an autonomous or semi-autonomous machine might operate.
[0033] In some embodiments, a remote control or teleoperation system can be used to remotely operate or control a vehicle or other machine. For example, the systems and methods described herein can be used to identify lane lines, road boundary lines, longitudinal features, potential paths, etc., which can be included in a visualization or mapping of the environment to assist a remote operator in controlling autonomous or semi-autonomous machines through environmental control (or providing road signs or other control or navigation instructions).
[0034] In some examples, one or more machine learning models described herein (e.g., deep neural networks, language models, LLMs, VLMs, multimodal language models, perceptual models, tracking models, fusion models, transformer models, diffusion models, encoder-only models, decoder-only models, encoder-decoder models, neural rendering field (NERF) models, etc.) can be packaged as microservices, such as inference microservices (e.g., NVIDIA NIM), which may include containers (e.g., operating system (OS) level virtualization packages) that may include an application programming interface (API) layer, a server layer, a runtime layer, and / or a model "engine". For example, an inference microservice might include the container itself and one or more models (e.g., weights and biases). In some instances (e.g., where one or more machine learning models are small enough (e.g., have a sufficiently few parameters)), one or more models may be included within the container itself. In other instances (e.g., where one or more models are large), one or more models may be hosted / stored in the cloud (e.g., in a data center) and / or may be hosted on the ground and / or at the edge (e.g., on a local server or computing device, but outside the container). In these embodiments, one or more models can be accessed via one or more APIs, such as REST APIs. Thus, and in some embodiments, one or more machine learning models described herein can be deployed as inference microservices to accelerate the deployment of one or more models on any cloud, data center, or edge computing system while ensuring data security. For example, an inference microservice may include one or more APIs, pre-configured containers for simplified deployment, optimized inference engines (e.g., using standardized AI model deployment to build execution software such as NVIDIA's Triton Inference Server) and / or one or more APIs for high-performance deep learning inference, which may include inference runtime and model optimizations for low latency and high throughput for production applications (such as NVIDIA's TensorRT) and / or enterprise management data for telemetry (e.g., including identity, metrics, health checks, and / or monitoring). One or more machine learning models described herein can be included as part of a microservice along with acceleration infrastructure capable of deployment using a single command and / or orchestrated and automatically scaled on the acceleration infrastructure (e.g., on a single device up to data center scale). Thus, inference microservices may include machine learning models (e.g., optimized for high-performance inference), inference runtime software that implements one or more machine learning models and provides outputs / responses to inputs (e.g., user queries, prompts, etc.), and enterprise management software that provides health checks, identity and / or other monitoring.In some embodiments, the inference microserver may include software that performs in-situ replacement and / or updates to one or more machine learning models. When a replacement or update is performed, the software performing the replacement / update can maintain the user configurations of the inference runtime software and the enterprise management software.
[0035] The systems and methods described herein can be, but are not limited to, non-autonomous vehicles or machines, semi-autonomous vehicles or machines (e.g., in one or more adaptive driver assistance systems (ADAS)), autonomous vehicles or machines, manned and unmanned robots or robotic platforms, warehouse vehicles, off-road vehicles, vehicles coupled to one or more trailers, airships, boats, shuttles, emergency response vehicles, motorcycles, electric or motorized bicycles, aircraft, engineering vehicles, underwater vehicles, drones, and / or other vehicle types. Furthermore, the systems and methods described herein can be used for a variety of purposes, by way of example but not limited to, systems for: machine control, machine motion, machine driving, synthetic data generation, model training, perception, augmented reality, virtual reality, mixed reality, robotics, safety and supervision, simulation and digital twins, autonomous or semi-autonomous machine applications, deep learning, environmental simulation, object or actor simulation and / or digital twins, data center processing, conversational AI, optical transport simulation (e.g., ray tracing, path tracing, etc.), collaborative content creation for 3D assets, cloud computing, and / or any other suitable application.
[0036] The disclosed embodiments can comprise a variety of different systems, such as automotive systems (e.g., control systems for autonomous or semi-autonomous machines, perception systems for autonomous or semi-autonomous machines), systems implemented using robots, aviation systems, medical systems, boating systems, smart zone monitoring systems, systems for performing deep learning operations, systems for performing simulation operations, systems for performing digital twin operations, systems implemented using edge devices, systems implementing language models (such as Large Language Models (LLM), Visual Language Models (VLM), and / or multimodal language models), systems implementing one or more multimodal language models, systems using or deploying one or more inference microservices, systems containing one or more machine learning models and operating system-level virtualization packages (e.g., containers) deployed in services or microservices, systems containing one or more virtual machines (VMs), systems for performing synthetic data generation operations, systems implemented at least partially in a data center, systems for performing conversational AI operations, systems for performing optical transmission simulations, systems for performing collaborative content creation for 3D assets, systems for performing generative AI operations, systems implemented at least partially using cloud computing resources, and / or other types of systems.
[0037] Reference Figure 1 , Figure 1Example data flow diagrams of process 100 for locating a machine using a specific path tracker according to some embodiments of this disclosure are shown. It should be understood that such and other arrangements described herein are merely examples. Other arrangements and elements (e.g., machines, interfaces, functions, sequences, functional groups, etc.) may be used to supplement or replace the arrangements and elements shown, and some elements may be omitted entirely. Furthermore, many of the elements described herein are functional entities that may be implemented as discrete or distributed components, or in combination with other components, and in any suitable combination and location. The various functions described herein as being performed by entities can be performed by hardware, firmware, and / or software. For example, various functions can be performed using a processor that executes instructions stored in memory. In some embodiments, the systems, methods, and processes described herein may use... Figure 10A-10D Example of autonomous vehicles 1000 Figure 11 Example computing device 1100 and / or Figure 12 The example data center 1200 uses components, features, and / or functions similar to those of other components, features, and / or functions to perform the same operation.
[0038] In additional or alternative components, process 100 can be implemented using path localization system 102 and drive stack component 104. Path localization system 102 may include tracker initializer 106, one or more trackers 108A-108N (where “N” can represent any number), scoring component 110, and selection component 112. Drive stack component 104 may include mapping component 114, path prediction component 116, and refinement component 118.
[0039] In summary, the process 100 may include: a path localization system 102 receiving input data 120, which may include one or more of the following: sensor data 122, map data 124, motion data 126, perception data 128, predicted path data 130, and / or localization data 132. A tracker initializer 106 may use the input data 120 to initialize one or more trackers 108A-108N (hereinafter collectively referred to as "trackers 108") for tracking one or more candidate poses 134A-134N (where "N" can represent any number). A scoring unit 110 may use the input data 120 to evaluate one or more candidate poses 134A-134N (hereinafter collectively referred to as "candidate poses 134") and calculate one or more scores corresponding to the candidate poses 134. Based on one or more scores, a selection unit 112 may select the candidate pose from the candidate poses 134 that most closely corresponds to the machine's actual pose, and / or select one of the trackers 108 that is tracking the candidate pose. Using the selected tracker and / or candidate poses, the path localization system 102 can output localization data 132, which can indicate the machine's position, pose, etc., regarding a specific road segment in the environment. The drive stack component 104 can use the localization data 132 to perform various operations. For example, the mapping component 114 can use the localization data 132 to determine if map data is valid, the path prediction component 116 can use the localization information 132 to predict the machine's path (e.g., the most likely path), and / or the refinement component 118 can use the localization signal 132, etc., to update or refine the machine's trajectory, update or refine the machine's operating parameters (e.g., maximum / minimum speed constraints, etc.).
[0040] As shown, input data 120 may include sensor data 122. Sensor data 122 may correspond to one or more modalities of sensor data, such as LiDAR data generated using one or more LiDAR sensors, RADAR data generated using one or more RADAR sensors, image data generated using one or more image sensors (e.g., cameras), ultrasonic data generated using one or more ultrasonic sensors, GPS data generated using one or more GPS sensors, inertial data generated using one or more inertial measurement units (IMUs), or any other type of sensor data. In some examples, sensor data 122 may include raw sensor data, preprocessed sensor data, and / or processed sensor data. For example, sensor data 122 may be captured in one format (e.g., RCCB, RCCC, RBGC, etc.) and then converted (e.g., during sensor data preprocessing) to another format. In some other examples, sensor data 122 may be provided as input to a sensor data or image data preprocessor (not shown) to generate preprocessed sensor data. For example, in the context of image data, many types of images or formats can be used as input, such as compressed images (e.g., in Joint Image Experts Group (JPEG), Red-Green-Blue (RGB), or Luminosity / Chromaticity (YUV) formats), compressed images as frames derived from compressed video formats (e.g., H.264 / Advanced Video Coding (AVC), H.265 / High Efficiency Video Coding (HEVC), VP8, VP9, Open Media Video Consortium 1 (AV1), Multifunction Video Coding (VVC), or any other video compression standard), and raw images (e.g., raw images derived from Red-Blue (RCCB), Red-Blue (RCCC), or other types of imaging sensors).
[0041] In some examples, map data 124 may include SD map data representing a navigation map or SD map of the environment (e.g., a map with less detail compared to an HD map, such as a map that does not include identifiers or indicators of how many lanes are on each road, classifiers for each lane, etc.). Map data 124 may indicate the topology of at least a portion of the road network that the machine is using. For example, map data 124 may indicate the approximate location of intersections of road segments in an environment where road segments are separated from each other. Map data 124 may also indicate the approximate angle of a road segment, the length of a road segment, the start and end points of a road segment, the geometry of a road segment, the curvature of a road segment, etc. In some instances, map data 124 may be continuously updated based on the machine's location. That is, the machine may continuously receive updated map data 124 for relevant portions of the map corresponding to the area near the machine's location. For example, map data 124 may represent a map of one or more road segments 100 meters ahead of the machine, the machine's current road segment, and a map of the next intersection and subsequent road segments ahead of the machine, the current road segment only, etc.
[0042] In some examples, motion data 126 may include relative motion data and / or global motion data. For instance, relative motion data may indicate one or more rotations and / or one or more translations of the machine relative to its current pose, orientation, and / or position. That is, relative motion data may indicate the machine's tracking path, but the orientation of the tracking path may be based on the machine's pose. References herein Figures 4A-4C The examples illustrate and describe the relative motion of the machine in more detail. In contrast, global motion data can indicate the machine's motion relative to a coordinate system, such as a GPS coordinate system that includes latitude and longitude. Global motion data can be used to represent the machine's path using a series of latitude and longitude points and the machine's orientation or attitude. Thus, GPS data, IMU data, and odometry data collected at multiple timestamps can be used to determine the machine's global motion data or global motion path.
[0043] Perception data 128 can indicate the perceived location of various features in the environment. A perception system (not shown) can be used to process at least sensor data 122 and / or possibly any other data described herein to determine or generate perception data 128. For example, the perception system can receive sensor data 122 (e.g., LiDAR data, radar data, image data, etc.) and process the sensor data 122 to determine the location of objects (e.g., vehicles, pedestrians, obstacles, etc.) and / or other features (e.g., road surface, curb, lane markings, pavement markings, etc.) in the environment. In some examples, the perception system may include one or more machine learning models for predicting the location of objects, features, etc. Thus, sensor data 122 can be (e.g., raw or pre-processed) applied to machine learning models of the perception system, which can process and analyze sensor data 122 to perceive objects, features, etc., in the environment surrounding the machine. In some examples, perception data 128 can indicate the curvature of a road segment, the location of road segment intersections, road network geometry, or any other attribute or feature associated with a driving surface. The sensing data 128 can also indicate the machine's location regarding the road network, road segments, lanes, etc.
[0044] Predicted path data 130 can indicate a predicted or most likely path for the machine. For example, the predicted path can indicate the route the machine is likely to use after passing an intersection. In some examples, the predicted path of the machine can be determined based on data indicating the intention of the machine's occupant. Such data indicating the occupant's intention may include, but is not limited to, the machine's turn signals, the machine's yaw rate, the machine's predetermined or set route, or the machine's destination. The occupant's intention can indicate the machine's path. Further details regarding the use of data indicating the occupant's intention to determine a machine's predicted path are described in U.S. Patent Application No. 18 / 811368, filed August 21, 2024, the entire contents of which are incorporated herein by reference and used for all purposes.
[0045] The path localization system 102 can use the techniques disclosed herein to determine the location data 132. That is, the location data 132 output by the path localization system 102 can be fed back into the path localization system 103 as input for future iterations. The location data 132 can indicate the machine's position and / or attitude in the environment. For example, the location data 132 can indicate the machine's approximate GPS position and the machine's attitude (e.g., orientation and heading) relative to a map of the environment. Therefore, in some examples, the location data 132 can indicate which road segment the machine is operating on and / or the machine's location relative to the start and end points of the road segment.
[0046] In some instances, location data 132 may be unavailable. For example, and as further described in detail herein, in some instances, the confidence level of the machine's location along a certain road segment may be below a threshold confidence level. That is, in such scenarios, one or more components of the path localization system 102 and / or its components may be unable to determine whether the machine is located along a first path / segment or a second path / segment. In these scenarios, the path localization system 102 may avoid outputting location data 132, or the location data 132 may indicate that the machine's location cannot be determined or has been determined with low confidence. In this way, other systems or components of the machine (e.g., drive stack component 104) can adjust their behavior and output accordingly.
[0047] As described herein, process 100 may include: tracker initializer 106 initializing tracker 108 for tracking candidate poses 134 of the machine. In some instances, tracker initializer 106 may initialize tracker 108 at least based on determining that the machine has traversed an intersection. An intersection may correspond to an area of an environment where two or more road segments meet or intersect. The complexity of the intersection may vary in various instances, from simple T-junctions or Y-junctions to more complex designs (such as roundabouts or multi-lane grade-separated intersections). Thus, an intersection may include multiple options for the machine's path. For example, the intersection may correspond to the intersection of a first road segment, a second road segment, a third road segment, etc. Therefore, at the intersection, the machine can choose one of the first road segment, the second road segment, the third road segment, etc., as part of its path to traverse.
[0048] In some instances, tracker initializer 106 may determine the existence of an intersection and / or the machine's position relative to the intersection (e.g., whether the machine is about to pass through, is passing through, or has already passed through the intersection) at least based on the tracking machine's position with respect to map data 124. For example, tracker initializer 106 may have initialized a tracker to track the machine's position relative to a previous segment of the road network (e.g., the segment the machine was used to navigate to the intersection) before tracker initializer 106 determines that the machine has passed through the intersection. The tracker may have already tracked the machine's offset on the previous segment, and once the offset indicates that the machine is no longer on the segment and has moved to a subsequent segment, tracker initializer 106 may have been invoked to determine whether any new tracker needs to be initialized. For example, if the machine has passed through an intersection with at least two options for the path used by the machine, tracker initializer 106 may initialize a new tracker 108.
[0049] In some examples, the trackers described herein (such as tracker 108) can be implemented using a Kalman filter. A Kalman filter can track a machine's candidate pose using state variables that are updated every frame, every other frame, etc. In some instances, the state variables may include at least a first state variable corresponding to a road segment identifier, a second state variable corresponding to the machine's offset relative to the road segment (e.g., the machine's position relative to the start or end point of the road segment), a third state variable corresponding to the confidence level associated with the offset, and possibly other state variables.
[0050] for example, Figure 2 An example architecture of a tracker 202, which can be used to track the position and / or attitude of a machine according to some embodiments of the present disclosure, is shown. Tracker 202 may correspond to one or more trackers 108. That is, similar to tracker 202, tracker 108 may include a Kalman filter architecture that includes at least a processing model 204 and a measurement model 206. Processing model 204 may be configured to generate one or more predicted states 208 (e.g., predicted values of state variables) based on one or more previous states 210 of tracker 202. Measurement model 206 may use input data 214 (which may be related to…) Figure 1 (Corresponding to input data 120 in the example) to refine one or more predicted states 208 and output one or more updated states 212. That is, measurement model 206 can use input data 214 to refine one or more values of state variables predicted by processing model 204 from one or more previous states 210. In some examples, path localization system 102 can use one or more updated states 212 to determine localization data 132.
[0051] Return to reference Figure 1 For example, the process 100 may include: a tracker initializer 106 initializing one or more trackers 108 to track one or more candidate poses 134 of the machine (also referred to herein as “candidate locations” and / or “possible locations”) along one or more subsequent road segments away from the intersection. For instance, if the intersection has three subsequent road segments, the tracker initializer 106 may initialize two additional trackers for tracking two candidate poses 134 of the machine along two of the subsequent road segments away from the intersection (e.g., road segments away from the end of the first road segment and / or the start of the second road segment), while the original tracker may continue tracking the candidate poses of the machine along another subsequent road segment.
[0052] for example, Figure 3 Example visualizations of candidate poses that can be determined and tracked by a machine, according to some embodiments of this disclosure, are shown. Figure 3In the example, machine 302 (which may correspond to autonomous vehicle 1000) is approaching an intersection 304 comprising multiple subsequent road segments, and the predicted path 306 of machine 302 is that the machine will turn right at the intersection 304 onto the first subsequent road segment 308A. According to the technology of this disclosure, based on the machine traversing the intersection 304, tracker initializer 106 can initialize multiple trackers along candidate poses of the subsequent road segments. For example, a first tracker can be initialized to track a first candidate pose 310A of machine 302 along the first subsequent road segment 308A, a second tracker can be initialized to track a second candidate pose 310B of machine 302 along the second subsequent road segment 308B, and a third tracker can be initialized to track a third candidate pose 310C of machine 302 along the third subsequent road segment 308C. Candidate poses 310A-310C can be associated with... Figure 1 The candidate pose 134 in the example corresponds to this. In some examples, because the first subsequent road segment 308A corresponds to the predicted path 306 of machine 302, the first tracker for the first candidate pose 310A may not need to be initialized. Instead, in at least some examples, the tracker may be the same tracker used to track the position of machine 302 along the road segment leading to the intersection 304.
[0053] Return to reference Figure 1 For example, the process 100 may include: a scoring component 110 calculating a score for each of the trackers 108 and / or candidate poses 134 of the machine after the intersection. These scores can indicate which candidate pose among the candidate poses 134 corresponds to the actual position of the machine. In some examples, the scoring component 110 may calculate multiple scores for each of the trackers 108 and / or candidate poses 134 over a period of time (e.g., calculating a new score for each tracker at each time step, round, etc.) and aggregate the scores for each tracker and / or candidate pose to determine the highest score. In some examples, the score value may be calculated using the same or similar measurement (e.g., input data 120) used to update the state of the Kalman filter of tracker 108.
[0054] for example, Figures 4A-4C Some embodiments of the present disclosure are shown in relation to the present disclosure. Figure 3 The example is related to scoring the candidate poses of the machine. To explain... Figures 4A-4C Assuming in Figure 3 In the example, machine 302 has actually traveled the predicted path 306 and turned right at the intersection 304 to continue on the first subsequent road segment 308A.
[0055] Reference Figure 4AThe diagram illustrates the relative motion 402 and global motion 404 of machine 302 with respect to a first candidate pose 310A along a first subsequent road segment 308A. As shown, the geometry of the relative motion 402 and the geometry of the global motion 404 are aligned with the geometry of the road segment and the previous position 406 of machine 302. In this scenario, when the alignment correlation between the relative motion 402, the global motion 404, and the road geometry between the previous position 406 and the candidate pose of machine 302 is strong, the scoring unit 110 can calculate a relatively high score for the first candidate pose 310A.
[0056] Now, referring to Figure 4B The diagram illustrates the relative motion 402 and global motion 404 of machine 302 with respect to a second candidate pose 310B along a second subsequent road segment 308B. As shown, when the relative motion 402 is aligned with the second candidate pose 310B, the geometry of the relative motion 402 deviates from the previous position 406 of machine 302. Additionally, because the global motion 404 is about a coordinate system (e.g., a GPS coordinate system), the geometry / orientation of the global motion 404 can remain constant regardless of the candidate pose. In other words, the global motion 404 can be superimposed on the map at one or more of its GPS locations and does not need to be determined relative to the pose of machine 302 as the relative motion 402 does. Thus, as shown, the geometry of the global motion 404 deviates from the second candidate pose 310B. Figure 4B In such scenarios as illustrated in the example, when the alignment correlation between the relative motion 402 of machine 302, the global motion 404, and the road geometry between the previous position 406 and the second candidate pose 310B is weak, the scoring component 110 may calculate a relatively low score for the second candidate pose 310B, which may indicate that the actual pose of machine 302 may not correspond to the second candidate pose 310B.
[0057] Now, referring to Figure 4C The diagram illustrates the relative motion 402 and global motion 404 of machine 302 with respect to a third candidate pose 310C along a third subsequent segment 308C. As shown, when relative motion 402 aligns with the third candidate pose 310C, the geometry of relative motion 402 deviates from the previous position 406 of machine 302. Additionally, as shown, the geometry of global motion 404 fails to align with the third candidate pose 310C. Figure 4CIn such scenarios as illustrated in the example, when the alignment correlation between the relative motion 402 of machine 302, global motion 404, and the road geometry between the previous position 406 and the third candidate pose 310C is weak, the scoring component 110 can calculate a relatively low score for the third candidate pose 310C, which may indicate that the actual pose of machine 302 may not correspond to the third candidate pose 310C.
[0058] Return to reference Figure 1 For example, the process 100 may include: selection component 112 determining which candidate pose among candidate poses 134 corresponds to the machine's actual pose, and selecting one tracker from trackers 108 to track the machine's pose or position along the current road segment. In some examples, the dominant candidate pose may be selected based on the score of the candidate poses 134 tracked by each tracker in trackers 108. Additionally or alternatively, the dominant candidate pose may be selected based on the geometry of the machine's relative motion and / or global motion, where the alignment between the geometry of the machine's previous position and the candidate pose is greater than an alignment threshold amount. For example, if the tip and tail of the machine's relative motion and / or global motion match or align with the machine's previous position and the machine's candidate pose, the candidate pose may be selected as the machine's actual pose, without necessarily calculating a score.
[0059] In some examples, selection component 112 can evaluate the score calculated by scoring component 110 to determine the tracker or candidate pose with the highest score, and whether that score is above a threshold compared to one or more other scores. For example, and referring to... Figure 3 For example, selection component 112 may determine that one or more scores associated with the first candidate pose 310A are greater than the scores associated with the second candidate pose 310B and the third candidate pose 310C. In this example, selection component 112 may determine that the position / pose of machine 302 corresponds to the first candidate pose 310A based at least on the first candidate pose 310A having the highest score. Additionally, in some examples, selection component 112 may determine whether one or more differences between the score of the first candidate pose 310A and the scores of the second and third candidate poses reach or exceed a threshold, and determining that the position / pose of machine corresponds to the first candidate pose 310A may also be based on one or more differences reaching or exceeding that threshold.
[0060] In some examples, while continuously recalculating and / or aggregating scores, selection component 112 can evaluate scores over a period of time, and once a score associated with one of the candidate poses 134 is determined to be significantly larger than other scores by a threshold, selection component 112 can set the tracker as the dominant candidate pose and remove / terminate other trackers for other candidate poses. For example, in Figure 3 In the example, if the first candidate pose 310A is the dominant pose, the selection component 112 can terminate the trackers associated with the second candidate pose 310B and the third candidate pose 310C. In some examples, if the machine travels beyond a threshold distance and / or beyond a threshold time period across multiple candidate poses and trackers, and none of the scores reach or exceed the threshold, the selection component 112 may mark the localization result as unavailable.
[0061] In some examples, path localization system 102 may use a selected tracker to track the machine's position, and path localization system 101 may generate localization data 132 based on the output of the selected tracker. In some examples, localization data 132 may indicate a specific road segment that the machine is using after an intersection, the machine's position along that road segment (e.g., the machine's offset along the road segment relative to the start or end point of the road segment), confidence associated with the localization result, etc. In some examples, localization data 132 may be sent to one or more drive stack components 104, and drive stack component 104 may use localization data 132 as input to perform various processes associated with the machine.
[0062] For example, mapping unit 114 can use location data 132 to determine whether to update the cached map of the environment. For example, as described above and herein, when a machine traverses an environment, it can continuously receive updated map data representing portions of the map (e.g., navigation map, SD map, etc.) for use in traversing the environment. Thus, the machine can receive and temporarily store relevant portions of the map, such as map portions corresponding to the machine's current and / or future location. In some examples, once a map portion is deemed valid (e.g., for the machine's path), mapping unit 114 can cache the map portion until a new map portion is received and is also deemed valid. Thus, mapping unit 114 can use location data 132 to determine whether the newly received map portion is valid and / or whether to update the cached map.
[0063] In these instances, if the newly received map portion does not include the road segments currently being used by the machine in the positioning data 132, the mapping unit 114 can determine that the newly received map portion is invalid and avoid updating the cached map. That is, because the new map portion is invalid, the mapping unit 114 may cause the machine to continue relying on the cached map for navigation. As another example, if the new map portion does not include road segments corresponding to the machine's predicted path or the machine's most likely path, the mapping unit 114 can avoid updating the cached map. However, if the new map portion includes all relevant road segments, the mapping unit 114 can determine that the new map portion is valid and allow the cached map to be updated to the newly received map.
[0064] for example, Figures 5A-5C The illustration shows an example of a cache map associated with determining whether to update the environment, according to some embodiments of the present disclosure. First, refer to... Figure 5A The cached map 502 may include multiple road segments 504A-504E representing a portion of the road network's topology. When a new map 508A is received, the mapping unit 114 can determine that the new map 508A is invalid because the new map 508B does not include road segment 504C corresponding to the predicted path 506 of machine 302 and / or the location of machine 302 along road segment 504C. Similarly, in Figure 5B In the process, when a new map 508B is received, even if the new map 508B includes the current road segment 504A where machine 302 is located, the mapping unit 114 can also determine that the new map 508B is invalid because it does not include the road segment 504C corresponding to the predicted path 506 of machine 302. However, in Figure 5C When a new map 508C is received, the mapping unit 114 can determine that the new map 508C is valid because it includes road segments 504B and 504E corresponding to the predicted path 508 of machine 301, as well as road segment 504B where machine 302 is currently located (as indicated in the positioning data 132). In this example, the mapping unit 114 can update the cached map 502 to correspond to the new map 508C.
[0065] Return to reference Figure 1For example, the process 100 may include: a path prediction component 116 using location data 132 to determine a predicted path for the machine. For instance, the path prediction component 116 may use the location data 132 to determine the current road segment the machine is traversing, and then, based on that current road segment and, in some instances, based on additional information (e.g., occupant intent data), determine the most probable path for the machine. Additionally, in some examples, a refinement component 118 may use the location data 132 to refine one or more of the machine's trajectories, one or more operational constraints of the machine, and / or any other operations performed by the machine. For example, the refinement component 118 may use the location data 132 to refine the machine's trajectory, or to refine the maximum speed under which the machine is allowed to operate. For example, if the location data 132 indicates that the machine is traversing a high-curvature road, the refinement component 118 may at least reduce the machine's maximum operating speed while the machine is traversing a high-curvature road.
[0066] Now, referring to Figure 6 , Figure 6 Examples of systems 600 according to some embodiments of the present disclosure, capable of performing one or more of the processes described herein, are illustrated. As shown, system 602 (which may represent and / or include one or more example computing devices 1100 and / or example data centers 1200) may include one or more processors 604 (which may be similar to and / or include CPU 1106 and / or GPU 1108) and memory 606 (which may be similar to and / or include memory 1104). For example, memory 606 may store one or more path localization systems 102, including a tracker initializer 106, one or more trackers 108A-108N, a scoring unit 110 and / or a selection unit 112, and a drive stack unit 104, which may include a mapping unit 114, a path prediction unit 116, and / or a refinement unit 118. Additionally, one or more processors 604 may implement one or more path localization systems 102, which include a tracker initializer 106, one or more trackers 108A-108N, a scoring unit 110 and / or a selection unit 112, and / or a drive stack unit 104, which includes a mapping unit 114, a path prediction unit 116, and / or a refinement unit 118, to perform one or more processes described herein.
[0067] Now, referring to Figure 7-9Each block of methods 700, 800, and 900 described herein includes a computational process that can be executed using any combination of hardware, firmware, and / or software. For example, various functions can be performed by a processor executing instructions stored in memory. These methods can also be embodied as computer-usable instructions stored on a computer storage medium. These methods can be provided by standalone applications, services (standalone or combined with other managed services), or plug-ins to managed services or other products, to name a few. Additionally, regarding... Figure 1 Methods 700, 800, and 900 are described with examples. However, these methods may be implemented additionally or alternatively by any one of the systems or any combination of systems, including but not limited to the systems described herein.
[0068] Figure 7 This is a flowchart illustrating an example of a method for locating a machine after an intersection using a path-specific tracker, according to some embodiments of the present disclosure. At block B702, the method 700 may include updating the tracking position of the machine. For example, a first tracker may use input data 120 to update the tracking position of the machine. In some examples, the tracker may correspond to tracker 202 and include a Kalman filter architecture that includes at least a processing model 204 and a measurement model 206. To update the tracking position of the machine, tracker 202 may use processing model 204 to generate one or more predicted states 208 based on one or more previous states 210, and measurement model 206 may use input data 214 (e.g., sensor data, perception data, map data, etc.) to refine one or more predicted states 208 and output one or more updated states 212. One or more updated states 212 may correspond to or indicate the tracking position of the machine. In some examples, updating the machine's tracking position may include: updating the state variable of the machine's offset relative to the current road segment using a Kalman filter, updating the road segment identifier state variable, or updating the confidence state variable, the confidence associated with the offset, and / or the road segment identifier.
[0069] At box B704, method 700 may include determining whether the machine has passed the intersection. For example, based on the updated tracking position of the machine, path localization system 102 and / or tracker initializer 106 may determine whether the machine has passed the intersection. In some examples, it may be determined whether the machine has passed the intersection if the Kalman filter's segment identifier state variable changes from a first segment identifier (associated with the previous segment) to a second segment identifier (associated with its subsequent segment) and / or if the offset value has moved beyond the end of the previous segment. If it is determined at box B704 that the machine has passed the intersection, method 700 proceeds to box B706. Otherwise, if the machine has not yet passed the intersection, method 700 continues back to box B702.
[0070] At box B706, the method 700 may include: initializing additional trackers. For example, tracker initializer 106 may initialize one or more trackers 108 under one or more candidate poses 134 of the machine. In some examples, tracker initializer 106 may initialize additional trackers for each subsequent road segment. For example, although the intersection of a four-way stop may include or be associated with a total of four road segments, the four-way stop intersection may include three subsequent road segments because the machine can use one of the road segments to approach the intersection. Thus, in some examples, tracker initializer 106 may initialize a tracker on each of the three subsequent road segments. Additionally or alternatively, in some instances, since the initial tracker used to approach the intersection along a non-subsequent road segment to track the machine may automatically move to one of the subsequent road segments, the tracker may only need to initialize a new tracker on two of the three subsequent road segments. Thus, in some examples, the tracker initializer can initialize fewer trackers than the total number of subsequent road segments.
[0071] At box B708, the method 700 may include scoring candidate poses. For example, scoring component 110 may use input data 120 to calculate a score for candidate pose 134. The score may indicate which candidate pose corresponds to the machine's actual pose or position, which may further indicate which subsequent road segment the machine is traversing. In some examples, scoring component 110 may calculate candidate poses based on the machine's relative motion or global motion, as described herein. Figures 4A-4CAs described. Additionally or alternatively, the scoring unit 110 may calculate candidate scores based on perception data 128, which may indicate lane markings and / or other road features. Additionally or alternatively, the scoring unit 110 may calculate candidate postures by comparing the machine's yaw rate with the orientation or geometry of the road segment. For example, if the machine's yaw rate is 30 degrees and a subsequent segment in a subsequent road segment also has a yaw rate of 30 degrees, the scoring unit 110 may calculate a score indicating a high probability that the machine is operating on that road segment.
[0072] At box B710, method 700 may include: determining the tracker / candidate pose with the highest score. For example, selection component 112 may determine the tracker with the highest score for tracker 108 and / or the pose with the highest score for candidate pose 134. In some examples, selection component 112 may determine the tracker / pose with the highest score over a period of time by aggregating the scores of the corresponding trackers / poses in multiple iterations of scoring component 110, which calculates scores. For example, scoring component 110 may calculate scores at each timestamp, each time new input data 120 is received, and / or each round, and selection component 112 may aggregate the scores to select the tracker / candidate pose with the highest score over that period of time.
[0073] At box B712, method 700 may include: determining one or more differences between the highest score and one or more other scores. For example, selection component 112 may determine one or more differences between the highest score and one or more other scores, such that the selected tracker / candidate pose is determined with high determinism. In some examples, selection component 112 may determine the difference between the highest score and the next highest score, without calculating multiple differences between each score. Additionally or alternatively, selection component 112 may calculate the differences between each score.
[0074] At box B714, method 700 may include determining whether one or more differences have reached or exceeded a threshold. For example, selection component 112 may determine whether one or more differences have reached or exceeded a threshold. In some instances, whether one or more differences have reached or exceeded a threshold may have a time factor, and one or more differences may need to reach or exceed the threshold within a threshold time period. In various examples, the threshold may be user-defined or determined by selection component 112 in real-time or near real-time. Thus, selection component 112 may have the ability to change and adjust the threshold as it deems appropriate. If one or more differences are determined to have reached or exceeded a threshold at box B714, method 700 may proceed to box B716. Otherwise, if one or more differences have not reached or exceeded the threshold, method 700 returns to box B708, and scores may be recalculated, re-aggregated, and re-evaluated.
[0075] At box B716, method 700 may include determining whether the current tracker is a new tracker (e.g., a tracker initialized after the intersection). That is, in some instances, selection component 112 may determine whether the dominant tracker / candidate pose (e.g., the highest-scoring tracker exceeding a threshold and / or exceeding a threshold for a certain period of time) is the original tracker tracking the machine's position toward the intersection, or whether the tracker is a new tracker initialized by tracker initializer 106 in response to the machine passing through the intersection. If the tracker is determined to be a new tracker at box B716, method 700 proceeds to box B718. Otherwise, method 700 proceeds to box B720.
[0076] At box B718, method 700 may include: reinitializing a new tracker in place of the previous tracker. For example, path localization system 102 may replace the original tracker that tracks the location of the machine toward the intersection with a new dominant tracker initialized by tracker initializer 106 in response to the machine passing through the intersection.
[0077] At box B720, method 700 may include terminating one or more other trackers. For example, path localization system 102 or tracker initializer 106 may terminate one or more non-dominant trackers or trackers not selected as having the highest scores. Thus, once the machine's position / attitude is known or determined with high confidence, path localization system 102 can track the machine's position using only one tracker. However, if the machine crosses another intersection again, method 700 can be repeated from box 706, and more trackers can be initialized.
[0078] Now, referring to Figure 8 , Figure 8 This is a flowchart illustrating an example of a method 800 for determining the location of a machine based on scoring multiple candidate locations, according to some embodiments of the present disclosure. At block B802, the method 800 may include: determining multiple candidate locations of the machine along multiple road segments included in a map of the environment. For example, tracker initializer 106, tracker 108, and / or scoring unit 110 may determine multiple candidate poses 134 of the machine along multiple road segments included in a map of the environment represented using map data 124.
[0079] At block B804, method 800 may include: calculating multiple scores based at least on sensor data indicating at least one tracking path of the machine, the multiple scores indicating which of a plurality of candidate positions corresponds to the position of the machine. For example, scoring unit 110 may use input data 120 to calculate multiple scores indicating which of a plurality of candidate poses 134 corresponds to the position of the machine.
[0080] At box B806, the method 800 may include: determining, at least based on the aggregation of multiple scores over a period of time, that the machine's location corresponds to a first candidate location among multiple candidate locations set along a first road segment of multiple road segments. For example, scoring component 110 and / or selection component 112 may aggregate multiple scores over a period of time, and selection component 112 may determine that the machine's location corresponds to a first candidate pose among multiple candidate poses 134 set along a first road segment of multiple road segments.
[0081] At block B808, method 800 may include performing one or more operations associated with the machine in the environment, based at least on tracking the machine's position along the first road segment. For example, drive stack component 104 may perform one or more operations associated with the machine in the environment, based at least on positioning data 132 indicating the machine's position along the first road segment.
[0082] Figure 9 This is a flowchart illustrating an example of a method 900 for determining a road segment that a machine is using after passing an intersection, according to some embodiments of the present disclosure. At block B902, the method 900 may include obtaining multiple possible positions of the machine along multiple road segments, at least based on determining that the machine has passed through an intersection associated with multiple road segments. For example, scoring unit 110 may obtain multiple candidate poses 134 indicating multiple possible positions of the machine along multiple road segments.
[0083] At box B904, method 900 may include: determining the machine's position along a first segment of a plurality of road segments based at least on one or more tracking movements of the machine relative to a plurality of possible positions. For example, selection component 112 may determine the machine's position along the first segment of a plurality of road segments. Selection component 112 may select one of the trackers 108 corresponding to the machine's position.
[0084] At box B906, method 900 may include performing one or more machine-related operations based at least on the machine's position along the first road segment. For example, drive stack component 104 may perform one or more machine-related operations based at least on the location data 132 indicating the machine's position along the first road segment.
[0085] Example autonomous vehicles
[0086] Figure 10A This is an illustration of an example autonomous vehicle 1000 according to some embodiments of the present disclosure. The autonomous vehicle 1000 (or, alternatively, referred to herein as “vehicle 1000”) may include, but is not limited to, passenger vehicles such as cars, trucks, buses, first response vehicles, shuttle buses, electric or motorized bicycles, motorcycles, fire trucks, police vehicles, ambulances, boats, construction vehicles, underwater vessels, robotic vehicles, drones, aircraft, vehicles coupled to trailers (e.g., semi-trailers for hauling goods) and / or other types of vehicles (e.g., driverless and / or vehicles accommodating one or more passengers). Autonomous vehicles are typically described according to the levels of automation defined by the National Highway Traffic Safety Administration (NHTSA), a division of the U.S. Department of Transportation, and the Society of Automotive Engineers (SAE) in its "Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles" (Standard No. J3016-201806, published June 15, 2018; Standard No. J3016-201609, published September 30, 2016; and previous and future versions of this standard). Vehicle 1000 may be able to achieve one or more functions that meet Level 3-5 of the autonomous driving level. Vehicle 1000 may be able to achieve one or more functions that meet Level 1-5 of the autonomous driving level. For example, depending on the embodiment, vehicle 1000 may be able to achieve driver assistance (Level 1), partial automation (Level 2), conditional automation (Level 3), high automation (Level 4), and / or full automation (Level 5). The term “autonomy” as used herein can include any and / or all types of autonomy of the vehicle 1000 or other machines, such as full autonomy, high autonomy, conditional autonomy, partial autonomy, provision of auxiliary autonomy, semi-autonomy, primary autonomy or other specified autonomy.
[0087] Vehicle 1000 may include components such as chassis, body, wheels (e.g., 2, 4, 6, 8, 18, etc.), tires, axles, and other vehicle components. Vehicle 1000 may include a propulsion system 1050, such as an internal combustion engine, a hybrid power plant, an all-electric motor, and / or another type of propulsion system. Propulsion system 1050 may be connected to the drivetrain of vehicle 1000, which may include a transmission, to enable propulsion of vehicle 1000. Propulsion system 1050 may be controlled in response to receiving a signal from throttle / accelerator 1052.
[0088] A steering system 1054, which may include a steering wheel, can be used to steer the vehicle 1000 (e.g., along a desired path or route) when the propulsion system 1050 is operating (e.g., when the vehicle is in motion). The steering system 1054 may receive signals from the steering actuator 1056. For fully automatic (level 5) functionality, the steering wheel may be optional.
[0089] The brake sensor system 1046 can be used to operate the vehicle brakes in response to receiving a flag from the brake actuator 1048 and / or the brake sensor.
[0090] It may include one or more System-on-Chip (SoC) 1004 ( Figure 10C One or more controllers 1036, including one or more GPUs, may provide (e.g., indicating commands) flags to one or more components and / or systems of vehicle 1000. For example, one or more controllers may send flags to operate vehicle brakes via one or more brake actuators 1048, to operate steering system 1054 via one or more steering actuators 1056, and to operate propulsion system 1050 via one or more throttles / accelerators 1052. One or more controllers 1036 may include one or more onboard (e.g., integrated) computing devices (e.g., supercomputers) that process sensor flags and output operating commands (e.g., flags indicating commands) to enable autonomous driving and / or human-assisted driving of vehicle 1000. One or more controllers 1036 may include a first controller 1036 for autonomous driving functions, a second controller 1036 for functional safety functions, a third controller 1036 for artificial intelligence functions (e.g., computer vision), a fourth controller 1036 for infotainment functions, a fifth controller 1036 for redundancy in emergency situations, and / or other controllers. In some examples, a single controller 1036 may handle two or more of the above functions, two or more controllers 1036 may handle a single function, and / or any combination thereof.
[0091] One or more controllers 1036 may provide indications for controlling one or more components and / or systems of vehicle 1000 in response to sensor data (e.g., sensor inputs) received from one or more sensors. Sensor data may be received from, for example, but not limited to, Global Navigation Satellite System (“GNSS”) sensors 1058 (e.g., Global Positioning System sensors), RADAR sensors 1060, ultrasonic sensors 1062, LIDAR sensors 1064, inertial measurement unit (IMU) sensors 1066 (e.g., accelerometers, gyroscopes, magnetic compasses, magnetometers, etc.), microphones 1096, stereo cameras 1068, wide-angle cameras 1070 (e.g., fisheye cameras), infrared cameras 1072, surround cameras 1074 (e.g., 360-degree cameras), long-range and / or medium-range cameras 1098, speed sensors 1044 (e.g., for measuring the rate of vehicle 1000), vibration sensors 1042, steering sensors 1040, braking sensors (e.g., as part of braking sensor system 1046), and / or other sensor types.
[0092] One or more of the controllers 1036 may receive inputs (e.g., represented by input data) from the instrument group 1032 of the vehicle 1000 and provide outputs (e.g., represented by output data, display data, etc.) via the human-machine interface (HMI) display 1034, auditory markers, speakers, and / or via other components of the vehicle 1000. These outputs may include, for example, vehicle speed, rate, time, map data (e.g., [missing information]). Figure 10C Information such as high-definition (“HD”) map 1022, location data (e.g., the location of vehicle 1000 on the map), direction, and the location of other vehicles (e.g., occupying a grid), as well as information about objects and their states perceived by controller 1036, etc. For example, HMI display 1034 may display information about the existence of one or more objects (e.g., street signs, warning signs, traffic light changes, etc.) and / or information about driving maneuvers that the vehicle has made, is making, or will make (e.g., changing lanes now, leaving 34B in two miles, etc.).
[0093] Vehicle 1000 also includes a network interface 1024, which can communicate via one or more networks using one or more wireless antennas 1026 and / or a modem. For example, network interface 1024 may be able to communicate via Long Term Evolution (“LTE”), Wideband Code Division Multiple Access (“WCDMA”), Universal Mobile Telecommunications System (“UMTS”), Global System for Mobile Communications (“GSM”), IMT-CDMA Multicarrier (“CDMA2000”), etc. One or more wireless antennas 1026 may also enable communication between objects in the environment (e.g., vehicles, mobile devices, etc.) using one or more local area networks such as Bluetooth, Bluetooth Low Energy (“LE”), Z-Wave, ZigBee, etc., and / or one or more low-power wide area networks (LPWAN) such as LoRaWAN, SigFox, etc.
[0094] Figure 10B For use in accordance with some embodiments of this disclosure Figure 10A This is an example of the camera position and field of view of an autonomous vehicle 1000. The camera and its respective field of view are an example embodiment and are not intended to be limiting. For example, additional and / or replaceable cameras may be included, and / or these cameras may be located at different positions on the vehicle 1000.
[0095] The camera type used for the camera may include, but is not limited to, a digital camera suitable for use with components and / or systems of vehicle 800. The camera may operate at Automotive Safety Integrity Level (ASIL) B and / or another ASIL. The camera type may have any image capture rate, such as 60 frames per second (fps), 120 fps, 240 fps, etc., depending on the embodiment. The camera may be able to use a rolling shutter, a global shutter, another type of shutter, or a combination thereof. In some examples, the color filter array may include a red-white-white-white (RCCC) color filter array, a red-white-white-blue (RCCB) color filter array, a red-blue-green-white (RBGC) color filter array, a Foveon X3 color filter array, a Bayer sensor (RGGB) color filter array, a monochrome sensor color filter array, and / or another type of color filter array. In some embodiments, a sharp-pixel camera, such as one having an RCCC, RCCB, and / or RBGC color filter array, may be used in efforts to improve light sensitivity.
[0096] In some examples, one or more of the cameras can be used to perform advanced driver assistance system (ADAS) functions (e.g., as part of a redundant or fail-safe design). For example, a multi-function monocular camera can be installed to provide functions including lane departure warning, traffic sign assistance, and intelligent headlight control. One or more of the cameras (e.g., all cameras) can simultaneously record and provide image data (e.g., video).
[0097] One or more of the cameras can be mounted in mounting components such as custom-designed (3D-printed) parts to cut off stray light and reflections from inside the vehicle (e.g., reflections from dashboard panels reflected in the windshield mirror) that may interfere with the camera's image data capture capabilities. Regarding wing mirror mounting components, these components can be custom-3D printed so that the camera mounting plate matches the shape of the wing mirror. In some examples, one or more cameras can be integrated into the wing mirror. For side-view cameras, one or more cameras can also be integrated into the four pillars at each corner of the cab.
[0098] A camera with a field of view that includes the environment in front of the vehicle 1000 (e.g., a front-facing camera) can be used for surround view to help identify forward paths and obstacles, and, with the assistance of one or more controllers 1036 and / or control SoCs, to provide information crucial for generating an occupancy grid and / or determining a preferred vehicle path. The front-facing camera can be used to perform many of the same ADAS functions as LiDAR, including emergency braking, pedestrian detection, and collision avoidance. The front-facing camera can also be used in ADAS functions and systems, including Lane Departure Warning (“LDW”), Autonomous Cruise Control (“ACC”), and / or other functions such as traffic sign recognition.
[0099] A variety of cameras can be used in front-facing configurations, including, for example, monocular camera platforms that include complementary metal-oxide-semiconductor (“CMOS”) color imagers. Another example could be a wide-angle camera 1070, which can be used to perceive objects entering the field of view from the periphery (such as pedestrians, traffic at intersections, or bicycles). Although Figure 10B The middle image shows only one wide-angle camera, but any number (including zero) of wide-angle cameras 1070 can exist on vehicle 1000. Furthermore, any number of remote cameras 1098 (e.g., long-view stereo camera pairs) can be used for depth-based object detection, especially for objects for which neural networks have not yet been trained. Remote cameras 1098 can also be used for object detection and classification, as well as basic object tracking.
[0100] Any number of stereo cameras 1068 may also be included in a front-mounted configuration. In at least one embodiment, one or more stereo cameras 1068 may include an integrated control unit that includes a scalable processing unit that can provide a multi-core microprocessor and programmable logic (“FPGA”) with an integrated controller area network (“CAN”) or Ethernet interface on a single chip. Such a unit can be used to generate a 3D map of the vehicle environment, including distance estimates for all points in the image. Alternative stereo cameras 1068 may include a compact stereo vision sensor that may include two camera lenses (one on each side) and an image processing chip that can measure the distance from the vehicle to a target object and use the generated information (e.g., metadata) to activate autonomous emergency braking and lane departure warning functions. Other types of stereo cameras 1068 may be used in addition to those described herein or alternatively.
[0101] Cameras (e.g., side-view cameras) with a field of view including the side of the vehicle 1000 can be used for surround view, providing information for creating and updating occupancy grids and generating side-impact collision warnings. For example, surround camera 1074 (e.g., ... Figure 10B The four surround cameras 1074 shown can be mounted on the vehicle 1000. The surround cameras 1074 can include a wide-angle camera 1070, a fisheye camera, a 360-degree camera, and / or similar devices. Four examples are provided; the four fisheye cameras can be positioned at the front, rear, and sides of the vehicle. In an alternative arrangement, the vehicle can use three surround cameras 1074 (e.g., left, right, and rear) and can utilize one or more other cameras (e.g., forward-facing cameras) as a fourth surround-view camera.
[0102] A camera (e.g., a rear-view camera) having a field of view that includes the environment behind the vehicle 1000 can be used for assisted parking, surround view, rear collision warning, and creating and updating occupancy grids. A wide variety of cameras can be used, including but not limited to those also suitable as front cameras as described herein (e.g., long-range and / or mid-range camera 1098, stereo camera 1068, infrared camera 1072, etc.).
[0103] Figure 10C For use in accordance with some embodiments of this disclosure Figure 10AThe example autonomous vehicle 1000 is illustrated in the block diagram of an example system architecture. It should be understood that this arrangement, and other arrangements described herein, are merely illustrative. Other arrangements and elements (e.g., machines, interfaces, functions, sequences, functional groupings, etc.) may be used in addition to or in place of those shown, and some elements may be omitted entirely. Furthermore, many of the elements described herein are functional entities, which may be implemented as discrete or distributed components or in combination with other components, and in any suitable combination and location. The various functions described herein as being performed by these entities can be implemented via hardware, firmware, and / or software. For example, the various functions can be implemented by a processor executing instructions stored in memory.
[0104] Figure 10C Each component, feature, and system in vehicle 1000 is illustrated as being connected via bus 1002. Bus 1002 may include a Controller Area Network (CAN) data interface (or, alternatively, referred to herein as the "CAN bus"). CAN may be a network within vehicle 1000 used to assist in controlling various features and functions of vehicle 1000, such as braking, acceleration, steering, windshield wipers, and other driving functions. The CAN bus may be configured to have dozens or even hundreds of nodes, each with its own unique identifier (e.g., CAN ID). The CAN bus can be read to find steering wheel angle, ground speed, engine speed per minute (RPM), button positions, and / or other vehicle status indicators. The CAN bus may be ASIL B compliant.
[0105] Although bus 1002 is described herein as a CAN bus, this is not intended to be limiting. For example, FlexRay and / or Ethernet may be used in addition to or alternatively to a CAN bus. Furthermore, although bus 1002 is represented by a single line, this is not intended to be limiting. For example, any number of buses 1002 may exist, which may include one or more CAN buses, one or more FlexRay buses, one or more Ethernet buses, and / or one or more other types of buses using different protocols. In some examples, two or more buses 1002 may be used to perform different functions and / or may be used for redundancy. For example, a first bus 1002 may be used for a collision avoidance function, and a second bus 1002 may be used for driving control. In any example, each bus 1002 may communicate with any component of vehicle 1000, and two or more buses 1002 may communicate with the same component. In some examples, each SoC 1004, each controller 1036, and / or each computer within the vehicle may have access to the same input data (e.g., input from sensors in the vehicle 1000) and may be connected to a common bus such as the CAN bus.
[0106] Vehicle 1000 may include one or more controllers 1036, such as those described herein. Figure 10A The controllers described herein. Controller 1036 can be used for a wide variety of functions. Controller 1036 can be coupled to any other different components and systems of vehicle 1000 and can be used for the control of vehicle 1000, artificial intelligence of vehicle 1000, infotainment and / or similar functions of vehicle 1000.
[0107] Vehicle 1000 may include one or more System-on-Chip (SoC) 1004. SoC 1004 may include CPU 1006, GPU 1008, processor 1010, cache 1012, accelerator 1014, data storage 1016, and / or other components and features not shown. SoC 1004 can be used to control vehicle 1000 in a wide variety of platforms and systems. For example, one or more SoCs 1004 may be combined with an HD map 1022 in a system (e.g., the system of vehicle 1000), the HD map being accessible from one or more servers (e.g., via a network interface 1024). Figure 10D One or more servers (1078) receive map refresh and / or updates.
[0108] CPU 1006 may include CPU clusters or CPU complexes (or, alternatively, referred to herein as "CCPLEX"). CPU 1006 may include multiple cores and / or L2 cache. For example, in some embodiments, CPU 1006 may include eight cores in a coherent multiprocessor configuration. In some embodiments, CPU 1006 may include four dual-core clusters, each cluster having a dedicated L2 cache (e.g., 2MB L2 cache). CPU 1006 (e.g., CCPLEX) may be configured to support simultaneous cluster operation, such that any combination of clusters of CPU 1006 can be active at any given time.
[0109] CPU 1006 can implement power management capabilities including one or more of the following features: automatic clock gating of hardware blocks when idle to save dynamic power; clock gating of each core when the core is not actively executing instructions due to the execution of WFI / WFE instructions; independent power gating of each core; independent clock gating of each core cluster when all cores are clock-gated or power-gated; and / or independent power gating of each core cluster when all cores are power-gated. CPU 1006 can further implement enhanced algorithms for managing power states, wherein allowed power states and desired wake-up times are specified, and the hardware / microcode determines the optimal power state to enter for the core, cluster, and CCPLEX. The processing core can support simplified power state entry sequences in software, with this work offloaded to the microcode.
[0110] GPU 1008 may include an integrated GPU (or, alternatively, referred to herein as an "iGPU"). GPU 1008 may be programmable and efficient for parallel workloads. In some examples, GPU 1008 may use an enhanced tensor instruction set. GPU 1008 may include one or more streaming microprocessors, wherein each streaming microprocessor may include an L1 cache (e.g., an L1 cache with at least 96KB of storage capacity), and two or more of these streaming microprocessors may share an L2 cache (e.g., an L2 cache with 512KB of storage capacity). In some embodiments, GPU 1008 may include at least eight streaming microprocessors. GPU 1008 may use a computation application programming interface (API). Furthermore, GPU 1008 may use one or more parallel computing platforms and / or programming models (e.g., NVIDIA's CUDA).
[0111] In automotive and embedded applications, the GPU 1008 can be power-optimized for optimal performance. For example, the GPU 1008 can be fabricated on FinFETs. However, this is not intended to be limiting, and the GPU 1008 can be fabricated using other semiconductor manufacturing processes. Each streaming microprocessor can combine several mixed-precision processing cores divided into multiple blocks. For example, and without limitation, 64 PF32 cores and 32 PF64 cores can be divided into four processing blocks. In such an example, each processing block can be allocated 16 FP32 cores, 8 FP64 cores, 16 INT32 cores, two mixed-precision NVIDIA Tensor cores for deep learning matrix arithmetic, an L0 instruction cache, a warp scheduler, a dispatch unit, and / or a 64KB register file. Furthermore, the streaming microprocessor can include independent parallel integer and floating-point data paths to leverage the mixture of computation and addressing computations to provide efficient execution of workloads. Streaming microprocessors may include independent thread scheduling capabilities to allow for finer-grained synchronization and cooperation between parallel threads. Streaming microprocessors may include combined L1 data caches and shared memory units to improve performance while simplifying programming.
[0112] The GPU 1008 may include, in some examples, a High Bandwidth Memory (HBM) and / or a 16GB HBM2 memory subsystem providing a peak memory bandwidth of approximately 900GB / s. In some examples, in addition to HBM memory or alternatively, Synchronous Graphics Random Access Memory (SGRAM), such as Generation 5 Graphics Double Data Rate Synchronous Random Access Memory (GDDR5), may be used.
[0113] The GPU 1008 may include unified memory technology, which includes access counters to allow memory pages to be migrated more precisely to the processors that access them most frequently, thereby improving the efficiency of shared memory ranges between processors. In some examples, Address Translation Service (ATS) support can be used to allow the GPU 1008 to directly access the CPU 1006 page tables. In such examples, when the GPU 1008 Memory Management Unit (MMU) experiences a miss, the address translation request can be transferred to the CPU 1006. In response, the CPU 1006 can look up the virtual-physical mapping for the address in its page tables and transfer the translation back to the GPU 1008. Thus, unified memory technology can allow a single unified virtual address space for the memory of both the CPU 1006 and the GPU 1008, simplifying GPU 1008 programming and porting applications to the GPU 1008.
[0114] In addition, the GPU 1008 may include access counters that track how frequently the GPU 1008 accesses the memory of other processors. Access counters can help ensure that memory pages are moved to the physical memory of the processor that accesses those pages most frequently.
[0115] SoC 1004 may include any number of caches 1012, including those described herein. For example, cache 1012 may include an L3 cache available to both CPU 1006 and GPU 1008 (e.g., it is connected to both CPU 1006 and GPU 1008). Cache 1012 may include a write-back cache, which can track the state of rows, for example, using a cache coherence protocol (e.g., MEI, MESI, MSI, etc.). Depending on the embodiment, the L3 cache may include 4MB or more, but a smaller cache size may also be used.
[0116] SoC 1004 may include an arithmetic logic unit (ALU) that can be utilized in processing of any of the various tasks or operations performed on vehicle 1000, such as processing a DNN. Furthermore, SoC 1004 may include a floating-point unit (FPU) (or other mathematical coprocessor or digital coprocessor type) for performing mathematical operations within the system. For example, SoC 1004 may include one or more FPUs integrated as execution units within CPU 1006 and / or GPU 1008.
[0117] SoC 1004 may include one or more accelerators 1014 (e.g., hardware accelerators, software accelerators, or combinations thereof). For example, SoC 1004 may include a hardware accelerator cluster, which may include optimized hardware accelerators and / or large on-chip memory. This large on-chip memory (e.g., 4MB SRAM) can enable the hardware accelerator cluster to accelerate neural networks and other computations. The hardware accelerator cluster can be used to secondary GPU 1008 and offload some tasks from GPU 1008 (e.g., freeing up more cycles of GPU 1008 to perform other tasks). As an example, accelerator 1014 may be used for targeted workloads (e.g., perceptrons, convolutional neural networks (CNNs), etc.) that are stable enough to be easily controlled for acceleration. When used herein, the term "CNN" can include all types of CNNs, including region-based or region convolutional neural networks (RCNNs) and fast RCNNs (e.g., for object detection).
[0118] Accelerator 1014 (e.g., a hardware accelerator cluster) may include a Deep Learning Accelerator (DLA). The DLA may include one or more Tensor Processing Units (TPUs) that can be configured to provide an additional 10 trillion operations per second for deep learning applications and inference. The TPU may be an accelerator configured to perform image processing functions (e.g., for CNNs, RCNNs, etc.) and optimized for performing image processing functions. The DLA may be further optimized for a specific set of neural network types and floating-point operations as well as inference. The DLA is designed to provide higher performance per millimeter than a general-purpose GPU and significantly outperform CPUs. The TPU can perform several functions, including single-instance convolution functions, support for INT8, INT16, and FP16 data types for both features and weights, and post-processor functions.
[0119] DLA can execute neural networks, especially CNNs, quickly and efficiently on processed or unprocessed data for any function across a wide variety of applications, such as, but not limited to: CNNs for object recognition and detection using data from camera sensors; CNNs for distance estimation using data from camera sensors; CNNs for emergency vehicle detection and recognition using data from microphones; CNNs for face recognition and vehicle owner recognition using data from camera sensors; and / or CNNs for safety and / or safety-related events.
[0120] The DLA can perform any function of the GPU 1008, and by using inference accelerators, for example, a designer can target either the DLA or the GPU 1008 for any function. For instance, a designer can focus the CNN processing and floating-point operations on the DLA and leave other functions to the GPU 1008 and / or other accelerators 1014.
[0121] Accelerator 1014 (e.g., a hardware accelerator cluster) may include a programmable vision accelerator (PVA), which may alternatively be referred to herein as a computer vision accelerator. The PVA may be designed and configured to accelerate computer vision algorithms for advanced driver assistance systems (ADAS), autonomous driving, and / or augmented reality (AR) and / or virtual reality (VR) applications. The PVA can provide a balance between performance and flexibility. For example, each PVA may include, for example, but not limited to, any number of reduced instruction set computer (RISC) cores, direct memory access (DMA), and / or any number of vector processors.
[0122] RISC cores can interact with image sensors (such as the image sensor of any camera described herein), image label processors, and / or similar objects. Each of these RISC cores may include any amount of memory. Depending on the embodiment, the RISC core may use any of several protocols. In some examples, the RISC core may execute a real-time operating system (RTOS). RISC cores may be implemented using one or more integrated circuit devices, application-specific integrated circuits (ASICs), and / or memory devices. For example, a RISC core may include an instruction cache and / or tightly coupled RAM.
[0123] DMA enables PVA components to access system memory independently of the CPU 1006. DMA can support any number of features to provide optimizations to the PVA, including but not limited to support for multidimensional addressing and / or circular addressing. In some examples, DMA can support addressing in up to six or more dimensions, which can include block width, block height, block depth, horizontal block step, vertical block step, and / or depth step.
[0124] A vector processor can be a programmable processor designed to efficiently and flexibly execute programming for computer vision algorithms and provide tag processing capabilities. In some examples, a PVA may include a PVA core and two vector processing subsystem partitions. The PVA core may include a processor subsystem, one or more DMA engines (e.g., two DMA engines), and / or other peripherals. The vector processing subsystem may operate as the main processing engine of the PVA and may include a vector processing unit (VPU), an instruction cache, and / or vector memory (e.g., VMEM). The VPU core may include a digital tag processor, such as, for example, a Single Instruction Multiple Data (SIMD) or Very Long Instruction Word (VLIW) digital tag processor. The combination of SIMD and VLIW can enhance throughput and speed.
[0125] Each of the vector processors may include an instruction cache and may be coupled to dedicated memory. As a result, in some examples, each of the vector processors may be configured to execute independently of other vector processors. In other examples, the vector processors included in a particular PVA may be configured to employ data parallelism. For example, in some embodiments, multiple vector processors included in a single PVA may execute the same computer vision algorithm, but on different regions of an image. In other examples, vector processors included in a particular PVA may execute different computer vision algorithms simultaneously on the same image, or even different algorithms on a sequence of images or portions of an image. Among other things, any number of PVAs may be included in a cluster of hardware accelerators, and any number of vector processors may be included in each of these PVAs. Furthermore, the PVA may include additional error correction code (ECC) memory to enhance overall system security.
[0126] Accelerator 1014 (e.g., a hardware accelerator cluster) may include an on-chip computer vision network and SRAM to provide high-bandwidth, low-latency SRAM for accelerator 1014. In some examples, on-chip memory may include at least 4MB of SRAM consisting of, for example, but not limited to, eight field-configurable memory blocks, accessible by both PVA and DLA. Each pair of memory blocks may include an Advanced Peripheral Bus (APB) interface, configuration circuitry, a controller, and a multiplexer. Any type of memory may be used. PVA and DLA may access memory via a backbone that provides high-speed memory access to PVA and DLA. The backbone may include (e.g., using an APB) an on-chip computer vision network that interconnects PVA and DLA to memory.
[0127] On-chip computer vision networks can include an interface that determines whether both the PVA and DLA provide a ready and valid flag before transmitting any control flags / addresses / data. Such an interface can provide separate phases and channels for transmitting control flags / addresses / data, as well as burst communication for continuous data transmission. This type of interface can conform to ISO 26262 or IEC 615010 standards, but other standards and protocols can also be used.
[0128] In some examples, SoC 1004 may include, for example, a real-time ray tracing hardware accelerator as described in U.S. Patent Application No. 16 / 101,232, filed August 10, 2018. This real-time ray tracing hardware accelerator can be used to quickly and efficiently determine the location and extent of objects (e.g., within a world model) to generate real-time visual simulations for use in RADAR sign interpretation, sound propagation synthesis and / or analysis, SONAR system simulation, general wave propagation simulation, comparison with LiDAR data for localization and / or other functional purposes, and / or other uses. In some embodiments, one or more Tree Traversal Units (TTUs) may be used to perform one or more ray tracing-related operations.
[0129] Accelerator 1014 (e.g., a cluster of hardware accelerators) has broad applications in autonomous driving. PVAs can be programmable vision accelerators used in critical processing stages of ADAS and autonomous vehicles. PVAs are well-suited to algorithmic domains requiring predictable processing, low power, and low latency. In other words, PVAs perform well in semi-dense or dense rule computation, even on small datasets requiring predictable runtimes with low latency and low power. Therefore, in the context of platforms for autonomous vehicles, PVAs are designed to run classical computer vision algorithms because they are efficient in object detection and integer arithmetic.
[0130] For example, according to one embodiment of this technology, PVA is used to perform computer stereo vision. In some examples, semi-global matching-based algorithms may be used, but this is not intended to be limiting. Many applications for Level 3-5 autonomous driving require instantaneous motion estimation / stereo matching (e.g., from moving structures, pedestrian recognition, lane detection, etc.). PVA can perform computer stereo vision functions on input from two monocular cameras.
[0131] In some examples, PVA can be used to perform intensive optical flow, providing processed RADAR data from the raw RADAR data (e.g., using 4D Fast Fourier Transform). In other examples, PVA is used for time-of-flight depth processing, which, for example, involves processing raw time-of-flight data to provide processed time-of-flight data.
[0132] DLA can be used to run any type of network to enhance control and driving safety, including, for example, neural networks that output a confidence metric for each object detection. Such a confidence value can be interpreted as a probability or as providing a relative “weight” for each detection compared to other detections. This confidence value allows the system to make further decisions about which detections should be considered true positives rather than false positives. For example, the system can set a threshold for the confidence and only consider detections exceeding the threshold as true positives. In an Automatic Emergency Braking (AEB) system, false positives can cause the vehicle to automatically perform emergency braking, which is clearly undesirable. Therefore, only the most confident detections should be considered as triggers for AEB. DLA can run neural networks to regress the confidence values. The neural network can take at least some subset of parameters as its input, such as bounding box dimensions, ground plane estimates obtained (e.g. from another subsystem), outputs from inertial measurement unit (IMU) sensors 1066 related to the orientation and distance of vehicle 1000, 3D position estimates of objects obtained from the neural network and / or other sensors (e.g., LiDAR sensor 1064 or RADAR sensor 1060), etc.
[0133] SoC 1004 may include one or more data storage units 1016 (e.g., memory). Data storage units 1016 may be on-chip memory of SoC 1004, which may store neural networks to be executed on the GPU and / or DLA. In some examples, for redundancy and security, data storage units 1016 may be large enough to store multiple instances of the neural network. Data storage units 1012 may include L2 or L3 cache units 1012. References to data storage units 1016 may include references to memory associated with PVA, DLA, and / or other accelerators 1014 as described herein.
[0134] SoC 1004 may include one or more processors 1010 (e.g., embedded processors). Processor 1010 may include a startup and power management processor, which may be a dedicated processor and subsystem for handling startup power and management functions, as well as safety implementation. The startup and power management processor may be part of the SoC 1004 startup sequence and may provide runtime power management services. The startup power and management processor may provide clock and voltage programming, auxiliary system low-power state transitions, SoC 1004 thermal and temperature sensor management, and / or SoC 1004 power state management. Each temperature sensor may be implemented as a ring oscillator whose output frequency is proportional to the temperature, and SoC 1004 may use the ring oscillator to detect the temperature of CPU 1006, GPU 1008, and / or accelerator 1014. If it is determined that the temperature exceeds a threshold, the startup and power management processor may enter a temperature fault routine and place SoC 1004 into a lower power state and / or place vehicle 1000 into a driver-safe parking mode (e.g., safely stop vehicle 1000).
[0135] The processor 1010 may also include a set of embedded processors that can be used as an audio processing engine. The audio processing engine can be an audio subsystem that allows for full hardware support for multi-channel audio via multiple interfaces, as well as a wide and flexible range of audio I / O interfaces. In some examples, the audio processing engine is a dedicated processor core with a digital flag processor and dedicated RAM.
[0136] The processor 1010 may also include an always-on-processor engine that can provide the necessary hardware features to support low-power sensor management and wake-up use cases. This always-on-processor engine may include a processor core, tightly coupled RAM, support for peripherals (such as timers and interrupt controllers), various I / O controller peripherals, and routing logic.
[0137] The processor 1010 may also include a security cluster engine, which comprises a dedicated processor subsystem for handling security management for automotive applications. The security cluster engine may include two or more processor cores, tightly coupled RAM, support for peripheral devices (e.g., timers, interrupt controllers, etc.), and / or routing logic. In secure mode, the two or more cores may operate in lockstep mode and function as a single core with comparison logic that detects any differences between their operations.
[0138] The processor 1010 may also include a real-time camera engine, which may include a dedicated processor subsystem for handling real-time camera management.
[0139] The processor 1010 may also include a high dynamic range flag processor, which may include an image flag processor, which is a hardware engine that is part of the camera processing pipeline.
[0140] Processor 1010 may include a video image compositer, which may be (e.g., implemented on a microprocessor) a processing block, implementing video post-processing functions required by the video playback application to generate the final image for the player window. The video image compositer may perform lens distortion correction on the wide-angle camera 1070, the surround camera 1074, and / or the in-cabin monitoring camera sensor. The in-cabin monitoring camera sensor is preferably monitored by a neural network running on another instance of an advanced SoC, configured to recognize in-cabin events and respond accordingly. The in-cabin system may perform lip reading to activate mobile phone services and make calls, dictate emails, change vehicle destinations, activate or change the vehicle's infotainment system and settings, or provide voice-activated web browsing. Some functions are only available to the driver when the vehicle is operating in autonomous mode and are disabled in other situations.
[0141] Video image compositers can include enhanced temporal denoising for both spatial and temporal noise reduction. For example, in the case of motion in the video, denoising appropriately weights spatial information, reducing the weight of information provided by neighboring frames. In cases where the image or part of the image does not contain motion, the temporal denoising performed by the video image compositer can use information from previous images to reduce noise in the current image.
[0142] The video image compositer can also be configured to perform stereo correction on input stereo camera frames. When the operating system desktop is in use and the GPU 1008 does not need to continuously render new surfaces, the video image compositer can be further used for user interface components. Even when the GPU 1008 is powered on and active, performing 3D rendering, the video image compositer can be used to offload the GPU 1008 to improve performance and responsiveness.
[0143] SoC 1004 may also include a Mobile Industry Processor Interface (MIPI) camera serial interface, a high-speed interface, and / or a video input block that can be used for camera and related pixel input functions for receiving video and input from a camera. SoC 1004 may also include an input / output controller that can be software-controlled and can be used to receive I / O flags not submitted to a specific role.
[0144] SoC 1004 may also include a wide range of peripheral interfaces to enable communication with peripherals, audio codecs, power management and / or other devices. SoC 1004 can be used to process data from cameras and sensors (e.g., LIDAR sensor 1064, RADAR sensor 1060, etc., which can be connected via Gigabit Multimedia Serial Link and Ethernet), data from bus 1002 (e.g., vehicle 1000 speed, steering wheel position, etc.), and data from GNSS sensor 1058 (connected via Ethernet or CAN bus). SoC 1004 may also include a dedicated high-performance, high-capacity memory controller, which may include its own DMA engine and can be used to free up CPU 1006 from routine data management tasks.
[0145] The SoC 1004 can be an end-to-end platform with a flexible architecture spanning Automation Levels 3-5, providing a comprehensive functional safety architecture that leverages and efficiently utilizes computer vision and ADAS technologies for diversity and redundancy, along with deep learning tools to deliver a flexible and reliable driving software stack. The SoC 1004 can be faster, more reliable, and even more energy- and space-efficient than conventional systems. For example, when combined with the CPU 1006, GPU 1008, and data storage 1016, the accelerator 1014 can provide a fast and efficient platform for Level 3-5 autonomous vehicles.
[0146] Therefore, this technology offers capabilities and functionalities that cannot be achieved through conventional systems. For example, computer vision algorithms can be executed on CPUs, which can be configured using high-level programming languages such as C to execute a wide variety of processing algorithms across a diverse range of visual data. However, CPUs often cannot meet the performance requirements of many computer vision applications, such as those related to execution time and power consumption. In particular, many CPUs cannot execute complex object detection algorithms in real time, which is a requirement for automotive ADAS applications and practical Level 3-5 autonomous vehicles.
[0147] In contrast to conventional systems, the techniques described in this paper, by providing CPU complexes, GPU complexes, and hardware accelerator clusters, allow multiple neural networks to be executed simultaneously and / or sequentially, and the results combined to achieve Level 3–5 autonomous driving capabilities. For example, a CNN executed on a DLA or dGPU (e.g., GPU 1020) could include text and word recognition, allowing a supercomputer to read and understand traffic signs, including those for which neural networks have not yet been specifically trained. The DLA could also include a neural network capable of recognizing, interpreting, and providing semantic understanding of the signs, and passing that semantic understanding to a path planning module running on the CPU complex.
[0148] As another example, multiple neural networks can operate simultaneously, as required for Level 3, 4, or 5 driving. For instance, a warning sign consisting of "Caution: Flashing lights indicate icy conditions," along with a light, can be interpreted independently or jointly by several neural networks. The sign itself can be recognized as a traffic sign by a deployed first neural network (e.g., a trained neural network), and the text "Flashing lights indicate icy conditions" can be interpreted by a deployed second neural network that informs the vehicle's path planning software (preferably executing on a CPU complex) that icy conditions exist when the flashing lights are detected. The flashing lights can be identified by a deployed third neural network operating across multiple frames, informing the vehicle's path planning software of the presence (or absence) of the flashing lights. All three neural networks can operate simultaneously, for example, within a DLA and / or on a GPU 1008.
[0149] In some examples, the CNN used for facial recognition and owner identification can use data from camera sensors to identify the presence of an authorized driver and / or owner of vehicle 1000. A processing engine always on the sensors can be used to unlock the vehicle and turn on the lights when the owner approaches the driver's door, and in safe mode, to disable the vehicle when the owner leaves. In this way, SoC 1004 provides security against theft and / or carjacking.
[0150] In another example, the CNN used for emergency vehicle detection and identification can use data from microphone 1096 to detect and identify emergency vehicle siren. In contrast to conventional systems that use a general classifier to detect siren and manually extract features, SoC 1004 uses a CNN to classify environmental and urban sounds as well as visual data. In a preferred embodiment, the CNN running on the DLA is trained to identify the relative shut-off rate of emergency vehicles (e.g., by using the Doppler effect). The CNN can also be trained to identify emergency vehicles specific to the localized area in which the vehicle operates, as identified by GNSS sensor 1058. Thus, for example, when operating in the EU, the CNN will seek to detect EU siren, and when operating in the US, the CNN will seek to identify siren only in North America. Once an emergency vehicle is detected, with the assistance of ultrasonic sensor 1062, the control program can be used to execute emergency vehicle safety routines, causing the vehicle to slow down, pull over to the side of the road, stop, and / or idle until the emergency vehicle passes.
[0151] The vehicle may include a CPU 1018 (e.g., a discrete CPU or dCPU) that can be coupled to the SoC 1004 via a high-speed interconnect (e.g., PCIe). The CPU 1018 may include, for example, an X106 processor. The CPU 1018 can be used to perform any of a wide variety of functions, including, for example, arbitrating the results of potential inconsistencies between ADAS sensors and the SoC 1004, and / or monitoring the status and health of the controller 1036 and / or the infotainment SoC 1030.
[0152] Vehicle 1000 may include a GPU 1020 (e.g., a discrete GPU or dGPU) that can be coupled to SoC 1004 via a high-speed interconnect (e.g., NVIDIA's NVLINK). GPU 1020 may provide additional artificial intelligence capabilities, for example by executing redundant and / or different neural networks, and can be used to train and / or update neural networks based at least in part on inputs (e.g., sensor data) from sensors of vehicle 1000.
[0153] Vehicle 1000 may also include a network interface 1024, which may include one or more wireless antennas 1026 (e.g., one or more wireless antennas for different communication protocols, such as cellular antennas, Bluetooth antennas, etc.). Network interface 1024 can be used to enable wireless connectivity via the Internet to the cloud (e.g., with server 1078 and / or other network devices), with other vehicles, and / or with computing devices (e.g., passenger client devices). For communication with other vehicles, a direct link can be established between the two vehicles, and / or an indirect link can be established (e.g., across networks and via the Internet). A direct link can be provided using a vehicle-to-vehicle communication link. The vehicle-to-vehicle communication link can provide vehicle 1000 with information about vehicles approaching vehicle 1000 (e.g., vehicles in front, to the side, and / or behind vehicle 1000). This functionality can be part of vehicle 1000's cooperative adaptive cruise control function.
[0154] Network interface 1024 may include a SoC that provides modulation and demodulation functions and enables controller 1036 to communicate via a wireless network. Network interface 1024 may include an RF front-end for up-conversion from baseband to RF and down-conversion from RF to baseband. Frequency conversion can be performed using known processes and / or using a superheterodyne process. In some examples, the RF front-end functionality may be provided by a separate chip. The network interface may include wireless functions for communication via LTE, WCDMA, UMTS, GSM, CDMA2000, Bluetooth, Bluetooth LE, Wi-Fi, Z-Wave, ZigBee, LoRaWAN, and / or other wireless protocols.
[0155] Vehicle 1000 may also include data storage 1028, which may include off-chip (e.g., off-chip SoC 1004) storage devices. Data storage 1028 may include one or more storage elements, including RAM, SRAM, DRAM, VRAM, flash memory, hard disk, and / or other components and / or devices capable of storing at least one bit of data.
[0156] Vehicle 1000 may also include a GNSS sensor 1058. The GNSS sensor 1058 (e.g., GPS, assisted GPS sensor, differential GPS (DGPS) sensor, etc.) is used for auxiliary mapping, sensing, occupancy grid generation, and / or path planning functions. Any number of GNSS sensors 1058 can be used, including, for example, but not limited to, GPS using a USB connector with an Ethernet-to-serial (RS-232) bridge.
[0157] Vehicle 1000 may also include a RADAR sensor 1060. The RADAR sensor 1060 can be used by vehicle 1000 for remote vehicle detection even in dark and / or inclement weather conditions. The RADAR functional safety level may be ASIL B. The RADAR sensor 1060 can use CAN and / or bus 1002 (e.g., to transmit data generated by the RADAR sensor 1060) for control and access to object tracking data, and in some examples, Ethernet access for accessing raw data. A wide variety of RADAR sensor types can be used. For example, and without limitation, the RADAR sensor 1060 can be adapted for front, rear, and side RADAR use. In some examples, a pulse Doppler RADAR sensor is used.
[0158] The RADAR sensor 1060 can include different configurations, such as long-range with a narrow field of view, short-range with a wide field of view, short-range side coverage, etc. In some examples, the long-range RADAR can be used for adaptive cruise control functions. The long-range RADAR system can provide a wide field of view (e.g., within 250m) achieved through two or more independent scans. The RADAR sensor 1060 can help distinguish between stationary and moving objects and can be used by ADAS systems for emergency braking assist and forward collision warning. The long-range RADAR sensor can include a single-site multi-mode RADAR with multiple (e.g., six or more) fixed RADAR antennas and high-speed CAN and FlexRay interfaces. In an example with six antennas, the four central antennas can create a focused beam pattern designed to record the vehicle 1000's surroundings at higher rates with minimal traffic interference from adjacent lanes. The other two antennas can extend the field of view, enabling rapid detection of vehicles entering or leaving the vehicle 1000's lane.
[0159] As an example, a mid-range RADAR system can include a range of up to 1060m (front) or 80m (rear) and a field of view of up to 42 degrees (front) or 1050 degrees (rear). Short-range RADAR systems can include, but are not limited to, RADAR sensors designed to be mounted at both ends of the rear bumper. When mounted at both ends of the rear bumper, such a RADAR sensor system can create two beams that continuously monitor blind spots behind and beside the vehicle.
[0160] Short-range RADAR systems can be used in ADAS systems for blind spot detection and / or lane change assistance.
[0161] Vehicle 1000 may also include ultrasonic sensors 1062. Ultrasonic sensors 1062, which can be positioned at the front, rear, and / or sides of vehicle 1000, can be used for parking assistance and / or creating and updating occupancy grids. A wide variety of ultrasonic sensors 1062 can be used, and different ultrasonic sensors 1062 can be used for different detection ranges (e.g., 2.5m, 4m). Ultrasonic sensors 1062 can operate at functional safety level ASIL B.
[0162] Vehicle 1000 may include a LIDAR sensor 1064. The LIDAR sensor 1064 may be used for object and pedestrian detection, emergency braking, collision avoidance, and / or other functions. The LIDAR sensor 1064 may be of functional safety level ASIL B. In some examples, vehicle 1000 may include multiple LIDAR sensors 1064 (e.g., two, four, six, etc.) that can use Ethernet (e.g., to provide data to a Gigabit Ethernet switch).
[0163] In some examples, the LiDAR sensor 1064 may be able to provide a list of objects and their distances within a 360-degree field of view. Commercially available LiDAR sensors 1064 may have an advertising range of, for example, approximately 1000m, with an accuracy of 2cm-3cm, and support for 1000Mbps Ethernet connectivity. In some examples, one or more non-protruding LiDAR sensors 1064 may be used. In such examples, the LiDAR sensor 1064 may be implemented as a small device that can be embedded in the front, rear, sides, and / or corners of a vehicle 1000. In such examples, the LiDAR sensor 1064 may provide a horizontal field of view of up to 120 degrees and a vertical field of view of 35 degrees, even for low-reflectivity objects, with a range of 200m. Front-mounted LiDAR sensors 1064 may be configured for a horizontal field of view between 45 degrees and 135 degrees.
[0164] In some examples, LIDAR technologies such as 3D flash LIDAR can also be used. 3D flash LIDAR uses a flash of laser light as the emission source to illuminate the vehicle's surroundings up to approximately 200 meters. A flash LIDAR unit includes a receiver that records the laser pulse propagation time and reflected light on each pixel, which in turn corresponds to the range from the vehicle to the object. Flash LIDAR allows for the generation of highly accurate and distortion-free images of the surrounding environment using each laser flash. In some examples, four flash LIDAR sensors can be deployed, one on each side of the vehicle. Available 3D flash LIDAR systems include solid-state 3D staring array LIDAR cameras (e.g., non-scanning LIDAR devices) without moving parts other than a fan. Flash LIDAR devices can use 5 nanosecond Class I (eye-safe) laser pulses per frame and can capture reflected laser light in the form of a 3D range point cloud and co-registered intensity data. By using a flash LIDAR, and because a flash LIDAR is a solid-state device with no moving parts, the LIDAR sensor 1064 is less susceptible to motion blur, vibration, and / or shock.
[0165] The vehicle may also include an IMU sensor 1066. In some examples, the IMU sensor 1066 may be located at the center of the rear axle of the vehicle 1000. The IMU sensor 1066 may include, for example, but not limited to, an accelerometer, a magnetometer, a gyroscope, a magnetic compass, and / or other sensor types. In some examples, such as in a six-axis application, the IMU sensor 1066 may include an accelerometer and a gyroscope, while in a nine-axis application, the IMU sensor 1066 may include an accelerometer, a gyroscope, and a magnetometer.
[0166] In some embodiments, the IMU sensor 1066 can be implemented as a miniature, high-performance GPS-assisted inertial navigation system (GPS / INS) that combines a microelectromechanical system (MEMS) inertial sensor, a high-sensitivity GPS receiver, and an advanced Kalman filter algorithm to provide estimates of position, velocity, and attitude. Thus, in some examples, the IMU sensor 1066 can enable the vehicle 1000 to estimate heading by directly observing and correlating velocity changes from GPS to the IMU sensor 1066 without input from a magnetic sensor. In some examples, the IMU sensor 1066 and the GNSS sensor 1058 can be combined into a single integrated unit.
[0167] The vehicle may include a microphone 1096 placed in and / or around the vehicle 1000. Among other things, the microphone 1096 may be used for emergency vehicle detection and identification.
[0168] The vehicle may also include any number of camera types, including stereo camera 1068, wide-angle camera 1070, infrared camera 1072, surround camera 1074, long-range and / or mid-range camera 1098, and / or other camera types. These cameras can be used to capture image data around the entire perimeter of the vehicle 1000. The types of cameras used depend on the embodiment and the requirements of the vehicle 1000, and any combination of camera types can be used to provide the necessary coverage around the vehicle 1000. Furthermore, the number of cameras may vary depending on the embodiment. For example, the vehicle may include six cameras, seven cameras, ten cameras, twelve cameras, and / or another number of cameras. As an example and without limitation, these cameras may support Gigabit Multimedia Serial Link (GMSL) and / or Gigabit Ethernet. Each of the cameras is described herein with respect to... Figure 10A and Figure 10B It was described in more detail.
[0169] Vehicle 1000 may also include vibration sensor 1042. Vibration sensor 1042 can measure vibrations of vehicle components such as axles. For example, changes in vibration can indicate changes in the road surface. In another example, when two or more vibration sensors 1042 are used, differences between vibrations can be used to determine friction or slippage on the road surface (e.g., when there is a vibration difference between the powered drive shaft and the free-rotating shaft).
[0170] Vehicle 1000 may include ADAS system 1038. In some examples, ADAS system 1038 may include SoC. ADAS system 1038 may include autonomous / adaptive / automatic cruise control (ACC), cooperative adaptive cruise control (CACC), forward collision warning (FCW), automatic emergency braking (AEB), lane departure warning (LDW), lane keeping assist (LKA), blind spot warning (BSW), rear cross traffic warning (RCTW), collision warning system (CWS), lane centering (LC) and / or other features and functions.
[0171] The ACC system can use a RADAR sensor 1060, a LIDAR sensor 1064, and / or a camera. The ACC system can include longitudinal ACC and / or lateral ACC. Longitudinal ACC monitors and controls the distance to vehicles immediately in front of vehicle 1000 and automatically adjusts the vehicle speed to maintain a safe distance. Lateral ACC performs distance holding and, if necessary, advises vehicle 1000 to change lanes. Lateral ACC is associated with other ADAS applications such as LCA and CWS.
[0172] CACC uses information from other vehicles, which can be received indirectly from other vehicles via a wireless link or through a network connection (e.g., via the Internet) through network interface 1024 and / or wireless antenna 1026. Direct links can be provided by vehicle-to-vehicle (V2V) communication links, while indirect links can be infrastructure-to-vehicle (I2V) communication links. Typically, the V2V communication concept provides information about vehicles immediately ahead (e.g., vehicles immediately in front of vehicle 1000 and in the same lane), while the I2V communication concept provides information about traffic further ahead. A CACC system can include either or both I2V and V2V information sources. Given information about vehicles ahead of vehicle 1000, CACC can be more reliable, and it has the potential to improve traffic flow and reduce road congestion.
[0173] The Forward-Warping (FCW) system is designed to alert the driver to hazards, enabling the driver to take corrective action. The FCW system uses a front-facing camera and / or RADAR sensor 1060 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback such as displays, speakers, and / or vibrating components. The FCW system can provide warnings in the form of, for example, audible, visual, haptic, and / or rapid braking pulses.
[0174] An AEB (Autonomous Emergency Braking) system detects an impending forward collision with another vehicle or other object and can automatically apply the brakes if the driver does not take corrective action within a specified time or distance parameter. The AEB system can use a front-facing camera and / or RADAR sensor 1060 coupled to a dedicated processor, DSP, FPGA, and / or ASIC. When the AEB system detects a hazard, it typically first alerts the driver to take corrective action to avoid a collision, and if the driver does not take corrective action, the AEB system can automatically apply the brakes to attempt to prevent or at least mitigate the effects of the predicted collision. The AEB system may include technologies such as dynamic brake support and / or collision approach braking.
[0175] The Lane Departure Warning (LDW) system provides visual, auditory, and / or tactile warnings, such as steering wheel or seat vibrations, to alert the driver when the vehicle crosses lane markings. When the driver indicates intentional lane departure, the LDW system is deactivated by activating the turn sign. The LDW system can utilize a front-facing camera coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback such as displays, speakers, and / or vibrating components.
[0176] The LKA system is a variation of the LDW system. If vehicle 1000 begins to leave the lane, the LKA system provides steering input or braking to correct vehicle 1000.
[0177] The BSW system detects and warns the driver of vehicles in the vehicle's blind spot. The BSW system can provide visual, auditory, and / or tactile alerts to indicate that merging or changing lanes is unsafe. The system can provide additional warnings when the driver uses turn signs. The BSW system can utilize a rear-facing camera and / or RADAR sensor 1060 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback such as displays, speakers, and / or vibrating components.
[0178] RCTW systems can provide visual, auditory, and / or tactile notifications when an object is detected outside the range of a rear-view camera while the vehicle is reversing. Some RCTW systems include AEB (Autonomous Emergency Braking) to ensure the application of the vehicle's brakes to avoid a collision. RCTW systems may use one or more rear-view RADAR sensors 1060 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback such as displays, speakers, and / or vibrating components.
[0179] Conventional ADAS systems can be prone to false positives, which can be annoying and distracting for the driver, but typically not catastrophic, as ADAS systems alert the driver and allow them to determine whether a safe condition truly exists and take appropriate action. However, in an autonomous vehicle 1000, in the event of conflicting results, the vehicle 1000 itself must decide whether to heed the results from the main computer or auxiliary computer (e.g., the first controller 1036 or the second controller 1036). For example, in some embodiments, ADAS system 1038 may be a backup and / or auxiliary computer for providing perception information to a backup computer rationality module. The backup computer rationality monitor may run redundant and diverse software on hardware components to detect faults in perception and dynamic driving tasks. Outputs from ADAS system 1038 may be provided to a supervisory MCU. If the outputs from the main computer and the auxiliary computer conflict, the supervisory MCU must determine how to reconcile the conflict to ensure safe operation.
[0180] In some examples, the master computer can be configured to provide a confidence score to the supervisory MCU, indicating the master computer's confidence level in the selected result. If the confidence score exceeds a threshold, the supervisory MCU can follow the master computer's direction regardless of whether the auxiliary computer provides conflicting or inconsistent results. If the confidence score does not meet the threshold and the master and auxiliary computers indicate different results (e.g., conflict), the supervisory MCU can arbitrate between these computers to determine the appropriate result.
[0181] The supervisory MCU can be configured to run a neural network trained and configured to determine, at least in part, the conditions under which the auxiliary computer provides a false alarm, based on outputs from both the host and auxiliary computers. Thus, the neural network in the supervisory MCU can learn when the output of the auxiliary computer can be trusted and when it cannot. For example, when the auxiliary computer is a RADAR-based FCW system, the neural network in the supervisory MCU can learn when the FCW system is identifying a metallic object that is not actually dangerous, such as a drain grid or manhole cover that triggers an alarm. Similarly, when the auxiliary computer is a camera-based LDW system, the neural network in the supervisory MCU can learn to ignore the LDW when a cyclist or pedestrian is present and lane departure is actually the safest strategy. In embodiments that include a neural network running on the supervisory MCU, the supervisory MCU may include at least one of a DLA or GPU suitable for running the neural network using associated memory. In a preferred embodiment, the supervisory MCU may include a component of SoC 1004 and / or be included as a component of SoC 1004.
[0182] In other examples, ADAS system 1038 may include an auxiliary computer that performs ADAS functions using conventional computer vision rules. This allows the auxiliary computer to use classic computer vision rules (if-then), and the presence of neural networks in the supervising MCU can improve reliability, safety, and performance. For example, diverse implementations and intentional non-identity make the entire system more fault-tolerant, especially for failures caused by software (or software-hardware interface) functionality. For instance, if a software vulnerability or bug exists in the software running on the host computer and non-identical software code running on the auxiliary computer provides the same overall result, the supervising MCU can be more confident that the overall result is correct and that the vulnerability in the software or hardware on the host computer does not cause a substantial error.
[0183] In some examples, the output of the ADAS system 1038 can be fed to the perception block and / or the dynamic driving task block of the main computer. For example, if the ADAS system 1038 issues a forward collision warning because an object is immediately in front, the perception block can use this information when identifying the object. In other examples, the assistance computer can have its own neural network, which is trained and thus reduces the risk of false positives as described herein.
[0184] Vehicle 1000 may also include an infotainment SoC 1030 (e.g., an in-vehicle infotainment system (IVI)). Although illustrated and described as an SoC, the infotainment system may not be an SoC and may include two or more discrete components. The infotainment SoC 1030 may include a combination of hardware and software that can be used to provide vehicle 1000 with audio (e.g., music, personal digital assistant, navigation instructions, news, radio, etc.), video (e.g., TV, movies, streaming media, etc.), telephone (e.g., hands-free calling), network connectivity (e.g., LTE, WiFi, etc.) and / or information services (e.g., navigation system, rear parking assistance, radio data system, vehicle-related information such as fuel level, total coverage distance, brake fuel level, fuel level, door opening / closing, air filter information, etc.). For example, the infotainment SoC 1030 may include a radio, disc player, navigation system, video player, USB and Bluetooth connectivity, in-vehicle computer, in-vehicle entertainment, WiFi, steering wheel audio controls, hands-free voice controls, head-up display (HUD), HMI display 1034, telematics device, control panel (e.g., for controlling and / or interacting with various components, features, and / or systems) and / or other components. The infotainment SoC 1030 may further be used to provide information (e.g., visual and / or auditory) to users of the vehicle, such as information from the ADAS system 1038, autonomous driving information such as planned vehicle maneuvers, trajectories, surrounding environment information (e.g., intersection information, vehicle information, road information, etc.), and / or other information.
[0185] The infotainment SoC 1030 may include GPU functionality. The infotainment SoC 1030 can communicate with other devices, systems, and / or components of the vehicle 1000 via bus 1002 (e.g., CAN bus, Ethernet, etc.). In some examples, the infotainment SoC 1030 may be coupled to a supervisory MCU, allowing the GPU of the infotainment system to perform some autonomous driving functions in the event of a failure of the main controller 1036 (e.g., the primary and / or backup computer of the vehicle 1000). In such an example, the infotainment SoC 1030 may place the vehicle 1000 into a driver-safe parking mode as described herein.
[0186] Vehicle 1000 may also include instrument cluster 1032 (e.g., digital instrument panel, electronic instrument cluster, digital instrument panel, etc.). Instrument cluster 1032 may include a controller and / or a supercomputer (e.g., a discrete controller or supercomputer). Instrument cluster 1032 may include a set of instruments such as speedometer, fuel level, oil pressure, tachometer, odometer, turn indicator, shift position indicator, seatbelt warning light, parking brake warning light, engine malfunction indicator, airbag (SRS) system information, lighting controls, safety system controls, navigation information, etc. In some examples, information may be displayed and / or shared between infotainment SoC 1030 and instrument cluster 1032. In other words, instrument cluster 1032 may be included as part of infotainment SoC 1030, or vice versa.
[0187] Figure 10D For cloud-based servers and according to some embodiments of this disclosure Figure 10A This is a schematic diagram of a system for communication between example autonomous vehicles 1000. System 1076 may include server 1078, network 1090, and vehicles including vehicle 1000. Server 1078 may include multiple GPUs 1084(A)-1084(H) (collectively referred to herein as GPU 1084), PCIe switches 1082(A)-1082(H) (collectively referred to herein as PCIe switch 1082), and / or CPUs 1080(A)-1080(B) (collectively referred to herein as CPU 1080). GPU 1084, CPU 1080, and PCIe switches may be interconnected with high-speed interconnects and / or PCIe connections 1086, such as, but not limited to, NVLink interface 1088 developed by NVIDIA. In some examples, GPU 1084 is connected via NVLink and / or NVSwitch SoC, and GPU 1084 and PCIe switch 1082 are connected via PCIe interconnect. Although the diagram illustrates eight GPUs 1084, two CPUs 1080, and two PCIe switches, it is not intended to be limiting. Depending on the embodiment, each of the servers 1078 may include any number of GPUs 1084, CPUs 1080, and / or PCIe switches. For example, each of the servers 1078 may include eight, sixteen, thirty-two, and / or more GPUs 1084.
[0188] Server 1078 can receive image data from vehicles via network 1090, representing images of unexpected or altered road conditions, such as recently initiated roadworks. Server 1078 can also transmit neural network 1092, updated neural network 1092, and / or map information 1094, including information about traffic and road conditions, to vehicles via network 1090. Updates to map information 1094 may include updates to HD map 1022, such as information about construction sites, potholes, bends, floods, or other obstacles. In some examples, neural network 1092, updated neural network 1092, and / or map information 1094 may have been generated from new training and / or data received from any number of vehicles in the environment, and / or based on experience gained from training performed at a data center (e.g., using server 1078 and / or other servers).
[0189] Server 1078 can be used to train machine learning models (e.g., neural networks) based on training data. Training data can be generated by the vehicle and / or generated in a simulation (e.g., using a game engine). In some examples, the training data is labeled (e.g., where the neural network benefits from supervised learning) and / or undergoes other preprocessing, while in other examples, the training data is not labeled and / or preprocessed (e.g., where the neural network does not require supervised learning). Training can be performed according to any one or more categories of machine learning techniques, including but not limited to categories such as: supervised training, semi-supervised training, unsupervised training, self-learning, reinforcement learning, joint learning, transfer learning, feature learning (including principal component and cluster analysis), multilinear subspace learning, manifold learning, representation learning (including alternative dictionary learning), rule-based machine learning, anomaly detection, and any variations or combinations thereof. Once the machine learning model is trained, it can be used by the vehicle (e.g., transmitted to the vehicle via network 1090), and / or the machine learning model can be used by server 1078 to remotely monitor the vehicle.
[0190] In some examples, server 1078 can receive data from a vehicle and apply that data to a state-of-the-art real-time neural network for real-time intelligent inference. Server 1078 may include a deep learning supercomputer powered by GPU 1084 and / or a dedicated AI computer, such as the DGX and DGX Station machines developed by NVIDIA. However, in some examples, server 1078 may include a deep learning infrastructure in a data center that uses only CPU power.
[0191] The deep learning infrastructure of server 1078 may be capable of rapid real-time inference and can be used to assess and verify the health status of the processor, software, and / or associated hardware in vehicle 1000. For example, the deep learning infrastructure may receive periodic updates from vehicle 1000, such as image sequences and / or objects located in those image sequences that vehicle 1000 has already located (e.g., via computer vision and / or other machine learning object classification techniques). The deep learning infrastructure may run its own neural network to identify objects and compare them with objects identified by vehicle 1000. If the results do not match and the infrastructure concludes that the AI in vehicle 1000 has malfunctioned, then server 1078 may transmit a flag to vehicle 1000, instructing vehicle 1000's fail-safe computer to take control, notify passengers, and complete a safe stopping operation.
[0192] For inference, server 1078 may include GPU 1084 and one or more programmable inference accelerators (such as NVIDIA's TensorRT). The combination of a GPU-powered server and inference acceleration enables real-time response. In other examples, such as where performance is less critical, CPU, FPGA, and other processor-powered servers can be used for inference.
[0193] Example computing device
[0194] Figure 11 This is a block diagram of an example computing device 1100 suitable for implementing some embodiments of the present disclosure. The computing device 1100 may include an interconnect system 1102 directly or indirectly coupled to the following devices: memory 1104, one or more central processing units (CPUs) 1106, one or more graphics processing units (GPUs) 1108, a communication interface 1110, input / output (I / O) ports 1112, input / output components 1114, a power supply 1116, one or more presentation components 1118 (e.g., one or more displays), and one or more logic units 1120. In at least one embodiment, one or more computing devices 1100 may include one or more virtual machines (VMs), and / or any component thereof may include virtual components (e.g., virtual hardware components). For a non-limiting example, one or more GPUs 1108 may include one or more vGPUs, one or more CPUs 1106 may include one or more vCPUs, and / or one or more logic units 1120 may include one or more virtual logic units. Thus, one or more computing devices 1100 may include discrete components (e.g., a full GPU dedicated to computing device 1100), virtual components (e.g., a portion of the GPU dedicated to computing device 1100), or a combination thereof.
[0195] although Figure 11 The various blocks are shown as connected via interconnect system 1102 using lines, but this is not intended to be limiting and is merely for clarity. For example, in some embodiments, presentation component 1118 (such as a display device) may be considered I / O component 1114 (e.g., if the display is a touchscreen). As another example, CPU 1106 and / or GPU 1108 may include memory (e.g., memory 1104 may represent a storage device other than the memory of GPU 1108, CPU 1106, and / or other components). In other words, Figure 11 The computing devices described are for illustrative purposes only. No distinction is made between this category of devices or systems such as “workstations,” “servers,” “laptop computers,” “desktop computers,” “tablet computers,” “client devices,” “mobile devices,” “handheld devices,” “game consoles,” “electronic control units (ECUs),” “virtual reality systems,” and / or other device or system types, as all are considered within this category. Figure 11 Within the scope of computing devices.
[0196] Interconnect system 1102 may represent one or more links or buses, such as address buses, data buses, control buses, or combinations thereof. Interconnect system 1102 may include one or more bus or link types, such as Industry Standard Architecture (ISA) bus, Extended Industry Standard Architecture (EISA) bus, Video Electronics Standards Association (VESA) bus, Peripheral Component Interconnect (PCI) bus, Fast Peripheral Component Interconnect (PCIe) bus, and / or another type of bus or link. In some embodiments, there is a direct connection between components. As an example, CPU 1106 may be directly connected to memory 1104. Further, CPU 1106 may be directly connected to GPU 1108. In cases where there is a direct or point-to-point connection between components, interconnect system 1102 may include a PCIe link to perform the connection. In these examples, a PCI bus is not required to be included in computing device 1100.
[0197] The memory 1104 may include any computer-readable medium from a variety of computer-readable media. A computer-readable medium may be any available medium accessible by the computing device 1100. Computer-readable media may include volatile and non-volatile media, as well as removable and non-removable media. By way of example and not limitation, computer-readable media may include computer storage media and communication media.
[0198] Computer storage media may include volatile and non-volatile media and / or removable and non-removable media implemented with any method or technology for storing information such as computer-readable instructions, data structures, program modules and / or other data types. For example, memory 1104 may store computer-readable instructions (e.g., representing one or more programs and / or one or more program elements, such as an operating system). Computer storage media may include, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic tape cassettes, magnetic tape, disk storage devices or other magnetic storage devices, or any other medium that can be used to store the desired information and is accessible by computing device 1100. As used herein, computer storage media does not include the markings themselves.
[0199] Computer storage media may embody computer-readable instructions, data structures, program modules, and / or other data types in modulated data flags, such as carrier waves or other transmission mechanisms, and include any information delivery medium. The term "modulated data flag" may refer to a flag that sets or alters one or more characteristics in a manner that encodes information in the flag. By way of example and not limitation, computer storage media may include wired media (such as wired networks or direct wired connections) and wireless media (such as acoustic, RF, infrared, and other wireless media). Any combination of the above should also be included within the scope of computer-readable media.
[0200] CPU 1106 may be configured to execute at least some of computer-readable instructions to control one or more components of computing device 1100 to perform one or more of the methods and / or processes described herein. Each CPU 1106 may contain one or more cores (e.g., one, two, four, eight, twenty-eight, seventy-two, etc.) capable of handling numerous software threads simultaneously. CPU 1106 may contain any type of processor and may contain different types of processors depending on the type of computing device 1100 implemented (e.g., processors with fewer cores for mobile devices and processors with more cores for servers). For example, depending on the type of computing device 1100, the processor may be an advanced RISC machine (ARM) processor implemented using Reduced Instruction Set Computing (RISC) or an x86 processor implemented using Complex Instruction Set Computing (CISC). In addition to one or more microprocessors or secondary coprocessors (such as math coprocessors), computing device 1100 may also include one or more CPUs 1106.
[0201] In addition to or in lieu of one or more CPUs 1106, one or more GPUs 1108 may be configured to execute at least some of computer-readable instructions to control one or more components of computing device 1100 to perform one or more of the methods and / or processes described herein. One or more GPUs 1108 may be integrated GPUs (e.g., with one or more CPUs 1106) and / or one or more GPUs 1108 may be discrete GPUs. In embodiments, one or more GPUs 1108 may be coprocessors of one or more CPUs 1106. GPUs 1108 may be used by computing device 1100 to render graphics (e.g., 3D graphics) or perform general-purpose computing. For example, GPUs 1108 may be used for general-purpose computing on a GPU (GPGPU). GPUs 1108 may include hundreds or thousands of cores capable of handling hundreds or thousands of software threads simultaneously. GPUs 1108 may produce pixel data of an output image in response to rendering commands (e.g., rendering commands received from CPUs 1106 via a host interface). GPU 1108 may include graphics memory (e.g., display memory) for storing pixel data or any other suitable data (e.g., GPGPU data). Display memory may be included as part of memory 1104. GPU 1108 may include two or more GPUs operating in parallel (e.g., via links). The links may directly connect the GPUs (e.g., using NVLINK) or may connect the GPUs via a switch (e.g., using NVSwitch). When combined, each GPU 1108 may produce pixel data or GPGPU data for different portions of the output or for different outputs (e.g., a first GPU for a first image and a second GPU for a second image). Each GPU may include its own memory or may share memory with other GPUs.
[0202] In addition to or in lieu of CPU 1106 and / or GPU 1108, logic unit 1120 may be configured to execute at least some of computer-readable instructions to control one or more components of computing device 1100 to perform one or more of the methods and / or processes described herein. In embodiments, one or more CPUs 1106, one or more GPUs 1108, and / or one or more logic units 1120 may execute any combination of methods, processes, and / or portions thereof discretely or jointly. One or more logic units 1120 may be a portion of one or more CPUs 1106 and / or GPUs 1108 and / or integrated into one or more CPUs 1106 and / or GPUs 1108, and / or one or more logic units 1120 may be discrete components or otherwise external to CPUs 1106 and / or GPUs 1108. In an embodiment, one or more of the logic units 1120 may be coprocessors of one or more of the CPU 1106 and / or one or more of the GPU 1108.
[0203] Examples of logic unit 1120 include one or more processing cores and / or components thereof, such as data processing unit (DPU), tensor core (TC), tensor processing unit (TPU), pixel vision core (PVC), vision processing unit (VPU), graphics processing cluster (GPC), texture processing cluster (TPC), streaming multiprocessor (SM), tree lateral unit (TTU), artificial intelligence accelerator (AIA), deep learning accelerator (DLA), arithmetic logic unit (ALU), application-specific integrated circuit (ASIC), floating-point unit (FPU), input / output (I / O) element, peripheral component interconnect (PCI) or fast peripheral component interconnect (PCIe) element, etc.
[0204] Communication interface 1110 may include one or more receivers, transmitters, and / or transceivers enabling computing device 1100 to communicate with other computing devices via electronic communication networks (including wired and / or wireless communications). Communication interface 1110 may include components and functions for enabling communication over any of a plurality of different networks, such as wireless networks (e.g., Wi-Fi, Z-Wave, Bluetooth, Bluetooth LE, ZigBee, etc.), wired networks (e.g., communication over Ethernet or wirelessband), low-power wide area networks (e.g., LoRaWAN, SigFox, etc.), and / or the Internet. In one or more embodiments, one or more logic units 1120 and / or communication interface 1110 may include one or more data processing units (DPUs) to directly transmit (e.g., to memory) data received via a network and / or via interconnect system 1102 to one or more GPUs 1108.
[0205] I / O port 1112 enables computing device 1100 to be logically coupled to other devices including I / O component 1114, one or more presentation components 1118, and / or other components, some of which may be built into (e.g., integrated into) computing device 1100. Illustrative I / O component 1114 includes microphones, mice, keyboards, joysticks, game pads, game controllers, satellite dish antennas, scanners, printers, wireless devices, etc. I / O component 1114 can provide a natural user interface (NUI) that processes air gestures, voice, or other physiological input generated by the user. In some cases, input may be transmitted to appropriate network elements for further processing. The NUI can implement any combination of voice recognition, pen recognition, facial recognition, biometric recognition, on-screen and near-screen gesture recognition, air gestures, head and eye tracking, and touch recognition (as described in more detail below) associated with the display of computing device 1100. Computing device 1100 may include depth cameras for gesture detection and recognition, such as stereo camera systems, infrared camera systems, RGB camera systems, touchscreen technology, and combinations thereof. Additionally, computing device 1100 may include an accelerometer or gyroscope (e.g., as part of an inertial measurement unit (IMU)) that enables motion detection. In some examples, computing device 1100 may use the output of the accelerometer or gyroscope to render immersive augmented reality or virtual reality.
[0206] The power supply 1116 may include a hardwired power supply, a battery power supply, or a combination thereof. The power supply 1116 may provide power to the computing device 1100 so that the components of the computing device 1100 can operate.
[0207] The presentation component 1118 may include a display (e.g., a monitor, touchscreen, television screen, head-up display (HUD), other display types, or combinations thereof), speakers, and / or other presentation components. The presentation component 1118 may receive data from other components (e.g., GPU 1108, CPU 1106, etc.) and output the data (e.g., as images, videos, sounds, etc.).
[0208] Example Data Center
[0209] Figure 12 An example data center 1200 that may be used in at least one embodiment of this disclosure is shown. The data center 1200 may include a data center infrastructure layer 1210, a framework layer 1220, a software layer 1230, and / or an application layer 1240.
[0210] like Figure 12 As shown, the data center infrastructure layer 1210 may include a resource coordinator 1210, grouped computing resources 1214, and node computing resources (“nodes CRs”) 1216(1)-1216(N), where “N” represents any complete positive integer. In at least one embodiment, nodes CRs 1216(1)-1216(N) may include, but are not limited to, any number of central processing units (“CPUs”) or other processors (including DPUs, accelerators, field-programmable gate arrays (FPGAs), graphics processors or graphics processing units (GPUs), etc.), memory devices (e.g., dynamic read-only memory), storage devices (e.g., solid-state or disk drives), network input / output (“NW I / O”) devices, network switches, virtual machines (“VMs”), power modules and / or cooling modules, etc. In some embodiments, one or more node CRs from nodes CRs 1216(1)-1216(N) may correspond to servers having one or more of the aforementioned computing resources. In addition, in some embodiments, nodes CRs 1216(1)-12161(N) may include one or more virtual components, such as vGPU, vCPU, etc., and / or one or more nodes CRs 1216(1)-1216(N) may correspond to virtual machines (VMs).
[0211] In at least one embodiment, the grouped computing resources 1214 may include individual groups of node CRs 1216 housed within one or more racks (not shown), or multiple racks housed within a data center at different geographical locations (also not shown). Individual groups of node CRs 1216 within the grouped computing resources 1214 may include grouped computing, networking, memory, or storage resources that can be configured or allocated to support one or more workloads. In at least one embodiment, several node CRs 1216, including CPUs, GPUs, and / or other processors, may be grouped within one or more racks to provide computing resources to support one or more workloads. One or more racks may also include any number of power modules, cooling modules, and / or network switches in any combination.
[0212] Resource coordinator 1222 may be configured or otherwise control one or more nodes CRs 1216(1)-1216(N) and / or grouped computing resources 1214. In at least one embodiment, resource coordinator 1222 may include a Software Design Infrastructure (“SDI”) management entity for data center 1200. Resource coordinator 1222 may include hardware, software, or some combination thereof.
[0213] In at least one embodiment, such as Figure 12 As shown, framework layer 1220 may include a job scheduler 1232, a configuration manager 1234, a resource manager 1236, and / or a distributed file system 1238. Framework layer 1220 may include a framework for software 1232 supporting software layer 1230 and / or one or more applications 1242 of application layer 1240. Software 1232 or application 1242 may respectively contain web-based service software or applications, such as those provided by Amazon Web Services, Google Cloud, and Microsoft Azure. Framework layer 1220 may be, but is not limited to, free and open-source software web application frameworks (such as Apache Spark) that can utilize distributed file system 1238 for large-scale data processing (e.g., "big data"). TM(Hereinafter referred to as "Spark") is a type of resource manager. In at least one embodiment, job scheduler 1232 may include Spark drivers to facilitate the scheduling of workloads supported by different layers of data center 1200. Configuration manager 1234 may be able to configure different layers, such as software layer 1230 and framework layer 1220 (which includes Spark and distributed file system 1238 for supporting large-scale data processing). Resource manager 1236 may be able to manage compute resources mapped to or allocated to clusters of distributed file system 1238 and job scheduler 1232 or to support clusters of distributed file system 1238 and job scheduler 1232. In at least one embodiment, clusters or groups of compute resources may include grouped compute resources 1214 in data center infrastructure layer 1210. Resource manager 1236 may coordinate with resource coordinator 1210 to manage these mapped or allocated compute resources.
[0214] In at least one embodiment, the software 1232 included in software layer 1230 may include software used in at least a portion of the nodes CRs 1216(1)-1216(N), the grouped computing resources 1214, and / or the distributed file system 1238 of framework layer 1220. One or more types of software may include, but are not limited to, internet web search software, email virus scanning software, database software, and streaming video content software.
[0215] In at least one embodiment, the application 1242 included in the application layer 1240 may include one or more types of applications used at least in part by nodes CRs 1216(1)-1216(N), grouped computing resources 1214, and / or the distributed file system 1238 of the framework layer 1220. One or more types of applications may include, but are not limited to, any number of genomics applications, cognitive computing, and machine learning applications, including training or inference software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), and / or other machine learning applications used in combination with one or more embodiments.
[0216] In at least one embodiment, any of the configuration manager 1234, resource manager 1236, and resource coordinator 1210 can implement any number and type of self-modification actions based on any amount and type of data acquired in any technically feasible manner. Self-modification actions can free data center operators of data center 1200 from making potentially poor configuration decisions and may prevent underutilization and / or poor performance of the data center.
[0217] According to one or more embodiments described herein, data center 1200 may include tools, services, software, or other resources to train one or more machine learning models or to use one or more machine learning models to predict or infer information. For example, one or more machine learning models may be trained by using the software and / or computing resources described above with respect to data center 1200 to compute weight parameters according to a neural network architecture. In at least one embodiment, a trained or deployed machine learning model corresponding to one or more neural networks may be used to infer or predict information using the resources described above with respect to data center 1200 by using weight parameters computed through one or more training techniques (such as, but not limited to, those described herein).
[0218] In at least one embodiment, the data center 1200 may use a CPU, application-specific integrated circuit (ASIC), GPU, FPGA, and / or other hardware (or corresponding virtual computing resources) to perform training and / or inference using the aforementioned resources. Furthermore, one or more of the software and / or hardware resources described above may be configured to allow a user to train or perform services that infer information, such as image recognition, speech recognition, or other artificial intelligence services.
[0219] Example network environment
[0220] A network environment suitable for implementing embodiments of this disclosure may include one or more client devices, servers, network-attached storage (NAS), other backend devices, and / or other device types. Client devices, servers, and / or other device types (e.g., each device) may be... Figure 11 It is implemented on one or more instances of one or more computing devices 1100—for example, each device may include similar components, features, and / or functions of one or more computing devices 1100. Furthermore, in the case of implementing backend devices (e.g., servers, NAS, etc.), the backend devices may be included as part of a data center 1200, examples of which are described herein. Figure 12 To describe in more detail.
[0221] Components of a network environment can communicate with each other via a network, which can be wired, wireless, or both. A network can include one or more networks. For example, a network can include one or more wide area networks (WANs), one or more local area networks (LANs), one or more public networks (such as the Internet and / or the Public Switched Telephone Network (PSTN)), and / or one or more private networks. Where the network includes a wireless telecommunications network, components such as base stations, communication towers, or even access points (and other components) can provide wireless connectivity.
[0222] A compatible network environment may include one or more peer-to-peer network environments (in which case the server may not be included in the network environment) and one or more client-server network environments (in which case one or more servers may be included in the network environment). In a peer-to-peer network environment, the functionality described herein for the server can be implemented on any number of client devices.
[0223] In at least one embodiment, the network environment may include one or more cloud-based network environments, distributed computing environments, combinations thereof, etc. The cloud-based network environment may include a framework layer, a job scheduler, a resource manager, and a distributed file system implemented on one or more servers, which may include one or more core network servers and / or edge servers. The framework layer may include a framework supporting software at the software layer and / or application at the application layer. The software or application may respectively include network-based service software or applications. In embodiments, one or more client devices may use the network-based service software or applications (e.g., by accessing the service software and / or applications via one or more application programming interfaces (APIs)). The framework layer may be, but is not limited to, a free and open-source software network application framework that can use a distributed file system for large-scale data processing (e.g., "big data").
[0224] A cloud-based network environment can provide cloud computing and / or cloud storage for any combination of the computing and / or data storage functions (or one or more portions thereof) described herein. Any of these different functions can be distributed across multiple locations from a central or core server (e.g., distributed across one or more data centers at the state, region, country, global, etc.). The core server may assign at least a portion of the functionality to the edge server if the connection to the user (e.g., a client device) is relatively close to the edge server. A cloud-based network environment can be private (e.g., limited to a single organization), public (e.g., available to many organizations), and / or a combination thereof (e.g., a hybrid cloud environment).
[0225] One or more client devices may include the information described in this article. Figure 11At least some of the components, features, and functions of one or more example computing devices 1100 described. By way of example and not limitation, the client device may be implemented as a personal computer (PC), laptop computer, mobile device, smartphone, tablet computer, smartwatch, wearable computer, personal digital assistant (PDA), MP3 player, virtual reality headset, global positioning system (GPS) or device, video player, camera, surveillance equipment or system, vehicle, ship, spacecraft, virtual machine, drone, robot, handheld communication device, hospital equipment, gaming device or system, entertainment system, vehicle computer system, embedded system controller, remote control, electrical appliance, consumer electronics device, workstation, edge device, any combination of these depicted devices, or any other suitable device.
[0226] This disclosure can be described in the general context of machine-usable instructions or computer code, including computer-executable instructions such as program modules, which are executed by a computer or other machine such as a personal digital assistant or other handheld device. Typically, a program module, including routines, programs, objects, components, data structures, etc., refers to code that performs a specific task or implements a specific abstract data type. This disclosure can be practiced in a wide variety of system configurations, including handheld devices, consumer electronics, general-purpose computers, more specialized computing devices, etc. This disclosure can also be practiced in distributed computing environments in which tasks are performed by remote processing devices linked via a communication network.
[0227] As used herein, the phrase "and / or" relating to two or more elements should be interpreted as referring to only one element or a combination of elements. For example, "element A, element B, and / or element C" can include only element A, only element B, only element C, element A and element B, element A and element C, element B and element C, or element A, B, and C. Furthermore, "at least one of element A or element B" can include at least one of element A, at least one of element B, or at least one of element A and at least one of element B. Further, "at least one of element A and element B" can include at least one of element A, at least one of element B, or at least one of element A and at least one of element B.
[0228] This document describes in detail the subject matter of this disclosure to meet statutory requirements. However, the description itself is not intended to limit the scope of this disclosure. Rather, the discloser has envisioned that the claimed subject matter may be embodied in other ways to include steps different from or similar combinations of steps described herein in conjunction with other current or future techniques. Moreover, although the terms "step" and / or "block" may be used herein to imply different elements of the method employed, these terms should not be construed as suggesting any particular order among or between the various steps disclosed herein, unless the order of the steps is explicitly described.
[0229] Example paragraph
[0230] A. A method comprising: determining multiple candidate locations of a machine along a map of an environment comprising multiple road segments corresponding to multiple options of a path of the machine at one or more intersections; calculating multiple scores based at least on sensor data indicating at least a tracking path of the machine, the multiple scores indicating which of the multiple candidate locations corresponds to a location of the machine; determining, at least based on the aggregation of the multiple scores over a period of time, that the location of the machine corresponds to a first candidate location among the multiple candidate locations, the first candidate location being set along a first road segment among the multiple road segments; and performing one or more operations associated with the machine in the environment based at least on tracking the location of the machine along the first road segment.
[0231] B. The method as described in paragraph A, further comprising: obtaining an updated version of the map of the environment; and determining, at least based on the machine's location corresponding to the first candidate location, whether the updated version of the map is effective for tracking the path of the machine along the first road segment, wherein performing the one or more operations associated with the machine is also based at least on whether the updated version of the map is effective.
[0232] C. The method as described in paragraph A or B, wherein calculating the plurality of scores comprises: using one or more heuristic scoring functions to calculate one or more scores for each of the plurality of candidate positions; and aggregating the one or more scores for each candidate position over the time period.
[0233] D. The method as described in any of paragraphs A and C, further comprising at least one of the following: determining a relative trajectory of the machine based at least on the sensor data, the sensor data indicating at least one or more rotations and one or more translations of the machine between the one or more timestamps; or determining a global trajectory of the machine based at least on the sensor data, the sensor data indicating at least one or more position coordinates and one or more orientations of the machine at the one or more timestamps, wherein the tracking path of the machine corresponds to at least one of the relative trajectory or the global trajectory.
[0234] E. The method as described in any of paragraphs AD, wherein the calculation of the plurality of scores is based on at least one or more of the following: a plurality of curvatures associated with the plurality of road segments; a yaw rate associated with the machine; or a predicted path of the machine.
[0235] F. The method as described in any of paragraphs AE, further comprising: determining that a first score associated with the first candidate location is greater than one or more second scores associated with one or more second candidate locations by more than a threshold, wherein the determination that the location of the machine corresponds to the first candidate location is based at least on the first score being greater than the one or more second scores by more than the threshold.
[0236] G. The method as described in any of paragraphs AF, further comprising: determining a predicted path for the machine prior to the one or more intersections; determining that the path of the machine is different from the predicted path of the machine, at least based on the fact that the position of the machine corresponds to a first candidate position set along the first road segment; and initialing one or more Kalman filters for the tracking of the machine's position along the first road segment, at least based on the fact that the path is different from the predicted path.
[0237] H. A system comprising: one or more processors configured to: obtain multiple possible positions of the machine along the multiple road segments, at least based on determining that the machine has passed through an intersection associated with multiple road segments; determine the position of the machine along a first road segment of the multiple road segments, at least based on one or more tracking movements of the machine relative to the multiple possible positions; and perform one or more operations associated with the machine, at least based on the position of the machine along the first road segment.
[0238] I. The system as described in paragraph H, wherein the one or more processors are further configured to: at one or more first moments, calculate one or more first scores associated with the plurality of possible locations; and at one or more second moments, calculate one or more second scores associated with the plurality of possible locations; wherein the determination of the machine's location is further based at least on the aggregation of the one or more first scores and the one or more second scores.
[0239] J. The system as described in any of paragraphs HI, wherein the calculation of at least one of the one or more first scores or the one or more second scores is based at least on: the tracking path of the machine relative to one or more coordinate systems; the attitude of the machine; the yaw rate of the machine; multiple curvatures associated with the plurality of road segments; multiple lanes associated with the plurality of road segments; or surface markings associated with the plurality of road segments.
[0240] K. The system as described in any of paragraphs HJ, wherein the one or more processors are further configured to: initialize a plurality of trackers for the plurality of road segments to track the plurality of possible locations of the machine; and terminate one or more of the plurality of trackers used to track one or more possible locations of the plurality of possible locations along one or more of the plurality of road segments, at least based on the determination that the location of the machine is along the first road segment.
[0241] L. The system as described in any of paragraphs HK, wherein the one or more processors are further configured to: initialize a Kalman filter to track one or more state variables indicating the location of the machine along the first road segment, the one or more state variables including at least one of the following: an identifier corresponding to the first road segment; the location of the machine relative to at least one point along the first road segment; or a confidence score corresponding to the location of the machine.
[0242] M. The system as described in any of paragraphs HL, wherein the one or more processors are further configured to: determine that a first score associated with a first possible location among the plurality of possible locations is greater than one or more second scores associated with one or more second possible locations among the plurality of possible locations; and determine that the location of the machine corresponds to the first possible location, at least based on the first score being greater than the one or more second scores.
[0243] N. The system as described in any of paragraphs HM, wherein the one or more processors are further configured to: determine whether one or more differences between the first score and the one or more second scores reach or exceed a threshold, wherein the determination that the position of the machine corresponds to the first possible position is also based at least on the one or more differences reaching or exceeding the threshold.
[0244] O. The system as described in any of paragraphs HN, wherein the one or more processors are further configured to: determine, at least based on map data representing the environment, that the machine has traversed the intersection; obtain an updated version of the map data; and determine, at least based on the location of the machine along the first road segment, whether the updated version of the map data is valid.
[0245] P. The system as described in any of paragraphs HO, wherein the one or more processors are further configured to: initialize one or more Kalman filters to track one or more positions of the machine relative to one or more road segments using one or more state variables; calculate one or more updated state variables of the one or more Kalman filters, the one or more updated state variables indicating one or more updated positions of the machine, based at least on at least one of sensor data or perception data; and use the one or more updated positions of the machine to determine that the machine has traversed the intersections associated with the plurality of road segments.
[0246] Q. A system as described in any of paragraphs HP, wherein the system comprises at least one of: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing one or more analog operations; a system for performing one or more digital twin operations; a system for performing an optical transmission model; a system for performing collaborative content creation for 3D assets; a system for performing one or more deep learning operations; a system implemented using an edge device; a system implemented using a robot; a system for performing one or more generative AI operations; a system for performing operations using a large language model; a system for performing operations using one or more visual language models (VLMs); a system for performing operations using one or more multimodal language models; a system for implementing one or more machine learning models as inference microservices using one or more operating system (OS) level virtualization packages; a system for performing one or more conversational AI operations; a system for generating synthetic data; a system for presenting at least one of virtual reality content, augmented reality content, or mixed reality content; a system comprising one or more virtual machines (VMs); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources.
[0247] R. One or more processors, the one or more processors comprising: a processing circuitry system for evaluating one or more path localization algorithms within a simulated environment rendered using one or more optical transport simulation algorithms, the one or more path localization algorithms determining the virtual machine's position in the simulated environment by initializing one or more trackers to track one or more possible positions of the virtual machine along the one or more road segments after the virtual machine has traversed one or more road segments from one or more intersections where they branch off.
[0248] S. One or more processors as described in any of paragraphs R, wherein the simulation is generated at least in part using a 3D content collaboration platform for 3D assets.
[0249] T. One or more processors as described in any of paragraphs RS, wherein the 3D content collaboration platform for 3D assets uses generic scene descriptor (USD) data to manage one or more attributes of the simulation environment associated with the simulation.
Claims
1. A method, the method comprising: Determine multiple candidate locations for the machine in a map of the environment, where multiple road segments correspond to multiple options for the machine's path at one or more intersections; Multiple scores are calculated based at least on sensor data that indicates at least the tracking path of the machine, the multiple scores indicating which of the multiple candidate locations corresponds to the location of the machine; Based at least on the aggregation of the multiple scores over a period of time, it is determined that the location of the machine corresponds to a first candidate location among the multiple candidate locations, the first candidate location being set along a first road segment among the multiple road segments; as well as At least based on tracking the machine's location along the first road segment, one or more operations associated with the machine are performed in the environment.
2. The method according to claim 1, further comprising: Obtain an updated version of the map of the environment; as well as At least based on the machine's location corresponding to the first candidate location, it is determined whether the updated version of the map is effective for tracking the machine's path along the first road segment. The execution of the one or more operations associated with the machine is also based at least on whether the updated version of the map is valid.
3. The method according to claim 1, wherein, The calculation of the plurality of scores includes: Use one or more heuristic scoring functions to compute one or more scores for each of the plurality of candidate locations; and Aggregate the one or more scores for each candidate position within the time period.
4. The method according to claim 1, further comprising: At least one of the following: The relative trajectory of the machine is determined based at least on the sensor data, which indicates at least one or more rotations and one or more translations of the machine between the one or more timestamps. or The global trajectory of the machine is determined based at least on the sensor data, which indicates at least one or more location coordinates and one or more orientations of the machine at one or more timestamps. The tracking path of the machine corresponds to at least one of the relative trajectory or the global trajectory.
5. The method according to claim 1, wherein, The calculation of the plurality of scores is also based on at least one or more of the following: Multiple curvatures associated with the multiple road segments; Yaw rate associated with the machine; or The predicted path of the machine.
6. The method according to claim 1, further comprising: The first score associated with the first candidate position is determined to be greater than one or more second scores associated with one or more second candidate positions by a threshold value. The determination that the machine's position corresponds to the first candidate position is based at least on the first score being greater than the threshold by one or more second scores.
7. The method according to claim 1, further comprising: Determine the machine's predicted path before the one or more intersections; Based at least on the fact that the machine's location corresponds to the first candidate location set along the first road segment, it is determined that the machine's path is different from the machine's predicted path; as well as At least based on the fact that the path is different from the predicted path, one or more Kalman filters are initialized for the machine to track the position along the first road segment.
8. A system comprising: One or more processors are used for: Based at least on determining that the machine has passed through intersections associated with multiple road segments, multiple possible positions of the machine along the multiple road segments are obtained; The position of the machine is determined along a first road segment of the plurality of road segments based at least on one or more tracking movements of the machine relative to the plurality of possible locations; At least based on the machine's position along the first road segment, one or more operations associated with the machine are performed.
9. The system of claim 8, wherein the one or more processors are further configured to: At one or more first moments, calculate one or more first scores associated with the plurality of possible locations; and At one or more second time points, calculate one or more second scores associated with the plurality of possible locations; in, The determination of the machine's location is also based at least on the aggregation of one or more first scores and one or more second scores.
10. The system of claim 9, wherein the calculation of at least one of the one or more first scores or the one or more second scores is further based on at least: The tracking path of the machine relative to one or more coordinate systems; The posture of the machine; The yaw rate of the machine; Multiple curvatures associated with the multiple road segments; Multiple lanes associated with the aforementioned multiple road segments; or Surface markings associated with the plurality of road segments.
11. The system of claim 8, wherein the one or more processors are further configured to: For the multiple road segments, initialize multiple trackers to track the multiple possible locations of the machine; and At least based on the determination of the machine's location along the first road segment, terminate one or more of the multiple trackers used to track one or more possible locations along one or more of the multiple road segments and one or more second road segments.
12. The system of claim 8, wherein the one or more processors are further configured to: Initialize a Kalman filter to track one or more state variables indicating the machine's position along the first road segment, said one or more state variables including at least one of the following: The identifier corresponding to the first road segment; The machine's position relative to at least one point along the first road segment; or The confidence score corresponding to the location of the machine.
13. The system of claim 8, wherein the one or more processors are further configured to: Determine that a first score associated with a first possible position among the plurality of possible positions is greater than one or more second scores associated with one or more second possible positions among the plurality of possible positions; and The machine's position is determined to correspond to the first possible position based at least on the first score being greater than one or more of the second scores.
14. The system of claim 13, wherein the one or more processors are further configured to: Determine whether one or more differences between the first score and the one or more second scores reach or exceed a threshold. in, The determination that the machine's position corresponds to the first possible position is also based at least on one or more differences reaching or exceeding the threshold.
15. The system of claim 8, wherein the one or more processors are further configured to: Based at least on map data representing the environment, it is determined that the machine traversed the intersection; Obtain an updated version of the map data; and The validity of the updated version of the map data is determined at least based on the machine's location along the first road segment.
16. The system of claim 8, wherein the one or more processors are further configured to: Initialize one or more Kalman filters to track one or more positions of the machine relative to one or more road segments using one or more state variables; Based on at least one of sensor data or perceived data, calculate one or more updated state variables of the one or more Kalman filters, the one or more updated state variables indicating one or more updated positions of the machine; and Using one or more updated locations of the machine, it is determined that the machine has traversed the intersections associated with the plurality of road segments.
17. The system of claim 8, wherein the system comprises at least one of the following: Control systems for autonomous or semi-autonomous machines; Sensing systems for autonomous or semi-autonomous machines; A system for performing one or more simulation operations; A system for performing one or more digital twin operations; A system for implementing optical transmission models; A system for performing collaborative content creation for 3D assets; A system for performing one or more deep learning operations; Systems implemented using edge devices; Systems implemented using robots; A system for performing one or more generative AI operations; A system for performing operations using a large language model; A system for performing operations using one or more visual language models (VLMs); A system for performing operations using one or more multimodal language models; A system that uses one or more operating system OS-level virtualization packages to implement one or more machine learning models as inference microservices; A system for performing one or more conversational AI operations; A system for generating synthetic data; A system for presenting at least one of virtual reality content, augmented reality content, or mixed reality content; A system containing one or more virtual machines (VMs); A system that is at least partially implemented in a data center; or A system that utilizes cloud computing resources at least in part.
18. One or more processors, said one or more processors comprising: A processing circuit system is configured to evaluate one or more path localization algorithms within a simulated environment rendered using one or more optical transmission simulation algorithms, the one or more path localization algorithms determining the virtual machine's position in the simulated environment by initializing one or more trackers to track one or more possible locations of the virtual machine along the one or more road segments after the virtual machine has traversed one or more road segments from one or more intersections where they branch off.
19. One or more processors according to claim 18, wherein, The simulation was generated at least in part using a 3D content collaboration platform for 3D assets.
20. One or more processors according to claim 19, wherein, The 3D content collaboration platform for 3D assets uses Universal Scene Descriptor (USD) data to manage one or more attributes of the simulation environment associated with the simulation.
Citation Information
Patent Citations
Method for programmable timeouts of tree traversal mechanisms in hardware
US10885698B2
Path prediction for autonomous or semi-autonomous systems and applications
US20260054737A1