Ground surface estimation using stereoscopic imaging for autonomous and semi-autonomous systems and applications
By using nonlinear optimization and stereo imaging techniques to correct biases in LiDAR data and generate surface parallax fields, the problem of insufficient accuracy in road profile estimation was solved, enabling safe and comfortable driving of autonomous vehicles in complex environments.
Patent Information
- Application Number
- CN202510785905.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-12-19
- Filing Date
- 2025-06-12
- Publication Date
- 2025-12-12
AI Technical Summary
Existing technologies have limited accuracy in estimating road surface contours, especially at long distances and in adverse weather conditions, which leads to poor handling of autonomous vehicles on uneven roads and affects vehicle stability.
By employing nonlinear optimization and stereo imaging techniques, and through bias correction of LiDAR data and generation of surface disparity fields, the accuracy of road profile estimation is improved. This includes self-motion compensation, stereo image processing, and constrained nonlinear hierarchical optimization to generate a surface disparity field for refined estimation.
It improves the accuracy and robustness of road surface contour estimation, enhances the safety and comfort of autonomous vehicles, effectively avoids obstacles and adjusts the suspension system to compensate for road surface irregularities, providing a smooth driving experience.
Smart Images

Figure CN121120913A_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] This application is a continuation to U.S. Application No. 18 / 987,171, filed December 19, 2024, which claims the benefit of U.S. Provisional Application No. 63 / 659,173, filed June 12, 2024. The entire contents of each of the foregoing applications are incorporated herein by reference. Background Technology
[0003] Designing a system capable of autonomous, safe, and comfortable driving without supervision is extremely difficult. An autonomous vehicle should at least be able to behave like a focused driver, utilizing a perception and motion system with an incredible ability to identify and react to dynamic and static hazards in complex environments, thereby navigating along its path in the surrounding three-dimensional (3D) environment. The ability to estimate the road surface is often crucial for autonomous driving perception systems. For example, the estimated ground surface can be used for tasks such as identifying navigable space (e.g., road surfaces), detecting obstacles on the road, adjusting suspension or other vehicle components for smoother driving, and estimating obstacle heights, to name just a few.
[0004] Therefore, various autonomous vehicles and advanced driver assistance functions rely on ground or road surface estimation. However, the execution of these functions is limited by the accuracy of the estimated road surface. For example, subtle changes in road height can be used to optimize suspension settings, but when the estimated road height is not accurate enough, small but far-reaching changes in the road profile may go undetected, resulting in poor handling on uneven or rough surfaces and potentially compromising vehicle stability, especially at higher speeds.
[0005] Conventional techniques for estimating road surfaces have limited accuracy. Detecting road surface contours at distances from vehicles is particularly challenging due to the inherent limitations of current sensor technology. LiDAR (Light Detection and Ranging) uses laser pulses to create a detailed 3D representation of the environment, typically providing good accuracy at close range, but producing sparse data points at greater distances. The sparsity of LiDAR data can be further limited by weather conditions (e.g., resulting in few or no measurements on wet surfaces), leading to incomplete and less reliable representations. Radar (RADAR), on the other hand, uses radio waves to detect objects and measure distances, operating effectively at long ranges and in various weather conditions, but its lower resolution makes it difficult to estimate surfaces with sufficient accuracy. Conventional camera-only solutions can provide high-resolution images of the surrounding environment, but these solutions struggle to estimate distances to surfaces with adequate accuracy, especially for longer measurement ranges.
[0006] Therefore, surface estimation techniques need to be improved. Summary of the Invention
[0007] Embodiments of this disclosure relate to ground surface estimation using local surface fitting, bias correction, stereo imaging, and / or ground parallax for autonomous and semi-autonomous systems and applications.
[0008] In some embodiments, nonlinear optimization can be used to estimate three-dimensional (3D) surface structures (e.g., road surface profiles) by fitting height values to (e.g., accumulated, bias-corrected) LiDAR detection results (e.g., sampled in local regions along one or more predicted trajectories). For example, LiDAR data can be generated using one or more LiDAR sensors of a self-driving machine as it navigates through its environment. A 3D representation of the ground or road surface can be estimated based on the LiDAR data, and the ground or road surface profile along one or more predicted trajectories (e.g., wheel tracks) can be estimated based on the LiDAR data and the estimated ground or road surface. For example, in some embodiments, LiDAR data (e.g., detected 3D point clouds) can be self-motor compensated, measurement bias corrected, accumulated, and sampled along one or more predicted trajectories, and nonlinear optimization can be used to fit the height of each trajectory point to the height of the corresponding sampled point. In this way, the obtained road surface profile (e.g., modeled along the wheel track) can be provided to the adaptive suspension control system to adjust the damping characteristics of the suspension system, thereby counteracting depressions (e.g., potholes, ditches, etc.) or bumps (e.g., speed bumps, metal plates, etc.) represented in the road surface profile.
[0009] In some embodiments, LiDAR measurement biases, such as range-dependent height offsets and / or reflectivity-dependent height offsets, can be estimated during offline operations, and the measured LiDAR height can be compensated for by eliminating these biases. To estimate range-dependent height bias, observed height values representing fixed locations (e.g., patches on the ground) in the accumulated LiDAR point cloud can be binned by measurement distance, and the height bias or offset for each range bin can be calculated based on the difference between the combined height values for that bin and a specified ground truth height. To estimate reflectivity-dependent height bias, one or more locations in the accumulated LiDAR point cloud representing sufficiently large reflectivity variations (e.g., patches on the ground) can be identified (e.g., a local neighborhood with high reflectivity paired with a low-reflectivity area of asphalt pavement, such as a local neighborhood of road markings), and observed height values measured from approximately the same distance can be binned based on the measured reflectivity. In this way, observed height values from each reflectance band can be combined (e.g., median height), and the height deviation or offset for each reflectance band can be calculated based on the difference between the combined height value of that band and a specified true height. The estimated deviation can be stored in any suitable manner (e.g., in one or more lookup tables indexed by distance and / or reflectance), and can be compensated for for LiDAR points measured in-line operation by looking up and subtracting the distance-related height deviation corresponding to the measured distance, and / or by looking up and subtracting the reflectance-related height deviation corresponding to the measured reflectance.
[0010] In some embodiments, 3D surface structures can be modeled as disparity fields, and a constrained nonlinear hierarchical optimization can be used to generate a surface disparity field representing surfaces (e.g., ground) in the environment. This constrained nonlinear hierarchical optimization processes stereo image data and iteratively refines the estimated surface disparity values based on weights that guide optimization toward expected surface values (e.g., ground, road). More specifically, as the ego machine traverses the environment, it can generate stereo image pairs using one or more stereo cameras. Each stereo image pair can be used to generate a stereo disparity field including stereo disparity values (also known as stereo parallax). The surface disparity field representing surfaces (e.g., ground) in the environment can also be generated by iteratively refining the estimated disparity values using a constrained nonlinear optimization process customized with one or more weights to directly solve for the disparity field of the surface (e.g., ground) (ground disparity field). The resulting surface (e.g., ground) parallax field can be used for a variety of downstream tasks, such as obstacle detection, airworthiness space segmentation, self-motion refinement, and / or generating estimated surface profiles.
[0011] Therefore, the techniques described in this paper can be used to estimate the 3D structure of surfaces such as ground or road surfaces, detect obstacles in the environment and / or detect airworthy space, and can provide a representation of the detection results to the autonomous vehicle driving stack to enable safe and comfortable planning and control of autonomous vehicles. Attached Figure Description
[0012] The following will describe in detail, with reference to the accompanying drawings, the system and method for ground surface estimation using local surface fitting, bias correction, stereo imaging and / or ground parallax for autonomous and semi-autonomous systems and applications, wherein:
[0013] Figure 1 This is a data flow diagram illustrating an example surface estimation pipeline according to some embodiments of the present disclosure;
[0014] Figure 2 This illustrates distance-related height deviations in LiDAR sensor data according to some embodiments of the present disclosure;
[0015] Figure 3 This illustrates reflectivity-related height deviations in LiDAR sensor data according to some embodiments of this disclosure;
[0016] Figure 4 An example process for estimating height deviations in LiDAR sensor data according to some embodiments of this disclosure is shown;
[0017] Figure 5Example techniques for sampling LiDAR detection results along a predicted trajectory, according to some embodiments of the present disclosure, are shown;
[0018] Figure 6 An example road surface profile along a two-wheel track is shown according to some embodiments of the present disclosure;
[0019] Figure 7 This is a data flow diagram illustrating an example surface disparity estimation pipeline according to some embodiments of the present disclosure;
[0020] Figure 8 Example stereo parallax pyramids and downsampling strategies according to some embodiments of this disclosure are shown;
[0021] Figure 9 An example asymmetric measurement deviation cost function according to some embodiments of this disclosure is shown;
[0022] Figure 10 Some example techniques for guiding ground disparity estimation according to some embodiments of this disclosure are shown;
[0023] Figure 11 This is a flowchart illustrating a method for surface estimation based at least on fitting height values in one or more local neighborhoods according to some embodiments of the present disclosure;
[0024] Figure 12 This is a flowchart illustrating a method for generating bias-corrected LiDAR detection results according to some embodiments of the present disclosure;
[0025] Figure 13 This is a flowchart illustrating a method for generating a surface disparity field representing estimated disparity values of surfaces in an environment, according to some embodiments of the present disclosure;
[0026] Figure 14 This is a flowchart illustrating a method for controlling one or more operations of a self-machine based at least on a surface parallax field, according to some embodiments of the present disclosure;
[0027] Figure 15A These are illustrations of example autonomous vehicles according to some embodiments of the present disclosure;
[0028] Figure 15B According to some embodiments of this disclosure Figure 15A Examples of camera positions and fields of view for autonomous vehicles;
[0029] Figure 15C According to some embodiments of this disclosure Figure 15A A block diagram of an example system architecture for an example autonomous vehicle;
[0030] Figure 15D This is according to some embodiments of the present disclosure for use in one or more cloud-based servers and Figure 15A A system diagram illustrating communication between autonomous vehicles;
[0031] Figure 16 This is a block diagram of an example computing device suitable for implementing some embodiments of the present disclosure; and
[0032] Figure 17 This is a block diagram of an example data center suitable for implementing some embodiments of the present disclosure. Detailed Implementation
[0033] Systems and methods for ground surface estimation using local surface fitting, bias correction, stereo imaging, and / or ground parallax are disclosed for autonomous and semi-autonomous systems and applications. In some embodiments, nonlinear optimization can be used to estimate three-dimensional (3D) surface structures (e.g., road surface profiles) by fitting height values to (e.g., accumulated, bias-corrected) LiDAR detection results (e.g., sampled in local regions along one or more predicted trajectories). In some embodiments, the 3D surface structures can be modeled as parallax fields, and constrained nonlinear hierarchical optimization can be used to generate surface parallax fields representing surfaces (e.g., ground) in the environment by processing stereo image data and iteratively refining the estimated surface parallax values based on weights guiding the optimization to expected surface values (e.g., ground, road). This technique can be used by autonomous vehicles, semi-autonomous vehicles, robots, and / or other object or machine types to estimate the 3D surface structure of other components of airworthy spaces or environments, and / or to detect and avoid potential obstacles based on the estimated 3D surface structure.
[0034] Although this disclosure may relate to example autonomous or semi-autonomous vehicles or machines 1500 (in this document, alternatively referred to as "vehicle 1500" or "self-machine 1500"), examples of which relate to Figures 15A to 15DThe description herein is provided, but is not intended to be limiting. For example, the systems and methods described herein may be used, but are not limited to, non-autonomous vehicles or machines, semi-autonomous vehicles or machines (e.g., in one or more advanced driver assistance systems (ADAS)), autonomous vehicles or machines, manned and unmanned robots or robotic platforms, warehouse vehicles, off-road vehicles, vehicles coupled to one or more trailers, aircraft, ships, shuttles, emergency response vehicles, motorcycles, electric or motorized bicycles, aircraft, engineering vehicles, trains, underwater vehicles, drones, and other remotely controlled vehicles and / or other vehicle types. Furthermore, although this disclosure may describe road surface estimation for autonomous driving, this is not intended to be limiting, and the systems and methods described herein may be used in augmented reality, virtual reality, mixed reality, robotics, security and supervision, autonomous or semi-autonomous machine applications, and / or any other technical field where surface estimation can be used.
[0035] In some embodiments, as the ego machine traverses an environment, LiDAR data can be generated using one or more LiDAR sensors of the ego machine. A 3D representation of the ground or road surface can be estimated based on the LiDAR data, and a ground or road surface profile along one or more predicted trajectories (e.g., wheel ruts) can be estimated based on the LiDAR data and the estimated ground or road surface. For example, in some embodiments, the LiDAR data (e.g., detected 3D point clouds) can be self-motor compensated, accumulated, and sampled along one or more predicted trajectories, and nonlinear optimization can be used to fit the height of each trajectory point to the height of the corresponding sampled point. Thus, the resulting road surface profile (e.g., modeled along wheel ruts) can be provided to an adaptive suspension control system to adjust the damping characteristics of the suspension system to counteract depressions (e.g., potholes) or bumps (e.g., speed bumps) represented in the road surface profile.
[0036] One possible approach to improve the accuracy and robustness of the estimated road surface profile is to create a sufficiently dense point cloud for the optimization step, such that each point in the road surface profile has a minimum number of contributing LiDAR points. For a single LiDAR frame (e.g., a point cloud generated from a single LiDAR spin), the achievable density on the road surface is often too low for robust estimation. Therefore, in some embodiments, a self-motion compensation process can be used to accumulate multiple frames of LiDAR data, identifying transformations that map point clouds from consecutive LiDAR frames to a common coordinate system, and the accuracy of this transformation can be improved using multi-spin (or point cloud) registration (e.g., iterative nearest point (ICP), point-to-surface matching). In some embodiments, to improve the accuracy of the registration process (and thus the accuracy of the estimated surface), the LiDAR points can be segmented into points belonging to a static reference surface (e.g., ground, vegetation, buildings) based on their height above the estimated ground surface (e.g., filtering out points within height ranges such as 10 cm to 3 m above the estimated ground surface). The resulting segmented point cloud is then registered with the segmented LiDAR points. This segmentation essentially eliminates a large number of potential outliers that do not have a direct correspondence between consecutive LiDAR frames because they are likely to be moving. Therefore, registering these segmented point clouds should improve the accuracy and speed of the point cloud registration process (and the accuracy of the estimated road surface profile). In some embodiments, the multi-spin registration process may primarily address inaccuracies in pitch angle estimation, as this is typically the most dynamic self-attitude parameter. Therefore, in some embodiments, the registration process can be simplified to estimating or refining the pitch angle difference between (e.g., the segmented point cloud) and (e.g., a known, estimated) reference surface (e.g., the ground).
[0037] Registration of multiple (e.g., segmented) LiDAR spins offers several benefits. For example, it effectively enables fine-tuning of relative transformations between LiDAR frames, compensating for potentially inaccurate ego-motion estimations (e.g., due to high-dynamic events such as hitting speed bumps or potholes, sudden braking, or acceleration). Registering multiple LiDAR frames to a common coordinate system (typically the ego-motion pose of the previous frame) also increases sampling density on the ground surface. Therefore, refining ego-motion estimation by registering segmented LiDAR spins can improve the accuracy and density of LiDAR data, which in turn enhances the accuracy of downstream tasks.
[0038] Depending on downstream applications, using raw LiDAR measurements (or even motion-compensated LiDAR measurements) may not achieve the desired accuracy in estimating road surface profiles because measurements using LiDAR sensors include systematic measurement biases that affect measured distance and height values. One type of LiDAR measurement bias is distance-related height shift caused by the divergence or spread of the LiDAR beam, manifesting as a positive deviation in detected height or z-coordinate (the point appears higher than the true value) and a negative deviation in distance or x-coordinate (the point appears closer to the vehicle itself). Another prominent LiDAR measurement bias is reflectivity-related offset. This bias is caused by the stronger echo of bright / reflective objects compared to darker objects and has a similar effect to distance-related height shift (e.g., estimated distance is shorter than its true value, estimated height is higher than its true value), but with a different amplitude. Assuming the road surface is approximately flat, this reflectivity-related bias primarily manifests as a height shift. Measurement accuracy can be improved by estimating the corresponding bias amplitude and by eliminating bias-compensated LiDAR points in the measurements.
[0039] In the example bias estimation process, one or more data acquisition vehicles can be used to generate and accumulate individual LiDAR measurements. To estimate the distance-dependent height bias, observed height values representing fixed locations (e.g., blocks on the ground) in the accumulated LiDAR point cloud can be binned according to measurement distance. Due to the nature of distance-dependent height bias, points observed from closer distances and steeper angles of incidence should be less affected by distance-dependent height bias than points observed from farther distances and flatter angles of incidence. Therefore, observed height values measured from closer distances should be more accurate than those measured from farther distances. Thus, observed height values from each bin distance segment (e.g., in one-meter bins) can be combined or aggregated (e.g., by taking the median height). In some embodiments, the height bias or offset of each distance bin can be calculated based on the difference between the combined height value of the bin and a specified true height. For example, the combined height value corresponding to the nearest measurement distance segment (e.g., within a specified measurement distance, such as one meter) can be used as the true value, or the true height can be calculated as a weighted median to assign higher weight to closer measurements.
[0040] To estimate reflectance-related height bias, one or more locations (e.g., multiple patches on the ground) representing sufficiently large reflectance variations in the accumulated LiDAR point cloud can be identified (e.g., a local neighborhood of high reflectance paired with a low-reflectance area of asphalt surface, such as a local neighborhood of road markings). These bins can then be formed based on the measured reflectance from observed height values measured at approximately the same distance. The observed height values within each bin's reflectance segment can then be combined (e.g., taking the median height), and the height bias or offset for each reflectance bin can be calculated based on the difference between the combined height values of that bin and a specified true height. The true height can be obtained from points in the accumulated LiDAR point cloud measured within a specified measurement distance (e.g., the median height of a point), or it can be calculated by compensating for the observed height values using a corresponding distance-related height bias. In some embodiments, instead of calculating the distance-related height bias and the reflectivity-related height bias in separate processes, both can be estimated by sampling the accumulated LiDAR point cloud, binning the observed height values into distance and reflectivity bins, and utilizing joint two-dimensional (2D) nonlinear optimization to calculate the two biases using the observed height values measured within a specified measurement distance as the true height.
[0041] Therefore, distance-related height deviations can be calculated (e.g., offline) for individual distance buckets of arbitrarily specified size, and reflectance-related height deviations can be calculated (e.g., offline) for individual reflectance buckets of arbitrarily specified size, and the estimated deviations can be stored in any suitable manner (e.g., stored in one or more lookup tables indexed by distance and / or reflectance). Thus, LiDAR points measured in-process can be compensated for by looking up and subtracting the distance-related height deviation corresponding to the measured distance, and / or by looking up and subtracting the reflectance-related height deviation corresponding to the measured reflectance. Correcting measurement deviations of the measured LiDAR points improves their accuracy, as well as the accuracy of the resulting estimated (e.g., ground or road) surface.
[0042] Returning to the example online procedure, by performing measurement bias correction, segmentation (e.g., segmentation only into ground points), and joint registration on the input LiDAR points, the predicted trajectory can be estimated by extrapolating the state of a vehicle steering model (e.g., an Ackermann steering model). An arbitrary number of 3D points can be sampled along each wheel trajectory (e.g., using certain placeholder heights, such as z=0), and the surface profile along each wheel trajectory can be estimated by sampling those LiDAR points located at 3D positions within a specified (e.g., orientation-dependent) 3D radius from the wheel rut point. With sufficient sampling density, the number of sampled LiDAR points per wheel rut point should be variable, and optimized height values can be fitted to the sampled LiDAR points for each wheel rut point using nonlinear optimization methods (e.g., classical second-order least squares solvers such as Levenberg-Marquardt) or (e.g., for real-time systems) first-order gradient descent (e.g., with a specified number of iterations)). Robust cost functions (e.g., Cauchy, Huber, etc.) can be used to mitigate the impact of anomalous observations (e.g., points whose height values differ significantly from the current estimates). In some embodiments, optimization can be simplified to 1D optimization, where the predicted trajectory and associated LiDAR points can be unfolded and represented in a height-distance (z / d) space (e.g., where z represents the estimated height of the contour points and d represents the distance to the ego vehicle on the unfolded trajectory). In some embodiments, one or more parameters of the cost function can be set based on a ground-value noise level estimate (e.g., the standard deviation of the z-residuals in a given ground patch can be averaged over any number of ground patches and used to adjust σ in the Cauchy loss function or δ in the Huber loss function).
[0043] Therefore, the optimization step can generate an optimized height (z) value for each wheel track point, and the resulting surface profile (e.g., ground, road) can be provided to the autonomous vehicle's driving stack for safe and comfortable planning and control. For example, the autonomous vehicle can navigate to avoid detected hazards or road bumps (e.g., depressions, potholes), adjust the vehicle's suspension system to match the detected road profile (e.g., by compensating for bumps in the road), and / or apply early acceleration or deceleration based on the approximate road surface gradient in the detected road profile. Any of these functions should contribute to enhanced safety, extended vehicle lifespan, improved energy efficiency, and / or a smooth driving experience.
[0044] In some embodiments, as the ego machine navigates its environment, it can generate stereo image pairs using one or more stereo cameras. Each stereo image pair can be used to generate a stereo disparity field including stereo disparity values (also known as stereo parallax). A surface disparity field representing a surface (e.g., ground) in the environment can be generated by iteratively refining the estimated disparity values using a constrained nonlinear optimization process customized with one or more weights to directly solve for the surface (e.g., ground) disparity field (ground disparity field). The resulting surface (e.g., ground) disparity field can be used for various downstream tasks, such as obstacle detection, airworthiness space segmentation, ego motion refinement, and / or generating estimated surface profiles.
[0045] More specifically, in some embodiments, surface structures, such as the height profile of the ground or road surface, can be estimated from stereo image pairs using nonparametric models. Standard stereo matching techniques typically attempt to generate high-fidelity reconstructions of all parts of the observed scene. Traditionally, estimation of specific geometric entities (e.g., ground surfaces) is performed by upscaling the stereo parallax field to 3D using triangulation. In contrast, some embodiments perform this estimation directly in the parallax field space. Unlike off-the-shelf stereo methods, prior knowledge of surface geometry (e.g., ground) can be directly embedded into a cost function used to derive the weights of the optimization algorithm. Some embodiments enhance the robustness of the estimation process by enforcing local smoothing of the estimated surface field (e.g., assuming no large height discontinuities on the ground, road, or other navigable surfaces). In embodiments where both stereo and ground parallax fields are available, the detection of obstacles on the road surface is significantly simplified.
[0046] In an example estimation of the ground disparity field, the stereo disparity field (which may also be referred to as a disparity image) can be progressively downsampled to form a pyramid of stereo disparity layers. Ground disparity estimation can begin with the coarsest pyramid layer and initialize the ground disparity using the stereo disparity from the coarsest pyramid layer. An iterative process can be used to generate and iteratively refine the estimated ground disparity values using a constrained hierarchical optimization that minimizes a cost function defining one or more weights. For example, the optimization process can be constrained using: measurement deviation weights that penalize measured disparity values (derived from the stereo image) that deviate from the estimated disparity values (causing the optimization to converge to smaller disparities at the ground layer); weights that emphasize disparity values based on proximity to the predicted ego trajectory (causing the optimization to focus on areas that may be part of the ground, road, or other navigable surfaces); weights that deemphasize disparity values below the detected horizon based on proximity to the detected horizon (since in various embodiments, disparity values at the horizon should be zero, while disparity values above the horizon should contribute nothing); and / or other terms. In this way, the refined parallax can be upsampled and passed to the next higher resolution pyramid layer. This process can be repeated until the highest resolution layer is reached and refined, ultimately producing a parallax field representing the Earth's surface (the ground parallax field).
[0047] There may be several variations. For example, in some embodiments, each stereo image can be iteratively downsampled to derive an image pyramid for each stereo image, performing stereo matching in the coarsest layer, and performing ground disparity estimation by iteratively refining the estimated ground disparity values using weights that emphasize disparity values corresponding to higher intensity gradients in the stereo image (prompting optimization to focus on areas that may be part of a road). This may involve looking up corrected image data, but since the source data is already incorporated into the optimization process, the accuracy and robustness of the estimation should be improved. Therefore, the refined ground disparity values can be upsampled and passed to the next higher-resolution pyramid layer, and this process (looking up intensity gradients from the corresponding layers of the image pyramid of the stereo image to generate corresponding weights for each refined layer) can be repeated until the highest resolution layer is reached and refined.
[0048] In another example, each stereo image can be iteratively downscaled to generate an image pyramid for each stereo image, stereo matching is performed in the coarsest layer, and ground disparity estimation is performed using measurement deviation weights derived from the difference between the estimated ground disparity values and the disparity values in the coarse disparity image. The coarse disparity image can be refined using optical flow (which compensates for calibration inaccuracies), then the coarse disparity image and the coarse ground disparity field are upsampled and passed to the next higher-resolution pyramid layer, and this process is repeated (using measurement deviation weights derived from the difference between the upsampled ground disparity field and the upsampled refined disparity image) until the highest resolution layer is reached and refined. These are just a few examples; other variations can be implemented within the scope of this disclosure.
[0049] Therefore, the resulting ground parallax field can be used to perform various tasks. Taking obstacle detection as an example, the difference between the stereo parallax field and the ground parallax field can be used to detect objects. For example, parallax values can be upscaled to 3D (e.g., converted to distance values, which can then be backprojected into 3D space) to derive corresponding height values, and a distance-dependent threshold height is applied to the difference between the stereo parallax values and the ground parallax values to detect obstacles based on their estimated height above the ground surface. In some embodiments, a corresponding threshold can be applied directly in the parallax space by applying a distance-dependent threshold parallax difference. (Distance correlation can be used to compensate for the decrease in parallax and detection height as scene depth increases.) Therefore, if the parallax in the ground parallax image is greater than the parallax in the stereo parallax image by more than a threshold amount, an object may be present in the corresponding region, and obstacles on the ground surface can be detected based on calculating the difference between the ground parallax and the stereo parallax and applying a specified threshold to that difference. In some embodiments, pixels that meet the detection threshold can be grouped into clusters, and clusters with a threshold size and / or a specified shape are considered as detected objects. In some embodiments, detected objects can be tracked and / or evaluated to confirm that they appeared in a threshold number of frames prior to confirmation of detection. Therefore, the (confirmed) object detection results can be passed to one or more downstream components to trigger one or more corresponding responses (e.g., path planning, emergency braking, etc.).
[0050] Taking the segmentation of an airworthiness space as an example, regions in the ground parallax field where the stereo parallax and ground parallax are within a specified threshold (or regions generated by subtracting stereo parallax from ground parallax) can be classified as ground. A representation of the airworthiness space can be generated by radially projecting 2D rays from a reference point (e.g., vehicle location, nearest ground location) in different directions into the ground parallax field (or difference image) towards a first location where a parallax difference exceeding a specified threshold occurs. In this example, each ray continues to propagate until it encounters an obstacle (indicated by the threshold parallax difference) or a road boundary. Regions where the rays do not encounter any obstacles or boundaries during their journey can be classified as airworthiness or drivable space. The points where the rays intersect with obstacles or boundaries can be used to generate 2D contour lines depicting the boundaries of the airworthiness space. Therefore, a representation of the airworthiness space (e.g., backprojecting the 2D contour lines into 3D space) can be provided to one or more downstream components to trigger one or more corresponding responses (e.g., path planning, emergency braking, etc.).
[0051] Taking ego-motion refinement as an example, the ground disparity field can be used to compensate for ego-motion caused by highly dynamic attitude changes. Real-time ego-motion is often limited by the use of high-frequency signals, such as inertial measurement and wheel odometers (primarily due to real-time constraints and the required update frequency), although spin-to-spin registration of point clouds can fail or lead to limited accuracy in cases involving highly dynamic motion. Therefore, ground surface observations over consecutive frames (e.g., ground disparity fields) can be used to refine the ego-motion estimate (transformation). More specifically, the ground disparity field estimated for consecutive frames can be boosted (e.g., converted to distance values and back-projected into 3D space), and then the resulting 3D point clouds are registered (e.g., using iterative nearest-point registration) to estimate the relative transformation between the 3D point clouds, and this relative transformation is used to refine the initial ego-motion transformation generated by ego-motion compensation.
[0052] Taking surface profile estimation as an example, the ground disparity field can be upscaled to 3D (e.g., by converting it to distance values and backpropagating it to 3D space), and the resulting upscaled point cloud can be interpreted as a surface model. In some embodiments, the upscaled point cloud can be sampled along one or more predicted trajectories, and nonlinear optimization can be used to fit the height of each trajectory point to the height of the corresponding sampled point to generate the surface profile.
[0053] Therefore, the techniques described herein can be used to estimate the 3D structure of surfaces (e.g., ground or road surface), detect obstacles in the environment, and / or detect airworthy space. The representation of the detection results can be provided to the autonomous vehicle's driving stack to enable safe and comfortable planning and control of the autonomous vehicle. For example, the autonomous vehicle can navigate to avoid detected obstacles or bumps (e.g., depressions, potholes) on the road, adjust the vehicle's suspension system to match the detected road profile (e.g., by compensating for bumps on the road), and / or apply early acceleration or deceleration based on the approximate surface slope in the detected road profile. Any of these functions can be used to enhance safety, extend vehicle lifespan, improve energy efficiency, and / or provide a smooth driving experience.
[0054] refer to Figure 1 , Figure 1 This is an example surface estimation pipeline 100 based on some embodiments of this disclosure. It should be understood that such and other arrangements described herein are merely illustrative examples. Other arrangements and elements (e.g., machines, interfaces, functions, sequences, functional groups, etc.) may be used in addition to the arrangements and elements shown, and certain elements may be omitted entirely. Furthermore, many of the elements described herein are functional entities that can be implemented as discrete or distributed components, or in combination with other components, and can be implemented in any suitable combination and location. The various functions performed by the entities described herein can be performed by hardware, firmware, and / or software. For example, various functions can be implemented by a processor executing instructions stored in memory. In some embodiments, the systems, methods, and processes described herein can be used with… Figures 15A to 15D Example of autonomous vehicle 1500 Figure 16 Example computing devices 1600 and / or Figure 17 The components, features, and / or functions of the example data center 1700 are similar to those of other components, features, and / or functions used in the implementation.
[0055] In some examples, the machine learning models described herein (e.g., deep neural networks, language models, LLMs, VLMs, multimodal language models, perceptual models, tracking models, fusion models, transformer models, diffusion models, encoder-only models, decoder-only models, encoder-decoder models, neural rendering field (NERF) models, etc.) can be packaged as microservices (e.g., inference microservices (e.g., NVIDIA NIM)), which may include containers (e.g., operating system (OS) level virtualization packages) that may include an application programming interface (API) layer, a server layer, a runtime layer, and / or a model "engine". For example, an inference microservice may include the container itself and the model (e.g., weights and biases). In some cases, such as when the machine learning model is small enough (e.g., has a sufficiently small number of parameters), the model may be included within the container itself. In other examples (e.g., when the model is large), the model may be hosted / stored in the cloud (e.g., in a data center) and / or may be hosted on-premises and / or at the edge (e.g., on a local server or computing device, but outside the container). In these embodiments, the model may be accessed via one or more APIs (e.g., REST APIs). Therefore, in some embodiments, the machine learning models described herein can be deployed as inference microservices to accelerate model deployment on any cloud, data center, or edge computing system while ensuring data security. For example, an inference microservice may include one or more APIs, pre-configured containers for simplified deployment, an optimized inference engine (e.g., built using standardized AI model deployment and execution software, such as NVIDIA's Triton Inference Server, and / or one or more APIs for high-performance deep learning inference, which may include inference runtime and model optimizations to provide low latency and high throughput for production applications, such as NVIDIA's TensorRT), and / or enterprise management data for telemetry (e.g., including identity, metrics, health checks, and / or monitoring). The machine learning models described herein may be included as part of a microservice along with an acceleration infrastructure capable of deployment using a single command and / or orchestration and autoscaling using a container orchestration system on the acceleration infrastructure (e.g., reaching data center scale on a single device). Therefore, the inference microservice may include a machine learning model (e.g., a model optimized for high-performance inference), inference runtime software for executing the machine learning model and providing output / response to inputs (e.g., user queries, prompts, etc.), and enterprise management software for providing health checks, identity verification, and / or other monitoring. In some embodiments, the inference microservice may include software for in-situ replacement and / or updating of the machine learning model. During replacement or updating, the software performing the replacement / update may maintain the user configurations of the inference runtime software and the enterprise management software.
[0056] In some embodiments, the systems and methods described herein can be performed using simulated data (e.g., simulated sensor data from simulated sensors of virtual machines or simulated machines) in a simulated environment (e.g., NVIDIA DriveSIM, NVIDIA ISAAC GYM, NVIDIA ISAAC SIM, etc.). For example, simulated sensor data can be used (e.g., processed using one or more machine learning models, neural networks, etc.) to perform the operations described herein, and this information can be used to perform operations associated with virtual machines in the environment (e.g., control, navigation, planning, etc.). These simulated operations can be used to test the performance of underlying algorithms, systems, and / or processes before deploying them to the real world. In some cases, simulation can be used to generate synthetic training data, e.g., training data containing regions of interest and / or subregions of interest from the simulation. In some embodiments, other methods besides simulation or as an alternative to simulation can be used to generate synthetic training data. For example, synthetic training data can be generated using neural rendering fields (NERF), Gaussian sputtering techniques, diffusion models, electrostatic models (e.g., Poisson flow generative models (PFGM)), etc. Then, synthetic training data (in addition to or as a substitute for real-world data) can be processed to determine geometry, curvature, semantic information, classification information, and / or other information related to features of interest, such as lines, longitudinal features (e.g., poles), and / or other features within a driving environment, warehouse, etc. In any example, such as in examples using a simulated environment for testing, validation, training, etc., one or more optical propagation algorithms (e.g., ray tracing and / or path tracing algorithms) can be used to render or otherwise generate the simulated environment and / or associated training data. In some embodiments, simulated environments and / or one or more objects, features, or components of a simulated environment can be generated or managed within a 3D content collaboration platform (e.g., NVIDIA's OMNIVERSE) for industrial digitization, generative physics AI, and / or other use cases, applications, or services. For example, a content collaboration platform or system may include a system that uses generic scene descriptors (USD) (e.g., OpenUSD) data to manage objects, features, scenes, etc., in simulated environments, digital environments, etc. The platform may include realistic physics simulations, such as using NVIDIA's PhysX SDK, to simulate real physics and physical interactions with simulations hosted on the platform. This platform can integrate OpenUSD and ray tracing / path tracing / light transport simulations (such as NVIDIA's RTX rendering technology) into software tools and simulation workflows for building, training, deploying, or testing AI systems, such as systems for testing, validating, training (e.g., machine learning models, neural networks, etc.) and / or other tasks related to automobiles, robots, machines, or other applications.
[0057] In some embodiments, remote operation or remote control of a vehicle or other machine may be performed using a remote control or remote operating system. For example, the systems and methods described herein may be used to identify road surface information, which may be included in the visualization or mapping of the environment, to assist a remote operator in controlling an autonomous or semi-autonomous machine, or to provide indications of waypoints or other control or navigation for the autonomous or semi-autonomous machine through the environment.
[0058] In some embodiments, the systems and methods described herein can be deployed in robotic applications. For example, a robot or robotic system may include one or more onboard processors (e.g., CPU, GPU, hardware-based deep learning accelerator (DLA), hardware-based programmable vision accelerator (PVA) – which may include one or more vector processing units (VPU), direct memory access (DMA) systems and / or pixel processing engines (PPE), hardware-based optical flow accelerators (OFA), SoCs, etc.) and memory and / or storage devices (e.g., for storing control algorithms, sensor data, and one or more machine learning models). The robotic system can use these processors to execute one or more machine learning models (e.g., language models) that allow the robotic system to perform complex tasks autonomously or semi-autonomously, such as interacting with and / or manipulating static and / or dynamic objects, or navigating its environment using sensors such as cameras, LiDAR, RADAR, and ultrasonic sensors. The system can use sensor fusion techniques to combine data from multiple sensors (e.g., cameras, infrared, LiDAR, radar, accelerometers) to create a comprehensive model of the robot's surrounding environment. This data can be processed locally on the robot or sent to a remote server for computationally intensive tasks such as 3D mapping or SLAM (Simultaneous Localization and Mapping). In one or more embodiments, data from individual robots (e.g., sensor data, task status, or environmental conditions) can be uploaded to the cloud, where a centralized AI model can analyze the data and distribute optimized commands to the entire robot fleet. In some embodiments, the machine learning models described herein (e.g., language models, VLM, LLM, MMLM, diffusion models, NeRF models, DNN, etc.) can be used to allow robots to perceive and reason about their environment and / or communicate with one or more other robots and / or people in the environment. In some embodiments, robots can communicate with one or more locally hosted servers / computing devices and / or one or more remotely located servers / computing devices (e.g., in one or more data centers) using one or more network interface cards (NICs) and / or data processing units (DPUs).
[0059] While this article may describe examples using machine learning models (e.g., neural networks), this is not intended to be limiting. For example, but not limited to, any of the various machine learning models and / or neural networks described herein can include any type of machine learning model, such as those using linear regression, logistic regression, decision trees, support vector machines (SVM), Naive Bayes, k-nearest neighbors (Knn), K-means clustering, random forests, dimensionality reduction algorithms, gradient boosting algorithms, neural networks (e.g., autoencoder neural networks, artificial neural networks (ANN), convolutional neural networks (CNN), recurrent neural networks (RNN), perceptrons, long / short-term memory (LSTM) networks, multilayer perceptron (MLP) networks, deep stacked networks (DSN), generative pre-trained (GPT) models or networks, feedforward networks, radial basis function ANNs, self-organizing maps (SOM), Kohonen... Machine learning models including mapping, Hopfield networks, Boltzmann machines, deep belief neural networks, deconvolutional neural networks, generative adversarial networks (GANs), liquid machines, modular neural networks, sequence-to-sequence models, networks using transformer architectures, diffusion models (e.g., diffusion probability models, score-based generative models, etc.), neural rendering field (NeRF) models, Kolmogorov-Arnold networks (KANs), models with encoder-only architectures, models with decoder-only architectures, models with encoder-decoder architectures, generative machine learning models, language models, large language models (LLMs), visual language models (VLMs), multimodal language models (MMLMs), etc., and / or other types of machine learning models.
[0060] exist Figure 1 In the illustrated embodiment, the surface estimation pipeline 100 uses LiDAR data 101 and ego motion data 102 (e.g., using ego machines (e.g.)). Figures 15A to 15DThe 3D surface structure of the estimated surface 180 is detected using sensors generated by the autonomous vehicle 1500. In the example overview, the motion compensation component 105 can apply motion compensation to the LiDAR data 101 using self-motion data 102, and the surface estimation component 140 can accumulate the resulting motion-compensated LiDAR data 110 into a sampling queue 175, sample the accumulated LiDAR data in the sampling queue 175 (e.g., along one or more predicted trajectories generated by the path generator 130), and use nonlinear optimization to fit the height value of each trajectory point to the height of the corresponding sampling point. In some embodiments, the surface estimation component 140 can improve the accuracy of the fitted height values by applying bias correction and / or applying self-motion refinement to the measured height values (e.g., the height values of LiDAR data 101, motion-compensated LiDAR data 110, etc.) to register continuous (e.g., segmented) point clouds and refine the accumulated LiDAR data in the sampling queue 175. In this way, the resulting estimated surface 180 (e.g., a road surface profile modeled along the wheel tracks of the self-machine) can be provided to the control component 190 of the self-machine, such as an adaptive suspension control system, which uses the estimated surface 180 to adjust the damping characteristics of the self-machine suspension system to counteract the indentations (e.g., potholes) or bumps (e.g., speed bumps) represented in the estimated surface 180.
[0061] More specifically, in some embodiments, self-machines (e.g., Figures 15A to 15D The autonomous vehicle (1500) can be equipped with one or more LiDAR sensors (e.g., Figure 15A The LiDAR sensor 1564 is used to generate LiDAR data 101 (e.g., when the self-machine is moving in its environment). Self-motion data 102, representing the self-movement of the self-machine, can be recorded using any known technology (e.g., using an inertial measurement unit (IMU), a global positioning system (GPS), etc.). Sensor data from any given sensor can be generated at any frame rate, synchronized with or otherwise correlated with sensor data from other sensors, and processed by the surface estimation pipeline 100 at any frame rate. Figure 1 The implementation shown is merely an example, and other embodiments may additionally or alternatively rely on other types of sensor data, such as RADAR data, sonar data, depth data, and / or other types.
[0062] Motion compensation component 105 can apply motion compensation to LiDAR data 101 from any number of LiDAR sensors and / or scans (or spins) using self-motion data 102 from the self-machine, to transform the raw LiDAR distance measurement results into a generic spatial representation (e.g., an aggregated LiDAR data frame representing a scene in the environment). For example, a LiDAR sensor may generate a representation of its surroundings at a certain rate (e.g., ten times per second, resulting in ten spins per second). However, a LiDAR sensor typically does not complete a full spin at a single timestamp. Instead, it typically rotates substantially continuously, requiring a duration (e.g., 100 milliseconds) to complete one spin. Therefore, each LiDAR point in the spin can be recorded at a unique timestamp, reflecting the sensor's continuous rotation. Thus, motion compensation component 105 can address this time offset problem by adjusting the spatial position of the LiDAR points using self-motion data 102 (e.g., typically provided by onboard sensors at a frequency of 100 Hz). This process can correct for the motion of the ego machine during each (e.g., 100 ms) capture period, aligning points in the spins to a single reference timestamp. For example, if the ego machine moves between generating the first and last point in the spins, the ego motion compensation component 105 can compute and apply the necessary transformations to interpret that motion, thereby effectively eliminating the ego machine's ego motion from the LiDAR data 101. Thus, ego motion compensation can generate a number of LiDAR spins per second (e.g., 10), each of which can be corrected to reflect a consistent spatial configuration at a common timestamp. Therefore, the motion compensation component 105 and / or the surface estimation component 140 can store the motion-compensated LiDAR data 110 in the sampling queue 175 and / or periodically update the motion-compensated LiDAR data 110 (e.g., such that the sampling queue 175 effectively stores a cumulative point cloud accumulated over a sliding window of a certain number of frames and updates it at any suitable frame rate, etc.).
[0063] At a high level, the surface estimation component 140 can generate an estimated surface 180 by sampling LiDAR data (e.g., in one or more local regions) along one or more predicted trajectory pairs (e.g., motion-compensated, accumulated) and using nonlinear optimization to fit height values to a set of sampled heights for each sampled trajectory point. Figure 1In the example shown, the surface estimation component 140 includes: a bias correction component 145 for applying bias correction to measured height values (e.g., measured height values of LiDAR data in sampling queue 175); a self-motion refinement component 150 for registering continuous (e.g., segmented) point clouds and refining the accumulated LiDAR data in sampling queue 175; a sampling component 165 for sampling (e.g., accumulated, bias-corrected, registered) LiDAR data in sampling queue 175 along one or more predicted trajectories (e.g., generated by path generator 130); and a fitting component 170 for fitting the height value of each trajectory point to the height of the corresponding sampling point using nonlinear optimization.
[0064] In some embodiments, the deviation correction component 145 applies deviation correction by removing measurement bias or offset from the measured height values (e.g., measured height values of LiDAR data in sampling queue 175, or measured height values at any other point in surface estimation pipeline 100). For example, distance-related height biases can be pre-calculated for various distance barrels of arbitrarily specified sizes, reflectivity-related height biases can be pre-calculated for various reflectivity barrels of arbitrarily specified sizes, and the estimated biases can be stored in any suitable form as LiDAR bias data 160 (e.g., stored in one or more lookup tables indexed by distance and / or reflectivity). Thus, the deviation correction component 145 can compensate for (e.g., measured, motion-compensated, registered) the height values of LiDAR points by looking up and subtracting the distance-related height bias corresponding to the measured distance, and / or looking up and subtracting the reflectivity-related height bias corresponding to the measured reflectivity.
[0065] Figure 2 The diagram illustrates distance-dependent altitude deviations in LiDAR sensor data. Due to the divergence of the beam emitted by LiDAR sensor 210, a portion of the beam's divergence cone (shown as dashed line 210) typically strikes the seaworthy surface earlier than the ideal beam (shown as dashed line 220). Depending on the strength of the returned signal, conventional LiDAR sensors report a measured 3D position that is closer and higher than the actual position (shown as points 230 and 240). The magnitude of this effect increases with distance.
[0066] Figure 3The figure illustrates a reflectivity-related height bias in LiDAR sensor data. More specifically, it depicts a scenario where a LiDAR beam emitted by a LiDAR sensor (not shown) (shown by a divergence cone 320 corresponding to an ideal beam 310) strikes a navigable surface 330 at an angle. For low-reflectivity surfaces (e.g., dark asphalt or concrete surfaces), the return signal 340 typically arrives at the LiDAR sensor detection threshold 360 later than the return signal 350 produced by high-reflectivity surfaces (e.g., surfaces with road markings, metal manhole covers). This results in a bias where the measured height coordinates of the reflective surface (e.g., point 370) are reported as higher than the actual surface height.
[0067] To correct for one or more of these measurement biases, in some embodiments, one or more data collection vehicles may be equipped with one or more LiDAR sensors (e.g., a single, roof-mounted 360° field-of-view LiDAR scanner; a forward-facing, grille-mounted or windshield-mounted long-range LiDAR sensor, etc.), and the LiDAR sensors of the data collection vehicles can be used to generate and accumulate individual LiDAR measurements. Depending on the desired use case, the environment and / or scenario can be selected or specified to cover a range of conditions, terrain, weather conditions, time of day, traffic density, and / or road type to ensure comprehensive data collection. In some embodiments, any suitable computing device (e.g., Figure 16 Computing device 1600, Figure 17 The bias estimation component, performed by computing devices in a data center 1700, uses any known self-motion compensation and / or point cloud registration techniques to perform self-motion compensation on the point cloud and / or register point clouds from multiple LiDAR spin or scans to each other to increase point density and generate an accumulated LiDAR point cloud.
[0068] To estimate distance-related height bias, the bias estimation component can identify fixed locations (e.g., blocks on the ground) represented in the accumulated LiDAR point cloud, bin the observed height values at the fixed locations according to the measured distance, combine or aggregate the height values in each bin (or bucket), and calculate the distance-related height bias by subtracting the true height from the combined height values of a given bin. Figure 4 An example process for estimating height deviations in LiDAR sensor data according to some embodiments of this disclosure is illustrated. Taking distance-related height deviations as an example, the deviation estimation component can convert height measurements representing common local neighborhoods in the accumulated point cloud (in...) Figure 4 (shown as white circles) are distributed into distance bins (e.g., of any suitable bin size, such as one meter). In this example, Figure 4The boxes on the left represent closer measurement distances (e.g., measurements taken closer to a common local neighborhood), while Figure 4 The box on the right represents a more distant measurement (e.g., a measurement taken at a greater distance from a common local neighborhood).
[0069] Therefore, the bias estimation component can use any suitable metric to combine or aggregate the height values within each bin (e.g., by calculating the median of all measurements within the bin). Figure 4 (Shown as black squares). Using the median height combination measurement results should be able to identify outliers (e.g., Figure 4 Robust height estimation is produced in the case of outliers shown in the middle box (x and x+2). Note that the actual number of measurements may be significantly higher than the actual number. Figure 4 The number shown (e.g., approximately several thousand measurements).
[0070] Therefore, the bias estimation component can calculate the height bias for any given distance box by subtracting the true height from the combined height value of the given box. Due to the nature of distance-dependent height bias, points observed from closer distances and steeper angles of incidence should be less affected by distance-dependent height bias than points observed from farther distances and gentler angles of incidence; therefore, observed height values measured from closer distances should be more accurate than observed height values measured from farther distances. Thus, the bias estimation component can use the combined height value corresponding to the nearest measurement distance band (e.g., within a specified measurement distance, such as one meter) as the true height for that local neighborhood. In some embodiments, the bias estimation component can calculate the true height by taking the weighted median of all height measurements (or a subset of height measurements), where closer measurements are given higher weights. These are merely examples, and other variations can be implemented within the scope of this disclosure.
[0071] Therefore, the deviation estimation component can calculate the distance-related height deviation for any number of range buckets. In some embodiments, the deviation estimation component can calculate multiple sets of distance-related height deviations based on different local neighborhoods in the accumulated point cloud, and can combine (e.g., average) the deviations of common range buckets. Therefore, the deviation estimation component can store the obtained correction values (e.g., stored in a 2D lookup table indexed by the measured distance, as...). Figure 1 At least a portion of the LiDAR bias data 160 (etc.) is used to facilitate efficient access during surface estimation.
[0072] In addition to estimating distance-related height bias, or as an alternative, the bias estimation component can also estimate reflectance-related height bias. For example, the bias estimation component can identify one or more locations (e.g., patches on the ground) in an accumulated LiDAR point cloud that have at least a threshold amount of reflectance variation (e.g., a local neighborhood with high reflectance paired with a low-reflectance area of asphalt pavement, such as a local neighborhood of road markings). Therefore, the bias estimation component can bin observation height values measured from approximately the same distance based on measured reflectance, combine or aggregate the height values in each bin (or bucket), and calculate the reflectance-related height bias by subtracting the true height from the combined height values of a given bin. Figure 4 The illustration shows an example binning technique for reflectivity-related height bias. The bias estimation component can represent the height measurements of any number of identified local neighborhoods in the accumulated point cloud (in... Figure 4 (shown as white circles) are distributed into reflectance boxes (e.g., with any suitable box size, such as 10 boxes of size 0.1). For example, Figure 4 The height measurement results shown can represent points measured from approximately the same distance, where Figure 4 The box on the left represents a lower measured reflectance value, while Figure 4 The box on the right represents a larger reflectance value. Therefore, the bias estimation component can use any suitable metric (e.g., by calculating the median of all measurements in the box). Figure 4 (Shown as black squares) to combine or aggregate height values within each box. Similar to the previous example of distance-related height deviation estimation, the number of actual measurements used to estimate reflectance-related height deviations may be significantly higher than... Figure 4 The quantity shown.
[0073] Therefore, the bias estimation component can calculate the height bias for any given reflectivity bin by subtracting the true height from the combined height values of a given bin. In some embodiments, the bias estimation component can use the combined (e.g., median) height of points in a local neighborhood in the accumulated LiDAR point cloud measured within a specified measurement distance (e.g., the measurement distance corresponding to the nearest measurement distance bin used to estimate the distance-related height offset) as the true height of that local neighborhood, or the bias estimation component can calculate the true height by compensating the observed height value with the corresponding distance-related height bias (e.g., subtracting the bias corresponding to the measurement distance of the given measurement). In some embodiments, instead of calculating the distance-related height bias and reflectivity-related height bias in separate processes, the bias estimation component can estimate both by sampling the accumulated LiDAR point cloud, binning the observed height values into distance and reflectivity bins, and utilizing joint 2D nonlinear optimization to calculate both biases using the observed height values measured within the specified measurement distance as the true height. These are just a few examples, and other variations can be implemented within the scope of this disclosure.
[0074] Therefore, the deviation estimation component can calculate reflectance-related height deviations for any number of reflectance buckets. In some embodiments, the deviation estimation component can calculate multiple sets of reflectance-related height deviations based on different local neighborhoods in the accumulated point cloud, and can combine (e.g., average) the deviations of common reflectance buckets. Therefore, the deviation estimation component can store the obtained correction values (e.g., stored in a 2D lookup table indexed by reflectance, as...). Figure 1 At least a portion of the LiDAR bias data 160 (etc.) is used to facilitate efficient access during surface estimation.
[0075] In addition to estimating height deviation, or as a substitute, by any suitable computing device (e.g., Figure 16 The computing device 1600 in the middle Figure 17 The noise estimation component, which operates in computing devices (such as those in a data center 1700), can estimate the true noise level and use the estimated true noise level to set... Figure 1 The fitting component 170 uses one or more parameters of the cost function (explained in more detail below). For example, the noise estimation component can identify local neighborhoods in the accumulated point cloud (e.g., circular blocks with a diameter of one meter), fit a second-order quadratic polynomial to the height measurement results in each local neighborhood, and calculate the true noise level as the standard deviation of the residuals (e.g., taking the average over any number of ground blocks).
[0076] Therefore, return Figure 1The bias correction component 145 can compensate for (e.g., measured, motion-compensated, registered) the height value of a LiDAR point by finding and subtracting a height bias related to distance corresponding to the measured distance, and / or finding and subtracting a height bias related to reflectance corresponding to the measured reflectance. While the various embodiments described herein contemplate correcting LiDAR measurement biases for surface estimation purposes, bias-corrected LiDAR data can be used for any suitable task, such as augmented reality, virtual reality, mixed reality, robotics, security and supervision, autonomous or semi-autonomous machine applications, and / or any other task in the technological space.
[0077] continue Figure 1 In the example shown, in some embodiments, the self-motion refinement component 150 registers continuous (e.g., segmented) point clouds to generate refined self-motion estimates, and uses the refined self-motion estimates to refine the accumulated LiDAR data in the sampling queue 175. Typically, the self-motion refinement component 150 can use any known registration process (e.g., ICP, point-to-surface matching, etc.) to improve the accuracy of the transformation mapping point clouds in continuous LiDAR frames to a common coordinate system.
[0078] In some embodiments, the self-motion refinement component 150 may use the estimated ground surface model 120 as a reference surface to improve the accuracy of self-motion refinement. For example, the ground surface estimation component 115 may use any known technique to estimate a representation of the ground surface model 120 based on LiDAR data (e.g., LiDAR data 101, motion-compensated LiDAR data 110). Example ground surface estimation techniques include classifying and projecting predicted points on the ground onto a grid and smoothing the projected values (e.g., as described in U.S. Patent Application No. 17 / 992,569, Publication No. US20240028041A1) and / or based on image data (e.g., generating an estimated representation of the ground surface using 3D reconstruction). The ground surface may be defined as a smooth, continuous surface extending outward from the pedestal or contact point of the self-machine, and the ground surface estimation component 115 may focus on estimating a single ground surface model 120, thereby avoiding ambiguity when multiple non-intersecting surface models are found. The ground surface model 120 can be defined to cover the same distance range as the LiDAR data (e.g., up to 300 meters) and can provide robust estimates of areas in the input data that are occluded by certain non-ground objects (e.g., people, vehicles, and other objects on the road surface).
[0079] Therefore, the self-motion refinement component 150 may include a reference surface segmentation component 155, which segments the LiDAR points in the sampling queue 175 into points belonging to a static reference surface (e.g., ground, vegetation, buildings) based on the height of the LiDAR points above the estimated ground surface represented by the estimated ground surface model 120. The self-motion refinement component 150 may register the segmented point cloud. For example, the reference surface segmentation component 155 may filter out LiDAR points above the ground at heights above a threshold (e.g., 10 cm) and / or below a threshold (e.g., 3 m). In some embodiments that apply both low and high threshold heights above the ground, the resulting filtered point cloud may effectively ignore height bands where moving objects are expected (e.g., 10 cm to 3 m above the estimated ground surface), thereby substantially eliminating a large number of potential outliers that may be moving and have no direct correspondence between consecutive LiDAR frames. Therefore, the self-motion refinement component 150 can register the obtained segmented point cloud to improve self-motion estimation and refine the LiDAR data in the sampling queue 175.
[0080] In some embodiments, the self-motion refinement component 150 can simplify the registration process to estimating or refining the pitch angle difference between (e.g., segmented) point cloud and (e.g., a known, estimated) reference surface (e.g., the ground). In some embodiments, the self-motion refinement component 150 can use the estimated ground disparity field to refine the self-motion estimate, thereby improving the accuracy of the LiDAR data in the sampling queue 175 (as explained in more detail below). These are merely examples, and other variations can be implemented within the scope of this disclosure.
[0081] In some embodiments, sampling component 165 samples LiDAR data (e.g., accumulated, bias-corrected, registered) in sampling queue 175 along one or more predicted trajectories (e.g., generated by path generator 130). Depending on the downstream use case, different surface features and different portions of the surface may be of interest. For example, some embodiments may attempt to detect a representation of the height profile of a road or other surface along a tire track in front of a vehicle. Thus, path generator 130 may sample any number of 2D or 3D points along a predicted 2D or 3D trajectory (e.g., sampling 2D points along one or more tire tracks in a bird's-eye view and assigning a candidate height, e.g., zero, to each point). This is merely an example; other techniques for sampling candidate points on roads or other surfaces are also feasible (e.g., sampling a specified number of points from a specified plane (e.g., z=0 or some other specified area in the environment)).
[0082] Continuing with the example of sampling one or more 2D or 3D points along one or more predicted trajectories, the path generator 130 can identify the trajectory based on wheel angles 125°. For example, the path generator 130 can use an Ackerman steering model and wheel angles 125° to generate a representation of the 2D or 3D trajectory of one or more tires by simulating the vehicle's kinematic behavior based on the vehicle's steering geometry. The Ackerman model defines a geometry that determines the angles the wheels make when the vehicle turns, thus allowing the vehicle to follow a smooth trajectory. Using this model, the inner and outer wheels of the vehicle should follow circular paths with different radii centered on a common turning point when turning. Therefore, the wheel angles 125° of the left and right tires can be detected using steering angle sensors and / or wheel position sensors, and the path generator 130 can use the wheel angles 125° of the left and right tires and the Ackerman model to calculate a representation of the predicted trajectory for each tire, such as the turning radius and curvature of each path in a 2D (e.g., top-down) view of the vehicle's motion plane (e.g., the xy plane) (assuming a flat road surface). The path generator 130 can use a representation of each predicted trajectory to sample any number of 2D or 3D points along each trajectory (e.g., at regular intervals, logarithmic, etc.). In some embodiments that sample 2D trajectory points, the path generator 130 can assign an initial height (e.g., zero) to each 2D point to generate a corresponding 3D sampling location.
[0083] Therefore, sampling component 165 can sample LiDAR points in sampling queue 175 located within a specified (e.g., orientation-dependent) 3D radius of the 3D position of each predicted 3D trajectory point (e.g., predicted 3D trajectory points in each of one or more predicted ruts). For example, surface estimation component 140 can maintain a registered point cloud corresponding to a number (e.g., recently observed) spins in sampling queue 171, path generator 130 can periodically update one or more predicted trajectories (e.g., at any suitable frame rate), and sampling component 165 can project each predicted 3D trajectory point into the (e.g., most recent) registered point cloud represented in sampling queue 171 to identify the corresponding 3D sampling position. Thus, sampling component 165 can sample points in the registered point cloud located within a specified 3D radius of each 3D sampling position. Figure 5 Example techniques for sampling LiDAR detection results along a predicted trajectory 510 according to some embodiments of this disclosure are illustrated. More specifically, Figure 5The LiDAR measurements in height-distance (z / d) space are depicted (shown as white circles) (e.g., where z represents the estimated height of the contour point and d represents the distance from the ego vehicle on the unfolded trajectory). Therefore, sampling component 165 can identify LiDAR points in the local neighborhood of each 3D sampling location, and fitting component 170 can use these sampled height values to fit the corresponding height value in each local neighborhood. Figure 5 (shown as a black circle in the middle).
[0084] Typically, the size of the threshold 3D radius used by the sampling component 165 to sample height values from the registered point cloud can be selected to customize the resolution of the surface fitting process. For example, sampling from a large 3D radius around each 3D sampling location may reduce the resolution of the estimated surface 180 and its ability to represent subtle variations in surface profile height (e.g., relatively small cracks in the road surface may not be represented in the estimated surface 180). On the other hand, if the threshold 3D radius is too small, there may not be enough LiDAR points for robust estimation. Therefore, the threshold 3D radius can be customized to correspond to the target use case (e.g., detecting small holes or large objects on a road) and can be selected to balance accuracy and resolution. In some embodiments, the threshold 3D radius can be selected to correspond to the width of a wheel. For example, the wheel track may be a strip region that depends on the width of the wheel (e.g., 200 mm or more wide), and sampling over a region that extends wider than the wheel width may not be beneficial (e.g., some embodiments may use a threshold 3D radius that is half the width of the wheel). In some embodiments, the threshold 3D radius can be asymmetric, for example, using a smaller sampling rate in the direction of navigation to capture smaller fluctuations in the height profile in that direction (e.g., some embodiments may use a radius of + / - 50 mm in the driving direction and a radius of + / - 200 mm along the wheel width). These are just a few examples, and other variations can be implemented within the scope of this disclosure.
[0085] Therefore, the fitting component 170 can use nonlinear optimization (e.g., a first-order method (e.g., gradient descent) or a second-order solver (e.g., a classic second-order least squares solver, such as Levenberg-Marquardt)) to fit the height values to the LiDAR points sampled for each trajectory point. For example, a second-order solver typically uses analytical gradients, Jacobian matrices, or Hessian matrices to achieve higher efficiency, thus requiring fewer iterations to converge, while a first-order method typically omits the use of analytical derivatives, providing a simpler implementation, but potentially at the cost of slower convergence. In some embodiments, the nonlinear optimization technique can be selected based on the applicable processor (e.g., some embodiments may use gradient descent on a graphics processing unit (GPU) or a least squares solver on a central processing unit (CPU). In some embodiments, the optimization can be simplified to a 1D optimization that fits the height values to the sampled LiDAR points represented in height-distance (z / d) space.
[0086] Nonlinear optimization can use cost functions to mitigate the impact of outlier observations (e.g., LiDAR points representing detected artifacts, particles in ambient weather such as snowflakes or raindrops, or transient events such as condensate, dust particles, or exhaust gases from a nozzle plume), for example, the cost function of a loss function that aggregates the difference between predicted and observed values (e.g., Cauchy loss or Huber loss). In some embodiments, one or more parameters of the cost function can be set based on the ground truth noise level estimated by the noise estimation component described above (e.g., by adjusting σ in the Cauchy loss function or δ in the Huber loss function based on the ground truth noise level).
[0087] Therefore, the fitting component 170 can generate optimized height (z) values for each trajectory point, collectively forming a 3D surface structure (e.g., road surface profile) of the estimated surface 180. Figure 6 An example road surface profile 650 along two wheel tracks 640 or ruts according to some embodiments of the present disclosure is shown. This example shows a perspective image 610 of the scene, and corresponding top-down projections 620 and perspective projections 630 representing the accumulated LiDAR point cloud of the scene (calculated via proof-of-concept offline processing). In this example, the tracks 640 of the left and right wheels represent example predicted tracks and extend approximately 20 meters from the ego vehicle. The road surface profile 650 includes two sections corresponding to the left and right wheel tracks 640, where peaks represent speed bumps 660 detected ahead.
[0088] Therefore, return Figure 1The surface estimation component 140 can output a representation of the estimated surface 180, which can be used by the control component 190 of the ego machine to perform one or more operations. For example, the ego machine can be an autonomous or semi-autonomous machine, such as... Figures 15A to 15D The autonomous vehicle 1500 is included, and the control component 190 may include, for example, a controller 1536, an ADAS system 1538, an adaptive suspension control system, and / or an autonomous driving software stack (e.g., the software stack described in U.S. Patent Application Publication No. 20210026355A1) executing on one or more components of the vehicle 1500 (e.g., SoC 1504, CPU 1518, GPU 1520, etc.). Therefore, the surface estimation component 140 may provide an estimated surface 180 to the control component 190 (such as these), and the control component 190 may use any known technology to navigate, plan, or otherwise use the estimated surface 180 to perform one or more operations (e.g., avoiding obstacles or bumps, keeping a lane, changing lanes, merging, splitting lanes, adjusting the autonomous vehicle's suspension system to match the current surface profile, applying early acceleration or deceleration based on the approximate surface slope, mapping, etc.).
[0089] In some embodiments, the 3D surface structure of a surface (e.g., ground) in the environment can be modeled and detected as a parallax field. Figure 7 This is a data flow diagram illustrating an example surface disparity estimation pipeline 700 according to some embodiments of the present disclosure. At a high level, the surface disparity estimation pipeline 700 can process stereo image data (e.g., left image 705a and right image 705b, respectively) using constrained nonlinear hierarchical optimization and iteratively refine the surface disparity field 770 representing surfaces (e.g., the ground) in the environment based on weights that guide the optimization to expected values of the surfaces.
[0090] In the example overview, stereo preprocessor 710 can process left image 705a and right image 705b (e.g., generated using one or more stereo or non-stereo camera pairs from a self-machine) to generate stereo pair 715. Stereo matcher 720 can create stereo disparity field 725 from stereo pair 715, and surface disparity estimation component 730 can generate surface disparity field 770 representing surfaces (e.g., ground) in the environment by iteratively refining the estimated disparity values using a constrained nonlinear optimization process, which is customized with one or more weights to directly solve for surface disparity field 770 (e.g., ground disparity field). Therefore, surface disparity estimation component 730 can provide surface disparity field 770 to one or more downstream components 780 for various tasks, such as obstacle detection, segmentation of airspace, self-motion refinement, and / or generation of estimated surface profiles.
[0091] More specifically, in some embodiments, self-machines (e.g., Figures 15A to 15D The autonomous vehicle (1500) can be equipped with one or more stereo cameras (e.g., Figures 15A to 15D The autonomous vehicle 1500 has a stereo camera 1568, which can be used to generate left image 705a and right image 705b (e.g., when the autonomous machine is moving through an environment). In some embodiments, left image 705a and right image 705b can be generated using a pair of non-stereo cameras, such as two cameras mounted on the autonomous machine, separated by a fixed baseline distance (e.g., about 0.22 meters), and synchronized to capture images of the scene substantially simultaneously. A trigger synchronization mechanism can be used to ensure substantially simultaneous exposure and acquisition of the two images. Left image 705a and right image 705b can be generated, synchronized, or otherwise paired at any frame rate and processed by surface parallax estimation pipeline 700 at any frame rate.
[0092] The stereo preprocessor 710 can use any known stereo processing technique to prepare the left image 705a and right image 705b for parallax calculation, such as distortion correction (e.g., applying intrinsic camera calibration parameters to nonlinear distortion to correct lens distortion in the left image 705a and right image 705b), correction (e.g., geometrically aligning the left image 705a and right image 705b), scaling (e.g., adjusting the size or resolution of the left image 705a and / or right image 705b to make them correspond to each other) and / or other techniques. Therefore, the stereo preprocessor 710 can process the left image 705a and right image 705b into a corresponding stereo pair 715.
[0093] exist Figure 7In the example embodiment shown, the surface disparity estimation pipeline 700 may use a multi-stage approach to estimate a smooth surface model (e.g., for the ground). In an initial stage, the stereo matcher 720 may use any known stereo matching technique to estimate the stereo disparity field 725 at the highest image resolution (e.g., local matching, global matching, semi-global matching, matching techniques using machine learning models (e.g., convolutional neural networks (CNNs) or transformers, etc.). The stereo disparity field 725 typically represents the difference in the horizontal position of corresponding points in the left and right images of the stereo pair 715. From this high-resolution stereo disparity field 725, the stereo pyramid generator 735 of the surface disparity estimation component 730 may generate a stereo disparity pyramid 740 using a specified downsampling strategy. In a subsequent stage, the iterative disparity estimation component 745 of the surface disparity estimation component 730 may use hierarchical iterative global optimization to estimate the surface disparity field 770. This hierarchical optimization may be implemented using an image pyramid (e.g., using a scaling factor of 2, halving the resolution of each successive pyramid layer) and a (e.g., Gaussian) smoothing step. The number of pyramid layers may be specified based on the target resolution. For example, an 8-megapixel (MP) image can be downsized to approximately 3MP resolution, and the number of pyramid layers can be set to three, resulting in a top-level resolution of 0.5MP. While some embodiments estimate surface disparity values across multiple pyramid levels, this is not always necessary. For instance, the surface disparity estimation component 730 can generate a surface disparity field 770 by iteratively refining the stereo disparity field 725 to minimize a cost function, without downsampling either the stereo disparity field 725 or the surface disparity field 770 into the corresponding image pyramid.
[0094] continue Figure 7 In the example shown, stereo matcher 720 can estimate stereo disparity field 725 at the highest image resolution from stereo pair 715 using stereo matching techniques (e.g., a hardware-accelerated variant of semi-global matching (SGM)), and stereo pyramid generator 735 can successively downsample stereo disparity field 725 to create stereo disparity pyramid 740. In some embodiments, stereo pyramid generator 735 uses a downsampling scheme that emphasizes the importance of small disparity values (which represent distant objects). For example, stereo pyramid generator 735 can use a min-filter to combine four disparity values from higher pyramid levels into a single disparity, which propagates the minimum of the four disparity values to the next pyramid level. Figure 8 An example stereo parallax pyramid 810 and a downsampling strategy according to some embodiments of the present disclosure are illustrated. In this example, Figure 8The full resolution of the bottom of the stereo parallax pyramid 810 is shown, and each layer can be continuously downsampled to generate a next-resolution stereo parallax image for the next pyramid layer, until a certain specified number of pyramid layers is reached. Figure 8 The lower half illustrates an example downsampling strategy using a minimum filter, where the four disparity values 820, 830, 840, and 850 in the corresponding pixel or cell are downsampled into a single value by propagating their minimum value (disparity value 830 in this example). This is merely an example and variations can be implemented within the scope of this disclosure.
[0095] Therefore, the iterative disparity estimation component 745 can utilize an iterative process to generate and iteratively refine (e.g., smooth) estimated surface disparity values for a specified surface (e.g., ground) using a constrained hierarchical optimization that minimizes a cost function. In the example overview, the iterative disparity estimation component 745 can start with the coarsest resolution pyramid layer (e.g., the highest pyramid layer), initialize the estimated surface disparity values with stereo disparity values in the corresponding pyramid layer of the stereo disparity pyramid 740, and iteratively refine that layer using nonlinear optimization until a specified criterion is met. For subsequent layers, the iterative disparity estimation component 745 can initialize the estimated surface disparity values by upsampling the refined disparity from the previous layer (e.g., by simply copying pixels to generate four values in the target layer from one value in the source layer), iteratively refining those estimated surface disparity values, and repeating, moving towards layers with increasing resolution, until the base layer of the pyramid is reached. Therefore, the iterative disparity estimation component 745 can effectively estimate the surface disparity pyramid, which has different pyramid levels representing surface disparity fields with different resolutions, with the lowest level (highest resolution) serving as the surface disparity field 770.
[0096] More specifically, starting from the coarsest layer of the surface disparity pyramid, the iterative disparity estimation component 745 can initialize the estimated disparity values to the stereo disparity values in the corresponding pyramid layers of the stereo disparity pyramid 740, initialize the corresponding weights, and use these weights to iteratively smooth the estimated disparity values over any number of iterations. Typically, the estimated surface disparity values can be in the form of a field, image, or grid of estimated surface disparity values corresponding to the 2D views represented by the left image 705a and the right image 705b. In some embodiments, the iterative disparity estimation component 745 can smooth the estimated surface disparity values in the grid by minimizing a cost function (or minimizing an approximate cost function) that penalizes the deviation between the stereo disparity values measured from the stereo disparity pyramid 740 and the estimated surface disparity values, and / or penalizes the deviation between adjacent estimates. For example, a cell grid (e.g., an image or field) can be used to model the 3D structure of a surface (e.g., the ground), where each cell (or pixel) stores the estimated surface disparity value. Therefore, the iterative disparity estimation component 745 can generate an initialized grid (or field) of estimated surface disparity values and / or a corresponding grid (or field) of corresponding cell / pixel weights, and the iterative disparity estimation component 745 can iteratively smooth the grid of initialized estimated surface disparity values based on the stereo disparity values measured from the stereo disparity pyramid 740 and the weight grid (e.g., weight map).
[0097] In some embodiments, the iterative disparity estimation component 745 calculates estimated surface disparity values that minimize (or approximately minimize) a global cost function, such as:
[0098]
[0099] In this context, each cell or pixel p in the grid for estimating surface disparity values can be assigned a weight w. p (For example, weighting disparity values based on proximity to the predicted ego trajectory, and weighting disparity values below the detected horizon based on proximity to the detected horizon) and / or measurement deviation weights (e.g., denoted by data_cost), which penalize the deviation between the measured stereo disparity value (derived from left images 705a and 705b) and the estimated surface disparity value (causing optimization to converge to the smaller disparity at the ground level), and where each cell or pixel q in the neighborhood N of p can be assigned a smoothness weight w. sThis weighting accounts for the difference between the estimated value of a neighboring cell or pixel q and the estimated value of a cell or pixel p. In this example, the global cost function in Equation 1 includes: a measurement term that encourages the estimated value to approximate the measured value; and a smoothing term that penalizes large changes in the local neighborhood, thereby encouraging or forcing the smoothness of the estimated surface disparity field 770 (e.g., for modeling a typical road surface). Using a global cost function in an iterative approximation method can encourage the current estimated value to propagate to a broad (e.g., global) neighborhood and help avoid getting trapped in local minima.
[0100] Typically, smaller stereo parallax represents relatively distant objects or surfaces (and vice versa). Therefore, surface obstacles and the boundaries between obstacles and the surfaces they traverse (e.g., the ground) generally have larger stereo parallax than the surfaces themselves (the surfaces themselves should appear relatively distant and thus should be represented by relatively smaller stereo parallax). Thus, a measurement term can include a cost function that reduces the weight of larger estimated parallax values (e.g., obstacles) and / or encourages smaller estimated parallax values (e.g., the ground). For example, a measurement term can include a cost and / or weighting function that defines a measurement deviation weight that increases as the deviation between the measured parallax and the estimated parallax increases, and / or defines different measurement deviation weights for measurements above the estimated value (e.g., measurement parallax above the estimated surface) compared to measurements below the estimated value (e.g., measurement parallax below the estimated surface).
[0101] Figure 9 Example asymmetric measurement deviations from the cost function are shown according to some embodiments of this disclosure. Figure 9In this example, the x-axis represents the signed difference (or deviation) between the measured disparity value and the currently estimated disparity value, and the y-axis represents the example cost that can be assigned. In this example, the cost function assigns zero cost when the deviation is zero; for small deviations, the cost function assigns symmetrical costs; the greater the distance between the estimated value and the measurement, the higher the cost. This example cost function becomes asymmetric above certain thresholds for both positive and negative deviations. For example, if the measured disparity is much lower than the estimated disparity, these measurements are likely outliers (e.g., noise) below a surface (e.g., the ground), and their impact can be limited by assigning higher costs (e.g., L2 or Huber). Conversely, if the measured disparity is significantly higher than the estimated disparity, these measurements are likely above the estimated surface (e.g., the ground) and can be assigned a cost (e.g., Huber, Turkey, Cauchy) that is higher than measurements closer to the estimated surface but gradually decreases (rolls off), preventing measurements with increasing disparity from having an increasingly larger impact. The advantage of this gradual decrease is that disparity values measured significantly above the cutoff distance (e.g., above a certain threshold) can be assigned constant costs and zero derivatives, thus avoiding contributions to the estimated surface disparity when minimizing (or approximately minimizing) the cost function.
[0102] In some embodiments, (e.g., asymmetric) measurement deviations from the cost function are used (e.g., by...). Figure 7 The iterative disparity estimation component 745 uses a measurement deviation weight function (e.g., defined as the measurement deviation cost function divided by the derivative of the deviation, unless the deviation is approximately zero, in which case a fixed weight (e.g., 1)) to calculate the weight of a given cell or pixel. Therefore, in some embodiments, the iterative disparity estimation component 745 can assign a measurement deviation weight to each cell or pixel p in the representation of the estimated surface disparity values.
[0103] Additionally or alternatively, the iterative disparity estimation component 745 may assign to each cell or pixel p in the representation of the estimated surface disparity values a weight that emphasizes the disparity value based on its proximity to the predicted ego trajectory and / or a weight that weakens the disparity value below the detected horizon based on its proximity to the detected horizon.
[0104] Taking proximity to the predicted self-trajectories as an example, we can assume the ego machine lies on a trajectory along the surface being estimated. Therefore, weights can be assigned to emphasize cells or pixels close to the estimated trajectory. Typically, any suitable technique can be used to estimate the self-trajectories. Figure 7In the example shown, the self-machine may include corresponding sensors that generate odometer data 750, derived from measurements of the self-machine's wheel rotation and self-motion data 755 representing the self-machine's self-motion (e.g., recorded using an IMU, GPS, or any other known technology). The self-trajectory estimation component 760 can estimate a trajectory (e.g., represented as a series of discrete position states) based on the odometer data 750 using dead reckoning, and can correct for drift using a Kalman filter based on the self-motion data 755. This is merely an example; other ways of estimating the trajectory can be implemented within the scope of this disclosure. Thus, the self-trajectory estimation component 760 can provide the estimated trajectory to an iterative parallax estimation component 745, which can sample the trajectory as a line, back-project it into camera space, and assign weights based on proximity to the back-projected trajectory (e.g., assigning weights within a specified threshold distance of the back-projected trajectory, such as 1, decreasing the weights or assigning weights to zero as the distance to the back-projected trajectory increases). In some embodiments, the self-trajectory estimation component 760 may use the weight only when estimating the first layer of the surface parallax pyramid.
[0105] Taking weighted calculations based on the detected horizon as an example, typically, the disparity value at the horizon should be zero, and depending on the application, the navigated surface is unlikely to appear above the horizon in the image. Therefore, weights can be assigned based on the position relative to the detected horizon to de-emphasize cells or pixels (e.g., so that disparity values measured above the horizon do not affect the estimation). Generally, any suitable technique can be used to detect the horizon. Figure 7 In the example shown, the ego motion data 755 may include inertial measurements estimating the orientation of the ego machine, while the horizon estimation component 765 may use the corresponding pitch and roll relative to the horizontal plane to estimate the position of the horizon in camera space (e.g., by adjusting the vertical displacement of the horizon from the image center based on the pitch angle, and adjusting the tilt of the horizon based on the roll angle). Therefore, the ego trajectory estimation component 760 may provide a representation of the detected horizon to the iterative disparity estimation component 745, which may assign weights to weaken units or pixels on the detected horizon, assign higher or increased weights as the distance below the detected horizon increases, and / or assign weights to offset contributions above the detected horizon.
[0106] Therefore, weighting based on the estimated trajectory and / or detected horizon can constrain and encourage optimization to converge to the disparity value of a seaworthy surface (e.g., the ground). For example, Figure 10The region of interest 1010 is shown when estimating the ground disparity field representing a straight road. Horizon 1020 can be used to assign weights to constrained optimization to focus on regions below and / or further away from horizon 1020 (e.g., during the estimation of each layer), while the projected ego trajectory 1030 can be used to assign weights to constrained optimization to focus on cells or pixels close to ego trajectory 1040 (e.g., during the estimation of the coarsest layer). Figure 10 The lower half of the diagram shows an example road segment with vehicle 1050 and obstacle 1060. Disparity values are represented by arrows of corresponding length. For smooth surfaces such as this road segment, the variation in disparity values within the local neighborhood is typically small. An asymmetric cost function can be used to encourage optimization to converge to smaller disparity values, and a local smoothing prior can be used to enforce local consistency, thereby encouraging the solution to converge toward a smooth disparity field.
[0107] Back Figure 7 The iterative disparity estimation component 745 can generate one or more weights for each cell or pixel, and can combine different types of weights into composite weights (e.g., based on averaging, weighted averaging, multiplication, addition, etc.). In some embodiments, the iterative disparity estimation component 745 generates weighted values (e.g., a weighted grid of measured disparities, a weighted map of measurements) by applying weights (e.g., weights from a weighted grid) to measured disparity values.
[0108] Therefore, taking the estimation of the first level of the surface disparity pyramid as an example, according to the embodiment, the iterative disparity estimation component 745 can use the corresponding values from the stereo disparity pyramid 740 to initialize the grid (or field) of the estimated surface disparity values, can generate the grid (or field) of the corresponding unit / pixel weights and / or the grid (or field) of the corresponding weighted measured disparity, and can apply smoothing processing to iteratively refine the estimated surface disparity values (e.g., their grid) based on the weighted measurements (e.g., their grid) and / or weights (e.g., their grid). In some embodiments, the iterative disparity estimation component 745 employs a nonlinear optimization scheme that iteratively updates the estimated disparity values, for example, by applying a weighted convolution to a weighted grid of measurements and / or a grid of weights, generating an updated estimate by dividing each smoothed weighted measurement by its corresponding smoothed weight, and updating the weighted grid of measurements and / or the grid of weights, for example, by updating the measurement deviation weights and / or the corresponding combined weights using the updated estimate (e.g., in some embodiments using measurement deviation weights). The iterative disparity estimation component 745 may run a specified number of iterations, or may use some other suitable termination criterion. In some embodiments, the iterative disparity estimation component 745 may skip updating regions where the estimated disparity values have converged (e.g., within a threshold) and continue updating the remaining regions.
[0109] In some embodiments where the iterative disparity estimation component 745 operates on a multi-level hierarchical representation (e.g., an image pyramid), the iterative disparity estimation component 745 may initially apply smoothing at the coarsest level for an arbitrary number of iterations. When the iterative disparity estimation component 745 terminates smoothing at a particular level, it may upsample the smoothed and / or updated mesh using any known technique to generate a corresponding representation for the next level, and may generate corresponding weights and apply smoothing at that level. Therefore, this process can be repeated, for example, iteratively smoothing and then upsampling at each successive level until iterative smoothing at the original resolution is complete. Thus, the resulting (e.g., highest resolution) solution can serve as the surface disparity field 770.
[0110] exist Figure 7 In an example variant of the illustrated implementation, the stereo pyramid generator 735 can iteratively downsample the left and right images in the stereo pair 715 to derive an image pyramid for each stereo image; the stereo matcher 720 can perform stereo matching in the coarsest layer; and the iterative disparity estimation component 745 can iteratively refine the estimated surface disparity values using weights that emphasize disparity values corresponding to higher intensity gradient consistency in the stereo images of the stereo pair 715, thereby prompting optimization to focus on regions that may be part of a surface (e.g., a road). For example, the iterative disparity estimation component 745 can compute the intensity gradients of the two images, compare the gradients at corresponding pixel locations, and assign higher weights to units or pixels in the estimated disparity field that correspond to lower intensity gradient differences. The iterative disparity estimation component 745 can additionally or alternatively assign weights that emphasize disparity values based on proximity to the predicted ego trajectory, and / or de-emphasize disparity values below the detected horizon based on proximity to the detected horizon. Therefore, the iterative disparity estimation component 745 can iteratively refine and upsample using weights, and pass the refined ground disparity values to the next higher resolution pyramid layer. The iterative disparity estimation component 745 can repeat this process (e.g., calculate the intensity gradient from the corresponding layer of the image pyramid of the stereo image in the stereo pair 715 to generate the corresponding weight for each refinement layer) until the highest resolution layer is reached and refined.
[0111] exist Figure 7In another example variant of the implementation shown, the stereo pyramid generator 735 can iteratively downsample the left and right images in the stereo pair 715 to derive an image pyramid for each stereo image. The stereo matcher 720 can perform stereo matching in the coarsest layer. The iterative disparity estimation component 745 can iteratively refine the estimated surface disparity values in the coarsest layer using measurement deviation weights derived from the difference between the estimated surface disparity values and the disparity values in the coarse disparity image. The iterative disparity estimation component 745 can additionally or alternatively use weights based on proximity to the predicted ego trajectory to emphasize disparity values, and / or use weights based on proximity to the detected horizon to de-emphasize disparity values below the detected horizon. In some embodiments, the iterative disparity estimation component 745 (or some other component) may use any known technique to refine the coarse disparity image using optical flow (e.g., using motion vectors estimated by optical flow to correct inconsistencies or artifacts, compensate for calibration inaccuracies, etc.), then upsample the coarse disparity image and the coarse surface disparity field and pass them to the next higher-resolution pyramid layer, and repeat the process (e.g., using measurement deviation weights derived from the difference between the upsampled surface disparity field and the upsampled refined disparity image) until the highest resolution layer is reached and refined. These are merely examples, and other variations may be implemented within the scope of this disclosure.
[0112] In this way, the surface disparity estimation component 730 can generate a representation of the surface disparity field 770 (e.g., and the stereo disparity field 725) and provide it to the downstream component 780 for various tasks, such as obstacle detection, segmentation of airspace, self-motion refinement and / or generation of estimated surface profiles.
[0113] More specifically, in some embodiments, downstream component 780 may include an object detector that performs object detection based on the difference between surface parallax field 770 and stereo parallax field 725. For example, the object detector may upscale the parallax values to 3D (e.g., by converting the parallax values to distance values and then backprojecting the distance values into 3D space) to derive corresponding height values and apply a distance-dependent threshold height to the difference between the stereo parallax values and the surface parallax values to detect obstacles based on the height of the obstacle above the estimated surface. In some embodiments, the object detector may apply a corresponding threshold directly in the parallax space by applying a distance-dependent threshold parallax difference. Therefore, if the parallax in the surface parallax image is greater than the parallax in the stereo parallax image by more than a threshold amount, an object may be present in the corresponding region, and the object detector may detect obstacles on the surface by calculating the difference between the surface parallax and the stereo parallax and applying a specified threshold to that difference. In some embodiments, the object detector may group pixels that meet the detection threshold into clusters and identify clusters having a threshold size and / or a specified shape as detected objects. In some embodiments, the object detector can track and / or evaluate detected objects to confirm that they appear in a threshold number of frames prior to confirmation of detection. Therefore, the object detector can feed data into the control components of the ego machine (e.g., Figure 1 Control component 190 in Figures 15A to 15D The controller 1536 or ADAS system 1538, adaptive suspension control system, autonomous driving software stack, etc. of the vehicle 1500 provides a (verified) representation of object detection to trigger one or more corresponding responses (e.g., path planning, emergency braking, etc.).
[0114] In some embodiments, downstream component 780 may include an airworthiness space detector that generates a representation of the airworthiness space based on a surface parallax field 770 (e.g., and a stereo parallax field 725). For example, the airworthiness space detector may classify a region of the surface parallax field 770 (or a region of a difference image generated by subtracting the stereo parallax field 725 from the surface parallax field 770) as part of a surface, where the stereo parallax and surface parallax are within a specified threshold, and may generate a representation of the airworthiness (e.g., drivable) space by radially projecting 2D rays from a reference point (e.g., the location of the ego machine, the nearest surface location) in different directions into the surface parallax field 770 (or the difference image) to a first location (where a parallax difference (e.g., an obstacle) or surface boundary is encountered above a specified threshold). The airworthiness space detector may classify areas where the rays do not hit any obstacles or boundaries as airworthiness or drivable space, and may use the points where the rays intersect with obstacles or boundaries to generate a 2D profile depicting the boundaries of the airworthiness space. Therefore, the airworthiness space detector can provide a representation of the airworthiness space (e.g., backprojecting a 2D profile into 3D space) to the control components of the self-machine (e.g., Figure 1 Control component 190 in Figures 15A to 15D The vehicle 1500's controller 1536 or ADAS system 1538, adaptive suspension control system, autonomous driving software stack, etc., trigger one or more corresponding responses (e.g., path planning, emergency braking, etc.).
[0115] In some embodiments, the downstream component 780 may include a self-motion refiner that compensates for self-motion estimation based on a surface disparity field 770. More specifically, the self-motion refiner may perform self-motion estimation (transformation) using instances of the surface disparity field 770 estimated for consecutive frames, for example, by upscaling disparity values to 3D (e.g., converting disparity values to distance values and back-projecting them into 3D space) to generate a 3D point cloud corresponding to the surface disparity field 770 estimated for each frame in any number of frames. Therefore, the self-motion refiner may use any known technique (e.g., ICP, point-to-surface matching) to register the 3D point cloud to estimate the relative transformation between the 3D point clouds, and this relative transformation may be used to refine the initial self-motion transformation generated by self-motion compensation. In some embodiments, the self-motion refiner corresponds to Figure 1 At least a portion of the self-movement refinement component 150 in the middle.
[0116] In some embodiments, downstream component 780 may include a surface profile estimation component that estimates a surface profile based on a surface disparity field 770. For example, the surface profile estimation component may upscale the surface disparity field 770 to 3D to generate a corresponding 3D point cloud, and the resulting upscaled point cloud can be interpreted as a surface model. Therefore, the surface profile estimation component can generate a surface profile by sampling the upscaled point cloud (or multiple accumulated upscaled point clouds) along one or more predicted trajectories and using nonlinear optimization to fit the height of each trajectory point to the height of the corresponding sampled point. Thus, the surface profile estimation component can provide a representation of the surface profile to the control component of the self-machine (e.g., Figure 1 Control component 190 in Figures 15A to 15D The vehicle 1500's controller 1536 or ADAS system 1538, adaptive suspension control system, autonomous driving software stack, etc., trigger one or more corresponding responses (e.g., obstacle or bump avoidance, lane keeping, lane changing, lane merging, lane separation, adjusting the autonomous machine's suspension system to match the current surface profile, applying early acceleration or deceleration based on the approximate surface slope, map rendering, etc.). In some embodiments, the surface profile estimation component corresponds to... Figure 1 The surface estimation component 140 in the middle.
[0117] Now for reference Figures 11 to 14 Each block of methods 1100-1400 described herein includes a computational process that can be executed using any combination of hardware, firmware, and / or software. For example, various functions can be implemented by a processor executing instructions stored in memory. Methods 1100-1400 can also be embodied as computer-usable instructions stored on a computer storage medium. Methods 1100-1400 can be provided by standalone applications, standalone services, managed services (alone or in combination with other managed services), or plug-ins to other products, to name a few. Furthermore, as an example, this document regarding... Figure 1 Example surface estimation pipeline 100 or Figure 7 The example surface disparity estimation pipeline 700 describes methods 1100-1400. However, these methods may be performed additionally or alternatively by any single system or any combination of systems, including but not limited to the systems described herein.
[0118] Figure 11This is a flowchart illustrating a method 1100 for surface estimation based at least on fitting height values in one or more local neighborhoods, according to some embodiments of the present disclosure. Method 1100 includes, at block B1102, generating a three-dimensional (3D) representation of an estimated surface in the environment of the ego machine, based at least on fitting one or more height values to a set or more sets of LiDAR detection results sampled along one or more predicted trajectories of the ego machine. For example, regarding... Figure 1 The surface estimation pipeline 100 and surface estimation component 140 can accumulate motion-compensated LiDAR data 110 in sampling queue 175, sample the accumulated LiDAR data in sampling queue 175 along one or more predicted trajectories generated by path generator 130, and use nonlinear optimization to fit the height value of each trajectory point to the height of the corresponding sampling point.
[0119] Method 1100 at box B1104 includes: controlling one or more operations of the self-machine based at least on an estimated 3D representation of the surface. For example, regarding Figure 1 The surface estimation pipeline 100, the estimated surface 180 (e.g., road surface profile modeled along the wheel tracks of the self-machine) can be provided to the control component 190 of the self-machine, such as an adaptive suspension control system, which uses the estimated surface 180 to adjust the damping characteristics of the self-machine suspension system to counteract the dents (e.g., potholes) or bumps (e.g., speed bumps) represented in the estimated surface 180.
[0120] Figure 12 This is a flowchart illustrating a method 1200 for generating bias-corrected LiDAR detection results according to some embodiments of the present disclosure. Method 1200 includes, at block B1202, generating one or more LiDAR detection results using one or more LiDAR sensors from a self-generated machine. For example, regarding... Figure 1 Surface estimation pipeline 100, self-machine (e.g., Figures 15A to 15D The autonomous vehicle (1500) can be equipped with one or more LiDAR sensors (e.g., Figure 15A The LiDAR sensor 1564 is used to generate LiDAR data 101 (e.g., when the ego machine is moving around in the environment).
[0121] Method 1200 includes, at box B1204, finding one or more estimated height offsets corresponding to at least one of one or more measured distance values or one or more measured reflectance values from one or more LiDAR detection results; and at box B1206, generating one or more bias-corrected LiDAR detection results based at least on removing one or more estimated height offsets from one or more measured height values from one or more LiDAR detection results. For example, regarding Figure 1 The surface estimation pipeline 100 and the deviation correction component 145 can compensate for the height value of the LiDAR point (e.g., measured, motion-compensated, registered) by finding and subtracting the height deviation related to the distance corresponding to the measured distance, and / or by finding and subtracting the height deviation related to the reflectivity corresponding to the measured reflectivity.
[0122] Method 1200, at box B1208, includes: controlling one or more operations of the ego machine based at least on one or more bias-corrected LiDAR detection results. For example, regarding Figure 1 The surface estimation pipeline 100, surface estimation component 140, can sample bias-corrected LiDAR data (e.g., accumulated, registered) in sampling queue 175 along one or more predicted trajectories (e.g., generated by path generator 130), and use nonlinear optimization to fit the height value of each trajectory point to the height of the corresponding sampling point. The resulting estimated surface 180 (e.g., a road surface profile modeled along the wheel ruts of the self-machine) is provided to the self-machine's control component 190, such as an adaptive suspension control system, which uses the estimated surface 180 to adjust the damping characteristics of the self-machine's suspension system to counteract indentations (e.g., potholes) or bumps (e.g., speed bumps) represented in the estimated surface 180. While the various embodiments described herein contemplate correcting LiDAR measurement bias for surface estimation, bias-corrected LiDAR data can be used for any suitable task, such as augmented reality, virtual reality, mixed reality, robotics, safety and supervision, autonomous or semi-autonomous machine applications, and / or tasks used in any other technical field.
[0123] Figure 13 This is a flowchart illustrating a method 1300 for generating a surface disparity field representing estimated disparity values of surfaces in an environment, according to some embodiments of the present disclosure. Method 1300 includes, at block B1302, generating a surface disparity field representing estimated disparity values of surfaces in the environment, based at least on a representation of stereo image data corresponding to the environment of the self-machine using nonlinear hierarchical optimization processing. For example, regarding... Figure 7The surface disparity estimation pipeline 700 and the surface disparity estimation component 730 can generate a surface disparity field 770 representing a surface (e.g., ground) in the environment by iteratively refining the estimated disparity values using a constrained nonlinear hierarchical optimization, which is customized with one or more weights to directly solve the surface disparity field 770 (e.g., ground disparity field).
[0124] Method 1300 at box B1304 includes: controlling one or more operations of the ego machine based at least on the surface parallax field of the surface. For example, regarding Figure 7 The surface disparity estimation pipeline 700 and the surface disparity estimation component 730 can provide the surface disparity field 770 to one or more downstream components 780 for various tasks, such as obstacle detection, airworthiness space segmentation, self-motion refinement and / or generation of estimated surface profiles.
[0125] Figure 14 This is a flowchart illustrating a method 1400 for controlling one or more operations of a self-machine based at least on a surface disparity field, according to some embodiments of the present disclosure. Method 1400 includes, at block B1402, generating a surface disparity field representing estimated disparity values of surfaces in the environment, based at least on a representation of stereoscopic image data corresponding to the environment of the self-machine. For example, regarding... Figure 7 The surface disparity estimation pipeline 700 and the surface disparity estimation component 730 can generate a surface disparity field 770 representing a surface (e.g., ground) in the environment by iteratively refining the estimated disparity values using a constrained nonlinear hierarchical optimization, which is customized with one or more weights to directly solve the surface disparity field 770 (e.g., ground disparity field).
[0126] Method 1400, at box B1404, includes: controlling one or more operations of the ego machine based at least on the surface parallax field of the surface. For example, regarding Figure 7 The surface disparity estimation pipeline 700 and the surface disparity estimation component 730 can provide the surface disparity field 770 to one or more downstream components 780 for various tasks, such as obstacle detection, airspace segmentation, self-motion refinement and / or generation of estimated surface profiles.
[0127] The systems and methods described herein may be used, or in combination with, but are not limited to, non-autonomous vehicles or machines, semi-autonomous vehicles or machines (e.g., in one or more adaptive driver assistance systems (ADAS)), manned and unmanned robots or robotic platforms, warehouse vehicles, off-road vehicles, vehicles coupled to one or more trailers, aircraft, boats, shuttles, emergency response vehicles, motorcycles, electric or motorized bicycles, aircraft, engineering vehicles, trains, underwater vehicles, remotely controlled vehicles (e.g., drones) and / or other vehicle types. Furthermore, the systems and methods described herein can be used for a variety of purposes, such as, but not limited to: machine control, machine motion, machine driving, synthetic data generation, model training, perception, augmented reality, virtual reality, mixed reality, robotics, safety and supervision, simulation and digital twins, autonomous or semi-autonomous machine applications, deep learning, environmental simulation, object or actor simulation and / or digital twins, data center processing, conversational AI, optical transport simulation (e.g., ray tracing, path tracing, etc.), distributed or collaborative content creation of 3D assets (e.g., using generic scene descriptor (USD) data, such as OpenUSD and / or other data types), cloud computing, generative artificial intelligence (e.g., using one or more diffusion models, converter models, etc.), language model applications (e.g., large language models (LLM), visual language models (VLM), etc.) and / or any other suitable applications.
[0128] The disclosed embodiments may be included in a variety of different systems, such as automotive systems (e.g., control systems for autonomous or semi-autonomous machines, perception systems for autonomous or semi-autonomous machines), systems implemented using robots or robotic platforms, aviation systems, medical systems, rowing systems, smart area monitoring systems, systems for performing deep learning operations, systems for performing simulation operations (e.g., in driving or vehicle simulations, in robot simulations, in smart city or surveillance simulations, etc.), systems for performing digital twin operations (e.g., in conjunction with collaborative content creation platforms or systems, such as, but not limited to, NVIDIA's OMNIVERSE and / or other platforms, systems, or services using USD or OpenUSD data types), systems implemented using edge devices, systems containing one or more virtual machines (VMs), and systems for performing deep learning operations. Systems that perform synthetic data generation operations (e.g., using one or more neural rendering fields (NERF), Gaussian sputtering techniques, diffusion models, converter models, etc.), systems that are at least partially implemented in a data center, systems that perform conversational AI operations, systems that implement one or more language models (e.g., one or more large language models (LLM), one or more visual language models (VLM), one or more multimodal language models, etc.), systems that perform optical transmission simulations, systems that perform collaborative content creation of 3D assets (e.g., using generic scene descriptor (USD) data, such as OpenUSD, computer-aided design (CAD) data, 2D and / or 3D graphics or design data and / or other data types), systems that are at least partially implemented using cloud computing resources, and / or other types of systems.
[0129] In some embodiments, the systems and methods described herein can be performed in a simulated environment (e.g., NVIDIA's DriveSIM) using simulated data (e.g., simulated sensor data from simulated sensors of a simulated machine). For example, simulated (or virtual) LiDAR (e.g., accumulated, bias-corrected) data (e.g., representing a simulated environment, such as a highway or warehouse environment, from the perspective of one or more simulated sensors of a simulated ego machine) can be used to estimate 3D surface structures (e.g., road surface profiles) by fitting height values to LiDAR data (e.g., sampled in local regions along one or more predicted trajectories) using nonlinear optimization. In some embodiments, the 3D surface structure can be modeled as a parallax field, and a surface parallax field representing a simulated surface (e.g., ground) in the simulated environment can be generated using constrained nonlinear hierarchical optimization, which processes the simulated stereo image data and iteratively refines the estimated surface parallax values based on weights that guide the optimization to the desired surface values (e.g., ground, road). Thus, the estimated 3D surface structure can be used to control a simulated ego machine in a simulated environment. These simulation operations can be used to test the performance of the underlying algorithms, systems, and / or processes before deploying them to the real world. In some cases, simulations can be used to generate synthetic training data—for example, images of a simulated environment generated from the perspective of one or more simulated sensors of a simulated self-machine—and the synthetic training data (as a supplement to or alternative to real-world data) can be used to train a multimodal language model (e.g., VLM). In any example, such as when the simulated environment is used for testing, validation, training, etc., one or more optical transport algorithms (e.g., ray tracing and / or path tracing algorithms) can be used to render or otherwise generate the simulated environment and / or associated training data. In some embodiments, simulated environments and / or one or more of their objects, features, or components can be generated or managed within a three-dimensional (3D) content collaboration platform (e.g., NVIDIA's OMNIVERSE) for industrial digitization, generative physics AI, and / or other use cases, applications, or services. For example, the content collaboration platform or system may include a system that uses or develops generic scene descriptors (USD) (e.g., OpenUSD) data to manage objects, features, scenes, etc., in simulated environments, digital environments, etc. The platform may include realistic physical simulations, such as using NVIDIA's PhysX SDK, to simulate real physics and physical interactions with simulations hosted on the platform. This platform can integrate OpenUSD with ray tracing / path tracing / light transport simulations (such as NVIDIA's RTX rendering technology) into software tools and simulation workflows for building, training, deploying, or testing AI systems, such as systems for testing, validating, training (e.g., machine learning models, neural networks, etc.) and / or other tasks related to automobiles, robots, machines, or other applications.
[0130] Example autonomous vehicles
[0131] Figure 15A This is an illustration of an example autonomous or semi-autonomous vehicle or machine 1500 according to some embodiments of this disclosure. The autonomous or semi-autonomous vehicle or machine 1500 (which may be alternatively referred to herein as "vehicle 1500", "machine 1500", "self-vehicle 1500", "self-machine 1500", "robot 1500", etc.) may include, but is not limited to, passenger vehicles such as cars, trucks, buses, ambulances, shuttles, electric or motorized bicycles, motorcycles, fire trucks, police cars, ambulances, boats, engineering vehicles, underwater vehicles, robotic vehicles, drones, aircraft, vehicles coupled to trailers (e.g., semi-trailer tractors for hauling cargo) and / or another type of vehicle (e.g., driverless and / or vehicles that can accommodate one or more passengers). Autonomous vehicles are typically described according to the level of automation defined by the National Highway Traffic Safety Administration (NHTSA) of the U.S. Department of Transportation and the Society of Automotive Engineers (SAE) in their standard "Classification and Definition of Terms Related to Driving Automation Systems for Road Motor Vehicles" (Standard No. J3016-201806, published June 15, 2018; Standard No. J3016-201609, published September 30, 2016; and previous and future versions of this standard). Vehicle 1500 may be able to operate at one or more levels of autonomous driving, from Level 3 to Level 5. Vehicle 1500 may be able to operate at one or more levels of autonomous driving, from Level 1 to Level 5. For example, vehicle 1500 may be able to perform driver assistance (Level 1), partial automation (Level 2), conditional automation (Level 3), high automation (Level 4), and / or full automation (Level 5), depending on the implementation. The term “autonomy” as used herein can include any and / or all types of autonomy of the vehicle 1500 or other machines, such as full autonomy, high autonomy, conditional autonomy, partial autonomy, provision of auxiliary autonomy, semi-autonomy, primary autonomy or other specified autonomy.
[0132] Vehicle 1500 may include components such as chassis, body, wheels (e.g., 2, 4, 6, 8, 18, etc.), tires, axles, and other vehicle components. Vehicle 1500 may include a propulsion system 1550, such as an internal combustion engine, a hybrid power plant, an all-electric motor, and / or another type of propulsion system. Propulsion system 1550 may be connected to the drivetrain of vehicle 1500, which may include a transmission, to enable propulsion of vehicle 1500. Propulsion system 1550 may be controlled in response to receiving a signal from throttle / accelerator 1552.
[0133] A steering system 1554, which may include a steering wheel, can be used to steer the vehicle 1500 (e.g., along a desired path or route) when the propulsion system 1550 is operating (e.g., when the vehicle is in motion). The steering system 1554 may receive signals from the steering actuator 1556. For fully automatic (level 5) functionality, the steering wheel may be optional.
[0134] The brake sensor system 1546 can be used to operate the vehicle brakes in response to receiving signals from the brake actuator 1548 and / or the brake sensor.
[0135] It can include one or more System-on-Chip (SoC) 1504 ( Figure 15C One or more controllers 1536, including one or more GPUs, may provide signals (e.g., signals representing commands) to one or more components and / or systems of vehicle 1500. For example, one or more controllers may send signals to operate vehicle brakes via one or more brake actuators 1548, to operate steering system 1554 via one or more steering actuators 1556, and to operate propulsion system 1550 via one or more throttles / accelerators 1552. One or more controllers 1536 may include one or more onboard (e.g., integrated) computing devices (e.g., supercomputers) that process sensor signals and output operating commands (e.g., signals representing commands) to enable autonomous driving and / or assist a human driver in driving vehicle 1500. One or more controllers 1536 may include a first controller 1536 for autonomous driving functions, a second controller 1536 for functional safety functions, a third controller 1536 for artificial intelligence functions (e.g., computer vision), a fourth controller 1536 for infotainment functions, a fifth controller 1536 for redundancy in emergency situations, and / or other controllers. In some examples, a single controller 1536 can handle two or more of the functions described above, and two or more controllers 1536 can handle a single function, and / or any combination thereof.
[0136] One or more controllers 1536 may provide signals for controlling one or more components and / or systems of vehicle 1500 in response to sensor data (e.g., sensor inputs) received from one or more sensors. Sensor data may be received from, for example, but not limited to, Global Navigation Satellite System (“GNSS”) sensor 1558 (e.g., Global Positioning System sensor), RADAR sensor 1560, ultrasonic sensor 1562, LiDAR sensor 1564, Inertial Measurement Unit (IMU) sensor 1566 (e.g., accelerometer, gyroscope, magnetic compass, magnetometer, etc.), microphone 1596, stereo camera 1568, wide-angle camera 1570 (e.g., fisheye camera), infrared camera 1572, surround camera 1574 (e.g., 360-degree camera), long-range and / or medium-range camera 1598, speed sensor 1544 (e.g., for measuring the rate of vehicle 1500), vibration sensor 1542, steering sensor 1540, braking sensor (e.g., as part of braking sensor system 1546), one or more Occupant Monitoring System (OMS) sensor 1501 (e.g., one or more interior cameras) and / or other sensor types.
[0137] One or more of the controllers 1536 may receive inputs (e.g., represented by input data) from the instrument cluster 1532 of the vehicle 1500 and provide outputs (e.g., represented by output data, display data, etc.) via a human-machine interface (HMI) display 1534, an auditory signaling device, a speaker, and / or via other components of the vehicle 1500. These outputs may include information such as vehicle speed, rate, time, map data (e.g., [missing information]). Figure 15C Information such as high-definition (“HD”) maps 1522, location data (e.g., the location of vehicle 1500 on the map), orientation, and the location of other vehicles (e.g., occupying grids), as well as information about objects and their states perceived by controller 1536, etc. For example, HMI display 1534 may display information about the existence of one or more objects (e.g., street signs, warning signs, traffic light changes, etc.) and / or information about driving maneuvers that the vehicle has made, is making, or will make (e.g., changing lanes now, leaving 34B in two miles, etc.).
[0138] Vehicle 1500 further includes a network interface 1524, which can communicate via one or more networks using one or more wireless antennas 1526 and / or a modem. For example, network interface 1524 may be able to communicate via Long Term Evolution (“LTE”), Wideband Code Division Multiple Access (“WCDMA”), Universal Mobile Telecommunications System (“UMTS”), Global System for Mobile Communications (“GSM”), IMT-CDMA Multicarrier (“CDMA2000”), etc. One or more wireless antennas 1526 may also enable communication between objects in the environment (e.g., vehicles, mobile devices, etc.) using one or more local area networks such as Bluetooth, Bluetooth Low Energy (“LE”), Z-Wave, ZigBee, etc., and / or one or more low-power wide area networks (LPWANs such as LoRaWAN, SigFox, etc.).
[0139] Figure 15B For use in accordance with some embodiments of this disclosure Figure 15A This is an example of the camera position and field of view of an autonomous vehicle 1500. The camera and its respective field of view are an example embodiment and are not intended to be limiting. For example, additional and / or replaceable cameras may be included, and / or these cameras may be located at different positions on the vehicle 1500.
[0140] The camera type used for the camera may include, but is not limited to, a digital camera suitable for use with components and / or systems of vehicle 1500. The camera may operate at Automotive Safety Integrity Level (ASIL) B and / or another ASIL. The camera type may have any image capture rate, such as 60 frames per second (fps), 920 fps, 240 fps, etc., depending on the embodiment. The camera may be able to use a rolling shutter, a global shutter, another type of shutter, or a combination thereof. In some examples, the color filter array may include a red-transparent (RCCC) color filter array, a red-transparent-blue (RCCB) color filter array, a red-blue-green (RBGC) color filter array, a Foveon X3 color filter array, a Bayer sensor (RGGB) color filter array, a monochrome sensor color filter array, and / or another type of color filter array. In some embodiments, a high-resolution camera, such as a camera with RCCC, RCCB, and / or RBGC color filter arrays, may be used in efforts to improve light sensitivity.
[0141] In some examples, one or more of the cameras can be used to perform advanced driver assistance system (ADAS) functions (e.g., as part of a redundant or fail-safe design). For example, a multi-function monocular camera can be installed to provide functions including lane departure warning, traffic sign assistance, and intelligent headlight control. One or more of the cameras (e.g., all cameras) can simultaneously record and provide image data (e.g., video).
[0142] One or more of the cameras can be mounted in mounting components such as custom-designed (3D-printed) components to cut off stray light and reflections from inside the vehicle (e.g., reflections from the dashboard reflected in the windshield mirror) that may interfere with the camera's image data capture capabilities. Regarding the wing mirror mounting components, the wing mirror components can be custom-3D printed so that the camera mounting plate matches the shape of the wing mirror. In some examples, one or more cameras can be integrated into the wing mirror. For side-view cameras, one or more cameras can also be integrated into the four pillars at each corner of the cab.
[0143] A camera with a field of view that includes the environment in front of the vehicle 1500 (e.g., a front-facing camera) can be used for surround view to help identify forward paths and obstacles, and, with the assistance of one or more controllers 1536 and / or control SoCs, to provide information crucial for generating an occupancy grid and / or determining a preferred vehicle path. The front-facing camera can be used to perform many of the same ADAS functions as LiDAR, including emergency braking, pedestrian detection, and collision avoidance. The front-facing camera can also be used in ADAS functions and systems, including lane departure warning (“LDW”), autonomous cruise control (“ACC”), and / or other functions such as traffic sign recognition.
[0144] A variety of cameras can be used in front-facing configurations, including monocular camera platforms such as complementary metal-oxide-semiconductor (CMOS) color imagers. Another example could be a wide-angle camera 1570, which can be used to perceive objects entering the field of view from the periphery (such as pedestrians, traffic at intersections, or bicycles). Although Figure 15B The middle image shows only one wide-angle camera, but any number (including zero) of wide-angle cameras 1570 can be present on vehicle 1500. Furthermore, any number of remote cameras 1598 (e.g., long-view stereo camera pairs) can be used for depth-based object detection, especially for objects for which neural networks have not yet been trained. Remote cameras 1598 can also be used for object detection and classification, as well as basic object tracking.
[0145] Any number of stereo cameras 1568 may also be included in the front-mounted configuration. In at least one embodiment, one or more stereo cameras 1568 may include an integrated control unit that includes a scalable processing unit that can provide a multi-core microprocessor and programmable logic (“FPGA”) with an integrated controller area network (“CAN”) or Ethernet interface on a single chip. Such a unit can be used to generate a 3D map of the vehicle environment, including distance estimates for all points in the image. Alternative stereo cameras 1568 may include a compact stereo vision sensor that may include two camera lenses (one on each side) and an image processing chip that can measure the distance from the vehicle to a target object and use the generated information (e.g., metadata) to activate autonomous emergency braking and lane departure warning functions. Other types of stereo cameras 1568 may be used in addition to those described herein or alternatively.
[0146] Cameras with a field of view including the side portion of the vehicle 1500 (e.g., side-view cameras) can be used for surround view, providing information for creating and updating occupancy grids and generating side-impact collision warnings. For example, surround camera 1574 (e.g., ... Figure 15B The four surround cameras 1574 shown can be positioned on the vehicle 1500. The surround cameras 1574 can include a wide-angle camera 1570, a fisheye camera, a 360-degree camera, and / or the like. For example, four fisheye cameras can be positioned at the front, rear, and sides of the vehicle. In an alternative arrangement, the vehicle can use three surround cameras 1574 (e.g., left, right, and rear) and can utilize one or more other cameras (e.g., forward-facing cameras) as a fourth surround-view camera.
[0147] A camera with a field of view that includes the environment behind the vehicle 1500 (e.g., a rear-view camera) can be used for parking assistance, surround view, rear collision warning, and creating and updating occupancy grids. A wide variety of cameras can be used, including but not limited to those also suitable as front-facing cameras as described herein (e.g., long-range and / or mid-range camera 1598, stereo camera 1568, infrared camera 1572, etc.).
[0148] Cameras (e.g., one or more OMS sensors 1501) having a field of view of the interior environment of various parts of the vehicle cabin 1500 can be used as part of an Occupant Monitoring System (OMS), such as, but not limited to, a Driver Monitoring System (DMS). For example, OMS sensors (e.g., OMS sensor 1501) can be used (e.g., by controller 1536) to track the gaze direction, head posture, and / or blinking of occupants and / or drivers. This gaze information can be used to determine the level of attention of the occupant or driver (e.g., detect drowsiness, fatigue, and / or distraction), and / or to take responsive actions to prevent harm to the occupant or operator. In some embodiments, data from the OMS sensors can be used to implement gaze control operations triggered by the driver and / or non-driver occupants, such as, but not limited to, adjusting cabin temperature and / or airflow, opening and closing windows, controlling cabin lighting, controlling the entertainment system, adjusting rearview mirrors, adjusting seat position, and / or other operations. In some embodiments, the OMS can be used for applications such as determining when an object and / or occupant is left in the cabin (e.g., by detecting the presence of an occupant after the driver has left the vehicle).
[0149] Figure 15C For use in accordance with some embodiments of this disclosure Figure 15A The example autonomous vehicle 1500 is illustrated in the block diagram of an example system architecture. It should be understood that this arrangement, and other arrangements described herein, are merely illustrative. Other arrangements and elements (e.g., machines, interfaces, functions, sequences, functional groupings, etc.) may be used in addition to or in place of those shown, and some elements may be omitted entirely. Furthermore, many of the elements described herein are functional entities, which may be implemented as discrete or distributed components or in combination with other components, and in any suitable combination and location. The various functions described herein as being performed by these entities can be implemented in hardware, firmware, and / or software. For example, the various functions can be implemented by a processor executing instructions stored in memory.
[0150] Figure 15C Each component, feature, and system in vehicle 1500 is illustrated as being connected via bus 1502. Bus 1502 may include a Controller Area Network (CAN) data interface (or, alternatively, referred to herein as the "CAN bus"). CAN may be a network within vehicle 1500 used to assist in the control of various features and functions of vehicle 1500, such as the actuation of brakes, acceleration, braking, steering, windshield wipers, etc. CAN bus can be configured to have dozens or even hundreds of nodes, each with its own unique identifier (e.g., CAN ID). CAN bus can be read to find steering wheel angle, ground speed, engine speed per minute (RPM), button positions, and / or other vehicle status indicators. CAN bus may be ASIL B compliant.
[0151] Although bus 1502 is described herein as a CAN bus, this is not intended to be limiting. For example, FlexRay and / or Ethernet may be used in addition to or alternatively to a CAN bus. Furthermore, although bus 1502 is represented by a single line, this is not intended to be limiting. For example, any number of buses 1502 may exist, which may include one or more CAN buses, one or more FlexRay buses, one or more Ethernet buses, and / or one or more other types of buses using different protocols. In some examples, two or more buses 1502 may be used to perform different functions and / or may be used for redundancy. For example, a first bus 1502 may be used for a collision avoidance function, and a second bus 1502 may be used for drive control. In any example, each bus 1502 may communicate with any component of vehicle 1500, and two or more buses 1502 may communicate with the same component. In some examples, each SoC 1504, each controller 1536, and / or each computer within the vehicle may have access to the same input data (e.g., input from sensors in the vehicle 1500) and may be connected to a common bus such as the CAN bus.
[0152] Vehicle 1500 may include one or more controllers 1536, such as those described herein. Figure 15A The controllers described herein. Controller 1536 can be used for a wide variety of functions. Controller 1536 can be coupled to any other different components and systems of vehicle 1500 and can be used for the control of vehicle 1500, artificial intelligence of vehicle 1500, infotainment and / or the like for vehicle 1500.
[0153] Vehicle 1500 may include one or more System-on-Chip (SoC) 1504. SoC 1504 may include CPU 1506, GPU 1508, processor 1510, cache 1512, accelerator 1514, data storage 1516, and / or other components and features not shown. SoC 1504 can be used to control vehicle 1500 across a wide variety of platforms and systems. For example, one or more SoCs 1504 may be combined with an HD map 1522 in a system (e.g., the system of vehicle 1500), the HD map being accessible from one or more servers (e.g., via a network interface 1524). Figure 15D One or more servers (1578) receive map refresh and / or updates.
[0154] CPU 1506 may include CPU clusters or CPU complexes (or, alternatively, referred to herein as "CCPLEX"). CPU 1506 may include multiple cores and / or L2 cache. For example, in some embodiments, CPU 1506 may include eight cores in a coherent multiprocessor configuration. In some embodiments, CPU 1506 may include four dual-core clusters, each cluster having a dedicated L2 cache (e.g., 2MB L2 cache). CPU 1506 (e.g., CCPLEX) may be configured to support simultaneous cluster operation, such that any combination of clusters of CPU 1506 can be active at any given time.
[0155] The CPU 1506 can implement power management capabilities including one or more of the following features: automatic clock gating of hardware blocks when idle to conserve dynamic power; clock gating of each core when the core is not actively executing instructions due to the execution of WFI / WFE instructions; independent power gating of each core; independent clock gating of each core cluster when all cores are clock-gated or power-gated; and / or independent power gating of each core cluster when all cores are power-gated. The CPU 1506 can further implement enhanced algorithms for managing power states, where allowed power states and desired wake-up times are specified, and the hardware / microcode determines the optimal power state to enter for the core, cluster, and CCPLEX. The processing core can support simplified power state entry sequences in software, with this work offloaded to the microcode.
[0156] GPU 1508 may include an integrated GPU (or, alternatively, referred to herein as an "iGPU"). GPU 1508 may be programmable and efficient for parallel workloads. In some examples, GPU 1508 may use an enhanced tensor instruction set. GPU 1508 may include one or more streaming microprocessors, wherein each streaming microprocessor may include an L1 cache (e.g., an L1 cache with at least 96KB of storage capacity), and two or more of these streaming microprocessors may share an L2 cache (e.g., an L2 cache with 512KB of storage capacity). In some embodiments, GPU 1508 may include at least eight streaming microprocessors. GPU 1508 may use a computation application programming interface (API). Furthermore, GPU 1508 may use one or more parallel computing platforms and / or programming models (e.g., NVIDIA's CUDA).
[0157] In automotive and embedded applications, the GPU 1508 can be power-optimized for optimal performance. For example, the GPU 1508 can be fabricated on FinFETs. However, this is not intended to be limiting, and the GPU 1508 can be fabricated using other semiconductor manufacturing processes. Each streaming microprocessor can combine several mixed-precision processing cores divided into multiple blocks. For example, and without limitation, 64 PF32 cores and 32 PF64 cores can be divided into four processing blocks. In such an example, each processing block can be allocated 16 FP32 cores, 8 FP64 cores, 16 INT32 cores, two mixed-precision NVIDIA Tensor cores for deep learning matrix arithmetic, an L0 instruction cache, a warp scheduler, dispatch units, and / or a 64KB register file. Furthermore, the streaming microprocessor can include independent parallel integer and floating-point data paths to leverage the mixture of computation and addressing computations for efficient workload execution. Streaming microprocessors may include independent thread scheduling capabilities to allow for finer-grained synchronization and cooperation between parallel threads. Streaming microprocessors may include combined L1 data caches and shared memory units to improve performance while simplifying programming.
[0158] The GPU 1508 may include, in some examples, a high-bandwidth memory (HBM) and / or a 16GB HBM2 memory subsystem providing peak memory bandwidth of approximately 900GB / s. In some examples, in addition to HBM memory or alternatively, synchronous graphics random access memory (SGRAM), such as fifth-generation graphics double data rate synchronous random access memory (GDDR5), may be used.
[0159] The GPU 1508 may include unified memory technology, which includes access counters to allow memory pages to be migrated more precisely to the processors that access them most frequently, thereby improving the efficiency of shared memory ranges between processors. In some examples, Address Translation Service (ATS) support can be used to allow the GPU 1508 to directly access the CPU 1506 page tables. In such examples, when the GPU 1508 Memory Management Unit (MMU) experiences a miss, the address translation request can be transferred to the CPU 1506. In response, the CPU 1506 can look up the virtual-physical mapping for the address in its page tables and transfer the translation back to the GPU 1508. Thus, unified memory technology can allow a single unified virtual address space for the memory of both the CPU 1506 and the GPU 1508, simplifying GPU 1508 programming and porting applications to the GPU 1508.
[0160] In addition, the GPU 1508 may include access counters that track how frequently the GPU 1508 accesses the memory of other processors. These access counters can help ensure that memory pages are moved to the physical memory of the processor that accesses those pages most frequently.
[0161] SoC 1504 may include any number of caches 1512, including those described herein. For example, cache 1512 may include an L3 cache available to both CPU 1506 and GPU 1508 (e.g., it is connected to both CPU 1506 and GPU 1508). Cache 1512 may include a write-back cache, which can track the state of rows, for example, using a cache coherence protocol (e.g., MEI, MESI, MSI, etc.). Depending on the embodiment, the L3 cache may include 4MB or more, but a smaller cache size may also be used.
[0162] SoC 1504 may include an arithmetic logic unit (ALU), which can be used to perform processing of any of a variety of tasks or operations related to vehicle 1500—such as processing a DNN. Additionally, SoC 1504 may include a floating-point unit (FPU)—or other mathematical coprocessor or digital coprocessor type—for performing mathematical operations within the system. For example, SoC 1504 may include one or more FPUs integrated as execution units within CPU 1506 and / or GPU 1508.
[0163] SoC 1504 may include one or more accelerators 1514 (e.g., hardware accelerators, software accelerators, or a combination thereof). For example, SoC 1504 may include a hardware acceleration cluster, which may include optimized hardware accelerators and / or large on-chip memory. This large on-chip memory (e.g., 4MB SRAM) can enable the hardware acceleration cluster to accelerate neural networks and other computations. The hardware acceleration cluster can be used to supplement GPU 1508 and offload some tasks from GPU 1508 (e.g., freeing up more cycles of GPU 1508 to perform other tasks). As an example, accelerator 1514 can be used for targeted workloads (e.g., perception, convolutional neural networks (CNNs), etc.) that are stable enough to be easily controlled for acceleration. When used herein, the term "CNN" can include all types of CNNs, including region-based or region convolutional neural networks (RCNNs) and fast RCNNs (e.g., for object detection).
[0164] Accelerator 1514 (e.g., a hardware acceleration cluster) may include a Deep Learning Accelerator (DLA). The DLA may include one or more Tensor Processing Units (TPUs) that can be configured to provide an additional 10 trillion operations per second for deep learning applications and inference. The TPU may be an accelerator configured to perform image processing functions (e.g., for CNNs, RCNNs, etc.) and optimized for performing image processing functions. The DLA may be further optimized for a specific set of neural network types and floating-point operations and inference. The DLA is designed to provide higher performance per millimeter than a general-purpose GPU and significantly outperform CPUs. The TPU can perform several functions, including single-instance convolution functions, support for INT8, INT16, and FP16 data types for both features and weights, and post-processor functions.
[0165] DLA can execute neural networks, especially CNNs, quickly and efficiently on processed or unprocessed data for any function across a wide variety of applications, such as, but not limited to: CNNs for object recognition and detection using data from camera sensors; CNNs for distance estimation using data from camera sensors; CNNs for emergency vehicle detection and recognition using data from microphones; CNNs for face recognition and vehicle owner recognition using data from camera sensors; and / or CNNs for safety and / or safety-related events.
[0166] The DLA can perform any function of the GPU 1508, and by using inference accelerators, for example, designers can target either the DLA or the GPU 1508 for any function. For instance, a designer can focus the CNN processing and floating-point operations on the DLA and leave other functions to the GPU 1508 and / or other accelerators 1514.
[0167] Accelerator 1514 (e.g., a hardware acceleration cluster) may include a programmable vision accelerator (PVA), which may alternatively be referred to herein as a computer vision accelerator. The PVA may be designed and configured to accelerate computer vision algorithms for advanced driver assistance systems (ADAS), autonomous driving, and / or augmented reality (AR) and / or virtual reality (VR) applications. The PVA can provide a balance between performance and flexibility. For example, each PVA may include, for example, but not limited to, any number of reduced instruction set computer (RISC) cores, direct memory access (DMA), and / or any number of vector processors.
[0168] RISC cores can interact with image sensors (such as the image sensor of any camera described herein), image signal processors, and / or the like. Each of these RISC cores may include any amount of memory. Depending on the embodiment, the RISC core may use any of several protocols. In some examples, the RISC core may execute a real-time operating system (RTOS). RISC cores may be implemented using one or more integrated circuit devices, application-specific integrated circuits (ASICs), and / or memory devices. For example, a RISC core may include an instruction cache and / or tightly coupled RAM.
[0169] DMA enables PVA components to access system memory independently of the CPU 1506. DMA can support any number of features to provide optimizations to the PVA, including but not limited to support for multidimensional addressing and / or circular addressing. In some examples, DMA can support addressing in up to six or more dimensions, which can include block width, block height, block depth, horizontal block step, vertical block step, and / or depth step.
[0170] A vector processor can be a programmable processor designed to efficiently and flexibly execute programming for computer vision algorithms and provide signal processing capabilities. In some examples, a PVA may include a PVA core and two vector processing subsystem partitions. The PVA core may include a processor subsystem, one or more DMA engines (e.g., two DMA engines), and / or other peripherals. The vector processing subsystem may operate as the main processing engine of the PVA and may include a vector processing unit (VPU), an instruction cache, and / or a vector memory (e.g., a VMEM). The VPU core may include a digital signal processor, such as, for example, a Single Instruction Multiple Data (SIMD) or Very Long Instruction Word (VLIW) digital signal processor. The combination of SIMD and VLIW can enhance throughput and speed.
[0171] Each of the vector processors may include an instruction cache and may be coupled to dedicated memory. Consequently, in some examples, each of the vector processors may be configured to execute independently of other vector processors. In other examples, the vector processors included in a particular PVA may be configured to employ data parallelization. For example, in some embodiments, multiple vector processors included in a single PVA may execute the same computer vision algorithm, but on different regions of an image. In other examples, vector processors included in a particular PVA may execute different computer vision algorithms simultaneously on the same image, or even different algorithms on a sequence of images or portions of an image. Among other things, any number of PVAs may be included in a hardware acceleration cluster, and any number of vector processors may be included in each of these PVAs. Furthermore, the PVA may include additional error correction code (ECC) memory to enhance overall system security.
[0172] Accelerator 1514 (e.g., a hardware acceleration cluster) may include an on-chip computer vision network and SRAM to provide high-bandwidth, low-latency SRAM for accelerator 1514. In some examples, on-chip memory may include at least 4MB of SRAM consisting of, for example, but not limited to, eight field-configurable memory blocks accessible by both PVA and DLA. Each pair of memory blocks may include an Advanced Peripheral Bus (APB) interface, configuration circuitry, a controller, and a multiplexer. Any type of memory may be used. PVA and DLA may access memory via a backbone that provides high-speed memory access to PVA and DLA. The backbone may include (e.g., using an APB) an on-chip computer vision network interconnecting PVA and DLA to memory.
[0173] On-chip computer vision networks can include interfaces that ensure both the PVA and DLA provide ready and valid signals before transmitting any control signals / addresses / data. Such interfaces can provide separate phases and channels for transmitting control signals / addresses / data, as well as burst communication for continuous data transmission. This type of interface can conform to ISO 26262 or IEC 61508 standards, but other standards and protocols can also be used.
[0174] In some examples, the SoC 1504 may include, for example, a real-time ray tracing hardware accelerator as described in U.S. Patent No. 10,885,698, issued January 5, 2021. This real-time ray tracing hardware accelerator can be used to quickly and efficiently determine the location and extent of objects (e.g., within a world model) to generate real-time visualization simulations for RADAR signal interpretation, sound propagation synthesis and / or analysis, SONAR system simulation, general wave propagation simulation, comparison with LiDAR data for localization and / or other functional purposes, and / or for other uses. In some embodiments, one or more Tree Traversal Units (TTUs) may be used to perform one or more ray tracing-related operations.
[0175] Accelerators 1514 (e.g., hardware accelerator clusters) have broad applications in autonomous driving. PVAs can be programmable vision accelerators used in critical processing stages of ADAS and autonomous vehicles. PVAs are well-suited to algorithmic domains requiring predictable processing, low power, and low latency. Therefore, PVAs perform well in semi-dense or dense rule computation, even on small datasets requiring predictable runtimes with low latency and low power. Thus, in the context of platforms for autonomous vehicles, PVAs are designed to run classical computer vision algorithms because they are efficient in object detection and integer arithmetic.
[0176] For example, according to one embodiment of this technology, PVA is used to perform computer stereo vision. In some examples, semi-global matching-based algorithms may be used, but this is not intended to be limiting. Many applications for Level 3-5 autonomous driving require instantaneous motion estimation / stereo matching (e.g., from moving structures, pedestrian recognition, lane detection, etc.). PVA can perform computer stereo vision functions on input from two monocular cameras.
[0177] In some examples, PVA can be used to perform intensive optical flow, processing raw RADAR data (e.g., using 4D Fast Fourier Transform) to provide processed RADAR. In other examples, PVA is used for time-of-flight depth processing, which, for example, involves processing raw time-of-flight data to provide processed time-of-flight data.
[0178] DLA can be used to run any type of network to enhance control and driving safety, including, for example, neural networks that output a confidence metric for each object detection. Such a confidence value can be interpreted as a probability or as providing a relative “weight” for each detection compared to other detections. This confidence value allows the system to make further decisions about which detections should be considered true positives rather than false positives. For example, the system can set a threshold for the confidence and only consider detections exceeding the threshold as true positives. In an Automatic Emergency Braking (AEB) system, false positives can cause the vehicle to automatically perform emergency braking, which is clearly undesirable. Therefore, only the most confident detections should be considered as triggers for AEB. DLA can run neural networks to regress the confidence value. The neural network can take at least some subset of parameters as its input, such as bounding box dimensions, ground plane estimates obtained (e.g. from another subsystem), outputs from inertial measurement unit (IMU) sensor 1566 related to the orientation and distance of vehicle 1500, 3D position estimates of objects obtained from the neural network and / or other sensors (e.g., LiDAR sensor 1564 or RADAR sensor 1560), etc.
[0179] SoC 1504 may include one or more data storage units 1516 (e.g., memory). The data storage unit 1516 may be on-chip memory of SoC 1504, which may store neural networks to be executed on the GPU and / or DLA. In some examples, for redundancy and security, the data storage unit 1516 may be large enough to store multiple instances of the neural network. The data storage unit 1516 may include L2 or L3 cache 1512. References to the data storage unit 1516 may include references to memory associated with the PVA, DLA, and / or other accelerators 1514 as described herein.
[0180] SoC 1504 may include one or more processors 1510 (e.g., embedded processors). Processor 1510 may include a startup and power management processor, which may be a dedicated processor and subsystem for handling startup power and management functions, as well as safety implementation. The startup and power management processor may be part of the SoC 1504 startup sequence and may provide runtime power management services. The startup power and management processor may provide clock and voltage programming, auxiliary system low-power state transitions, SoC 1504 thermal and temperature sensor management, and / or SoC 1504 power state management. Each temperature sensor may be implemented as a ring oscillator whose output frequency is proportional to the temperature, and SoC 1504 may use the ring oscillator to detect the temperature of CPU 1506, GPU 1508, and / or accelerator 1514. If it is determined that the temperature exceeds a threshold, the startup and power management processor may enter a temperature fault routine and place SoC 1504 into a lower power state and / or place vehicle 1500 into a driver-safe parking mode (e.g., safely stop vehicle 1500).
[0181] The processor 1510 may further include a set of embedded processors that can be used as an audio processing engine. The audio processing engine can be an audio subsystem that allows for full hardware support for multi-channel audio via multiple interfaces, as well as a wide and flexible range of audio I / O interfaces. In some examples, the audio processing engine is a dedicated processor core with a digital signal processor and dedicated RAM.
[0182] The processor 1510 may further include an always-on-processor engine that can provide the necessary hardware features to support low-power sensor management and wake-up use cases. This always-on-processor engine may include a processor core, tightly coupled RAM, support for peripherals (such as timers and interrupt controllers), various I / O controller peripherals, and routing logic.
[0183] The processor 1510 may further include a security cluster engine, which includes a dedicated processor subsystem for handling security management for automotive applications. The security cluster engine may include two or more processor cores, tightly coupled RAM, support for peripheral devices (e.g., timers, interrupt controllers, etc.), and / or routing logic. In secure mode, the two or more cores may operate in lockstep mode and function as a single core with comparison logic that detects any differences between their operations.
[0184] The processor 1510 may further include a real-time camera engine, which may include a dedicated processor subsystem for handling real-time camera management.
[0185] The processor 1510 may further include a high dynamic range signal processor, which may include an image signal processor, which is a hardware engine that is part of the camera processing pipeline.
[0186] Processor 1510 may include a video image compositer, which may be (e.g., implemented on a microprocessor) a processing block, implementing video post-processing functions required by the video playback application to generate the final image for the player window. The video image compositer may perform lens distortion correction on the wide-angle camera 1570, the surround camera 1574, and / or the in-cabin monitoring camera sensor. The in-cabin monitoring camera sensor is preferably monitored by a neural network running on another instance of an advanced SoC, configured to recognize in-cabin events and respond accordingly. The in-cabin system may perform lip reading to activate mobile phone services and make calls, dictate emails, change vehicle destinations, activate or change the vehicle's infotainment system and settings, or provide voice-activated web browsing. Some functions are only available to the driver when the vehicle is operating in autonomous mode and are disabled in other situations.
[0187] Video image compositers can include enhanced temporal denoising for both spatial and temporal noise reduction. For example, in the case of motion in the video, denoising appropriately weights spatial information, reducing the weight of information provided by neighboring frames. In cases where the image or part of the image does not contain motion, the temporal denoising performed by the video image compositer can use information from previous images to reduce noise in the current image.
[0188] The video image compositer can also be configured to perform stereo correction on input stereo camera frames. When the operating system desktop is in use and the GPU 1508 does not need to continuously render new surfaces, the video image compositer can be further used for user interface components. Even when the GPU 1508 is powered on and performing 3D rendering, the video image compositer can be used to offload the GPU 1508 to improve performance and responsiveness.
[0189] The SoC 1504 may further include a Mobile Industry Processor Interface (MIPI) camera serial interface, a high-speed interface, and / or a video input block that can be used for camera and related pixel input functions for receiving video and input from a camera. The SoC 1504 may further include an input / output controller that can be software-controlled and can be used to receive I / O signals not assigned to a specific role.
[0190] SoC 1504 may further include a wide range of peripheral interfaces to enable communication with peripherals, audio codecs, power management and / or other devices. SoC 1504 can be used to process data from cameras and sensors (e.g., LiDAR sensor 1564, RADAR sensor 1560, etc., which can be connected via Gigabit Multimedia Serial Link and Ethernet), data from bus 1502 (e.g., vehicle 1500 speed, steering wheel position, etc.), and data from GNSS sensor 1558 (connected via Ethernet or CAN bus). SoC 1504 may further include a dedicated high-performance, high-capacity memory controller, which may include its own DMA engine, and which can be used to free up CPU 1506 from routine data management tasks.
[0191] The SoC 1504 can be an end-to-end platform with a flexible architecture spanning Automation Levels 3-5, providing a comprehensive functional safety architecture that leverages and efficiently utilizes computer vision and ADAS technologies for diversity and redundancy, along with deep learning tools to deliver a flexible and reliable driving software stack. The SoC 1504 can be faster, more reliable, and even more energy- and space-efficient than conventional systems. For example, when combined with the CPU 1506, GPU 1508, and data storage 1516, the accelerator 1514 can provide a fast and efficient platform for Level 3-5 autonomous vehicles.
[0192] Therefore, this technology offers capabilities and functionalities that cannot be achieved through conventional systems. For example, computer vision algorithms can be executed on CPUs, which can be configured using high-level programming languages such as C to execute a wide variety of processing algorithms across a diverse range of visual data. However, CPUs often cannot meet the performance requirements of many computer vision applications, such as those related to execution time and power consumption. In particular, many CPUs cannot execute complex object detection algorithms in real time, which is a requirement for automotive ADAS applications and practical Level 3-5 autonomous vehicles.
[0193] In contrast to conventional systems, the techniques described in this paper, by providing CPU complexes, GPU complexes, and hardware acceleration clusters, allow multiple neural networks to be executed simultaneously and / or sequentially, and the results to be combined to achieve Level 3–5 autonomous driving capabilities. For example, a CNN executed on a DLA or dGPU (e.g., GPU 1520) could include text and word recognition, allowing a supercomputer to read and understand traffic signs, including those for which neural networks have not yet been specifically trained. The DLA could further include a neural network capable of recognizing, interpreting, and providing semantic understanding of the signs, and passing that semantic understanding to a path planning module running on the CPU complex.
[0194] As another example, multiple neural networks can operate simultaneously, as required for Level 3, 4, or 5 driving. For instance, a warning sign consisting of "Caution: Flashing lights indicate icy conditions," along with a light, can be interpreted independently or jointly by several neural networks. The sign itself can be recognized as a traffic sign by a deployed first neural network (e.g., a trained neural network), and the text "Flashing lights indicate icy conditions" can be interpreted by a deployed second neural network, which informs the vehicle's path planning software (preferably executing on a CPU complex) that icy conditions exist when the flashing lights are detected. The flashing lights can be identified by a deployed third neural network operating across multiple frames, which informs the vehicle's path planning software of the presence (or absence) of the flashing lights. All three neural networks can operate simultaneously, for example, within a DLA and / or on a GPU 1508.
[0195] In some examples, the CNN used for facial recognition and owner identification can use data from camera sensors to identify the presence of an authorized driver and / or owner of vehicle 1500. A processing engine always on the sensors can be used to unlock the vehicle and turn on the lights when the owner approaches the driver's door, and in safe mode, to disable the vehicle when the owner leaves. In this way, SoC 1504 provides security against theft and / or carjacking.
[0196] In another example, the CNN used for emergency vehicle detection and identification can use data from microphone 1596 to detect and identify emergency vehicle siren. In contrast to conventional systems that use a general classifier to detect siren and manually extract features, SoC 1504 uses a CNN to classify environmental and urban sounds as well as visual data. In a preferred embodiment, the CNN running on the DLA is trained to identify the relative shut-off rate of emergency vehicles (e.g., by using the Doppler effect). The CNN can also be trained to identify emergency vehicles specific to the localized area in which the vehicle operates, as identified by GNSS sensor 1558. Thus, for example, when operating in Europe, the CNN will seek to detect European siren, and when operating in the United States, the CNN will seek to identify siren only in North America. Once an emergency vehicle is detected, with the assistance of ultrasonic sensor 1562, the control program can be used to execute emergency vehicle safety routines, causing the vehicle to slow down, pull over to the side of the road, stop, and / or idle until the emergency vehicle passes.
[0197] The vehicle may include a CPU 1518 (e.g., a discrete CPU or dCPU) that can be coupled to the SoC 1504 via a high-speed interconnect (e.g., PCIe). The CPU 1518 may include, for example, an x86 processor. The CPU 1518 can be used to perform any of a wide variety of functions, including, for example, arbitrating the results of potential inconsistencies between ADAS sensors and the SoC 1504, and / or monitoring the status and health of the controller 1536 and / or the infotainment SoC 1530.
[0198] Vehicle 1500 may include a GPU 1520 (e.g., a discrete GPU or dGPU) that can be coupled to SoC 1504 via a high-speed interconnect (e.g., NVIDIA's NVLINK). GPU 1520 may provide additional artificial intelligence capabilities, for example by executing redundant and / or different neural networks, and can be used to train and / or update neural networks based on inputs (e.g., sensor data) from sensors of vehicle 1500.
[0199] Vehicle 1500 may further include a network interface 1524, which may include one or more wireless antennas 1526 (e.g., one or more wireless antennas for different communication protocols, such as cellular antennas, Bluetooth antennas, etc.). Network interface 1524 can be used to enable wireless connectivity via the Internet to the cloud (e.g., with server 1578 and / or other network devices), with other vehicles, and / or with computing devices (e.g., passenger client devices). For communication with other vehicles, a direct link can be established between the two vehicles, and / or an indirect link can be established (e.g., across networks and via the Internet). A direct link can be provided using a vehicle-to-vehicle communication link. The vehicle-to-vehicle communication link can provide vehicle 1500 with information about vehicles approaching vehicle 1500 (e.g., vehicles in front, to the side, and / or behind vehicle 1500). This functionality may be part of vehicle 1500's cooperative adaptive cruise control function.
[0200] Network interface 1524 may include a SoC that provides modulation and demodulation functions and enables controller 1536 to communicate via a wireless network. Network interface 1524 may include an RF front-end for up-conversion from baseband to RF and down-conversion from RF to baseband. Frequency conversion can be performed using known processes and / or using a superheterodyne process. In some examples, the RF front-end functionality may be provided by a separate chip. The network interface may include wireless functions for communication via LTE, WCDMA, UMTS, GSM, CDMA2000, Bluetooth, Bluetooth LE, Wi-Fi, Z-Wave, ZigBee, LoRaWAN, and / or other wireless protocols.
[0201] Vehicle 1500 may further include data storage 1528, which may include off-chip (e.g., off-chip SoC 1504) storage devices. Data storage 1528 may include one or more storage elements, including RAM, SRAM, DRAM, VRAM, flash memory, hard disk, and / or other components and / or devices capable of storing at least one bit of data.
[0202] Vehicle 1500 may further include a GNSS sensor 1558. The GNSS sensor 1558 (e.g., GPS, assisted GPS sensor, differential GPS (DGPS) sensor, etc.) is used for auxiliary mapping, sensing, occupancy grid generation, and / or path planning functions. Any number of GNSS sensors 1558 can be used, including, for example, but not limited to, GPS using a USB connector with an Ethernet-to-serial (RS-232) bridge.
[0203] Vehicle 1500 may further include a RADAR sensor 1560. The RADAR sensor 1560 can be used by vehicle 1500 for remote vehicle detection even in dark and / or inclement weather conditions. The RADAR functional safety level can be ASIL B. The RADAR sensor 1560 can use CAN and / or bus 1502 (e.g., to transmit data generated using the RADAR sensor 1560) for control and access to object tracking data, and in some examples, Ethernet access for accessing raw data. A wide variety of RADAR sensor types can be used. For example, and without limitation, the RADAR sensor 1560 can be adapted for front, rear, and side RADAR use. In some examples, a pulse Doppler RADAR sensor is used.
[0204] The RADAR sensor 1560 can include different configurations, such as long-range with a narrow field of view, short-range with a wide field of view, short-range side coverage, etc. In some examples, the long-range RADAR can be used for adaptive cruise control functions. The long-range RADAR system can provide a wide field of view (e.g., within 250m) achieved through two or more independent scans. The RADAR sensor 1560 can help distinguish between stationary and moving objects and can be used by ADAS systems for emergency braking assistance and forward collision warning. The long-range RADAR sensor can include a single-site multi-mode RADAR with multiple (e.g., six or more) fixed RADAR antennas and high-speed CAN and FlexRay interfaces. In an example with six antennas, the four central antennas can create a focused beam pattern designed to record the vehicle 1500's surroundings at higher rates with minimal traffic interference from adjacent lanes. The other two antennas can extend the field of view, enabling rapid detection of vehicles entering or leaving the vehicle 1500's lane.
[0205] As an example, a mid-range RADAR system can include a range of up to 960m (front) or 80m (rear) and a field of view of up to 42 degrees (front) or 1550 degrees (rear). Short-range RADAR systems can include, but are not limited to, RADAR sensors designed to be mounted at both ends of the rear bumper. When mounted at both ends of the rear bumper, such a RADAR sensor system can create two beams that continuously monitor blind spots behind and beside the vehicle.
[0206] Short-range RADAR systems can be used in ADAS systems for blind spot detection and / or lane change assistance.
[0207] Vehicle 1500 may further include ultrasonic sensors 1562. Ultrasonic sensors 1562, which may be positioned at the front, rear, and / or sides of vehicle 1500, can be used for parking assistance and / or creating and updating occupancy grids. A wide variety of ultrasonic sensors 1562 can be used, and different ultrasonic sensors 1562 can be used for different detection ranges (e.g., 2.5m, 4m). Ultrasonic sensors 1562 can operate at functional safety level ASIL B.
[0208] Vehicle 1500 may include a LiDAR sensor 1564. The LiDAR sensor 1564 may be used for object and pedestrian detection, emergency braking, collision avoidance, and / or other functions. The LiDAR sensor 1564 may be of functional safety level ASIL B. In some examples, vehicle 1500 may include multiple LiDAR sensors 1564 (e.g., two, four, six, etc.) that can use Ethernet (e.g., to provide data to a Gigabit Ethernet switch).
[0209] In some examples, the LiDAR sensor 1564 may be able to provide a list of objects and their distances within a 360-degree field of view. Commercially available LiDAR sensors 1564 may have an advertising range of, for example, approximately 100m, with an accuracy of 2cm-3cm, and support for 100Mbps Ethernet connectivity. In some examples, one or more non-protruding LiDAR sensors 1564 may be used. In such examples, the LiDAR sensor 1564 may be implemented as a small device that can be embedded in the front, rear, sides, and / or corners of a vehicle 1500. In such examples, the LiDAR sensor 1564 may provide a horizontal field of view of up to 120 degrees and a vertical field of view of 35 degrees, even for low-reflectivity objects, with a range of 200m. Front-mounted LiDAR sensors 1564 may be configured for a horizontal field of view between 45 degrees and 135 degrees. Figure 15B Example long-range and short-range horizontal fields of view of the LiDAR sensor 1564 with an example mounting position above the windshield are shown, but other configurations (such as including a grille-mounted LiDAR sensor 1564) are also shown. Figure 15A Configurations such as (as shown) and / or roof-mounted LiDAR scanners (e.g., for data acquisition vehicles) are also possible.
[0210] In some examples, LiDAR technologies such as 3D flash LiDAR can also be used. 3D flash LiDAR uses flashes of laser light as the emission source to illuminate the vehicle's surroundings up to approximately 200 meters. A flash LiDAR unit includes a receiver that records the laser pulse propagation time and reflected light on each pixel, which in turn corresponds to the range from the vehicle to the object. Flash LiDAR allows for the generation of highly accurate and distortion-free images of the surrounding environment using each laser flash. In some examples, four flash LiDAR sensors can be deployed, one on each side of the vehicle. Available 3D flash LiDAR systems include solid-state 3D staring array LiDAR cameras (e.g., non-scanning LiDAR devices) without moving parts other than fans. Flash LiDAR devices can use 5 nanosecond Class I (eye-safe) laser pulses per frame and can capture reflected laser light in the form of a 3D range point cloud and co-registered intensity data. By using flash LiDAR, and because flash LiDAR is a solid-state device with no moving parts, the LiDAR sensor 1564 is less susceptible to motion blur, vibration, and / or shock.
[0211] The vehicle may further include an IMU sensor 1566. In some examples, the IMU sensor 1566 may be located at the center of the rear axle of the vehicle 1500. The IMU sensor 1566 may include, for example, but not limited to, an accelerometer, a magnetometer, a gyroscope, a magnetic compass, and / or other sensor types. In some examples, such as in a six-axis application, the IMU sensor 1566 may include an accelerometer and a gyroscope, while in a nine-axis application, the IMU sensor 1566 may include an accelerometer, a gyroscope, and a magnetometer.
[0212] In some embodiments, the IMU sensor 1566 can be implemented as a miniature, high-performance GPS-assisted inertial navigation system (GPS / INS) that combines a microelectromechanical system (MEMS) inertial sensor, a high-sensitivity GPS receiver, and an advanced Kalman filter algorithm to provide estimates of position, velocity, and attitude. Thus, in some examples, the IMU sensor 1566 can enable the vehicle 1500 to estimate heading without input from a magnetic sensor by directly observing and correlating velocity changes from GPS to the IMU sensor 1566. In some examples, the IMU sensor 1566 and the GNSS sensor 1558 can be combined into a single integrated unit.
[0213] The vehicle may include a microphone 1596 placed in and / or around the vehicle 1500. Among other things, the microphone 1596 may be used for emergency vehicle detection and identification.
[0214] The vehicle may further include any number of camera types, including stereo camera 1568, wide-angle camera 1570, infrared camera 1572, surround camera 1574, long-range and / or mid-range camera 1598, and / or other camera types. These cameras can be used to capture image data around the entire perimeter of the vehicle 1500. The types of cameras used depend on the embodiment and the requirements of the vehicle 1500, and any combination of camera types can be used to provide the necessary coverage around the vehicle 1500. Furthermore, the number of cameras may vary depending on the embodiment. For example, the vehicle may include six cameras, seven cameras, ten cameras, twelve cameras, and / or another number of cameras. As an example and without limitation, these cameras may support Gigabit Multimedia Serial Link (GMSL) and / or Gigabit Ethernet. Each of the cameras is described herein with respect to... Figure 15A and Figure 15B It was described in more detail.
[0215] Vehicle 1500 may further include vibration sensor 1542. Vibration sensor 1542 can measure vibrations of vehicle components such as axles. For example, changes in vibration can indicate changes in the road surface. In another example, when two or more vibration sensors 1542 are used, differences between vibrations can be used to determine friction or slippage on the road surface (e.g., when there is a vibration difference between a power drive shaft and a free-rotating shaft).
[0216] Vehicle 1500 may include ADAS system 1538. In some examples, ADAS system 1538 may include SoC. ADAS system 1538 may include autonomous / adaptive / automatic cruise control (ACC), cooperative adaptive cruise control (CACC), forward collision warning (FCW), automatic emergency braking (AEB), lane departure warning (LDW), lane keeping assist (LKA), blind spot warning (BSW), rear cross traffic warning (RCTW), collision warning system (CWS), lane centering (LC) and / or other features and functions.
[0217] The ACC system can use a RADAR sensor 1560, a LiDAR sensor 1564, and / or a camera. The ACC system can include longitudinal ACC and / or lateral ACC. Longitudinal ACC monitors and controls the distance to vehicles immediately in front of vehicle 1500 and automatically adjusts the vehicle speed to maintain a safe distance. Lateral ACC performs distance holding and, if necessary, advises vehicle 1500 to change lanes. Lateral ACC is associated with other ADAS applications such as LCA and CWS.
[0218] CACC uses information from other vehicles, which can be received indirectly from other vehicles via a wireless link or through a network connection (e.g., via the Internet) through network interface 1524 and / or wireless antenna 1526. Direct links can be provided by vehicle-to-vehicle (V2V) communication links, while indirect links can be infrastructure-to-vehicle (I2V) communication links. Typically, the V2V communication concept provides information about vehicles immediately ahead (e.g., vehicles immediately in front of vehicle 1500 and in the same lane), while the I2V communication concept provides information about traffic further ahead. A CACC system can include either or both of these I2V and V2V information sources. Given information about vehicles ahead of vehicle 1500, CACC can be more reliable, and it has the potential to improve traffic flow and reduce road congestion.
[0219] The Forward-Looking Warning (FCW) system is designed to alert the driver to hazards, enabling the driver to take corrective action. The FCW system uses a front-facing camera and / or RADAR sensor 1560 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback such as a display, speaker, and / or vibrating components. The FCW system can provide warnings in the form of, for example, audible, visual, haptic, and / or rapid braking pulses.
[0220] An AEB (Autonomous Emergency Braking) system detects an impending forward collision with another vehicle or other object and can automatically apply the brakes if the driver does not take corrective action within a specified time or distance parameter. The AEB system can use a front-facing camera and / or RADAR sensor 1560 coupled to a dedicated processor, DSP, FPGA, and / or ASIC. When the AEB system detects a hazard, it typically first alerts the driver to take corrective action to avoid a collision, and if the driver does not take corrective action, the AEB system can automatically apply the brakes to attempt to prevent or at least mitigate the effects of the predicted collision. The AEB system may include technologies such as dynamic brake support and / or collision proximity braking.
[0221] The Lane Departure Warning (LDW) system provides visual, auditory, and / or tactile warnings, such as steering wheel or seat vibrations, to alert the driver when the vehicle crosses lane markings. When the driver indicates intentional lane departure, the LDW system is deactivated by activating a turn signal. The LDW system can utilize a front-facing camera coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback such as a display, speaker, and / or vibrating components.
[0222] The LKA system is a variation of the LDW system. If vehicle 1500 begins to leave the lane, the LKA system provides steering input or braking to correct vehicle 1500.
[0223] The BSW system detects and warns the driver of vehicles in the vehicle's blind spot. The BSW system can provide visual, auditory, and / or tactile alerts to indicate that merging or changing lanes is unsafe. The system can provide additional warnings when the driver uses turn signals. The BSW system can utilize a rear-facing camera and / or RADAR sensor 1560 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback such as a display, speaker, and / or vibrating components.
[0224] The RCTW system can provide visual, auditory, and / or tactile notifications when an object is detected outside the range of the rear camera while the vehicle is reversing. Some RCTW systems include AEB to ensure the application of the vehicle's brakes to avoid a collision. The RCTW system may use one or more rear-mounted RADAR sensors 1560 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback such as a display, speaker, and / or vibrating components.
[0225] Conventional ADAS systems can be prone to false positives, which can be annoying and distracting for the driver, but typically not catastrophic, as the ADAS system alerts the driver and allows them to determine whether a safe condition truly exists and take appropriate action. However, in the autonomous vehicle 1500, in the event of conflicting results, the vehicle 1500 itself must decide whether to heed the results from the main computer or auxiliary computer (e.g., the first controller 1536 or the second controller 1536). For example, in some embodiments, the ADAS system 1538 may be a backup and / or auxiliary computer for providing perception information to a backup computer rationality module. The backup computer rationality monitor may run redundant and diverse software on hardware components to detect faults in perception and dynamic driving tasks. The output from the ADAS system 1538 may be provided to a supervisory MCU. If the outputs from the main computer and the auxiliary computer conflict, the supervisory MCU must determine how to reconcile the conflict to ensure safe operation.
[0226] In some examples, the master computer can be configured to provide a confidence score to the supervisory MCU, indicating the master computer's confidence level in the selected result. If the confidence score exceeds a threshold, the supervisory MCU can follow the master computer's direction regardless of whether the auxiliary computer provides conflicting or inconsistent results. If the confidence score does not meet the threshold and the master and auxiliary computers indicate different results (e.g., conflict), the supervisory MCU can arbitrate between these computers to determine the appropriate result.
[0227] The supervisory MCU can be configured to run a neural network trained and configured to determine the conditions under which the auxiliary computer provides a false alarm based on outputs from both the host and auxiliary computers. Thus, the neural network in the supervisory MCU can learn when the output of the auxiliary computer can be trusted and when it cannot. For example, when the auxiliary computer is a RADAR-based FCW system, the neural network in the supervisory MCU can learn when the FCW system is identifying a metallic object that is not actually dangerous, such as a drain grid or manhole cover that triggers an alarm. Similarly, when the auxiliary computer is a camera-based LDW system, the neural network in the supervisory MCU can learn to ignore the LDW when a cyclist or pedestrian is present and lane departure is actually the safest strategy. In embodiments that include a neural network running on the supervisory MCU, the supervisory MCU may include at least one of a DLA or GPU suitable for running the neural network using associated memory. In a preferred embodiment, the supervisory MCU may include components of and / or be included as components of the SoC 1504.
[0228] In other examples, ADAS system 1538 may include an auxiliary computer that performs ADAS functions using conventional computer vision rules. This allows the auxiliary computer to use classic computer vision rules (if-then), and the presence of neural networks in the supervising MCU can improve reliability, safety, and performance. For example, diverse implementations and intentional non-identity make the entire system more fault-tolerant, especially for failures caused by software (or software-hardware interface) functionality. For instance, if a software vulnerability or bug exists in the software running on the host computer and non-identical software code running on the auxiliary computer provides the same overall result, the supervising MCU can be more confident that the overall result is correct and that the vulnerability in the software or hardware on the host computer does not cause a substantial error.
[0229] In some examples, the output of ADAS system 1538 can be fed to the perception block and / or the dynamic driving task block of the main computer. For example, if ADAS system 1538 issues a forward collision warning because an object is immediately in front, the perception block can use this information when recognizing the object. In other examples, the assistance computer can have its own neural network, which is trained and thus reduces the risk of false positives as described herein.
[0230] Vehicle 1500 may further include an infotainment SoC 1530 (e.g., an in-vehicle infotainment system (IVI)). Although illustrated and described as an SoC, the infotainment system may not be an SoC and may include two or more discrete components. The infotainment SoC 1530 may include a combination of hardware and software that can be used to provide vehicle 1500 with audio (e.g., music, personal digital assistant, navigation instructions, news, radio, etc.), video (e.g., TV, movies, streaming media, etc.), telephone (e.g., hands-free calling), network connectivity (e.g., LTE, Wi-Fi, etc.) and / or information services (e.g., navigation system, rear parking assistance, radio data system, vehicle-related information such as fuel level, total coverage distance, brake fuel level, fuel level, door opening / closing, air filter information, etc.). For example, the infotainment SoC 1530 may include a radio, disc player, navigation system, video player, USB and Bluetooth connectivity, in-vehicle computer, in-vehicle entertainment, Wi-Fi, steering wheel audio controls, hands-free voice controls, head-up display (HUD), HMI display 1534, telematics device, control panel (e.g., for controlling and / or interacting with various components, features, and / or systems) and / or other components. The infotainment SoC 1530 may further be used to provide information (e.g., visual and / or auditory) to the vehicle's users, such as information from the ADAS system 1538, autonomous driving information such as planned vehicle maneuvers, trajectories, surrounding environment information (e.g., intersection information, vehicle information, road information, etc.), and / or other information.
[0231] The infotainment SoC 1530 may include GPU functionality. The infotainment SoC 1530 can communicate with other devices, systems, and / or components of the vehicle 1500 via bus 1502 (e.g., CAN bus, Ethernet, etc.). In some examples, the infotainment SoC 1530 may be coupled to a supervisory MCU, allowing the GPU of the infotainment system to perform some autonomous driving functions in the event of a failure of the main controller 1536 (e.g., the primary and / or backup computer of the vehicle 1500). In such an example, the infotainment SoC 1530 may place the vehicle 1500 into a driver-safe parking mode as described herein.
[0232] Vehicle 1500 may further include instrument cluster 1532 (e.g., digital instrument panel, electronic instrument cluster, digital instrument panel, etc.). Instrument cluster 1532 may include a controller and / or a supercomputer (e.g., a discrete controller or supercomputer). Instrument cluster 1532 may include a set of instruments such as speedometer, fuel level, oil pressure, tachometer, odometer, turn indicator, shift position indicator, seatbelt warning light, parking brake warning light, engine malfunction indicator, airbag (SRS) system information, lighting controls, safety system controls, navigation information, etc. In some examples, information may be displayed and / or shared between infotainment SoC 1530 and instrument cluster 1532. Therefore, instrument cluster 1532 may be included as part of infotainment SoC 1530, or vice versa.
[0233] Figure 15D For cloud-based servers and according to some embodiments of this disclosure Figure 15A The following is a system diagram illustrating communication between example autonomous vehicles 1500. System 1576 may include server 1578, network 1590, and vehicles including vehicle 1500. Server 1578 may include multiple GPUs 1584(A)-1584(H) (collectively referred to herein as GPU 1584), PCIe switches 1582(A)-1582(D) (collectively referred to herein as PCIe switch 1582), and / or CPUs 1580(A)-1580(B) (collectively referred to herein as CPU 1580). GPU 1584, CPU 1580, and PCIe switches may interconnect with high-speed interconnects and / or PCIe connections 1586, such as, but not limited to, NVLink interface 1588 developed by NVIDIA. In some examples, GPU 1584 is connected via NVLink and / or NVSwitch SoC, and GPU 1584 and PCIe switch 1582 are connected via PCIe interconnect. Although the diagram illustrates eight GPUs 1584, two CPUs 1580, and two PCIe switches, it is not intended to be limiting. Depending on the embodiment, each of the servers 1578 may include any number of GPUs 1584, CPUs 1580, and / or PCIe switches. For example, each of the servers 1578 may include eight, sixteen, thirty-two, and / or more GPUs 1584.
[0234] Server 1578 can receive image data from vehicles via network 1590, representing images of unexpected or changed road conditions such as recently commenced roadworks. Server 1578 can transmit neural network 1592, updated neural network 1592, and / or map information 1594, including information about traffic and road conditions, to vehicles via network 1590. Updates to map information 1594 may include updates to HD map 1522, such as information about construction sites, potholes, bends, floods, or other obstacles. In some examples, neural network 1592, updated neural network 1592, and / or map information 1594 may have been generated from new training and / or data received from any number of vehicles in the environment and / or based on experience from training performed at a data center (e.g., using server 1578 and / or other servers).
[0235] Server 1578 can be used to train machine learning models (e.g., neural networks) based on training data. Training data can be generated using vehicles and / or generated in simulations (e.g., using a game engine). In some examples, the training data is labeled (e.g., where the neural network benefits from supervised learning) and / or undergoes other preprocessing, while in other examples, the training data is not labeled and / or preprocessed (e.g., where the neural network does not require supervised learning). Training can be performed according to any one or more categories of machine learning techniques, including but not limited to: categories such as supervised training, semi-supervised training, unsupervised training, self-learning, reinforcement learning, joint learning, transfer learning, feature learning (including principal component and cluster analysis), multilinear subspace learning, manifold learning, representation learning (including alternative dictionary learning), rule-based machine learning, anomaly detection, and any variations or combinations thereof. Once the machine learning model is trained, it can be used by the vehicle (e.g., transmitted to the vehicle via network 1590), and / or the machine learning model can be used by server 1578 to remotely monitor the vehicle.
[0236] In some examples, server 1578 can receive data from vehicles and apply that data to state-of-the-art real-time neural networks for real-time intelligent inference. Server 1578 may include a deep learning supercomputer powered by GPU 1584 and / or a dedicated AI computer, such as the DGX and DGX station machines developed by NVIDIA. However, in some examples, server 1578 may include a deep learning infrastructure in a data center that uses only CPU power.
[0237] The deep learning infrastructure of server 1578 may be capable of rapid real-time inference and can be used to assess and verify the health status of the processor, software, and / or associated hardware in vehicle 1500. For example, the deep learning infrastructure may receive periodic updates from vehicle 1500, such as image sequences and / or objects located in those image sequences by vehicle 1500 (e.g., via computer vision and / or other machine learning object classification techniques). The deep learning infrastructure may run its own neural network to identify objects and compare them with objects identified by vehicle 1500. If the results do not match and the infrastructure concludes that the AI in vehicle 1500 has malfunctioned, then server 1578 may transmit a signal to vehicle 1500 instructing the vehicle's fail-safe computer to take control, notify passengers, and complete a safe stopping operation.
[0238] For inference, server 1578 may include GPU 1584 and one or more programmable inference accelerators (such as NVIDIA's TensorRT). The combination of a GPU-powered server and inference acceleration enables real-time response. In other examples, such as where performance is less critical, CPU, FPGA, and other processor-powered servers can be used for inference.
[0239] Reasoning and training logic
[0240] One or more embodiments may be implemented using inference and / or training logic for performing inference and / or training operations. Details regarding the inference and / or training logic are provided below.
[0241] In at least one embodiment, the inference and / or training logic may include, but is not limited to, code and / or data storage for storing forward and / or output weights and / or input / output data, and / or other parameters for configuring neurons or layers of a neural network trained and / or used for inference in one or more embodiments. In at least one embodiment, the training logic may include or be coupled to code and / or data storage for storing graph code or other software to control timing and / or sequence, wherein weight and / or other parameter information is loaded to configure logic including integer and / or floating-point units (collectively, an arithmetic logic unit (ALU)). In at least one embodiment, code (such as graph code) loads weight or other parameter information into the processor ALU based on the architecture of the neural network to which the code corresponds. In at least one embodiment, the code and / or data storage is stored during training and / or inference using one or more embodiments, incorporating weight parameters and / or input / output data of each layer of a neural network trained or used in one or more embodiments during forward propagation of input / output data and / or weight parameters. In at least one embodiment, any portion of the code and / or data storage may be included within other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory.
[0242] In at least one embodiment, any portion of the code and / or data storage may be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, the code and / or data storage may be cache memory, dynamic random access memory (“DRAM”), static random access memory (“SRAM”), non-volatile memory (e.g., flash memory), or other storage. In at least one embodiment, the choice of whether the code and / or data storage is internal or external to the processor, for example, or including DRAM, SRAM, flash memory, or some other storage type, may depend on the available on-chip versus off-chip storage, the latency requirements of the training and / or inference functions being performed, the batch size of the data used in the inference and / or training of the neural network, or some combination of these factors.
[0243] In at least one embodiment, the inference and / or training logic may include, but is not limited to, code and / or data storage for storing backpropagation and / or output weights and / or input / output data corresponding to neurons or layers of a neural network trained and / or used for inference in one or more embodiments. In at least one embodiment, during training and / or inference using one or more embodiments, the code and / or data storage is stored in conjunction with the weight parameters and / or input / output data of each layer of the neural network trained or used in one or more embodiments during backpropagation of the input / output data and / or weight parameters. In at least one embodiment, the inference and / or training logic may include or be coupled to code and / or data storage for storing graph code or other software to control timing and / or sequence, wherein weight and / or other parameter information is loaded to configure logic including integer and / or floating-point units (collectively, an arithmetic logic unit (ALU)). In at least one embodiment, code (such as graph code) loads weight or other parameter information into the processor ALU based on the architecture of the neural network to which the code corresponds. In at least one embodiment, any portion of the code and / or data storage may be included together with other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory. In at least one embodiment, any portion of the code and / or data storage may be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, the code and / or data storage may be cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other storage. In at least one embodiment, the choice of whether the code and / or data storage is internal or external to the processor, for example, including DRAM, SRAM, flash memory, or some other type of storage, may depend on the available on-chip versus off-chip storage, the latency requirements of the training and / or inference functions being performed, the batch size of the data used in the inference and / or training of the neural network, or some combination of these factors.
[0244] In at least one embodiment, the code and / or data storage, and the code and / or data storage itself, may be separate storage structures. In at least one embodiment, the code and / or data storage, and the code and / or data storage itself, may be the same storage structure. In at least one embodiment, the code and / or data storage, and the code and / or data storage itself, may be partially combined and partially separated. In at least one embodiment, the code and / or data storage, and any portion thereof, may be included together with other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory.
[0245] In at least one embodiment, the inference and / or training logic may include, but is not limited to, one or more arithmetic logic units (“ALUs”) (including integer and / or floating-point units) for performing logical and / or mathematical operations at least in part based on or instructed by training and / or inference code (e.g., graph code), the results of which may produce activations stored in activation storage (e.g., output values from layers or neurons within a neural network), which are functions of input / output and / or weight parameter data stored in code and / or data storage. In at least one embodiment, activations stored in activation storage are generated based on linear algebra and / or matrix-based mathematics performed by the ALU in response to execution instructions or other code, wherein weight values stored in and / or data storage serve as operands, and other values, such as bias values, gradient information, momentum values, or other parameters or hyperparameters, may be stored in code and / or data storage or other on-chip or off-chip storage.
[0246] In at least one embodiment, one or more processors or other hardware logic devices or circuits include one or more ALUs, while in another embodiment, one or more ALUs may be external to the processor or other hardware logic device or the circuitry using them (e.g., a coprocessor). In at least one embodiment, ALUs may be included within an execution unit of a processor, or otherwise included in an ALU bank accessible by the execution unit of the processor, which may be within the same processor or distributed among different processors of different types (e.g., a central processing unit, a graphics processing unit, a fixed-function unit, etc.). In at least one embodiment, code and / or data storage, code and / or data storage, and activation storage may share the processor or other hardware logic device or circuitry, while in another embodiment, they may be in different processors or other hardware logic devices or circuitry, or some combination of the same and different processors or other hardware logic devices or circuitry. In at least one embodiment, any portion of the activation storage may be included together with other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory. Furthermore, inference and / or training code may be stored together with other code accessible to the processor or other hardware logic or circuitry, and may be retrieved and / or processed using the processor’s fetch, decode, schedule, execute, exit, and / or other logic circuitry.
[0247] In at least one embodiment, the active memory may be a cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other memory. In at least one embodiment, the active memory may be entirely or partially located within or outside one or more processors or other logic circuits. In at least one embodiment, the choice of whether the active memory is internal or external to the processor, for example, or including DRAM, SRAM, flash memory, or certain other memory types, may depend on the availability of on-chip versus off-chip memory, latency requirements for performing training and / or inference functions, batch size of data used in inference and / or training the neural network, or some combination of these factors. In at least one embodiment, the inference and / or training logic may be used in conjunction with an application-specific integrated circuit (“ASIC”), such as those from Google. Processing unit, from Graphcore TM The inference processing unit (IPU) or from Intel. (e.g., "Lake Crest") processor. In at least one embodiment, the inference and / or training logic may be used in conjunction with central processing unit ("CPU") hardware, graphics processing unit ("GPU") hardware, or other hardware (e.g., field-programmable gate array ("FPGA")).
[0248] In at least one embodiment, the inference and / or training logic may include, but is not limited to, hardware logic, wherein computational resources, along with weight values or other information corresponding to one or more layers of neurons within a neural network, are used dedicatedly or otherwise exclusively. In at least one embodiment, the inference and / or training logic may be used in conjunction with an application-specific integrated circuit (ASIC), such as those from Google. Processing unit, from Graphcore TM The inference processing unit (IPU) or from Intel. (e.g., "Lake Crest") processor. In at least one embodiment, the inference and / or training logic may be used in conjunction with central processing unit (CPU) hardware, graphics processing unit (GPU) hardware, or other hardware (e.g., field-programmable gate array (FPGA)). In at least one embodiment, the inference and / or training logic includes, but is not limited to, code and / or data storage and code and / or data storage, which may be used to store code (e.g., graph code), weight values, and / or other information, including bias values, gradient information, momentum values, and / or other parameter or hyperparameter information. In at least one embodiment, each of the code and / or data storage and code and / or data storage is associated with a dedicated computing resource (e.g., computing hardware and computing hardware). In at least one embodiment, each of the computing hardware and computing hardware includes one or more ALUs that perform mathematical functions (e.g., linear algebraic functions) only on the information stored in the code and / or data storage and code and / or data storage, and the results are stored in active memory.
[0249] In at least one embodiment, each of the code and / or data storage and the corresponding computing hardware corresponds to a different layer of the neural network, such that an activation obtained from one storage / computation pair of the code and / or data storage and computing hardware is provided as input to the next storage / computation pair of the code and / or data storage and computing hardware to reflect the conceptual organization of the neural network. In at least one embodiment, each storage / computation pair may correspond to more than one neural network layer. In at least one embodiment, additional storage / computation pairs (not shown) may be included in the inference and / or training logic, either following or parallel to the storage / computation pair.
[0250] Example computing device
[0251] Figure 16This is a block diagram suitable for implementing some embodiments of the present disclosure of an example computing device 1600. The computing device 1600 may include an interconnect system 1602 directly or indirectly coupled to the following devices: memory 1604, one or more central processing units (CPUs) 1606, one or more graphics processing units (GPUs) 1608, a communication interface 1610, input / output (I / O) ports 1612, input / output components 1614, a power supply 1616, one or more presentation components 1618 (e.g., displays), and one or more logic units 1620. In at least one embodiment, the computing device 1600 may include one or more virtual machines (VMs), and / or any component thereof may include virtual components (e.g., virtual hardware components). For a non-limiting example, one or more GPUs 1608 may include one or more vGPUs, one or more CPUs 1606 may include one or more vCPUs, and / or one or more logic units 1620 may include one or more virtual logic units. Therefore, computing device 1600 may include discrete components (e.g., a complete GPU dedicated to computing device 1600), virtual components (e.g., a portion of the GPU dedicated to computing device 1600), or a combination thereof.
[0252] although Figure 16 The various blocks are shown connected via an interconnect system 1602 with wiring, but this is not intended to be limiting and is merely for clarity. For example, in some embodiments, a presentation component 1618, such as a display device, may be considered an I / O component 1614 (e.g., if the display is a touchscreen). As another example, the CPU 1606 and / or GPU 1608 may include memory (e.g., memory 1604 may represent a storage device other than the memory of the GPU 1608, CPU 1606, and / or other components). Therefore, Figure 16 The computing devices mentioned are merely illustrative. No distinction is made between categories such as "workstation," "server," "laptop," "desktop," "tablet," "client device," "mobile device," "handheld device," "game console," "electronic control unit (ECU)," "virtual reality system," and / or other device or system types, as all of these are considered within the same category. Figure 16 Within the scope of computing devices.
[0253] Interconnect system 1602 may represent one or more links or buses, such as address buses, data buses, control buses, or combinations thereof. Interconnect system 1602 may include one or more bus or link types, such as Industry Standard Architecture (ISA) bus, Extended Industry Standard Architecture (EISA) bus, Video Electronics Standards Association (VESA) bus, Peripheral Component Interconnect (PCI) bus, Peripheral Component Interconnect Fast (PCIe) bus, and / or another type of bus or link. In some embodiments, there is a direct connection between components. For example, CPU 1606 may be directly connected to memory 1604. Furthermore, CPU 1606 may be directly connected to GPU 1608. Where there is a direct or point-to-point connection between components, interconnect system 1602 may include a PCIe link to perform the connection. In these examples, a PCI bus is not required in computing device 1600.
[0254] The memory 1604 may include any of a wide variety of computer-readable media. Computer-readable media can be any available medium that can be accessed by the computing device 1600. Computer-readable media may include volatile and non-volatile media, as well as removable and non-removable media. For example and without limitation, computer-readable media may include computer storage media and communication media.
[0255] Computer storage media may include volatile and non-volatile media and / or removable and non-removable media, implemented in any way or by any method or technique for storing information such as computer-readable instructions, data structures, program modules, and / or other data types. For example, memory 1604 may store computer-readable instructions (e.g., representing programs and / or program elements, such as an operating system). Computer storage media may include, but is not limited to, RAM, ROM, EEPROM, flash memory or other storage technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage devices, magnetic tape cassettes, magnetic tape, disk storage devices or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by computing device 1600. As used herein, computer storage media does not include the signal itself.
[0256] Computer storage media may include computer-readable instructions, data structures, program modules, and / or other data types in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium. The term "modulated data signal" may refer to a signal whose characteristics are set or altered in a manner that encodes information into that signal. For example and without limitation, computer storage media may include wired media such as wired networks or direct wired connections, and wireless media such as sound, RF, infrared, and other wireless media. Any combination of the above should also be included within the scope of computer-readable media.
[0257] CPU 1606 may be configured to execute at least some of computer-readable instructions to control one or more components of computing device 1600 to perform one or more of the methods and / or processes described herein. Each of CPU 1606 may include one or more cores (e.g., one, two, four, eight, twenty-eight, seventy-two, etc.) capable of processing a large number of software threads simultaneously. CPU 1606 may include any type of processor and may include different types of processors depending on the type of computing device 1600 implemented (e.g., processors with fewer cores for mobile devices and processors with more cores for servers). For example, depending on the type of computing device 1600, the processor may be an advanced RISC machine (ARM) processor implemented using Reduced Instruction Set Computing (RISC) or an x86 processor implemented using Complex Instruction Set Computing (CISC). In addition to one or more microprocessors or supplementary coprocessors such as math coprocessors, computing device 1600 may also include one or more CPUs 1606.
[0258] In addition to or as a replacement for CPU 1606, one or more GPUs 1608 may be configured to execute at least some of computer-readable instructions to control one or more components of computing device 1600 to perform one or more of the methods and / or processes described herein. One or more GPUs 1608 may be integrated GPUs (e.g., having one or more CPUs 1606) and / or one or more GPUs 1608 may be discrete GPUs. In embodiments, one or more GPUs 1608 may be coprocessors of one or more CPUs 1606. Computing device 1600 may use GPUs 1608 to render graphics (e.g., 3D graphics) or perform general-purpose computing. For example, one or more GPUs 1608 may be used for general-purpose computing on a GPU (GPGPU). One or more GPUs 1608 may include hundreds or thousands of cores capable of processing hundreds or thousands of software threads simultaneously. GPUs 1608 may generate pixel data for outputting an image in response to rendering commands (e.g., rendering commands received from CPU 1606 via a host interface). GPU 1608 may include graphics memory, such as display memory, for storing pixel data or any other suitable data, such as GPGPU data. Display memory may be included as part of memory 1604. One or more GPUs 1608 may include two or more GPUs operating in parallel (e.g., via a link). The link may be directly connected to the GPUs (e.g., using NVLINK) or connected via a switch (e.g., using NVSwitch). When combined, each GPU 1608 may generate pixel data or GPGPU data for different portions of the output or different outputs (e.g., the first GPU for the first image, the second GPU for the second image). Each GPU may include its own memory or may share memory with other GPUs.
[0259] In addition to or as an alternative to CPU 1606 and / or GPU 1608, logic unit 1620 may be configured to execute at least some of computer-readable instructions to control one or more components of computing device 1600 to perform one or more of the methods and / or processes described herein. In embodiments, CPU 1606, GPU 1608, and / or logic unit 1620 may execute any combination of methods, processes, and / or portions thereof, discretely or jointly. One or more logic units 1620 may be part of and / or integrated into one or more of CPU 1606 and / or GPU 1608, and / or one or more logic units 1620 may be discrete components or otherwise separate from CPU 1606 and / or GPU 1608. In embodiments, one or more logic units 1620 may be coprocessors of one or more CPUs 1606 and / or one or more GPUs 1608.
[0260] Examples of logic unit 1620 include one or more processing cores and / or components thereof, such as data processing unit (DPU), tensor core (TC), tensor processing unit (TPU), pixel vision core (PVC), vision processing unit (VPU), graphics processing cluster (GPC), texture processing cluster (TPC), streaming multiprocessor (SM), tree traversal unit (TTU), artificial intelligence accelerator (AIA), deep learning accelerator (DLA), arithmetic logic unit (ALU), application-specific integrated circuit (ASIC), floating-point unit (FPU), input / output (I / O) element, peripheral component interconnect (PCI) or peripheral component interconnect fast (PCIe) element, etc.
[0261] Communication interface 1610 may include one or more receivers, transmitters, and / or transceivers that enable computing device 1600 to communicate with other computing devices via electronic communication networks, including wired and / or wireless communications. Communication interface 1610 may include components and functions that enable communication via any of several different networks, such as wireless networks (e.g., Wi-Fi, Z-Wave, Bluetooth, Bluetooth LE, ZigBee, etc.), wired networks (e.g., communication via Ethernet or InfiniBand), low-power wide area networks (e.g., LoRaWAN, SigFox, etc.), and / or the Internet. In one or more embodiments, logic unit 1620 and / or communication interface 1610 may include one or more data processing units (DPUs) to directly transmit data received via a network and / or via interconnect system 1602 to one or more GPUs 1608 (e.g., their memory).
[0262] I / O port 1612 enables computing device 1600 to be logically coupled to other devices, including I / O component 1614, presentation component 1618, and / or other components, some of which may be built into (e.g., integrated into) computing device 1600. Illustrative I / O component 1614 includes microphones, mice, keyboards, joysticks, game pads, game controllers, satellite dish antennas, scanners, printers, wireless devices, and so on. I / O component 1614 can provide a Natural User Interface (NUI) for processing user-generated air gestures, voice, or other physiological input. In some instances, the input may be transmitted to appropriate network elements for further processing. The NUI can implement any combination of voice recognition, stylus recognition, facial recognition, biometric recognition, on-screen and adjacent-screen gesture recognition, air gestures, head and eye tracking, and touch recognition associated with the display of computing device 1600 (described in more detail below). Computing device 1600 may include depth cameras such as stereo camera systems, infrared camera systems, RGB camera systems, touchscreen technology, and combinations thereof for gesture detection and recognition. In addition, the computing device 1600 may include an accelerometer or gyroscope that enables motion detection (e.g., as part of an inertial measurement unit (IMU)). In some examples, the output of the accelerometer or gyroscope may be used by the computing device 1600 to render immersive augmented reality or virtual reality.
[0263] Power supply 1616 may include hard-wired power supply, battery power supply, or a combination thereof. Power supply 1616 may supply power to computing device 1600 so that components of computing device 1600 can operate.
[0264] The presentation component 1618 may include a display (such as a monitor, touch screen, television screen, head-up display (HUD), other display types, or combinations thereof), speakers, and / or other presentation components. The presentation component 1618 may receive data from other components (such as GPU 1608, CPU 1606, DPU, etc.) and output that data (such as as images, videos, sounds, etc.).
[0265] Example Data Center
[0266] Figure 17 An example data center 1700 that may be used in at least one embodiment of this disclosure is shown. The data center 1700 may include a data center infrastructure layer 1710, a framework layer 1720, a software layer 1730, and / or an application layer 1740.
[0267] like Figure 17As shown, the data center infrastructure layer 1710 may include a resource coordinator 1712, grouped computing resources 1714, and node computing resources (“nodes CR”) 1716(1)-1716(N), where “N” represents any complete positive integer. In at least one embodiment, nodes CR 1716(1)-1716(N) may include, but are not limited to, any number of central processing units (“CPUs”) or other processors (including DPUs, accelerators, field-programmable gate arrays (FPGAs), graphics processors or graphics processing units (GPUs), etc.), memory devices (e.g., dynamic read-only memory), storage devices (e.g., solid-state or disk drives), network input / output (“NW I / O”) devices, network switches, virtual machines (“VMs”), power modules and / or cooling modules, etc. In some embodiments, one or more node CRs from nodes CR 1716(1)-1716(N) may correspond to servers having one or more of the aforementioned computing resources. In addition, in some embodiments, nodes CR1716(1)-17161(N) may include one or more virtual components, such as vGPU, vCPU, etc., and / or one or more nodes CRs1916(1)-1916(N) may correspond to virtual machines (VMs).
[0268] In at least one embodiment, the grouped computing resources 1714 may include individual groups of nodes CR1716 housed within one or more racks (not shown), or multiple racks housed within a data center at different geographical locations (also not shown). Individual groups of nodes CR1716 within the grouped computing resources 1714 may include grouped computing, networking, memory, or storage resources that can be configured or allocated to support one or more workloads. In at least one embodiment, several nodes CR1716, including CPUs, GPUs, DPUs, and / or other processors, may be grouped within one or more racks to provide computing resources to support one or more workloads. One or more racks may also include any number of power modules, cooling modules, and / or network switches in any combination.
[0269] Resource coordinator 1712 may be configured or otherwise control one or more nodes CR1716(1)-1716(N) and / or grouped computing resources 1714. In at least one embodiment, resource coordinator 1712 may include a Software Design Infrastructure (“SDI”) management entity for data center 1700. Resource coordinator 1712 may include hardware, software, or some combination thereof.
[0270] In at least one embodiment, such as Figure 17As shown, framework layer 1720 may include job scheduler 1733, configuration manager 1734, resource manager 1736, and / or distributed file system 1738. Framework layer 1720 may include a framework of software 1732 supporting software layer 1730 and / or one or more applications 1742 supporting application layer 1740. Software 1732 or application 1742 may respectively contain web-based service software or applications, such as those provided by Amazon Web Services, Google Cloud, and Microsoft Azure. Framework layer 1720 may be, but is not limited to, a free and open-source software web application framework (such as Apache Spark™ (hereinafter “Spark”)) that can leverage distributed file system 1738 for large-scale data processing (e.g., “big data”). In at least one embodiment, job scheduler 1733 may include Spark drivers to facilitate the scheduling of workloads supported by different layers of data center 1700. Configuration manager 1734 may be able to configure different layers, such as software layer 1730 and framework layer 1720 (which includes Spark and distributed file system 1738 for supporting large-scale data processing). Resource manager 1736 may be able to manage computing resources mapped to or allocated to clusters of distributed file system 1778 and job scheduler 1733 to support distributed file system 1738 and job scheduler 1733. In at least one embodiment, clustered or grouped computing resources may include grouped computing resources 1714 in data center infrastructure layer 1710. Resource manager 1736 may coordinate with resource coordinator 1712 to manage these mapped or allocated computing resources.
[0271] In at least one embodiment, the software 1732 included in software layer 1730 may include software used in at least a portion of nodes CR1716(1)-1716(N), grouped computing resources 1714, and / or the distributed file system 1738 of framework layer 1720. One or more types of software may include, but are not limited to, internet web search software, email virus scanning software, database software, and streaming video content software.
[0272] In at least one embodiment, the application 1742 included in the application layer 1740 may include one or more types of applications used at least in part by nodes CR1716(1)-1716(N), grouped computing resources 1714, and / or the distributed file system 1738 of the framework layer 1720. One or more types of applications may include, but are not limited to, any number of genomics applications, cognitive computing, and machine learning applications, including training or inference software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), and / or other machine learning applications used in combination with one or more embodiments.
[0273] In at least one embodiment, any of the configuration manager 1734, resource manager 1736, and resource coordinator 1712 can implement any number and type of self-modification actions based on any amount and type of data obtained in any technically feasible manner. Self-modification actions can free the data center operator of data center 1700 from making potentially poor configuration decisions and potentially avoid underutilization and / or poor performance of the data center.
[0274] According to one or more embodiments described herein, data center 1700 may include tools, services, software, or other resources to train one or more machine learning models or to use one or more machine learning models to predict or infer information. For example, one or more machine learning models may be trained by using the software and / or computing resources described above with respect to data center 1700 to compute weight parameters according to a neural network architecture. In at least one embodiment, a trained or deployed machine learning model corresponding to one or more neural networks may be used to infer or predict information using the resources described above with respect to data center 1700 by using weight parameters computed through one or more training techniques, such as, but not limited to, those described herein.
[0275] In at least one embodiment, the data center 1700 may use a CPU, application-specific integrated circuit (ASIC), GPU, FPGA, and / or other hardware (or corresponding virtual computing resources) to perform training and / or inference using the aforementioned resources. Furthermore, one or more of the software and / or hardware resources described above may be configured to allow a user to train or perform services that infer information, such as image recognition, speech recognition, or other artificial intelligence services.
[0276] Example network environment
[0277] A network environment suitable for implementing embodiments of this disclosure may include one or more client devices, servers, network-attached storage (NAS), other backend devices, and / or other device types. Client devices, servers, and / or other device types (e.g., each device) may be... Figure 16 This is implemented on one or more instances of computing devices 1600—for example, each device may include similar components, features, and / or functions of one or more computing devices 1600. Furthermore, in the case of implementing backend devices (e.g., servers, NAS, etc.), the backend devices may be included as part of a data center 1700, examples of which are described in this document. Figure 17 To describe in more detail.
[0278] Components of a network environment can communicate with each other via a network, which can be wired, wireless, or both. A network can include multiple networks or one of multiple networks. For example, a network can include one or more wide area networks (WANs), one or more local area networks (LANs), one or more public networks (such as the Internet and / or the Public Switched Telephone Network (PSTN)), and / or one or more private networks. Where the network includes a wireless telecommunications network, components such as base stations, communication towers, or even access points (and other components) can provide wireless connectivity.
[0279] A compatible network environment may include one or more peer-to-peer network environments (in which case the server may not be included in the network environment) and one or more client-server network environments (in which case one or more servers may be included in the network environment). In a peer-to-peer network environment, the functionality described herein for the server can be implemented on any number of client devices.
[0280] In at least one embodiment, the network environment may include one or more cloud-based network environments, distributed computing environments, combinations thereof, etc. The cloud-based network environment may include a framework layer, a job scheduler, a resource manager, and a distributed file system implemented on one or more servers, which may include one or more core network servers and / or edge servers. The framework layer may include a framework supporting software at the software layer and / or application at the application layer. The software or application may respectively include network-based service software or applications. In embodiments, one or more client devices may use the network-based service software or applications (e.g., by accessing the service software and / or applications via one or more application programming interfaces (APIs)). The framework layer may be, but is not limited to, a free and open-source software network application framework that can use a distributed file system for large-scale data processing (e.g., "big data").
[0281] A cloud-based network environment can provide cloud computing and / or cloud storage for any combination of the computing and / or data storage functions (or one or more portions thereof) described herein. Any of these different functions can be distributed across multiple locations from a central or core server (e.g., distributed across one or more data centers at the state, region, country, global, etc.). The core server may assign at least a portion of the functionality to the edge server if the connection to the user (e.g., a client device) is relatively close to the edge server. A cloud-based network environment can be private (e.g., limited to a single organization), public (e.g., available to many organizations), and / or a combination thereof (e.g., a hybrid cloud environment).
[0282] One or more client devices may include the information described in this article. Figure 16 At least some of the components, features, and functions of one or more example computing devices 1600 described. By way of example and not limitation, the client device may be implemented as a personal computer (PC), laptop computer, mobile device, smartphone, tablet computer, smartwatch, wearable computer, personal digital assistant (PDA), MP3 player, virtual reality headset, global positioning system (GPS) or device, video player, camera, surveillance equipment or system, vehicle, ship, spacecraft, virtual machine, drone, robot, handheld communication device, hospital equipment, gaming equipment or system, entertainment system, vehicle computer system, embedded system controller, remote control, electrical appliance, consumer electronics device, workstation, edge device, any combination of these depicted devices, or any other suitable device.
[0283] This disclosure can be described in the general context of machine-usable instructions or computer code, including computer-executable instructions such as program modules, which are executed by a computer or other machine such as a personal digital assistant or other handheld device. Typically, a program module, including routines, programs, objects, components, data structures, etc., refers to code that performs a specific task or implements a specific abstract data type. This disclosure can be practiced in a wide variety of system configurations, including handheld devices, consumer electronics, general-purpose computers, more specialized computing devices, etc. This disclosure can also be practiced in distributed computing environments where tasks are performed by remote processing devices linked via a communication network.
[0284] Other variations are within the spirit of this disclosure. Therefore, although the disclosed technology is readily adaptable to various modifications and alternative constructions, certain embodiments thereof are illustrated in the accompanying drawings and have been described in detail above. However, it should be understood that the disclosure is not intended to be limited to one or more specific forms disclosed, but rather, it is intended to cover all modifications, alternative constructions, and equivalents falling within the spirit and scope of this disclosure as defined in the appended claims.
[0285] Unless otherwise stated or obviously contradicted by the context, the terms “a,” “an,” and “the,” and similar pronouns, used in the context of describing the disclosed embodiments (particularly in the context of the appended claims), should be interpreted as encompassing both singular and plural forms, rather than as definitions of the terms. Unless otherwise stated, the terms “comprising,” “having,” “including,” and “containing” should be interpreted as open-ended terms (meaning “including, but not limited to”). The term “connection” (wherein it is not modified, it refers to a physical connection) should be interpreted as partially or wholly included, attached to, or connected together, even with some intervening elements. Unless otherwise indicated herein, references to numerical ranges herein are intended only as a way of abbreviating each individual value falling within that range, and each individual value is incorporated into the specification as if it were separately described herein. Unless otherwise indicated or contradicted by the context, the use of the terms “set” (e.g., “item set”) or “subset” should be interpreted as a non-empty set comprising one or more members. Furthermore, unless otherwise indicated or contradicted by the context, the term “subset” of the corresponding set does not necessarily mean an appropriate subset of the corresponding set, but rather that the subset and the corresponding set can be equal.
[0286] Unless otherwise explicitly stated or clearly contradicted by the context, connective phrases such as “at least one of A, B, and C” or “at least one of A, B, and C” are understood in the context to generally refer to items, terms, etc., which can be A or B or C, or any non-empty subset of the set A, B, and C. For example, in an illustrative example of a set with three members, the connective phrases “at least one of A, B, and C” and “at least one of A, B, and C” refer to any of the following sets: {A}, {B}, {C}, {A, B}, {A, C}, {B, C}, {A, B, C}. Therefore, such connective language is generally not intended to imply that some embodiments require the presence of each of at least one of A, at least one of B, and at least one of C. Additionally, unless otherwise stated or contradicted by the context, the term “multiple” indicates a plural state (e.g., “multiple items” indicates multiple items). Multiple means at least two items, but more can be indicated if explicitly stated or by the context. Furthermore, unless otherwise stated or clearly understood from the context, the phrase “based on” means “at least partially based on” rather than “based on only”.
[0287] Unless otherwise indicated herein or clearly contradicted by the context, the operations of the processes described herein may be performed in any suitable order. In at least one embodiment, processes such as those described herein (or variations thereof and / or combinations thereof) are executed under the control of one or more computer systems configured with executable instructions and are implemented as code (e.g., executable instructions, one or more computer programs, or one or more application programs) executed jointly by hardware or a combination thereof on one or more processors. In at least one embodiment, the code is stored on a computer-readable storage medium, for example, in the form of a computer program comprising a plurality of instructions executable by one or more processors. In at least one embodiment, the computer-readable storage medium is a non-transitory computer-readable storage medium that excludes transient signals (e.g., propagating transient electrical or electromagnetic transmissions) but includes non-transitory data storage circuitry (e.g., buffers, caches, and queues) within a transceiver of transient signals. In at least one embodiment, code (e.g., executable code or source code) is stored on a set of one or more non-transitory computer-readable storage media (or other memory for storing executable instructions) on which executable instructions are stored, which, when executed by one or more processors of a computer system (i.e., as a result of execution), cause the computer system to perform the operations described herein. In at least one embodiment, the set of non-transitory computer-readable storage media comprises a plurality of non-transitory computer-readable storage media, and one or more of the various non-transitory storage media lack the complete code, but the plurality of non-transitory computer-readable storage media collectively store the complete code. In at least one embodiment, the executable instructions are executed such that different instructions are executed by different processors; for example, the non-transitory computer-readable storage media store the instructions, and the main central processing unit (“CPU”) executes some instructions while the graphics processing unit (“GPU”) executes other instructions. In at least one embodiment, different components of the computer system have separate processors, and the different processors execute different subsets of the instructions.
[0288] Therefore, in at least one embodiment, the computer system is configured to implement one or more services that perform the processes described herein individually or collectively, and such a computer system is configured with suitable hardware and / or software to enable the performance of the operations. Furthermore, the computer system implementing at least one embodiment of this disclosure is a single device, and in another embodiment it is a distributed computer system comprising multiple devices operating differently, such that the distributed computer system performs the operations described herein, and that no single device performs all operations.
[0289] The use of any and all examples or exemplary language (e.g., “such as”) provided herein is intended only to better illustrate embodiments of this disclosure and does not impose a limitation on the scope of the disclosure unless otherwise required. No language in the specification should be construed as indicating that any unclaimed element is essential to the practice of the disclosure.
[0290] Unless otherwise expressly stated, it will be understood that throughout this specification, terms such as “processing,” “calculation,” “operation,” “determine,” etc., refer to the actions and / or processes of a computer or computing system or similar electronic computing device that manipulate and / or convert data represented as physical quantities (e.g., electronic quantities) in the registers and / or memory of the computing system into other data similarly represented as physical quantities in the memory, registers, or other such information storage, transmission, or display devices of the computing system.
[0291] Similarly, the term "processor" can refer to any device or part of a device that processes electronic data from registers and / or memory and converts that electronic data into other electronic data that can be stored in registers and / or memory. As a non-limiting example, a "processor" can be a CPU or a GPU. A "computing platform" can include one or more processors. As used herein, a "software" process can include, for example, software and / or hardware entities that perform work over time, such as tasks, threads, and intelligent agents. Likewise, each process can refer to multiple processes that execute instructions sequentially or in parallel, continuously or intermittently. The terms "system" and "method" are used interchangeably herein, provided that a system can embody one or more methods, and a method can be considered a system.
[0292] In this document, reference may be made to obtaining, acquiring, receiving, or inputting analog or digital data into a subsystem, computer system, or computer-implemented machine. Obtaining, acquiring, receiving, or inputting analog and digital data can be accomplished in various ways, such as by receiving data as a parameter to a function call or a call to an application programming interface (API). In at least one embodiment, the process of obtaining, acquiring, receiving, or inputting analog or digital data can be accomplished by transmitting data via a serial or parallel interface. In another embodiment, the process of obtaining, acquiring, receiving, or inputting analog or digital data can be accomplished by transmitting data from a providing entity to an acquiring entity via a computer network. Reference may also be made to providing, outputting, transmitting, sending, or presenting analog or digital data. In various examples, the process of providing, outputting, transmitting, sending, or presenting analog or digital data can be implemented by transmitting data as an input or output parameter to a function call, an API, or an inter-process communication mechanism.
[0293] While the discussion above illustrates example implementations of the described technologies, other architectures can be used to implement the described functionality and are intended to fall within the scope of this disclosure. Furthermore, although specific assignments of responsibilities have been defined above for discussion purposes, various functions and responsibilities can be assigned and divided in different ways depending on the circumstances.
[0294] Furthermore, although the subject matter has been described in language specific to structural features and / or methodological actions, it should be understood that the subject matter claimed in the appended claims is not necessarily limited to the specific features or actions described. Rather, specific features and actions are disclosed as exemplary forms for implementing the claims. The subject matter of this disclosure has been described in detail herein to satisfy statutory requirements. However, the description itself is not intended to limit the scope of this disclosure. Rather, the inventors have envisioned that the claimed subject matter may also be embodied in other ways to include steps different from or similar combinations of steps described herein in conjunction with other current or future techniques. Moreover, although the terms “step” and / or “block” may be used herein to imply different elements of the method employed, these terms should not be construed as implying any particular order among or between the various steps disclosed herein, unless the order of the steps is explicitly described.
[0295] Example text support
[0296] The disclosure of this application also includes the following numbered clauses:
[0297] Clause 1. One or more processors, comprising: processing circuitry for generating an estimated three-dimensional (3D) representation of a surface in the environment of the ego machine, based at least on a set or more sets of LiDAR detection results sampled in one or more local neighborhoods along one or more predicted trajectories of the ego machine by fitting one or more height values to them.
[0298] Clause 2. One or more processors as described in Clause 1, wherein the processing circuitry is further configured to: control one or more operations of the self-machine based at least on an estimated 3D representation of the surface.
[0299] Clause 3. One or more processors according to Clause 1 or 2, wherein the processing circuitry is further configured to: refine one or more self-motion transformations of the LiDAR detection results based at least on a segmented LiDAR point cloud representing one or more static reference surfaces in the environment.
[0300] Clause 4. One or more processors according to Clause 1 or 2, wherein the processing circuitry is further configured to: refine one or more self-motion transformations of the LiDAR detection results based at least on registration of LiDAR point clouds segmented at least based on estimated heights above the ground surface.
[0301] Clause 5. One or more processors according to Clause 1 or 2, wherein the processing circuitry is further configured to: refine one or more self-motion transformations of the LiDAR detection results based at least on a segmented LiDAR point cloud of one or more points in the height band above the ground surface estimated by registration removal.
[0302] Clause 6. One or more processors according to Clause 1 or 2, wherein the processing circuitry is further configured to: refine one or more self-motion transformations for aligning the LiDAR detection results, at least based on an estimated pitch angle relative to the estimated ground surface.
[0303] Clause 7. One or more processors according to Clause 1 or 2, wherein the processing circuitry is further configured to: sample a set of trajectory points along the one or more predicted trajectories, and sample the set or more sets of LiDAR detection results within one or more specified 3D radii of at least one individual trajectory point in the set of trajectory points.
[0304] Clause 8. One or more processors according to Clause 1 or 2, wherein the processing circuitry is further configured to: sample the set or more LiDAR detection results within a direction-dependent 3D radius of at least one individual trajectory point in the one or more predicted trajectories.
[0305] Clause 9. One or more processors according to Clause 1 or 2, wherein fitting the one or more height values applies nonlinear optimization to the observed height values of the set or more LiDAR detection results sampled in at least one individual local neighborhood among the one or more local neighborhoods.
[0306] Clause 10. One or more processors according to Clause 1 or 2, wherein fitting the one or more height values applies a one-dimensional 1D nonlinear optimization to the observed height values of the set or more LiDAR detection results sampled along the one or more predicted trajectories.
[0307] Clause 11. One or more processors according to Clause 1 or 2, wherein the one or more operations of the self-machine include at least one of: adjusting the suspension system, generating a path to avoid bumps, or applying acceleration or deceleration based at least on the estimated 3D representation of the surface.
[0308] Clause 12. One or more processors as described in Clause 1 or 2, wherein said one or more processors are included in at least one of: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing simulation operations; a system for performing digital twin operations; a system for performing optical transmission simulation; a system for performing collaborative content creation of 3D assets; a system for performing deep learning operations; a system for performing remote operations; a system for performing real-time streaming; a system for generating or presenting one or more of augmented reality content, virtual reality content, or mixed reality content; a system implemented using edge devices; a system implemented using robots; a system for performing conversational AI operations; a system implementing one or more language models; a system implementing one or more large language models (LLMs); a system implementing one or more visual language models (VLMs); a system implementing one or more multimodal language models; a system for generating synthetic data; a system for generating synthetic data using AI; a system for performing one or more generative AI operations; a system containing one or more virtual machines (VMs); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources.
[0309] Clause 13. A method comprising: generating a three-dimensional 3D representation of an estimated road surface in the environment of the ego machine, based at least on a set or more sets of LiDAR detections sampled in one or more local neighborhoods along one or more predicted trajectories of the ego machine; and
[0310] Clause 14. The method according to Clause 13 further includes: controlling one or more operations of the autonomous machine based at least on the estimated 3D representation of the road surface.
[0311] Clause 15. The method according to Clause 13 or 14 further comprises: sampling a set of trajectory points along the one or more predicted trajectories, and sampling the set or more sets of LiDAR detection results within one or more specified 3D radii of at least one individual trajectory point in the set of trajectory points.
[0312] Clause 16. The method according to Clause 13 or 14, wherein fitting the one or more height values involves applying nonlinear optimization to the observed height values of the set or more LiDAR detection results sampled in at least one individual local neighborhood among the one or more local neighborhoods.
[0313] Clause 17. The method according to Clause 13 or 14, wherein the method is performed by at least one of: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing simulated operations; a system for performing digital twin operations; a system for performing optical transmission simulation; a system for performing collaborative content creation of 3D assets; a system for performing deep learning operations; a system for performing remote operations; a system for performing real-time streaming; a system for generating or presenting one or more of augmented reality content, virtual reality content, or mixed reality content; a system implemented using an edge device; a system implemented using a robot; a system for performing conversational AI operations; a system implementing one or more language models; a system implementing one or more large language models (LLMs); a system implementing one or more visual language models (VLMs); a system implementing one or more multimodal language models; a system for generating synthetic data; a system for generating synthetic data using AI; a system for performing one or more generative AI operations; a system containing one or more virtual machines (VMs); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources.
[0314] Clause 18. A system comprising: one or more processors, said one or more processors being configured to: control one or more operations of an ego machine in a simulation rendered using one or more optical transmission simulation algorithms, based at least on an estimated three-dimensional 3D representation of a road surface in the simulation environment, said 3D representation of the road surface being generated based at least on a set or more sets of simulated LiDAR detection results sampled in one or more local neighborhoods along one or more predicted trajectories of said ego machine by fitting one or more height values to the ego machine.
[0315] Clause 19. The system pursuant to Clause 18, wherein the simulation is generated at least in part using one or more content creation applications of a 3D content collaboration platform for 3D assets.
[0316] Clause 20. The system pursuant to Clause 19, wherein the simulation environment is represented using the OpenUSD format in at least one of the one or more content creation applications.
[0317] Clause 21. The system according to Clause 18, wherein fitting the one or more height values is nonlinearly optimized and applied to the simulated height values of the set or more simulated LiDAR detection results sampled in at least one individual local neighborhood among the one or more local neighborhoods.
[0318] Clause 22. The system according to Clause 18, wherein at least one of the processors is implemented in at least one of a plurality of processing nodes in a data center and is accessible by one or more remote clients via at least one of an Application Programming Interface (API) or an application plugin.
[0319] Clause 23. One or more processors, comprising: processing circuitry for generating one or more LiDAR detection results using one or more LiDAR sensors of an autonomous machine.
[0320] Clause 24. One or more processors as described in Clause 23, wherein the processing circuitry is further configured to: find one or more estimated height offsets corresponding to at least one of one or more measured distance values or one or more measured reflectance values of the one or more LiDAR detection results.
[0321] Clause 25. One or more processors as described in Clause 23 or 24, wherein the processing circuitry is further configured to: generate one or more bias-corrected LiDAR detection results based at least on removing the one or more estimated height offsets from one or more measured height values of the one or more LiDAR detection results.
[0322] Clause 26. One or more processors as described in Clauses 23, 24 or 25, wherein the processing circuitry is further configured to: control one or more operations of the ego machine based at least on the one or more bias-corrected LiDAR detection results.
[0323] Clause 27. One or more processors as described in Clauses 23, 24, 25 or 26, wherein one or more estimated altitude offsets include one or more distance-related altitude offsets.
[0324] Clause 28. One or more processors as described in Clauses 23, 24, 25 or 26, wherein the one or more estimated height offsets include one or more reflectivity-related height offsets.
[0325] Clause 29. One or more processors as described in Clauses 23, 24, 25 or 26, wherein the processing circuitry is further configured to: locate the one or more estimated height offsets from one or more data structures, the one or more data structures being binned based at least on measured distances of the one or more estimated height offsets.
[0326] Clause 30. One or more processors as described in Clauses 23, 24, 25 or 26, wherein the processing circuitry is further configured to: locate the one or more estimated height offsets from one or more data structures, the one or more data structures being binned based at least on reflectivity of the one or more estimated height offsets.
[0327] Clause 31. One or more processors as described in Clauses 23, 24, 25 or 26, wherein the processing circuitry is further configured to: generate one or more distance-related height offsets of the one or more estimated height offsets by subtracting the true height of the local neighborhood from the total observed height of accumulated LiDAR measurements of the local neighborhood based on the measured distance bins.
[0328] Clause 32. One or more processors as described in Clauses 23, 24, 25 or 26, wherein the processing circuitry is further configured to: generate one or more reflectivity-related height offsets of the one or more estimated height offsets based at least on subtracting one or more true heights of the ground surface from the total height of accumulated LiDAR detection results based on measured reflectivity bins.
[0329] Clause 33. One or more processors as described in Clauses 23, 24, 25 or 26, wherein the processing circuitry is further configured to: generate the one or more estimated height offsets based at least on the total observed height of accumulated LiDAR detection results measured within a specified measurement distance, designated as the true height.
[0330] Clause 34. One or more processors as described in Clauses 23, 24, 25 or 26, wherein the processing circuitry is further configured to: generate an estimated three-dimensional (3D) representation of a surface in the environment based at least on the one or more bias-corrected LiDAR detection results.
[0331] Clause 35. One or more processors pursuant to Clauses 23, 24, 25, or 26, wherein said one or more processors are included in at least one of: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing analog operations; a system for performing digital twin operations; a system for performing optical transmission simulation; a system for performing collaborative content creation of 3D assets; a system for performing deep learning operations; a system for performing remote operations; a system for performing real-time streaming; and a system for generating or presenting one or more of augmented reality content, virtual reality content, or mixed reality content. Systems that include: systems implemented using edge devices; systems implemented using robots; systems for performing conversational AI operations; systems that implement one or more language models; systems that implement one or more large language models (LLMs); systems that implement one or more visual language models (VLMs); systems that implement one or more multimodal language models; systems for generating synthetic data; systems for generating synthetic data using AI; systems for performing one or more generative AI operations; systems containing one or more virtual machines (VMs); systems implemented at least partially in a data center; or systems implemented at least partially using cloud computing resources.
[0332] Clause 36. A method comprising: generating one or more LiDAR detection results using one or more LiDAR sensors.
[0333] Clause 37. The method according to Clause 36 further comprises: generating one or more bias-corrected LiDAR detection results by removing, at least based on, one or more estimated height biases corresponding to at least one of the one or more measured distance values or one or more measured reflectance values of the one or more LiDAR detection results from one or more measured height values of the one or more LiDAR detection results.
[0334] Clause 38. The method according to Clause 36 or 37, wherein the one or more estimated altitude deviations include one or more distance-related altitude deviations.
[0335] Clause 39. The method according to Clause 36 or 37, wherein the one or more estimated height deviations include one or more reflectivity-related height deviations.
[0336] Clause 40. The method according to Clause 36 or 37 further comprises: searching for the one or more estimated height deviations from one or more data structures, said one or more data structures binning the one or more estimated height deviations at least based on the measured distance.
[0337] Clause 41. The method according to Clause 36 or 37 further comprises: searching for the one or more estimated height deviations from one or more data structures, said one or more data structures binning the one or more estimated height deviations at least based on reflectivity.
[0338] Clause 42. The method according to Clause 36 or 37, wherein the method is performed by at least one of: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing simulated operations; a system for performing digital twin operations; a system for performing optical transmission simulation; a system for performing collaborative content creation of 3D assets; a system for performing deep learning operations; a system for performing remote operations; a system for performing real-time streaming; a system for generating or presenting one or more of augmented reality content, virtual reality content, or mixed reality content; a system implemented using an edge device; a system implemented using a robot; a system for performing conversational AI operations; a system implementing one or more language models; a system implementing one or more large language models (LLMs); a system implementing one or more visual language models (VLMs); a system implementing one or more multimodal language models; a system for generating synthetic data; a system for generating synthetic data using AI; a system for performing one or more generative AI operations; a system containing one or more virtual machines (VMs); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources.
[0339] Clause 43. A system comprising: one or more processors, said one or more processors configured to: control one or more operations of an ego machine in a simulated environment, based at least on one or more bias-corrected LiDAR detection results, in a simulation rendered using one or more optical transmission simulation algorithms, said one or more bias-corrected LiDAR detection results being generated at least based on removing one or more estimated height biases corresponding to at least one of one or more distance values or one or more reflectivity values of said one or more simulated LiDAR detection results from one or more height values generated using one or more simulated LiDAR sensors.
[0340] Clause 44. The system pursuant to Clause 43, wherein the simulation is generated at least in part using one or more content creation applications of a 3D content collaboration platform for three-dimensional (3D) assets.
[0341] Clause 45. The system according to Clause 44, wherein the simulation environment is represented using the OpenUSD format in at least one of the one or more content creation applications.
[0342] Clause 46. The system according to Clause 43, wherein at least one of the one or more processors is implemented in at least one of a plurality of processing nodes in a data center and is accessible by one or more remote clients via at least one of an application programming interface (API) or an application plugin.
[0343] Clause 47. One or more processors, including processing circuitry, the processing circuitry being configured to: generate a surface disparity field representing estimated disparity values of surfaces in the environment, based at least on a representation of stereo image data corresponding to the environment of the self-machine using nonlinear hierarchical optimization processing.
[0344] Clause 48. One or more processors as described in Clause 47, wherein the processing circuitry is further configured to: control one or more operations of the ego machine, at least based on the surface parallax field of the surface.
[0345] Clause 49. One or more processors as described in Clause 47 or 48, wherein the processing circuitry is further configured to: generate the surface disparity field based on at least one or more weights, said one or more weights causing the nonlinear hierarchical optimization to converge to a smaller disparity.
[0346] Clause 50. One or more processors as described in Clause 47 or 48, wherein the processing circuitry is further configured to: generate the surface disparity field using one or more weights, the one or more weights emphasizing disparity values based at least on proximity to the estimated trajectory of the ego machine.
[0347] Clause 51. One or more processors as described in Clause 47 or 48, wherein the processing circuitry is further configured to: generate the surface parallax field using one or more weights, said one or more weights weakening parallax values below the detected horizon based at least on proximity to the detected horizon.
[0348] Clause 52. One or more processors as described in Clause 47 or 48, wherein the processing circuitry is further configured to: generate the surface disparity field based at least on a measurement deviation weight, the measurement deviation weight penalizing the deviation between the stereo disparity value and the estimated disparity value of the surface.
[0349] Clause 53. One or more processors according to Clause 47 or 48, wherein the processing circuitry is further configured to: generate the surface parallax field based at least on the deviation between the stereo parallax value of the stereo parallax layer pyramid and the estimated parallax value of the surface.
[0350] Clause 54. One or more processors as described in Clause 47 or 48, wherein the processing circuitry is further configured to: generate the surface disparity field based on at least one or more weights, the one or more weights emphasizing disparity values corresponding to higher intensity gradient consistency in the stereo image data.
[0351] Clause 55. One or more processors as described in Clause 47 or 48, wherein the processing circuitry is further configured to: generate the surface disparity field based at least on a measurement deviation weight, the measurement deviation weight penalizing the deviation between the optically refined stereo disparity value and the estimated disparity value of the surface.
[0352] Clause 56. One or more processors according to Clause 47 or 48, wherein the processing circuitry is further configured to: generate the surface disparity field based at least on upsampling the optically refined stereo disparity values and the estimated disparity values of the surface in the nonlinear hierarchical optimization.
[0353] Clause 57. One or more processors as described in Clause 47 or 48, wherein said one or more processors are included in at least one of: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing analog operations; a system for performing digital twin operations; a system for performing optical transmission simulation; a system for performing collaborative content creation of 3D assets; a system for performing deep learning operations; a system for performing remote operations; a system for performing real-time streaming; a system for generating or presenting one or more of augmented reality content, virtual reality content, or mixed reality content; a system implemented using an edge device; a system implemented using a robot; a system for performing conversational AI operations; a system implementing one or more language models; a system implementing one or more large language models (LLMs); a system implementing one or more visual language models (VLMs); a system implementing one or more multimodal language models; a system for generating synthetic data; a system for generating synthetic data using AI; a system for performing one or more generative AI operations; a system containing one or more virtual machines (VMs); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources.
[0354] Clause 58. A method comprising:
[0355] At least based on the representation of stereo image data corresponding to the environment of the self-machine using nonlinear hierarchical optimization processing, a ground disparity field representing the estimated disparity values of the ground surface in the environment is generated.
[0356] Clause 59. The method according to Clause 58 further includes: controlling one or more operations of the ego machine based at least on the ground parallax field.
[0357] Clause 60. The method according to Clause 58 or 59 further comprises: generating the ground disparity field based on at least one or more weights, said one or more weights causing the nonlinear hierarchical optimization to converge to a smaller disparity.
[0358] Clause 61. The method according to Clause 58 or 59 further comprises: generating the ground disparity field using one or more weights, said one or more weights emphasizing disparity values based at least on proximity to the estimated trajectory of the ego machine.
[0359] Clause 62. The method according to Clause 58 or 59 further comprises: generating the ground disparity field using one or more weights, said one or more weights weakening disparity values below the detected horizon based at least on proximity to the detected horizon.
[0360] Clause 63. The method according to Clause 58 or 59, wherein the method is performed by at least one of: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing simulated operations; a system for performing digital twin operations; a system for performing optical transmission simulation; a system for performing collaborative content creation of 3D assets; a system for performing deep learning operations; a system for performing remote operations; a system for performing real-time streaming; a system for generating or presenting one or more of augmented reality content, virtual reality content, or mixed reality content; a system implemented using an edge device; a system implemented using a robot; a system for performing conversational AI operations; a system implementing one or more language models; a system implementing one or more large language models (LLMs); a system implementing one or more visual language models (VLMs); a system implementing one or more multimodal language models; a system for generating synthetic data; a system for generating synthetic data using AI; a system for performing one or more generative AI operations; a system containing one or more virtual machines (VMs); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources.
[0361] Clause 64. A system comprising: one or more processors configured to control one or more operations of an ego machine in a simulation rendered using one or more optical transmission simulation algorithms, based at least on a ground disparity field representing estimated disparity values of a ground surface in the simulation environment, the ground disparity field being generated at least based on a representation of stereo image data corresponding to the simulation environment using nonlinear hierarchical optimization processing.
[0362] Clause 65. The system pursuant to Clause 64, wherein the simulation is generated at least in part using one or more content creation applications of a 3D content collaboration platform for 3D assets.
[0363] Clause 66. The system according to Clause 65, wherein the simulation environment is represented using the OpenUSD format in at least one of the one or more content creation applications.
[0364] Clause 67. The system according to Clause 64, wherein the one or more processors are further configured to: generate the ground disparity field based on at least one or more weights, the one or more weights causing the nonlinear hierarchical optimization to converge to a smaller disparity.
[0365] Clause 68. The system according to Clause 64, wherein at least one of the processors is implemented in at least one of a plurality of processing nodes in a data center and is accessible by one or more remote clients via at least one of an Application Programming Interface (API) or an application plugin.
[0366] Clause 69. One or more processors, comprising: processing circuitry configured to: generate a surface disparity field representing estimated disparity values of surfaces in the environment, based at least on a representation of stereo image data corresponding to the environment of the self-machine.
[0367] Clause 70. One or more processors as described in Clause 69, wherein the processing circuitry is further configured to: control one or more operations of the ego machine based at least on the surface parallax field of the surface.
[0368] Clause 71. One or more processors according to Clause 69 or 70, wherein the one or more operations include: detecting one or more obstacles based at least on the difference between a boosted representation of the surface disparity field to which a distance-dependent threshold height is applied and a stereo disparity field corresponding to the representation of the stereo image data.
[0369] Clause 72. One or more processors according to Clause 69 or 70, wherein the one or more operations include: detecting one or more obstacles based at least on a distance-related threshold disparity difference between the surface disparity field and the stereo disparity field corresponding to the representation of the stereo image data.
[0370] Clause 73. One or more processors as described in Clause 69 or 70, wherein the one or more operations include: controlling the navigation of the ego machine based on at least one of: a) determining that one or more obstacles are represented by one or more clusters of the surface parallax field satisfying a specified threshold, or b) determining that at least one or more obstacles detected based on the surface parallax field appear in at least a threshold number of frames.
[0371] Clause 74. One or more processors according to Clause 69 or 70, wherein the one or more operations include: generating an aerospace-appropriate representation based at least on radially projecting two-dimensional (2D) rays from a reference point into a representation of the surface parallax field to one or more points corresponding to one or more parallax differences for at least a specified threshold.
[0372] Clause 75. One or more processors as described in Clause 69 or 70, wherein the one or more operations include: refining the ego-motion transformation of one or more estimates of aligned LiDAR detection results based at least on an enhanced representation of the surface disparity field registered from consecutive frames.
[0373] Clause 76. One or more processors as described in Clause 69 or 70, wherein the one or more operations include: generating an estimated three-dimensional (3D) representation of a surface in the environment, at least based on the surface parallax field.
[0374] Clause 77. One or more processors as described in Clause 69 or 70, wherein the one or more operations include: generating an estimated three-dimensional (3D) representation of the surface in the environment by sampling one or more detection results generated based on upscaling the surface disparity field to 3D, at least based on one or more predicted trajectories along the ego machine.
[0375] Clause 78. One or more processors as described in Clause 69 or 70, wherein said one or more processors are included in at least one of: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing analog operations; a system for performing digital twin operations; a system for performing optical transmission simulation; a system for performing collaborative content creation of 3D assets; a system for performing deep learning operations; a system for performing remote operations; a system for performing real-time streaming; a system for generating or presenting one or more of augmented reality content, virtual reality content, or mixed reality content; a system implemented using edge devices; a system implemented using robots; a system for performing conversational AI operations; a system implementing one or more language models; a system implementing one or more large language models (LLMs); a system implementing one or more visual language models (VLMs); a system implementing one or more multimodal language models; a system for generating synthetic data; a system for generating synthetic data using AI; a system for performing one or more generative AI operations; a system containing one or more virtual machines (VMs); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources.
[0376] Clause 79. A method comprising: controlling one or more operations of a self-operating machine in the environment based at least on a ground parallax field representing estimated parallax values of the ground surface in the environment.
[0377] Clause 80. The method according to Clause 79, wherein the one or more operations comprise controlling the navigation of the ego machine based on at least one of: a) determining that one or more obstacles are represented by one or more clusters of the ground parallax field satisfying a specified threshold, or b) determining that at least one or more obstacles detected based on the ground parallax field appear in at least a threshold number of frames.
[0378] Clause 81. The method according to Clause 79, wherein the one or more operations comprise: generating an airspace-appropriate representation based at least on radially projecting two-dimensional (2D) rays from a reference point into a representation of the ground parallax field to one or more points corresponding to one or more parallax differences for at least a specified threshold.
[0379] Clause 82. The method according to Clause 79, wherein the one or more operations comprise: refining the ego-motion transformation of one or more estimates of the aligned LiDAR detection results based at least on the enhanced representation of the ground disparity field registered from consecutive frames.
[0380] Clause 83. The method according to Clause 79, wherein the one or more operations include: generating an estimated three-dimensional (3D) representation of the ground surface in the environment, at least based on the ground parallax field.
[0381] Clause 84. The method according to Clause 79, wherein the method is performed by at least one of: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing simulated operations; a system for performing digital twin operations; a system for performing optical transmission simulation; a system for performing collaborative content creation of 3D assets; a system for performing deep learning operations; a system for performing remote operations; a system for performing real-time streaming; a system for generating or presenting one or more of augmented reality content, virtual reality content, or mixed reality content; a system implemented using an edge device; a system implemented using a robot; a system for performing conversational AI operations; a system implementing one or more language models; a system implementing one or more large language models (LLMs); a system implementing one or more visual language models (VLMs); a system implementing one or more multimodal language models; a system for generating synthetic data; a system for generating synthetic data using AI; a system for performing one or more generative AI operations; a system containing one or more virtual machines (VMs); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources.
[0382] Clause 85. A system comprising: one or more processors, said one or more processors being configured to: control one or more operations of an ego machine in a simulation environment, at least based on a ground parallax field representing estimated parallax values of the ground surface in the simulation environment, in a simulation rendered using one or more optical transmission simulation algorithms.
[0383] Clause 86. The system pursuant to Clause 85, wherein the simulation is generated at least in part using one or more content creation applications of a 3D content collaboration platform for three-dimensional (3D) assets.
[0384] Clause 87. The system according to Clause 86, wherein the simulation environment is represented using the OpenUSD format in at least one of the one or more content creation applications.
[0385] Clause 88. The system according to Clause 85, wherein the one or more processors are further configured to: generate the ground parallax field based at least on a representation of stereo image data representing the simulated environment.
[0386] Clause 89. The system according to Clause 85, wherein at least one of the processors is implemented in at least one of a plurality of processing nodes in a data center and is accessible by one or more remote clients via at least one of an application programming interface (API) or an application plugin.
Claims
1. One or more processors, including processing circuitry, said processing circuitry being used to: Based at least on the representation of stereo image data corresponding to the environment of the self-machine using nonlinear hierarchical optimization processing, a surface disparity field representing the estimated disparity values of surfaces in the environment is generated; and One or more operations of the ego machine are controlled, at least based on the surface parallax field of the surface.
2. The processors according to claim 1, wherein, The processing circuit is also configured to: generate the surface disparity field based on at least one or more weights, the one or more weights causing the nonlinear hierarchical optimization to converge to a smaller disparity.
3. The processors according to claim 1, wherein, The processing circuitry is further configured to: generate the surface disparity field using one or more weights, the one or more weights emphasizing the disparity values based at least on proximity to the estimated trajectory of the ego machine.
4. The processors according to claim 1, wherein, The processing circuit is further configured to: generate the surface disparity field using one or more weights, the one or more weights being configured to weaken disparity values below the detected horizon based at least on proximity to the detected horizon.
5. The processors according to claim 1, wherein, The processing circuit is further configured to: generate the surface disparity field based at least on a measurement deviation weight, the measurement deviation weight penalizing the deviation between the stereo disparity value and the estimated disparity value of the surface.
6. The processors according to claim 1, wherein, The processing circuit is also configured to generate the surface parallax field based at least on the deviation between the stereo parallax value of the stereo parallax layer pyramid and the estimated parallax value of the surface.
7. The processors according to claim 1, wherein, The processing circuit is further configured to: generate the surface disparity field based on at least one or more weights, the one or more weights emphasizing disparity values corresponding to higher intensity gradient consistency in the stereo image data.
8. The processors according to claim 1, wherein, The processing circuit is further configured to: generate the surface disparity field based at least on a measurement deviation weight, the measurement deviation weight penalizing the deviation between the optically refined stereo disparity value and the estimated disparity value of the surface.
9. The processors according to claim 1, wherein, The processing circuit is further configured to generate the surface disparity field based at least on upsampling the optically refined stereo disparity value and the estimated disparity value of the surface in the nonlinear hierarchical optimization.
10. The processors according to claim 1, wherein, The one or more processors are included in at least one of the following: Control systems for autonomous or semi-autonomous machines; Sensing systems for autonomous or semi-autonomous machines; A system used to perform simulation operations; Systems used to perform digital twin operations; A system for performing optical transmission simulation; A system for collaborative content creation of 3D assets; A system used to perform deep learning operations; A system used to perform remote operations; Systems used for performing real-time streaming; A system for generating or presenting one or more of augmented reality content, virtual reality content, or mixed reality content; Systems implemented using edge devices; Systems implemented using robots; Systems used to perform conversational AI operations; A system that implements one or more language models; A system that implements one or more large language model LLMs; A system that implements one or more Visual Language Models (VLMs); A system that implements one or more multimodal language models; A system for generating synthetic data; Systems for generating synthetic data using AI; A system for performing one or more generative AI operations; A system containing one or more virtual machines (VMs); A system that is at least partially implemented in a data center; or A system that utilizes cloud computing resources at least in part.
11. A method comprising: Based at least on the representation of stereo image data corresponding to the environment of the self-machine using nonlinear hierarchical optimization processing, a ground disparity field representing the estimated disparity values of the ground surface in the environment is generated; and One or more operations of the self-machine are controlled based at least on the ground parallax field.
12. The method of claim 11, further comprising: The ground disparity field is generated based on at least one or more weights, which cause the nonlinear hierarchical optimization to converge to a smaller disparity.
13. The method of claim 11, further comprising: The ground disparity field is generated using one or more weights, wherein the one or more weights emphasize the disparity values based at least on their proximity to the estimated trajectory of the ego machine.
14. The method of claim 11, further comprising: The ground disparity field is generated using one or more weights, which weaken the disparity values below the detected horizon based at least on their proximity to the detected horizon.
15. The method according to claim 11, wherein, The method is performed by at least one of the following: Control systems for autonomous or semi-autonomous machines; Sensing systems for autonomous or semi-autonomous machines; A system used to perform simulation operations; Systems used to perform digital twin operations; A system for performing optical transmission simulation; A system for collaborative content creation of 3D assets; A system used to perform deep learning operations; A system used to perform remote operations; Systems used for performing real-time streaming; A system for generating or presenting one or more of augmented reality content, virtual reality content, or mixed reality content; Systems implemented using edge devices; Systems implemented using robots; Systems used to perform conversational AI operations; A system that implements one or more language models; A system that implements one or more large language model LLMs; A system that implements one or more Visual Language Models (VLMs); A system that implements one or more multimodal language models; A system for generating synthetic data; Systems for generating synthetic data using AI; A system for performing one or more generative AI operations; A system containing one or more virtual machines (VMs); A system that is at least partially implemented in a data center; or A system that utilizes cloud computing resources at least in part.
16. A system comprising: One or more processors are configured to control one or more operations of an ego machine in a simulation rendered using one or more optical transmission simulation algorithms, based at least on a ground disparity field representing estimated disparity values of the ground surface in the simulation environment, the ground disparity field being generated at least based on a representation of stereo image data corresponding to the simulation environment using nonlinear hierarchical optimization processing.
17. The system according to claim 16, wherein, The simulation was generated, at least in part, using one or more content creation applications from a 3D content collaboration platform for 3D assets.
18. The system according to claim 17, wherein, The simulation environment is represented using the OpenUSD format in at least one of the one or more content creation applications.
19. The system according to claim 16, wherein, The one or more processors are further configured to: generate the ground disparity field based on at least one or more weights, the one or more weights causing the nonlinear hierarchical optimization to converge to a smaller disparity.
20. The system according to claim 16, wherein, At least one of the processors is implemented in at least one of the multiple processing nodes in the data center and can be accessed by one or more remote clients via at least one of an application programming interface (API) or an application plugin.
Citation Information
Patent Citations
Method for programmable timeouts of tree traversal mechanisms in hardware
US10885698B2
Deep neural network for segmentation of road scenes and animate object instances for autonomous driving applications
US20210026355A1
Ground surface estimation using depth information for autonomous systems and applications
US20240028041A1