Viewpoint adaptive perception of autonomous machines and applications using analog sensor data

By training the 3D perception network to adapt to the data configured by different camera devices, the performance degradation of multi-view or 3D perception on different models is solved, and the perception effect is achieved that performs well on different models.

CN120020794APending Publication Date: 2025-05-20NVIDIA CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411637211.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-05-31
Filing Date
2024-11-15
Publication Date
2025-05-20

AI Technical Summary

Technical Problem

The prior art is difficult to achieve effective multi-view or 3D perception on different vehicle models, especially when processing data configured by different camera devices, resulting in a degradation of perceived performance.

Method used

By training a 3D perception network using real source device data and simulated source device data and simulated target device data, the network is adapted to adapt to unavailable target device data.

Benefits of technology

A universal 3D perception network that performs well on different models is realized, avoiding the high cost and impractical requirements of collecting new data for each model and developing dedicated DNNs, improving detection accuracy and reducing the use of computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120020794A_ABST
    Figure CN120020794A_ABST
Patent Text Reader

Abstract

The invention relates to viewpoint adaptive perception of autonomous machines and applications using analog sensor data. Systems and methods related to perception of viewpoint adaptation for autonomous machines and applications are disclosed. The 3D perception network may be adapted to cope with unavailable target device data by training one or more layers of the 3D perception network using simulated source device data and simulated target device data as part of a trained network. Consistency loss comparing transformed feature maps extracted from simulated source device data and simulated target device data (from top to bottom) may be used to minimize differences across training channels. Accordingly, one or more paths through the training network may be designated as a 3D perceptual network, and target device data may be applied to the 3D perceptual network to perform one or more perceptual tasks.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This application claims the benefit of U.S. Provisional Application No. 63 / 600,567, filed on Nov. 17, 2023, the content of which is incorporated herein by reference in its entirety. BACKGROUND OF THE INVENTION

[0003] Designing a system for unsupervised, autonomous, and safe driving of vehicles is extremely difficult. An autonomous vehicle should be able to perform at least as a functional equivalent of a focused driver who utilizes a perception and actuation system with the amazing ability to identify moving and stationary obstacles in a complex environment and react to them to avoid collisions with other objects or structures along the vehicle's path. Thus, the ability to detect animate objects (such as cars, pedestrians, etc.) and other parts of the environment is typically crucial for an autonomous driving perception system. Many traditional perception methods rely on deep neural networks (DNNs) to evaluate images or other sensor data for detection tasks.

[0004] One of the most significant challenges in visual perception for autonomous driving stems from the diversity of camera rig configurations implemented on different vehicle models by different vehicle manufacturers (e.g., different numbers of cameras, fields of view, positions, orientations, etc.). Unlike two-dimensional (2D) computer vision where DNNs can provide some translational invariance and thus can be somewhat robust to certain changes in camera position, multi-view or three-dimensional (3D) perception (e.g., which relies on lifting 2D (e.g., image) features to 3D space) typically suffers from performance degradation when processing data generated using different camera rigs, especially when the road is non-planar. Thus, the diversity of camera rig configurations on different vehicle models (e.g., sedans, trucks, sport utility vehicles) limits the ability to develop a general DNN that performs well on different vehicle models. Ideally, the DNN would be trained using data generated by a data collection vehicle equipped with the same camera rig as the target vehicle model (e.g., the same number of cameras, viewpoints, camera parameters) to minimize the domain gap. However, this is often infeasible in the real world. For example, collecting new data and developing a dedicated DNN for each vehicle model is highly resource-intensive and impractical, and is usually impossible for new vehicle models. Thus, there is a need for improved perception techniques that can support multi-view or 3D perception. SUMMARY OF THE INVENTION

[0005] Embodiments of the present disclosure relate to viewpoint - adapted perception for autonomous machines and applications. For example, systems and methods are disclosed for adapting a 3D perception network to unavailable target rig data. In some embodiments, by using real source rig data as well as simulated source rig data and simulated target rig data to train one or more layers of a 3D perception network that is part of a training network, the 3D perception network can be adapted to unavailable target rig data. Feature statistics extracted from real source data can be used to transform features extracted from simulated data during training. The paths of real data and simulated data through the training network can be alternately trained on real data and simulated data to update the shared weights of different paths.

[0006] Additionally or alternatively, by using simulated source rig data and simulated target rig data (without real source rig data) to train one or more layers of a 3D perception network that is part of a training network, the 3D perception network can be adapted to unavailable target rig data. A consistency loss comparing (e.g., top - down) the transformed feature maps extracted from simulated source rig data and simulated target rig data can be used to minimize differences across training channels.

[0007] Thus, one or more paths through the training network can be designated as the 3D perception network, and target rig data can be applied to the 3D perception network to perform one or more perception tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] The present system and method for viewpoint - adapted perception for autonomous machines and applications are described in detail below with reference to the accompanying drawings, in which:

[0009] Figure 1 is a diagram of an example viewpoint adaptation system according to some embodiments of the present disclosure;

[0010] Figure 2 is a diagram of an example coarse viewpoint adaptation network according to some embodiments of the present disclosure;

[0011] Figure 3 is a diagram of an example fine - grained viewpoint adaptation network according to some embodiments of the present disclosure;

[0012] Figure 4 is a diagram of an example viewpoint - adapted network according to some embodiments of the present disclosure;

[0013] Figure 5is a data flow diagram showing a method of updating one or more 3D perception networks based at least on applying real 2D sensor data and simulated 2D sensor data to the one or more 3D perception networks according to some embodiments of the present disclosure;

[0014] Figure 6 is a data flow diagram showing a method of updating one or more 3D perception networks based at least on applying simulated 2D sensor data of a simulated source sensor device and a target sensor device according to some embodiments of the present disclosure;

[0015] Figure 7A is an illustration of an example autonomous vehicle according to some embodiments of the present disclosure;

[0016] Figure 7B is according to some embodiments of the present disclosure Figure 7A an example of the camera positions and fields of view of an example autonomous vehicle;

[0017] Figure 7C is according to some embodiments of the present disclosure Figure 7A a block diagram of an example system architecture of an example autonomous vehicle;

[0018] Figure 7D is a system schematic of the communication between a cloud-based server and Figure 7A an example autonomous vehicle according to some embodiments of the present disclosure;

[0019] Figure 8 is a block diagram of an example computing device suitable for implementing some embodiments of the present disclosure; and

[0020] Figure 9 is a block diagram of an example data center suitable for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION

[0021] Systems and methods related to view-point adaptive perception for autonomous machines and applications are disclosed. For example, to adapt a 3D perception network to unavailable target device data, inter-domain or coarse view-point adaptation can be performed by training the 3D perception network using real source device data as well as simulated source device data and simulated target device data, and / or inter-domain or coarse view-point adaptation can be performed by training the 3D perception network using simulated source device data and simulated target device data. Thus, the 3D perception network can be trained to perform (e.g., top-down) object detection to identify or detect objects (e.g., cars, trucks, pedestrians, cyclists, etc.) and / or parts of the environment from (e.g., perspective) target device data. The present technology can be used to generate view-point adaptive or device-invariant DNNs for autonomous vehicles, semi-autonomous vehicles, robots, and / or other object types.

[0022] Although the present disclosure may be described with respect to an exemplary autonomous or semi-autonomous vehicle or machine 700 (alternatively referred to herein as "vehicle 700" or "automated machine 700"), examples of which are described with respect to Figures 7A - 7D it is not intended to be limiting. For example, the systems and methods described herein may be used by, but are not limited to, non-autonomous vehicles or machines, semi-autonomous vehicles or machines (e.g., in one or more advanced driver assistance systems (ADAS)), autonomous vehicles or machines, manned and unmanned robots or robotic platforms, warehouse vehicles, off-road vehicles, vehicles coupled to one or more trailers, airships, boats, shuttles, emergency response vehicles, motorcycles, electric or motorized bicycles, airplanes, construction vehicles, trains, underwater vehicles, remotely operated vehicles (e.g., drones), and / or other vehicle types. Additionally, although the present disclosure may be described with respect to 3D perception for autonomous vehicles, it is not intended to be limiting, and the systems and methods described herein may be used in augmented reality, virtual reality, mixed reality, robotics, security and surveillance, autonomous or semi-autonomous machine applications, and / or any other technical field where 3D perception or reconstruction may be used.

[0023] At a high level, a DNN may be used to perform a type of 3D perception using 2D sensor data representing a 3D environment and may rely on 3D reconstruction techniques (e.g., 2D-to-3D transformation or 2D-3D fusion) to extract or infer 3D information from one or more 2D inputs. Some example tasks involving 2D-to-3D transformation include: depth estimation from a monocular image, where the DNN predicts 3D depth values from a 2D image; and view synthesis or multi-view reconstruction, where one or more 2D images are transformed into a 3D representation and / or some other 2D view (e.g., leveraging implicit 3D spatial relationships or representations). For example, the DNN may include multiple constituent DNNs or stages linked together to sequentially process different views of the 3D environment. Example multi-view Figure 3The D perception DNN may include: an input stage that extracts features from a first view (e.g., a perspective view) or multiple first views (e.g., applying images generated from different perspectives to different input channels); transformation to a second view (e.g., a top-down or bird's-eye view); and an output stage that operates in the second view (e.g., performs class segmentation and / or instance regression). Some example tasks that integrate or fuse information from both 2D and 3D sources include: pose estimation, where the DNN uses 2D projections of observations of an object to estimate its 3D pose; and augmented reality applications, where the DNN combines real-world 2D images or videos with virtual 3D objects for tasks such as object detection and tracking, pose estimation, semantic segmentation, depth estimation, or image-based rendering. The applicable sensor (e.g., camera) devices typically depend on the specific 3D reconstruction task.

[0024] In some embodiments, real source device data as well as simulated source device data and simulated target device data can be used to train the DNN to process target device data. For example, the feature extractor of the DNN can be used as a shared feature extractor shared across different channels or branches, or can be copied into multiple channels of the shared feature extractor, including a first channel for real source device data (e.g., images generated using an available camera device configuration), a second channel for simulated source device data (e.g., simulated from the perspective of a virtual camera device configuration corresponding to the available camera device configuration), and a third channel for simulated target device data (e.g., simulated from the perspective of the target device configuration). The output of each channel can be transformed from a first (e.g., perspective) 2D view to a second (e.g., top-down) 2D view and applied to a corresponding output (e.g., classification) channel. To address the domain gap challenge, during training, the feature statistics extracted from real source data can be used to transform the features extracted from simulated data. The paths of real data and simulated data through the resulting trained network can be alternately trained on real data and simulated data to update the shared weights of different paths. This type of training can be regarded as coarse viewpoint adaptation because the update can effectively train (update) the DNN to handle (e.g., process, coordinate, etc.) relatively large differences between the source viewpoint and the target viewpoint and the corresponding device configurations (e.g., relatively large changes in the number of cameras, field of view, position, or orientation, etc.).

[0025] Additionally or alternatively, simulated source device data and simulated target device data can be used to train (update) the DNN to process target device data. For example, the feature extractor of the DNN can be replicated into multiple channels of the feature extractor, including a first channel for simulating source device data (e.g., simulated from the perspective of a virtual camera device configuration corresponding to an available camera device configuration), and a second channel for simulating target device data (e.g., simulated from the perspective of the target device configuration). The output of the first channel can be transformed from a first (e.g., perspective) 2D view to a second (e.g., top-down) 2D view, applied to an output (e.g., classification) channel, and used to compute a corresponding (e.g., cross-entropy) loss. The output of the second channel can be transformed from a first (e.g., perspective) 2D view to a second (e.g., top-down) 2D view, compared with the transformed output of the first channel, and used to compute a (e.g., top-down view or bird's-eye view) consistency loss that minimizes the differences between the transformed features across channels. Thus, the resulting training network can be trained using pairwise simulated data that represents the same scene content simulated with different device configurations, where the consistency loss effectively enforces viewpoint invariance in the viewpoint transformation. This type of training can be considered fine viewpoint adaptation, as the update can effectively train the DNN to handle relatively small differences (e.g., relatively small changes in camera position or orientation) between the source viewpoint and the target viewpoint and the corresponding device configurations.

[0026] In some embodiments, a simulation system (e.g., NVIDIA DRIVE Sim TM 、Car Learning to Act (CARLA), Grand Theft Auto V) and / or a simulation dataset (e.g., SYNTHetic Images and ANnotations Collection (SYNTHIA), Virtual Karlsruhe Institute of Technology and Toyota Technological Institute (Virtual KITTI)) can be used to generate simulated device data. Taking DRIVE Sim as an example, a virtual source camera device representing an available source camera device (e.g., including six cameras) and a virtual target camera device representing a target camera device (e.g., including five or six cameras) can be generated in DRIVE Sim and used to generate a set of simulated data (e.g., images from each camera perspective, corresponding ground truth data such as 2D / 3D bounding boxes, depth maps, optical flow, segmentation masks, etc.) for each of a plurality of simulated scenarios and / or time slices. Thus, the source camera sensor device and the target camera sensor device can be simulated and used to generate simulated training data, and the simulated training data can be used to adapt the DNN from the source viewpoint to the target viewpoint.

[0027] Accordingly, the techniques described herein can be used to adapt a DNN to detect and classify animate objects and / or environmental portions from data (e.g., perspective images) generated using a target device different from the source device used to generate the training data, and these detections and classifications can be provided to an autonomous vehicle driving stack to enable safe planning and control of the autonomous vehicle. Different from conventional methods, the present technique can be used to train or update a 3D perception DNN that performs well in different device configurations and vehicle models, even when data from the target device configuration is not available. In addition, the present technique avoids the highly resource-intensive process of collecting new data and developing a dedicated DNN for each vehicle model. Therefore, the DNN trained using the present technique should produce improved detection accuracy on the unavailable target device configuration and significantly reduce the use of computing resources compared to the prior art.

[0028] Reference Figure 1 , Figure 1 FIG. 100 is an example viewpoint adaptation system 100 according to some embodiments of the present disclosure. It should be understood that such arrangements and other arrangements described herein are presented only as examples. Other arrangements and elements (e.g., machines, interfaces, functions, sequences, function groupings, etc.) may be used in addition to or in place of the shown arrangements and elements, and some elements may be omitted entirely. Moreover, many of the elements described herein are functional entities that may be implemented as discrete or distributed components or in combination with other components and may be implemented in any suitable combination and location. The various functions performed by the entities described herein may be executed by hardware, firmware, and / or software. For example, the various functions may be executed by a processor that executes instructions stored in a memory. In some embodiments, the systems, methods, and processes described herein may use components, features, and / or functions similar to those of Figures 7A - 7D the example autonomous vehicle 700 of Figure 8 FIG. 11, Figure 9 the example computing device 800 of

[0029] As a high-level overview, the view point adaptation system 100 can adapt a 3D perception network to handle unavailable target device data. For example, the neural network development component 101 can accept, specify, construct, or otherwise identify a coarse view point adaptation network 125 and / or a fine view point adaptation network 150, which includes one or more layers of the 3D perception network to be view point adapted. The training data generator 105 can generate or otherwise identify training data 175 for the coarse view point adaptation network 125 and / or the fine view point adaptation network 150, and the training component 110 can use the training data 175 to train or update the coarse view point adaptation network 125 and / or the fine view point adaptation network 150. After training, the neural network development component 101 and / or the training component 110 can specify one or more trained layers as the view point adapted network 170.

[0030] Generally, the view point adaptation system 100 and its components can be implemented using a deep learning platform, such as TensorFlow, PyTorch, Keras, MXNet, Caffe, Microsoft Cognitive Toolkit (CNTK), Fast.ai, or Theano. Thus, the view point adaptation system 100 can include a neural network development component 101 of the deep learning platform, which provides application programming interfaces (APIs) and libraries to define and create neural network architectures. For example, the neural network development component 101 can accept inputs that specify the number and configuration of network layers (such as convolutional, recurrent, etc.), activation functions, initialized network weights, and the like. Additionally, the view point adaptation system 100 can include a training component 110 of the deep learning platform, which provides configurable training algorithms and optimization techniques. For example, the training component 110 can accept inputs that specify loss functions, learning rates, batch sizes, optimization algorithms (such as stochastic gradient descent, Adam optimizer), regularization techniques, and other parameters or algorithms to customize the desired training process.

[0031] For example, the neural network development component 101 and / or the training component 110 can be used to create the view point adapted network 170. More specifically, the neural network development component 101 can use one or more layers to accept, specify, construct, or otherwise identify the coarse view point adaptation network 125 and / or the fine view point adaptation network 150, and the one or more layers, once trained by the training component 110, can be used as at least part of the view point adapted network 170.

[0032] More specifically, the neural network development component 101 may accept, specify, or otherwise identify the coarse viewpoint adaptation network 125, the fine viewpoint adaptation network 150, and / or the viewpoint adaptation network 170 as 3D perception networks that accept or otherwise process representations of 2D sensor data (e.g., image data) representing a 3D environment. For example, a 3D perception network may accept and process representations of sensor data generated or simulated using a real sensor device or a simulated sensor device of a machine (e.g., Figures 7A - 7D example autonomous vehicle 700) of the self. A 3D perception network may rely on 3D reconstruction (e.g., 2D to 3D transformation or 2D-3D fusion) to extract or infer 3D information from one or more 2D inputs. For example, a 3D perception network may perform depth estimation, pose estimation, object detection, 3D reconstruction, mapping or localization, 3D point classification, and / or other tasks, whether in a 3D view or in multiple 2D views.

[0033] Depending on the task and / or embodiment, a 3D perception network may use any known architecture and may include any number of input paths (e.g., heads, channels, layers, branches, etc.), transformations, fusion operations, and / or output paths (e.g., heads, channels, layers, branches, etc.). Taking an autonomous machine application as an example, a 3D perception network may include an input path that includes one or more layers (e.g., of a feature extractor) of each of one or more sensors (e.g., cameras) of an autonomous machine (e.g., Figures 7A - 7D example autonomous vehicle 700), such as six external cameras. A 3D perception network may include one or more common trunks that connect the input path to the corresponding output paths, and these common trunks perform classification and / or regression to detect and / or regress the characteristics of any number of classes of objects and / or environmental parts. For example, a 3D perception network may include an output path that includes one or more layers (e.g., classifiers) for each of one or more tasks and / or classes (e.g., free space segmentation, obstacle detection, scene parsing, detection of classes, such as vehicles, vulnerable road users, traffic signs, lane lines, drivable surfaces, its subclasses, etc.). Taking an example multi-view Figure 3 D perception network including one or more feature extractors, view transformers, and one or more classifiers as an example, the feature extractor may extract image features from images generated by the corresponding cameras using a camera device, the view transformer may transform the extracted image features from a first 2D (e.g., perspective) view to a second (e.g., top-down or bird's-eye) view, and the classifier may perform (e.g., in the top-down view or bird's-eye view) classification and / or segmentation in the second view. This is only an example, and other 3D perception networks may also be used in various implementations.

[0034] Continue with the multi-view example Figure 3 D perception network, multi-view Figure 3 One or more viewpoints of the D perception network can be adapted to one or more target viewpoints represented by the target sensor device by training (or fine-tuning, updating, etc.) one or more layers (e.g., one or more feature extractors, view transformers, and / or one or more classifiers) that are part of the coarse viewpoint adaptation network 125. In Figure 1 the example shown, the coarse viewpoint adaptation network 125 includes a shared feature extractor 130 corresponding to the feature extractor of the multi-view Figure 3 D perception network, a style injector 135, a view transformer 140 corresponding to the view transformer of the multi-view Figure 3 D perception network, and a shared classifier 145 corresponding to the classifier of the multi-view Figure 3 D perception network. Figure 2 is an illustration of an example coarse viewpoint adaptation network 200, which can be used as Figure 1 the coarse viewpoint adaptation network 125.

[0035] For example, Figure 1 the neural network development component 101 of Figure 3 can use the feature extractor of the multi-view Figure 1 D perception network to be viewpoint-adapted to a shared feature extractor 230 that is shared across different channels or branches (e.g., corresponding to Figure 1 the shared feature extractor 130 of

[0036] ), or the feature extractor can be copied into multiple channels of the shared feature extractor 230. The shared feature extractor 230 can include a first channel that receives real source device data 210 (e.g., an image generated using an available camera device configuration), a second channel that receives simulated source device data 215 (e.g., simulated from the perspective of a virtual camera device configuration corresponding to the available camera device configuration), and a third channel that receives simulated target device data 220 (e.g., simulated from the perspective of the target device configuration). The view transformer 240 (e.g., corresponding to Figure 1 the view transformer 140 of Figure 1The style injector 135) can extract one or more feature statistics from the features extracted from the real source device data 210 and can use them to transform the features extracted from the simulated source device data 215 and / or the simulated target device data 220. For example, given a pair of features extracted from the simulated data and the real data The style injector 235 can perform the transformation t(·,·), for example:

[0037]

[0038] where are the mean and standard deviation of the corresponding features calculated across the spatial dimensions:

[0039]

[0040] where C, H, and W represent the channel dimension, height, and width. The superscripts s and r represent simulated and real, respectively. F chw represents the element at position (c, h, w) in F, and ε can be set to 1x10 -5 . For example, the feature extractor can include any number of layers that generate any number of feature maps, and the style injector 235 can calculate statistics from the feature maps extracted from any layer in the path of the real data and inject the calculated statistics into the corresponding feature maps extracted from the corresponding layers in the path of the simulated data. This style injection can be used to minimize the style and content gap between the features extracted from the real path and the simulated path.

[0041] Figure 1 The training component 110 of can include a coarse viewpoint adaptation component 115 that alternately trains the path of the real source device data 210 with the path of the simulated source device data 215 and / or the simulated target device data 220 to update the shared weights of the different paths. For example, the training component 110 can alternately train the different paths of the real data and the simulated data by separately evaluating the performance of each path via inference and comparison with the ground truth, and then updating the shared weights based on the combined loss. In an example implementation involving segmentation, the following training objective can be used:

[0042]

[0043] where represents the predicted segmentation map, Y represents the ground truth segmentation map, and the superscripts rs, ss, and st represent real source, simulated source, and simulated target, respectively. Thus, the coarse viewpoint adaptation component 115 can use Equation 4 to calculate the cross-entropy loss based on the real source device data 210 can use Equation 5 to calculate the cross-entropy loss based on the simulated source device data 215 and the simulated target device data 220 and the loss can be calculated using Equation 6 This loss combines the losses based on real data and simulated data. The balancing hyperparameter λ aux can be set to 1. Thus, the coarse viewpoint adaptation component 115 can: alternate between, on the one hand, applying the real source device data 210 to the coarse viewpoint adaptation network 200 and, on the other hand, applying the simulated source device data 215 and the simulated target device data 220 to the coarse viewpoint adaptation network 200, calculate the combined loss, and update the coarse viewpoint adaptation network 200 accordingly.

[0044] Additionally or alternatively, continuing the example multi-view Figure 3 D perception network, the multi-view Figure 3 D perception network can be viewpoint-adapted to handle one or more target viewpoints represented by the target sensor device by training one or more of its layers (e.g., the feature extractor, the view transformer, and / or the classifier) that are part of the fine viewpoint adaptation network 150. In Figure 1 the example shown, the fine viewpoint adaptation network 150 includes: a feature extractor 155 that corresponds to the feature extractor of the multi-view Figure 3 D perception network and / or the shared feature extractor 130 of the coarse viewpoint adaptation network 125; a view transformer 160 that corresponds to the view transformer of the multi-view Figure 3 D perception network and / or the view transformer 140 of the coarse viewpoint adaptation network 125; and one or more classifiers 165 that correspond to one or more classifiers of the multi-view Figure 3 D perception network and / or one or more shared classifiers 145 of the coarse viewpoint adaptation network 125. Figure 3 is an illustration of an example fine viewpoint adaptation network 300 that can be used as Figure 1 the fine viewpoint adaptation network 150.

[0045] For example, Figure 1 the neural network development component 101 of Figure 3 can use the feature extractor of the multi-view Figure 1 D perception network to perform viewpoint adaptation in multiple channels of the feature extractor 355 (e.g., corresponding to Figure 1 the feature extractor 155) (e.g., corresponding to one or more paths through the shared feature extractor 130). The feature extractor 355 can include one or more first channels that receive the simulated source device data 315 (e.g., simulated from the perspective of a virtual camera device configuration corresponding to the available camera device configurations), and one or more second channels that receive the simulated target device data 320 (e.g., simulated from the perspective of the target device configuration). The view transformer 340 (e.g., corresponding to Figure 1 the view transformer 140 of Figure 2The view transformer 240) can use any known technique to transform the output of each input channel (e.g., the extracted image features) from a first (e.g., perspective) 2D view to a second (e.g., top-down or bird's-eye view) 2D view, and the view transformer 340 can apply the transformed features from the first channel to the corresponding output (e.g., classification) channel (e.g., corresponding to one or more shared classifiers 145).

[0046] Figure 1 The training component 110 can include a fine viewpoint adaptation component 120, which applies the simulated source device data 315 and the simulated target device data 320, calculates one or more losses, and updates the fine viewpoint adaptation network 300 accordingly. For example, the fine viewpoint adaptation component 120 can use the transformed feature maps to calculate a consistency loss (e.g., top-down view or bird's-eye view) that minimizes the difference between the transformed features across channels, such as:

[0047]

[0048] where are the transformed (e.g., top-down view or bird's-eye view) feature maps extracted from the simulated source device data 315 and the simulated target device data 320, and C, L, and W represent the number of channels, length, and width of the feature maps, respectively. In an example implementation involving segmentation, the following training objective can be used:

[0049]

[0050] where λ bev is a balancing hyperparameter that can be selected by design. Thus, the fine viewpoint adaptation component 120 can use the pairwise simulated data to train the fine viewpoint adaptation network 200, and the pairwise simulated data represents the same scene content simulated with different device configurations, where the consistency loss effectively enforces viewpoint invariance in the viewpoint transformation.

[0051] Therefore, returning Figure 1 , the training component 110 can train the coarse viewpoint adaptation component 115 and / or the fine viewpoint adaptation component 120. After training, the trained components can be used to create the viewpoint adaptation network 170. Figure 4 is an illustration of an example viewpoint adaptation network 400 according to some embodiments of the present disclosure. Figure 4 The viewpoint adaptation network 400 of Figure 1The viewpoint-adapted network 170. Whether using coarse viewpoint adaptation or fine viewpoint adaptation, the resulting viewpoint-adapted network 170 and / or 400 can be used to accurately process real target device data, even if its training may not use any real target device data.

[0052] In some embodiments and returning to Figure 1 the training data generator 105 can generate, specify, or otherwise determine training data 175 for the coarse viewpoint adaptation network 125 and / or the fine viewpoint adaptation network 150.

[0053] For example, one or more data collection machines (e.g., data collection vehicles) can be used to collect real source device data 180, such as Figures 7A - 7D the example autonomous vehicle 700. Any known technique can be used to annotate the real source device data 180. For example, objects can be classified and / or categorized, such as by labeling different parts of the real source device data 180 based on classes (e.g., for a landscape image, parts of the image (e.g., pixels or groups of pixels) can be labeled as cars, sky, trees, roads, buildings, water, waterfalls, vehicles, buses, trucks, sedans, etc.). Whether the annotations are generated automatically and / or manually, the annotations for the training data 175 (e.g., training images) can be generated in a drawing program (e.g., annotation program), a computer-aided design (CAD) program, a labeling program, other types of programs suitable for generating annotations, and / or can be hand-drawn. Generally, the annotations can be synthetically generated (e.g., generated from a computer model or rendering), real-world generated (e.g., designed and generated from real-world data), machine-automated (e.g., using feature analysis and learning to extract features from the data and then generate labels), manually annotated (e.g., an annotator or annotation expert defining the location of the labels), and / or combinations thereof (e.g., humans formulate one or more rules or annotation conventions and the machine generates the annotations).

[0054] Thus, the training data generator 105 can access the real source device data 180 (e.g., frames of image data generated using one or more data collection vehicles), and can generate or receive corresponding ground truth training data that matches the size and dimensions of the corresponding model inputs and outputs (e.g., classification data, such as a top-down segmentation map). Thus, the training data generator 105 can associate the corresponding real input training data frames (e.g., a set of real image data frames generated using a source sensor device) with the ground truth training data (e.g., the generated or annotated classification data, such as a top-down segmentation map).

[0055] In some embodiments, the training data generator 105 includes or interacts with a simulation system that generates simulated source device data 185 and / or simulated target device data 190. The type of simulation system can depend on the use case or application. For example, the training data generator 105 can include or use a physics-based simulation engine (such as Unreal Engine or Unity3D) to create a realistic virtual environment in which simulated objects can interact. Depending on the use case and the desired training data 175, a physics-based simulation engine can be used to simulate various scenarios, such as indoor or outdoor scenes and environments. In some embodiments, the training data generator 105 can include or use a domain-specific simulation platform. Depending on the implementation, the training data generator 105 can include or interact with a simulation system that simulates the training data 175 in any suitable simulation domain, such as simulations of vehicle environments, robotic environments, augmented and virtual reality (AR / VR) environments, industrial or manufacturing environments, medical or surgical environments, gaming environments, and / or other environments. Taking vehicle simulation as an example, the training data generator 105 can include or use a domain-specific simulation platform, such as NVIDIA Drive Sim or CARLA, which are designed specifically for autonomous driving simulation. These platforms provide detailed 3D environments, realistic vehicle models, and tools for generating diverse scenarios with different weather conditions, traffic patterns, and road layouts. In some embodiments, data augmentation can be applied in such simulation systems to create variations in lighting, texture, object position, and / or other factors.

[0056] Taking DRIVE Sim as an example, DRIVE Sim can be used to generate physically accurate simulated sensor data and ground truth labels to facilitate the training of autonomous vehicle perception models. Simulation systems like DRIVE Sim can support real-time rendering, hardware-in-the-loop (HIL) simulation, and a high level of photorealism. In autonomous driving, the value of simulated data (also known as synthetic data) stems largely from the difficulty of collecting and annotating real-world data, especially in the case of rare and critical events (such as uncommon objects, far-field objects, fatal accidents, and unusual driving scenarios). In cases where the target sensor device is unavailable, the training data generator 105 can use a simulation system to generate simulated target device data 190 that represents the target sensor device.

[0057] Continuing with DRIVE Sim as an example, DRIVE Sim can ingest a representation of a specified simulation (e.g., Universal Scene Description (USD)), which includes a scene script, source and target sensor devices, and a specified set of 3D assets (e.g., models of source and / or target vehicle types). For example, the scene script can be written in a High-Level Scene Description Language (HSDL), can define specified behaviors, events, and interactions in the simulation environment, and can be converted to a compatible format (e.g., USD). The script can outline domain-specific simulation scenarios, such as traffic patterns, pedestrian movement, weather conditions, road layouts, vehicle behavior, domain randomization, etc. In this way, the training data generator 105 can run the specified simulation in the simulation system and render a set of scenes, thereby generating a frame of simulated sensor data (e.g., simulated source device data 185, simulated target device data 190) from the perspective of each simulated sensor (e.g., including lens distortion) and specified ground truth data corresponding to the outputs of the coarse viewpoint adaptation network 125, the fine viewpoint adaptation network 150, and / or the viewpoint adaptation network 170 (e.g., 2D / 3D bounding boxes, depth maps, optical flow, segmentation masks).

[0058] For example, taking an example source sensor device including six cameras, an example target sensor device including six cameras, and an example 3D perception network (e.g., coarse viewpoint adaptation network 125, fine viewpoint adaptation network 150, viewpoint adaptation network 170) that performs top-down class segmentation at its output stage as an example, the training data generator 105 can generate training data 175 that includes simulated input data (e.g., simulated images of each of the six cameras) and simulated ground truth data (e.g., annotated 3D bounding boxes, annotated top-down view segmentation masks, etc.) for each sensor device and for each of a plurality of time slices representing different simulated 3D scenes. Thus, the training data generator 105 can associate the corresponding simulated frames of the input training data (e.g., a set of frames of simulated image data simulated from the perspective of the source and / or target sensor devices) with the ground truth training data (e.g., the generated or annotated classification data, such as a top-down segmentation map).

[0059] Now referring to Figure 5 and Figure 6 , each block of the methods 500 and 600 described herein includes a computational process that can be executed using any combination of hardware, firmware, and / or software. For example, various functions can be performed by a processor executing instructions stored in a memory. These methods can also be embodied as computer-usable instructions stored on a computer storage medium. These methods can be provided by a stand-alone application, a service, or a hosted service (independently or in combination with another hosted service) or a plug-in of another product, to name just a few. Additionally, by way of example, with respect to Figure 1The viewpoint adaptation system 100 to understand methods 500 and 600. However, these methods can be additionally or alternatively performed by any one system or any combination of systems, including but not limited to the systems described herein.

[0060] Figure 5 is a data flow diagram showing a method 500 of updating a 3D perception network based at least on applying real 2D sensor data and simulated 2D sensor data to the 3D perception network according to some embodiments of the present disclosure. At block B502, method 500 includes updating a three-dimensional (3D) perception network based at least on: applying real two-dimensional (2D) sensor data generated by a source sensor device of the machine to at least a portion of the 3D perception network; and applying simulated 2D sensor data simulated based at least on a simulated source sensor device and a simulated target sensor device to at least that portion of the 3D perception network. For example, regarding Figure 1 the training component 110 (e.g., the coarse viewpoint adaptation component 115) can alternate between applying real source device data 210 to the coarse viewpoint adaptation network 200 on the one hand and applying simulated source device data 215 and simulated target device data 220 to the coarse viewpoint adaptation network 200 on the other hand, calculate a combined loss, and update the coarse viewpoint adaptation network 200 accordingly. Thus, the coarse viewpoint adaptation component 115 can alternately train the path of real source device data 210 and the path of simulated source device data 215 and / or simulated target device data 220 to update the shared weights of different paths.

[0061] Figure 6 is a data flow diagram showing a method 600 of updating a 3D perception network based at least on applying simulated 2D sensor data of a simulated source sensor device and a target sensor device according to some embodiments of the present disclosure. At block B602, method 600 includes: updating a three-dimensional (3D) perception network based at least on applying simulated two-dimensional (2D) sensor data to at least a portion of the 3D perception network, the simulated two-dimensional (2D) sensor data being generated based at least on a simulated target sensor device and a source sensor device of the machine. For example, regarding Figure 1 the fine viewpoint adaptation component 120 can apply simulated source device data 185 and simulated target device data 190, calculate one or more losses, and update the fine viewpoint adaptation network 150 based on the one or more losses. For example, the fine viewpoint adaptation component 120 can calculate a consistency loss (e.g., a top-down view or a bird's-eye view) that minimizes the difference between the transformed features across channels using the transformed feature maps, and use the consistency loss to update the fine viewpoint adaptation network 150.

[0062] The systems and methods described herein can be used by non-autonomous vehicles, semi-autonomous vehicles (e.g., in one or more Adaptive Driving Assistance Systems (ADAS)), manned and unmanned robots or robotic platforms, warehouse vehicles, off-road vehicles, vehicles coupled to one or more trailers, aircraft, boats, shuttles, emergency response vehicles, motorcycles, electric or motorized bicycles, airplanes, engineering vehicles, trains, underwater vehicles, remotely controlled vehicles (e.g., drones) and / or other vehicle types, but are not limited thereto. Additionally, the systems and methods described herein can be used for various purposes, such as but not limited to, for machine control, machine motion, machine driving, synthetic data generation, model training, perception, augmented reality, virtual reality, mixed reality, robotics, security and surveillance, simulation and digital twins, autonomous or semi-autonomous machine applications, deep learning, environmental simulation, object or participant simulation and / or digital twins, data center processing, conversational AI (Artificial Intelligence), light transport simulation (e.g., ray tracing, path tracing, etc.), collaborative content creation of 3D assets, cloud computing, generative AI and / or any other suitable applications.

[0063] The disclosed embodiments can be incorporated in a variety of different systems, such as automotive systems (e.g., control systems of autonomous or semi-autonomous machines, perception systems of autonomous or semi-autonomous machines), systems implemented using robots, aviation systems, medical systems, navigation systems, intelligent area monitoring systems, systems performing deep learning operations, systems for performing simulation operations, systems for performing digital twin operations, systems implemented using edge devices, systems containing one or more virtual machines (VMs), systems for performing synthetic data generation operations, systems at least partially implemented in a data center, systems for performing conversational AI operations, systems implementing one or more language models (e.g., one or more large language models (LLMs)), systems for performing light transport simulation, systems for performing collaborative content creation of 3D assets, systems at least partially implemented using cloud computing resources and / or other types of systems.

[0064] Example autonomous vehicle

[0065] Figure 7AIllustrated is an example autonomous vehicle 700 in accordance with some embodiments of the present disclosure. The autonomous vehicle 700 (alternatively, referred to herein as "vehicle 700") can include, but is not limited to, passenger vehicles such as cars, trucks, buses, emergency vehicles, shuttle vehicles, electric or motorized bicycles, motorcycles, fire trucks, police cars, ambulances, boats, engineering vehicles, underwater vessels, robotic vehicles, drones, airplanes, vehicles coupled to trailers (e.g., semi-trailer trucks for hauling cargo) and / or other types of vehicles (e.g., driverless and / or accommodating one or more passengers). Autonomous vehicles are typically described according to the levels of automation defined by the National Highway Traffic Safety Administration (NHTSA), a division of the United States Department of Transportation, and the Society of Automotive Engineers (SAE) in "Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles" (Standard No. J3016-201806, released June 15, 2018; Standard No. J3016-201609, released September 30, 2016; and prior and future versions of the standard). The vehicle 700 is capable of implementing functions corresponding to one or more of Levels 3-5 of the autonomous driving level. The vehicle 700 is capable of implementing functions corresponding to one or more of Levels 1-5 of the autonomous driving level. For example, depending on the embodiment, the vehicle 700 is capable of implementing the capabilities of driver assistance (Level 1), partial automation (Level 2), conditional automation (Level 3), highly automation (Level 4), and / or full automation (Level 5). The term "autonomous" as used herein can include any and / or all types of autonomy of 700 or other machines, such as full autonomy, highly autonomy, conditional autonomy, partial autonomy, providing assisted autonomy, semi-autonomy, primary autonomy, or other names.

[0066] The vehicle 700 can include components such as a chassis, a body, wheels (e.g., 2, 4, 6, 8, 18, etc.), tires, axles, and other components of the vehicle. The vehicle 700 can include a propulsion system 750, such as an internal combustion engine, a hybrid power plant, a fully electric motor, and / or another type of propulsion system. The propulsion system 750 can be connected to a driveline of the vehicle 700 that can include a transmission to effect propulsion of the vehicle 700. The propulsion system 750 can be controlled in response to signals received from a throttle / accelerator 752.

[0067] A steering system 754 that may include a steering wheel can be used to steer a vehicle 700 (e.g., along a desired path or route) while a propulsion system 750 is operating (e.g., while the vehicle is in motion). The steering system 754 can receive a signal from a steering actuator 756. For fully autonomous (Level 5) functionality, the steering wheel can be optional.

[0068] A brake sensor system 746 can be used to operate vehicle brakes in response to receiving a signal from a brake actuator 748 and / or a brake sensor.

[0069] One or more controllers 736 that may include one or more system-on-chips (SoCs) 704 ( Figure 7C ) and / or one or more GPUs can provide signals (e.g., representing commands) to one or more components and / or systems of the vehicle 700. For example, one or more controllers can send signals to operate vehicle brakes via one or more brake actuators 748, operate the steering system 754 via one or more steering actuators 756, and operate the propulsion system 750 via one or more throttles / accelerators 752. One or more controllers 736 can include one or more on-board (e.g., integrated) computing devices (e.g., supercomputers) that process sensor signals and output operation commands (e.g., signals representing commands) to enable autonomous driving and / or assist a human driver in driving the vehicle 700. One or more controllers 736 can include a first controller 736 for autonomous driving functionality, a second controller 736 for functional safety functionality, a third controller 736 for artificial intelligence functionality (e.g., computer vision), a fourth controller 736 for infotainment functionality, a fifth controller 736 for redundancy in emergency situations, and / or other controllers. In some examples, a single controller 736 can handle two or more of the above functions, two or more controllers 736 can handle a single function, and / or any combination thereof.

[0070] One or more controllers 736 may provide signals for controlling one or more components and / or systems of vehicle 700 in response to sensor data (e.g., sensor inputs) received from one or more sensors. The sensor data may be received from, for example and without limitation, a Global Navigation Satellite System (“GNSS”) sensor 758 (e.g., a Global Positioning System sensor), a RADAR sensor 760, an ultrasonic sensor 762, a LiDAR sensor 764, an Inertial Measurement Unit (IMU) sensor 766 (e.g., an accelerometer, a gyroscope, a magnetic compass, a magnetometer, etc.), a microphone 796, a stereo camera 768, a wide-angle camera 770 (e.g., a fisheye camera), an infrared camera 772, a surround camera 774 (e.g., a 360-degree camera), a long-range and / or mid-range camera 798, a speed sensor 744 (e.g., for measuring the speed of vehicle 700), a vibration sensor 742, a steering sensor 740, a brake sensor (e.g., as part of a brake sensor system 546), one or more Occupant Monitoring System (OMS) sensors 701 (e.g., one or more interior cameras), and / or other sensor types.

[0071] One or more of the controllers 736 may receive inputs (e.g., represented by input data) from the instrument cluster 732 of vehicle 700 and provide outputs (e.g., represented by output data, display data, etc.) via a Human Machine Interface (HMI) display 734, an audible annunciator, a speaker, and / or via other components of vehicle 700. These outputs may include information such as vehicle speed, rate, time, map data (e.g., Figure 7C a high-definition (HD) map 722), position data (e.g., the position of vehicle 700 on a map, for example), direction, the positions of other vehicles (e.g., occupancy grid), information about objects and object states as perceived by the controller 736, and so on. For example, the HMI display 734 may display information about the presence of one or more objects (e.g., street signs, warning signs, traffic light changes, etc.) and / or information about driving maneuvers that the vehicle has made, is making, or will make (e.g., changing lanes now, exiting 34B in two miles, etc.).

[0072] Vehicle 700 further includes a network interface 724, which may communicate over one or more networks using one or more wireless antennas 726 and / or a modem. For example, network interface 724 may be capable of communicating via Long Term Evolution (LTE), Wideband Code Division Multiple Access (WCDMA), Universal Mobile Telecommunications System (UMTS), Global System for Mobile Communications (GSM), IMT-CDMA Multi-Carrier (CDMA2000), etc. One or more wireless antennas 726 may also be used to enable communication between objects (such as vehicles, mobile devices, etc.) in an environment using one or more local area networks such as Bluetooth, Bluetooth Low Energy (LE), Z-Wave, ZigBee, etc. and / or one or more low power wide area networks (LPWANs) such as LoRaWAN, SigFox, etc.

[0073] Figure 7B For an example autonomous vehicle 700 according to some embodiments of the present disclosure for Figure 7A an example camera location and field of view. The cameras and respective fields of view are one example embodiment and are not intended to be limiting. For example, additional and / or alternative cameras may be included, and / or these cameras may be located at different positions on vehicle 700.

[0074] The camera type for the cameras may include, but is not limited to, digital cameras that may be adapted to be used with components and / or systems of vehicle 700. The cameras may operate under Automotive Safety Integrity Level (ASIL) B and / or under another ASIL. The camera type may have any image capture rate, such as 60 frames per second (fps), 120 fps, 240 fps, etc., depending on the embodiment. The cameras may be capable of using a rolling shutter, a global shutter, another type of shutter, or a combination thereof. In some examples, the color filter array may include a Red Clear Clear Clear (RCCC) color filter array, a Red Clear Clear Blue (RCCB) color filter array, a Red Blue Green Clear (RBGC) color filter array, a Foveon X3 color filter array, a Bayer sensor (RGGB) color filter array, a monochrome sensor color filter array, and / or another type of color filter array. In some embodiments, clear pixel cameras such as those with RCCC, RCCB, and / or RBGC color filter arrays may be used in an effort to increase light sensitivity.

[0075] In some examples, one or more of the cameras may be used to perform Advanced Driver Assistance System (ADAS) functions (e.g., as part of a redundant or fail-safe design). For example, a multi-functional monocular camera may be installed to provide functions including lane departure warning, traffic sign assistance, and intelligent headlight control. One or more of the cameras (e.g., all cameras) may record and provide image data (e.g., video) simultaneously.

[0076] One or more of the cameras may be mounted in a mounting assembly such as a custom-designed (3D printed) component to cut off stray light and reflections from within the vehicle (such as reflections from the dashboard reflected in the windshield mirror) that may interfere with the image data capture capabilities of the camera. Regarding the wing mirror mounting assembly, the wing mirror assembly may be custom 3D printed such that the camera mounting plate matches the shape of the wing mirror. In some examples, one or more cameras may be integrated into the wing mirror. For side view cameras, one or more cameras may also be integrated into the four pillars at each corner of the cab.

[0077] Cameras having a field of view that includes an environmental portion in front of the vehicle 700 (such as a front camera) may be used for surround view to help identify forward paths and obstacles and to assist in providing information critical for generating an occupancy grid and / or determining a preferred vehicle path with the help of one or more controllers 736 and / or a control SoC. The front camera may be used to perform many of the same ADAS functions as LiDAR, including emergency braking, pedestrian detection, and collision avoidance. The front camera may also be used for ADAS functions and systems, including lane departure warning (LDW), adaptive cruise control (ACC), and / or other functions such as traffic sign recognition.

[0078] A variety of cameras may be used in a front-facing configuration, including, for example, a monocular camera platform including a complementary metal oxide semiconductor (CMOS) color imager. Another example may be a wide-angle camera 770, which may be used to sense objects entering the field of view from the periphery (such as pedestrians, intersection traffic, or bicycles). Although Figure 7B only one wide-angle camera is illustrated, any number (including zero) of wide-angle cameras 770 may be present on the vehicle 700. Additionally, any number of long-range cameras 798 (such as a long-view stereo camera pair) may be used for depth-based object detection, especially for objects for which a neural network has not been trained. The long-range cameras 798 may also be used for object detection and classification and basic object tracking.

[0079] Any number of stereo cameras 768 may also be included in a front-facing configuration. In at least one embodiment, one or more stereo cameras 768 may include an integrated control unit that includes a scalable processing unit that may provide a multi-core microprocessor and programmable logic (FPGA) with an integrated controller area network (CAN) or Ethernet interface on a single chip. Such a unit may be used to generate a 3D map of the vehicle environment, including distance estimates for all points in the image. Alternative stereo cameras 768 may include a compact stereo vision sensor that may include two camera lenses (one on the left and one on the right) and an image processing chip that may measure the distance from the vehicle to a target object and activate autonomous emergency braking and lane departure warning functions using the generated information (e.g., metadata). Other types of stereo cameras 768 may be used in addition to or alternatively to those described herein.

[0080] Cameras having a field of view that includes an environmental portion of a side of the vehicle 700 (e.g., side-view cameras) may be used for surround viewing, providing information used to create and update an occupancy grid and generate side-impact collision warnings. For example, surround cameras 774 (e.g., four surround cameras 774 as shown in Figure 7B may be disposed on the vehicle 700. The surround cameras 774 may include wide-angle cameras 770, fisheye cameras, 360-degree cameras, and / or the like. By way of example, four fisheye cameras may be disposed on the front, rear, and sides of the vehicle. In an alternative arrangement, the vehicle may use three surround cameras 774 (e.g., left, right, and rear) and may utilize one or more other cameras (e.g., a forward camera) as a fourth surround camera.

[0081] Cameras having a field of view that includes an environmental portion of the rear of the vehicle 700 (e.g., rear-view cameras) may be used for assisting with parking, surround viewing, rear collision warnings, and creating and updating an occupancy grid. A wide variety of cameras may be used, including but not limited to cameras that are also suitable as front cameras as described herein (e.g., long-range and / or mid-range cameras 798, stereo cameras 768, infrared cameras 772, etc.).

[0082] A camera (e.g., one or more OMS sensors 701) that views a portion of the interior environment within the passenger compartment of vehicle 700 can be used as part of an Occupant Monitoring System (OMS), such as, but not limited to, a Driver Monitoring System (DMS). For example, an OMS sensor (e.g., OMS sensor 701) can be used (e.g., by controller 736) to track the gaze direction, head pose, and / or blinks of an occupant and / or driver. This gaze information can be used to determine the attention level of the occupant or driver (e.g., detect drowsiness, fatigue, and / or distraction), and / or take responsive actions to prevent harm to the occupant or operator. In some embodiments, data from the OMS sensor can be used to implement gaze control operations triggered by a driver and / or non-driver passenger, such as, but not limited to, adjusting the passenger compartment temperature and / or airflow, opening and closing windows, controlling the passenger compartment lighting, controlling the entertainment system, adjusting the rearview mirror, adjusting the seat position, and / or other operations. In some embodiments, the OMS can be used in applications such as determining when an object and / or passenger remains within the passenger compartment (e.g., by detecting the presence of a passenger after the driver has left the vehicle).

[0083] Figure 7C For an example autonomous vehicle 700 in accordance with some embodiments of the present disclosure Figure 7A is a block diagram of an example system architecture. It should be understood that this and other arrangements described herein are presented only as examples. Other arrangements and elements (e.g., machines, interfaces, functions, orders, function groupings, etc.) can be used in addition to or instead of those shown, and some elements can be omitted entirely. Further, many of the elements described herein are functional entities that can be implemented as discrete or distributed components or in combination with other components, and in any suitable combination and location. The various functions described herein as being performed by entities can be implemented by hardware, firmware, and / or software. For example, the various functions can be implemented by a processor executing instructions stored in memory.

[0084] Figure 7C Each of the components, features, and systems in vehicle 700 is illustrated as being connected via a bus 702. The bus 702 can include a Controller Area Network (CAN) data interface (alternatively referred to herein as the "CAN bus"). The CAN can be a network within vehicle 700 that aids in controlling various features and functions of vehicle 700, such as driving of brakes, acceleration, braking, steering, windshield wipers, and so on. The CAN bus can be configured to have dozens or even hundreds of nodes, each with its own unique identifier (e.g., CAN ID). The CAN bus can be read to find the steering wheel angle, ground speed, engine revolutions per minute (RPM), button positions, and / or other vehicle state indicators. The CAN bus can be ASIL B compliant.

[0085] Although the bus 702 is described herein as a CAN bus, this is not intended to be limiting. For example, in addition to or alternatively to the CAN bus, FlexRay and / or Ethernet can be used. Further, although the bus 702 is shown as a single line, this is not intended to be limiting. For example, there can be any number of buses 702, which can include one or more CAN buses, one or more FlexRay buses, one or more Ethernet buses, and / or one or more other types of buses using different protocols. In some examples, two or more buses 702 can be used to perform different functions, and / or can be used for redundancy. For example, a first bus 702 can be used for a collision avoidance function, and a second bus 702 can be used for drive control. In any example, each bus 702 can communicate with any component of the vehicle 700, and two or more buses 702 can communicate with the same component. In some examples, each SoC 704, each controller 736, and / or each computer in the vehicle can have access to the same input data (e.g., input from sensors of the vehicle 700), and can be connected to a common bus such as a CAN bus.

[0086] Vehicle 700 can include one or more controllers 736, such as those described herein with respect to Figure 7A The controllers 736 can be used for a variety of functions. The controllers 736 can be coupled to any other different components and systems of the vehicle 700, and can be used for control of the vehicle 700, artificial intelligence of the vehicle 700, infotainment for the vehicle 700, and / or the like.

[0087] Vehicle 700 can include one or more system-on-chips (SoC) 704. The SoC 704 can include a CPU 706, a GPU 708, a processor 710, a cache 712, an accelerator 714, a data store 716, and / or other components and features not shown. In a variety of platforms and systems, the SoC 704 can be used to control the vehicle 700. For example, one or more SoC 704 can be combined with an HD map 722 in a system (e.g., a system of the vehicle 700), and the HD map can obtain map refreshes and / or updates via a network interface 724 from one or more servers (e.g., Figure 7D one or more servers 778 of

[0088] The CPU 706 may include a CPU cluster or a CPU complex (alternatively referred to herein as a "CCPLEX"). The CPU 706 may include multiple cores and / or L2 caches. For example, in some embodiments, the CPU 706 may include eight cores in a coherent multi-processor configuration. In some embodiments, the CPU 706 may include four dual-core clusters, each with a dedicated L2 cache (e.g., a 2MB L2 cache). The CPU 706 (e.g., CCPLEX) may be configured to support simultaneous cluster operation such that any combination of the clusters of the CPU 706 can be active at any given time.

[0089] The CPU 706 may implement power management capabilities including one or more of the following: each hardware block may automatically perform clock gating when idle to save dynamic power; each core clock may be gated when the core is not actively executing instructions due to the execution of WFI / WFE instructions; each core may be independently power gated; when all cores are clock gated or power gated, each core cluster may be independently clock gated; and / or when all cores are power gated, each core cluster may be independently power gated. The CPU 706 may further implement an enhanced algorithm for managing power states, where allowed power states and desired wake-up times are specified, and the hardware / microcode determines the optimal power state for the cores, clusters, and CCPLEX to enter. The processing cores may support a simplified power state entry sequence in software, and this task is offloaded to the microcode.

[0090] The GPU 708 may include an integrated GPU (alternatively referred to herein as an "iGPU"). The GPU 708 may be programmable and efficient for parallel workloads. In some examples, the GPU 708 may use an enhanced tensor instruction set. The GPU 708 may include one or more streaming microprocessors, where each streaming microprocessor may include an L1 cache (e.g., an L1 cache with at least 96KB of storage capacity), and two or more of these streaming microprocessors may share an L2 cache (e.g., an L2 cache with 512KB of storage capacity). In some embodiments, the GPU 708 may include at least eight streaming microprocessors. The GPU 708 may use a compute application programming interface (API). Additionally, the GPU 708 may use one or more parallel computing platforms and / or programming models (e.g., NVIDIA's CUDA).

[0091] In automotive and embedded use cases, the GPU 708 can be power optimized for best performance. For example, the GPU 708 can be fabricated on fin field-effect transistors (FinFETs). However, this is not intended to be restrictive, and the GPU 708 can be fabricated using other semiconductor manufacturing processes. Each streaming microprocessor can incorporate a number of mixed-precision processing cores divided into multiple blocks. For example and without limitation, 64 PF32 cores and 32 PF64 cores can be divided into four processing blocks. In such an example, each processing block can be allocated 16 FP32 cores, 8 FP64 cores, 16 INT32 cores, two mixed-precision NVIDIA tensor cores for deep learning matrix arithmetic, an L0 instruction cache, a warp scheduler, a dispatch unit, and / or a 64KB register file. Additionally, the streaming microprocessor can include independent parallel integer and floating-point data paths to enable efficient execution of workloads that utilize a mix of computational and addressing computations. The streaming microprocessor can include independent thread scheduling capabilities to allow for more fine-grained synchronization and cooperation between parallel threads. The streaming microprocessor can include a combined L1 data cache and shared memory unit to improve performance while simplifying programming.

[0092] The GPU 708 can include, in some examples, high-bandwidth memory (HBM) that provides a peak memory bandwidth of approximately 900GB / s and / or a 16GB HBM2 memory subsystem. In some examples, in addition to or alternatively to HBM memory, synchronous graphics random access memory (SGRAM), such as fifth-generation graphics double data rate synchronous random access memory (GDDR5), can be used.

[0093] The GPU 708 can include unified memory technology that includes access counters to allow memory pages to be more precisely migrated to the processors that most frequently access them, thereby improving the efficiency of the memory ranges shared between processors. In some examples, address translation service (ATS) support can be used to allow the GPU 708 to directly access the CPU 706 page tables. In such an example, when the GPU 708 memory management unit (MMU) experiences a miss, an address translation request can be transmitted to the CPU 706. In response, the CPU 706 can look up the virtual-physical mapping for the address in its page table and transmit the translation back to the GPU 708. In this way, the unified memory technology can allow for a single unified virtual address space for the memory of both the CPU 706 and the GPU 708, thereby simplifying GPU 708 programming and porting applications to the GPU 708.

[0094] In addition, the GPU 708 may include an access counter that can track how frequently the GPU 708 accesses the memory of other processors. The access counter can help ensure that memory pages are moved to the physical memory of the processor that most frequently accesses those pages.

[0095] The SoC 704 may include any number of caches 712, including those described herein. For example, the cache 712 may include an L3 cache (e.g., which is connected to both the CPU 706 and the GPU 708) that is available to both the CPU 706 and the GPU 708. The cache 712 may include a write-back cache that can track the state of lines, for example, by using a cache coherence protocol (e.g., MEI, MESI, MSI, etc.). Depending on the embodiment, the L3 cache may include 4MB or more, but smaller cache sizes may also be used.

[0096] The SoC 704 may include one or more arithmetic logic units (ALUs) that can be used to perform processing for any of a variety of tasks or operations regarding the vehicle 700, such as processing a DNN. In addition, the SoC 704 may include a floating point unit (FPU) or other math co-processor or digital co-processor type for performing mathematical operations within the system. For example, the SoC 704 may include one or more FPUs integrated within the execution units of the CPU 706 and / or the GPU 708.

[0097] The SoC 704 may include one or more accelerators 714 (e.g., hardware accelerators, software accelerators, or a combination thereof). For example, the SoC 704 may include a hardware acceleration cluster that may include optimized hardware accelerators and / or large on-chip memory. This large on-chip memory (e.g., 4MB SRAM) may enable the hardware acceleration cluster to accelerate neural networks and other computations. The hardware acceleration cluster can be used to supplement the GPU 708 and offload some of the tasks of the GPU 708 (e.g., freeing up more cycles of the GPU 708 for performing other tasks). As an example, the accelerator 714 may be used for targeted workloads that are stable enough to be easily accelerated (e.g., perception, convolutional neural networks (CNNs), etc.). When used herein, the term "CNN" may include all types of CNNs, including region-based or region convolutional neural networks (RCNNs) and fast RCNNs (e.g., for object detection).

[0098] The accelerator 714 (e.g., a hardware acceleration cluster) may include a deep learning accelerator (DLA). The DLA may include one or more tensor processing units (TPUs) that can be configured to provide an additional one trillion operations per second for deep learning applications and inference. The TPU may be an accelerator configured to perform image processing functions (e.g., for CNN, RCNN, etc.) and optimized for performing image processing functions. The DLA may be further optimized for a specific set of neural network types and floating-point operations and inference. The design of the DLA may provide higher performance per millimeter than a general-purpose GPU and far exceed the performance of a CPU. The TPU may perform several functions, including single-instance convolution functions, supporting INT8, INT16, and FP16 data types for both features and weights, and post-processor functions.

[0099] The DLA may quickly and efficiently execute neural networks, especially CNNs, for any of a variety of functions on processed or unprocessed data, such as, but not limited to: CNNs for object recognition and detection using data from camera sensors; CNNs for distance estimation using data from camera sensors; CNNs for emergency vehicle detection and identification and detection using data from microphones; CNNs for face recognition and vehicle owner recognition using data from camera sensors; and / or CNNs for security and / or safety-related events.

[0100] The DLA may perform any function of the GPU 708, and by using inference accelerators, for example, a designer may target the DLA or the GPU 708 for any function. For example, a designer may focus the processing and floating-point operations of a CNN on the DLA and leave other functions to the GPU 708 and / or other accelerators 714.

[0101] The accelerator 714 (e.g., a hardware acceleration cluster) may include a programmable vision accelerator (PVA), which may alternatively be referred to herein as a computer vision accelerator. The PVA may be designed and configured to accelerate computer vision algorithms for advanced driver assistance systems (ADAS), autonomous driving, and / or augmented reality (AR) and / or virtual reality (VR) applications. The PVA may provide a balance between performance and flexibility. For example, each PVA may include, for example and without limitation, any number of reduced instruction set computer (RISC) cores, direct memory access (DMA), and / or any number of vector processors.

[0102] The RISC cores can interact with image sensors (e.g., the image sensors of any of the cameras described herein), image signal processors, and / or the like. Each of these RISC cores can include any number of memories. Depending on the embodiment, the RISC cores can use any of several protocols. In some examples, the RISC cores can execute a real-time operating system (RTOS). The RISC cores can be implemented using one or more integrated circuit devices, application-specific integrated circuits (ASICs), and / or storage devices. For example, the RISC cores can include an instruction cache and / or tightly coupled RAM.

[0103] DMA can enable components of the PVA to access system memory independently of the CPU 706. DMA can support any number of features used to optimize the PVA, including but not limited to supporting multi-dimensional addressing and / or circular addressing. In some examples, DMA can support up to six or more dimensions of addressing, which can include block width, block height, block depth, horizontal block step, vertical block step, and / or depth step.

[0104] The vector processor can be a programmable processor that can be designed to efficiently and flexibly execute programming for computer vision algorithms and provide signal processing capabilities. In some examples, the PVA can include a PVA core and two vector processing subsystem partitions. The PVA core can include a processor subsystem, one or more DMA engines (e.g., two DMA engines), and / or other peripherals. The vector processing subsystem can operate as the main processing engine of the PVA and can include a vector processing unit (VPU), an instruction cache, and / or vector memory (e.g., VMEM). The VPU core can include a digital signal processor, such as, for example, a single instruction multiple data (SIMD), very long instruction word (VLIW) digital signal processor. The combination of SIMD and VLIW can enhance throughput and rate.

[0105] Each of the vector processors may include an instruction cache and may be coupled to dedicated memory. As a result, in some examples, each of the vector processors may be configured to execute independently of the other vector processors. In other examples, the vector processors included in a particular PVA may be configured to employ data parallelization. For example, in some embodiments, multiple vector processors included in a single PVA may execute the same computer vision algorithm, but on different regions of an image. In other examples, the vector processors included in a particular PVA may execute different computer vision algorithms simultaneously on the same image, or even execute different algorithms on sequential images or portions of an image. Among other things, any number of PVAs may be included in a hardware acceleration cluster, and any number of vector processors may be included in each of these PVAs. Additionally, a PVA may include additional error correction code (ECC) memory to enhance overall system security.

[0106] The accelerator 714 (e.g., a hardware acceleration cluster) may include an on-chip computer vision network and SRAM to provide high-bandwidth, low-latency SRAM for the accelerator 714. In some examples, the on-chip memory may include at least 4MB SRAM consisting of, for example and without limitation, eight field-configurable memory blocks, which may be accessed by both the PVA and the DLA. Each pair of memory blocks may include an advanced peripheral bus (APB) interface, configuration circuitry, a controller, and a multiplexer. Any type of memory may be used. The PVA and the DLA may access the memory via a backbone that provides high-speed memory access to the PVA and the DLA. The backbone may include, for example using the APB, an on-chip computer vision network that interconnects the PVA and the DLA to the memory.

[0107] The on-chip computer vision network may include an interface that determines that both the PVA and the DLA provide ready and valid signals before transmitting any control signals / address / data. Such an interface may provide separate phases and separate channels for transmitting control signals / address / data, as well as burst communication for continuous data transfer. This type of interface may comply with the ISO 26262 or IEC 61508 standards, but other standards and protocols may also be used.

[0108] In some examples, SoC 704 can include a real-time ray tracing hardware accelerator as described, for example, in U.S. Patent Application No. 16 / 101,232, filed on August 10, 2018. The real-time ray tracing hardware accelerator can be used to quickly and efficiently determine the position and extent of objects (e.g., within a world model) in order to generate a real-time visualization simulation for RADAR signal interpretation, for sound propagation synthesis and / or analysis, for SONAR system simulation, for general wave propagation simulation, for comparison with LiDAR data for positioning and / or other functional purposes, and / or for other uses. In some embodiments, one or more tree traversal units (TTUs) can be used to perform one or more ray tracing-related operations.

[0109] Accelerator 714 (e.g., a hardware accelerator cluster) has a wide range of autonomous driving uses. The PVA can be a programmable vision accelerator that can be used in key processing stages in ADAS and autonomous vehicles. The capabilities of the PVA are a good match for algorithm domains that require predictable processing, low power, and low latency. Thus, the PVA performs well on semi-dense or dense regular computations, even on small data sets that require predictable runtimes with low latency and low power. Thus, in the context of a platform for autonomous vehicles, the PVA is designed to run classical computer vision algorithms because they are effective in object detection and integer math operations.

[0110] For example, according to one embodiment of the technology, the PVA is used to perform computer stereo vision. In some examples, an algorithm based on semi-global matching can be used, but this is not intended to be limiting. Many applications for level 3 - 5 autonomous driving require instantaneous motion estimation / stereo matching (e.g., structure from motion, pedestrian recognition, lane detection, etc.). The PVA can perform computer stereo vision functions on inputs from two monocular cameras.

[0111] In some examples, the PVA can be used to perform dense optical flow. Process raw RADAR data (e.g., using a 4D fast Fourier transform) to provide processed RADAR. In other examples, the PVA is used for time-of-flight depth processing, which, for example, processes raw time-of-flight data to provide processed time-of-flight data.

[0112] DLA can be used to run any type of network to enhance control and driving safety, including, for example, a neural network that outputs a confidence metric for each object detection. Such confidence values can be interpreted as probabilities or as providing a relative "weight" of each detection compared to other detections. The confidence value enables the system to make further decisions about which detections should be considered true positive detections rather than false positive detections. For example, the system can set a threshold for the confidence and consider only detections that exceed the threshold as true positive detections. In an automatic emergency braking (AEB) system, false positive detections can cause the vehicle to automatically perform emergency braking, which is clearly undesirable. Therefore, only the most confident detections should be considered as triggers for AEB. DLA can run a neural network for regressing confidence values. The neural network can take as its input at least some subset of parameters, such as bounding box dimensions, a ground plane estimate obtained (e.g., from another subsystem), the output of an inertial measurement unit (IMU) sensor 766 related to the vehicle 700 orientation and distance, a 3D position estimate of an object obtained from a neural network and / or other sensors (such as a LiDAR sensor 764 or a RADAR sensor 760), etc.

[0113] The SoC 704 can include one or more data stores 716 (e.g., memory). The data store 716 can be on-chip memory of the SoC 704, which can store neural networks to be executed on the GPU and / or DLA. In some examples, for redundancy and safety, the data store 716 can be large enough in capacity to store multiple instances of the neural network. The data store 716 can include an L2 or L3 cache 712. References to the data store 716 can include references to memory associated with PVAs, DLAs, and / or other accelerators 714 as described herein.

[0114] The SoC 704 may include one or more processors 710 (e.g., embedded processors). The processor 710 may include a boot and power management processor, which may be a dedicated processor and subsystem for handling boot power and management functions as well as security implementation related. The boot and power management processor may be part of the SoC 704 boot sequence and may provide runtime power management services. The boot power and management processor may provide clock and voltage programming, auxiliary system low power state transitions, SoC 704 heat and temperature sensor management, and / or SoC 704 power state management. Each temperature sensor may be implemented as a ring oscillator whose output frequency is proportional to temperature, and the SoC 704 may use the ring oscillator to detect the temperature of the CPU 706, GPU 708, and / or accelerator 714. If it is determined that the temperature exceeds a threshold, then the boot and power management processor may enter a temperature fault routine and place the SoC 704 in a lower power state and / or place the vehicle 700 in a driver safety stop mode (e.g., safely stop the vehicle 700).

[0115] The processor 710 may further include a set of embedded processors that may be used as an audio processing engine. The audio processing engine may be an audio subsystem that allows for full hardware support for multi-channel audio over multiple interfaces and a wide and flexible range of audio I / O interfaces. In some examples, the audio processing engine is a dedicated processor core with a digital signal processor with dedicated RAM.

[0116] The processor 710 may further include an always-on processor engine, which may provide the necessary hardware features to support low power sensor management and wake-up use cases. The always-on processor engine may include a processor core, tightly coupled RAM, support peripherals (e.g., timers and interrupt controllers), various I / O controller peripherals, and routing logic.

[0117] The processor 710 may further include a security cluster engine, which includes a dedicated processor subsystem for handling security management of automotive applications. The security cluster engine may include two or more processor cores, tightly coupled RAM, support peripherals (e.g., timers, interrupt controllers, etc.), and / or routing logic. In the security mode, the two or more cores may operate in a lockstep mode and act as a single core with comparison logic for detecting any differences between their operations.

[0118] The processor 710 may further include a real-time camera engine, which may include a dedicated processor subsystem for handling real-time camera management.

[0119] The processor 710 may further include a high dynamic range signal processor, which may include an image signal processor, which is a hardware engine that is part of the camera processing pipeline.

[0120] The processor 710 may include a video image compositor that may be a processing block (e.g., implemented on a microprocessor) that implements the video post-processing functions required for a video playback application to generate the final image for the player window. The video image compositor may perform lens distortion correction on the wide-angle camera 770, the surround camera 774, and / or the in-cab monitoring camera sensor. The in-cab monitoring camera sensor is preferably monitored by a neural network running on another instance of the advanced SoC, configured to identify in-cab events and respond accordingly. The in-cab system may perform lip reading to activate mobile phone services and make calls, dictate emails, change the vehicle destination, activate or change the vehicle's infotainment system and settings, or provide voice-activated web surfing. Certain functions are only available to the driver when the vehicle is operating in autonomous mode and are disabled otherwise.

[0121] The video image compositor may include enhanced temporal noise reduction for spatial and temporal noise reduction. For example, in the case of motion in the video, the noise reduction appropriately weights the spatial information, reducing the weight of the information provided by neighboring frames. In the case where the image or a portion of the image does not include motion, the temporal noise reduction performed by the video image compositor may use information from a previous image to reduce the noise in the current image.

[0122] The video image compositor may also be configured to perform stereo correction on input stereo lens frames. When the operating system desktop is in use and the GPU 708 does not need to continuously render new surfaces, the video image compositor may be further used for user interface composition. Even when the GPU 708 is powered on and active, performing 3D rendering, the video image compositor may be used to relieve the burden on the GPU 708 to improve performance and responsiveness.

[0123] The SoC 704 may further include a Mobile Industry Processor Interface (MIPI) camera serial interface, a high-speed interface, and / or a video input block for receiving video and inputs from cameras and may be used for camera and related pixel input functions. The SoC 704 may further include an input / output controller that may be software-controlled and may be used to receive I / O signals not committed to a specific role.

[0124] The SoC 704 may further include a wide range of peripheral device interfaces to enable communication with peripheral devices, audio codecs, power management, and / or other devices. The SoC 704 may be used to process data from cameras (via gigabit multimedia serial link and Ethernet connections), sensors (such as LiDAR sensor 764, RADAR sensor 760, etc. that may be connected via Ethernet), data from bus 702 (such as the speed of vehicle 700, steering wheel position, etc.), and data from GNSS sensor 758 (connected via Ethernet or CAN bus). The SoC 704 may further include dedicated high-performance large-capacity storage controllers, which may include their own DMA engines and which may be used to free the CPU 706 from routine data management tasks.

[0125] The SoC 704 may be an end-to-end platform with a flexible architecture that spans automation levels 3 - 5, thus providing an integrated functional safety architecture for a platform that leverages and efficiently uses computer vision and ADAS technologies to achieve diversity and redundancy, along with deep learning tools to provide a flexible and reliable driving software stack. The SoC 704 may be faster, more reliable, and even more energy-efficient and space-efficient than conventional systems. For example, when combined with the CPU 706, GPU 708, and data storage 716, the accelerator 714 may provide a fast and efficient platform for level 3 - 5 autonomous vehicles.

[0126] Thus, this technology provides capabilities and functions that cannot be achieved by conventional systems. For example, computer vision algorithms may be executed on CPUs that may be configured using high-level programming languages such as the C programming language to perform various processing algorithms across a variety of visual data. However, CPUs often cannot meet the performance requirements of many computer vision applications, such as those related to execution time and power consumption. In particular, many CPUs cannot execute complex object detection algorithms in real time, which is a requirement for in-vehicle ADAS applications and for practical level 3 - 5 autonomous vehicles.

[0127] In contrast to conventional systems, the technology described herein allows multiple neural networks to be executed simultaneously and / or sequentially by providing a CPU complex, a GPU complex, and a hardware acceleration cluster, and combining the results to achieve level 3 - 5 autonomous driving functions. For example, a CNN executed on a DLA or a dGPU (such as GPU 720) may include text and word recognition, allowing a supercomputer to read and understand traffic signs, including signs for which the neural network has not been specifically trained. The DLA may further include a neural network capable of recognizing, interpreting, and providing semantic understanding of the signs and passing that semantic understanding to a path planning module running on the CPU complex.

[0128] As another example, multiple neural networks can operate simultaneously, such as required for level 3, 4, or 5 driving. For example, a warning sign consisting of "Caution: Flashing lights indicate icy conditions" together with the electric lights can be interpreted independently or jointly by several neural networks. The sign itself can be recognized as a traffic sign by a first neural network deployed (e.g., a trained neural network), and the text "Flashing lights indicate icy conditions" can be interpreted by a second neural network deployed, which informs the vehicle's path planning software (preferably executed on the CPU complex) that when the flashing lights are detected, there are icy conditions. The flashing lights can be recognized by operating a third neural network deployed over multiple frames, which informs the vehicle's path planning software of the presence (or absence) of the flashing lights. All three neural networks can operate simultaneously, for example, within the DLA and / or on the GPU 708.

[0129] In some examples, the CNNs for face recognition and owner recognition can use data from the camera sensors to recognize the presence of an authorized driver and / or owner of the vehicle 700. A processing engine that is always on the sensor can be used to unlock the vehicle and turn on the lights when the owner approaches the driver's door, and in a security mode, to disable the vehicle when the owner leaves the vehicle. In this way, the SoC 704 provides security against theft and / or carjacking.

[0130] In another example, the CNN for emergency vehicle detection and recognition can use data from the microphone 796 to detect and recognize an emergency vehicle siren. In contrast to conventional systems that use a general classifier to detect the siren and manually extract features, the SoC 704 uses a CNN to classify environmental and urban sounds as well as visual data. In a preferred embodiment, the CNN running on the DLA is trained to recognize the relative closing rate of an emergency vehicle (e.g., by using the Doppler effect). The CNN can also be trained to recognize emergency vehicles specific to the local area in which the vehicle is operating as recognized by the GNSS sensor 758. Thus, for example, when operating in Europe, the CNN will seek to detect European sirens, and when in the United States, the CNN will seek to recognize only North American sirens. Once an emergency vehicle is detected, with the assistance of the ultrasonic sensor 762, a control program can be used to execute an emergency vehicle safety routine to slow down the vehicle, pull over to the side of the road, stop the vehicle, and / or idle the vehicle until the emergency vehicle passes.

[0131] The vehicle may include a CPU 718 (e.g., a discrete CPU or dCPU) that can be coupled to the SoC 704 via a high-speed interconnect (e.g., PCIe). The CPU 718 may include, for example, an X86 processor. The CPU 718 can be used to perform any of a variety of functions, including, for example, arbitrating potentially inconsistent results between ADAS sensors and the SoC 704, and / or monitoring the status and health of the controller 736 and / or the infotainment SoC 730.

[0132] The vehicle 700 may include a GPU 720 (e.g., a discrete GPU or dGPU) that can be coupled to the SoC 704 via a high-speed interconnect (e.g., NVIDIA's NVLINK). The GPU 720 can provide additional artificial intelligence capabilities, for example, by executing redundant and / or different neural networks, and can be used to train and / or update neural networks based on inputs (e.g., sensor data) from the sensors of the vehicle 700.

[0133] The vehicle 700 may further include a network interface 724 that can include one or more wireless antennas 726 (e.g., one or more wireless antennas for different communication protocols, such as cellular antennas, Bluetooth antennas, etc.). The network interface 724 can be used to enable wireless connections to the cloud (e.g., to the server 778 and / or other network devices), to other vehicles, and / or to computing devices (e.g., the client devices of passengers) via the Internet. To communicate with other vehicles, a direct link can be established between the two vehicles, and / or an indirect link (e.g., across a network and via the Internet) can be established. The direct link can be provided using a vehicle-to-vehicle communication link. The vehicle-to-vehicle communication link can provide the vehicle 700 with information about vehicles approaching the vehicle 700 (e.g., vehicles in front of, to the side of, and / or behind the vehicle 700). This function can be part of the cooperative adaptive cruise control function of the vehicle 700.

[0134] The network interface 724 can include an SoC that provides modulation and demodulation functions and enables the controller 736 to communicate via a wireless network. The network interface 724 can include a radio frequency front end for upconversion from baseband to radio frequency and downconversion from radio frequency to baseband. The frequency conversion can be performed by a known process and / or can be performed using a super-heterodyne process. In some examples, the radio frequency front end functions can be provided by a separate chip. The network interface can include wireless capabilities for communicating via LTE, WCDMA, UMTS, GSM, CDMA2000, Bluetooth, Bluetooth LE, Wi-Fi, Z-Wave, ZigBee, LoRaWAN, and / or other wireless protocols.

[0135] Vehicle 700 may further include a data store 728 that may include off-chip (e.g., outside of SoC 704) storage devices. The data store 728 may include one or more storage elements including RAM, SRAM, DRAM, VRAM, flash memory, hard drives, and / or other components and / or devices that may store at least one bit of data.

[0136] Vehicle 700 may further include a GNSS sensor 758. The GNSS sensor 758 (e.g., GPS, assisted GPS sensor, differential GPS (DGPS) sensor, etc.) is used to assist mapping, perception, occupancy grid generation, and / or path planning functions. Any number of GNSS sensors 758 may be used, including for example and without limitation a GPS using a USB connector with an Ethernet to serial (RS-232) bridge.

[0137] Vehicle 700 may further include a RADAR sensor 760. The RADAR sensor 760 may be used by the vehicle 700 for remote vehicle detection even in dark and / or adverse weather conditions. The RADAR functional safety level may be ASIL B. The RADAR sensor 760 may use CAN and / or bus 702 (e.g., to transmit data generated by the RADAR sensor 760) for control as well as to access object tracking data, and in some examples accesses Ethernet to access raw data. A variety of RADAR sensor types may be used. For example and without limitation, the RADAR sensor 760 may be suitable for front, rear, and side RADAR use. In some examples, a pulsed Doppler RADAR sensor is used.

[0138] The RADAR sensor 760 may include different configurations, such as long range with a narrow field of view, short range with a wide field of view, short range side coverage, and so on. In some examples, long range RADAR may be used for adaptive cruise control functions. The long range RADAR system may provide a wide field of view (e.g., within 250 m) achieved through two or more independent scans. The RADAR sensor 760 may help distinguish between static and moving objects and may be used by the ADAS system for emergency braking assistance and forward collision warning. The long range RADAR sensor may include a single station multi-mode RADAR with multiple (e.g., six or more) fixed RADAR antennas and high-speed CAN and FlexRay interfaces. In an example with six antennas, the central four antennas may create a focused beam pattern that is designed to record the surroundings of the vehicle 700 at a higher rate with minimal traffic interference from adjacent lanes. The other two antennas may extend the field of view, making it possible to quickly detect vehicles entering or leaving the lane of the vehicle 700.

[0139] As an example, a mid-range RADAR system can include a range of up to 760 m (front) or 80 m (rear) and a field of view of up to 42 degrees (front) or 750 degrees (rear). A short-range RADAR system can include, but is not limited to, RADAR sensors designed to be mounted at both ends of the rear bumper. When mounted at both ends of the rear bumper, such a RADAR sensor system can create two beams that continuously monitor the rear and the blind spots alongside the vehicle.

[0140] The short-range RADAR system can be used in an ADAS system for blind spot detection and / or lane change assistance.

[0141] Vehicle 700 can further include ultrasonic sensors 762. Ultrasonic sensors 762 that can be placed at the front, rear, and / or sides of vehicle 700 can be used for parking assistance and / or creating and updating occupancy grids. A variety of ultrasonic sensors 762 can be used, and different ultrasonic sensors 762 can be used for different detection ranges (e.g., 2.5 m, 4 m). Ultrasonic sensors 762 can operate at the functional safety level of ASIL B.

[0142] Vehicle 700 can include a LiDAR sensor 764. The LiDAR sensor 764 can be used for object and pedestrian detection, emergency braking, collision avoidance, and / or other functions. The LiDAR sensor 764 can be at the functional safety level of ASIL B. In some examples, vehicle 700 can include multiple LiDAR sensors 764 (e.g., two, four, six, etc.) that can use Ethernet (e.g., to provide data to a gigabit Ethernet switch).

[0143] In some examples, the LiDAR sensor 764 may be capable of providing a list of objects and their distances for a 360-degree field of view. Commercially available LiDAR sensors 764 can have, for example, an advertised range of approximately 700 m, an accuracy of 2 cm - 3 cm, and support for a 700 Mbps Ethernet connection. In some examples, one or more non-protruding LiDAR sensors 764 can be used. In such examples, the LiDAR sensor 764 can be implemented as a small device that can be embedded into the front, rear, sides, and / or corners of vehicle 700. In such examples, the LiDAR sensor 764 can provide a field of view of up to 120 degrees horizontally and 35 degrees vertically, even for low-reflectivity objects, with a range of 200 m. The front-mounted LiDAR sensor 764 can be configured for a horizontal field of view between 45 degrees and 135 degrees.

[0144] In some examples, LiDAR technologies such as 3D flash LiDAR can also be used. 3D flash LiDAR uses the flash of a laser as the emission source to illuminate the vehicle's surroundings up to approximately 200m. The flash LiDAR unit includes a receiver that records the laser pulse transit time and the reflected light on each pixel, which in turn corresponds to the range from the vehicle to the object. Flash LiDAR can allow for the generation of highly accurate and distortion-free images of the surroundings using each laser flash. In some examples, four flash LiDAR sensors can be deployed, one on each side of the vehicle 700. Available 3D flash LiDAR systems include solid-state 3D staring array LiDAR cameras (e.g., non-scanning LiDAR devices) that have no moving parts other than a fan. The flash LiDAR device can use Class I (eye-safe) laser pulses of 5 nanoseconds per frame and can capture the reflected laser in the form of 3D range point clouds and co-registered intensity data. By using flash LiDAR, and since flash LiDAR is a solid-state device with no moving parts, the LiDAR sensor 764 can be less susceptible to motion blur, vibration, and / or shock.

[0145] The vehicle can further include an IMU sensor 766. In some examples, the IMU sensor 766 can be located at the center of the rear axle of the vehicle 700. The IMU sensor 766 can include, for example and without limitation, an accelerometer, a magnetometer, a gyroscope, a magnetic compass, and / or other sensor types. In some examples, such as in a six-axis application, the IMU sensor 766 can include an accelerometer and a gyroscope, while in a nine-axis application, the IMU sensor 766 can include an accelerometer, a gyroscope, and a magnetometer.

[0146] In some embodiments, the IMU sensor 766 can be implemented as a miniature high-performance GPS-aided inertial navigation system (GPS / INS) that combines microelectromechanical systems (MEMS) inertial sensors, a high-sensitivity GPS receiver, and an advanced Kalman filtering algorithm to provide estimates of position, velocity, and attitude. Thus, in some examples, the IMU sensor 766 can enable the vehicle 700 to estimate the heading without input from a magnetic sensor by directly observing the change in velocity from the GPS to the IMU sensor 766 and correlating it. In some examples, the IMU sensor 766 and the GNSS sensor 758 can be combined into a single integrated unit.

[0147] The vehicle can include a microphone 796 placed in and / or around the vehicle 700. Among other things, the microphone 796 can be used for emergency vehicle detection and identification.

[0148] The vehicle may further include any number of camera types, including a stereo camera 768, a wide-angle camera 770, an infrared camera 772, a surround camera 774, a long-range and / or mid-range camera 798, and / or other camera types. These cameras can be used to capture image data around the entire periphery of the vehicle 700. The camera types used depend on the embodiment and the requirements of the vehicle 700, and any combination of camera types can be used to provide the necessary coverage around the vehicle 700. Additionally, the number of cameras can vary according to the embodiment. For example, the vehicle may include six cameras, seven cameras, ten cameras, twelve cameras, and / or another number of cameras. As an example and without limitation, these cameras can support Gigabit Multimedia Serial Link (GMSL) and / or Gigabit Ethernet. Each of the cameras is described in more detail herein with respect to Figure 7A and Figure 7B More detailed descriptions are provided.

[0149] The vehicle 700 may further include a vibration sensor 742. The vibration sensor 742 can measure the vibration of components of the vehicle, such as the axle. For example, a change in vibration can indicate a change in the road surface. In another example, when two or more vibration sensors 742 are used, the difference between the vibrations can be used to determine the friction or slip of the road surface (e.g., when there is a vibration difference between a powered drive axle and a free-spinning axle).

[0150] The vehicle 700 may include an ADAS system 738. In some examples, the ADAS system 738 may include a SoC. The ADAS system 738 may include autonomous / adaptive / auto cruise control (ACC), cooperative adaptive cruise control (CACC), forward collision warning (FCW), automatic emergency braking (AEB), lane departure warning (LDW), lane keeping assist (LKA), blind spot warning (BSW), rear cross-traffic warning (RCTW), collision warning system (CWS), lane centering (LC), and / or other features and functions.

[0151] The ACC system can use RADAR sensors 760, LiDAR sensors 764, and / or cameras. The ACC system can include longitudinal ACC and / or lateral ACC. Longitudinal ACC monitors and controls the distance to the vehicle immediately in front of the vehicle 700 and automatically adjusts the vehicle speed to maintain a safe distance from the vehicle ahead. Lateral ACC performs distance keeping and, when necessary, advises the vehicle 700 to change lanes. Lateral ACC is related to other ADAS applications such as LCA and CWS.

[0152] The CACC uses information from other vehicles, which can be received indirectly from other vehicles via the network interface 724 and / or the wireless antenna 726 via a wireless link or through a network connection (e.g., via the Internet). The direct link can be provided by a vehicle-to-vehicle (V2V) communication link, while the indirect link can be an infrastructure-to-vehicle (I2V) communication link. Generally, the V2V communication concept provides information about the immediately preceding vehicle (e.g., a vehicle immediately in front of vehicle 700 and in the same lane), while the I2V communication concept provides information about traffic further ahead. The CACC system can include either or both of the I2V and V2V information sources. Given information about the vehicle in front of vehicle 700, the CACC can be more reliable, and it has the potential to improve the smoothness of traffic flow and reduce road congestion.

[0153] The FCW system is designed to alert the driver of a hazard so that the driver can take corrective action. The FCW system uses a front camera and / or RADAR sensor 760 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback such as a display, speaker, and / or vibration component. The FCW system can provide warnings in the form of, for example, audible, visual warnings, vibrations, and / or rapid braking pulses.

[0154] The AEB system detects an impending frontal collision with another vehicle or other object and can automatically apply the brakes if the driver does not take corrective action within specified time or distance parameters. The AEB system can use a front camera and / or RADAR sensor 760 coupled to a dedicated processor, DSP, FPGA, and / or ASIC. When the AEB system detects a hazard, it typically first alerts the driver to take corrective action to avoid the collision, and if the driver does not take corrective action, then the AEB system can automatically apply the brakes in an effort to prevent or at least mitigate the impact of the predicted collision. The AEB system can include technologies such as dynamic brake support and / or collision imminent braking.

[0155] The LDW system provides visual, audible, and / or tactile warnings such as steering wheel or seat vibrations to alert the driver when vehicle 700 crosses a lane marking. The LDW system is not activated when the driver indicates an intentional lane departure by activating the turn signal. The LDW system can use a front-side facing camera coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback such as a display, speaker, and / or vibration component.

[0156] The LKA system is a variant of the LDW system. If vehicle 700 starts to leave the lane, then the LKA system provides a steering input or braking to correct the vehicle 700.

[0157] The BSW system detects and warns the driver of vehicles in the vehicle's blind spot. The BSW system can provide visual, audible, and / or tactile alerts to indicate that merging or changing lanes is unsafe. The system can provide additional warnings when the driver uses the turn signal. The BSW system can use a rear-facing camera and / or RADAR sensor 760 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback such as a display, speaker, and / or vibrating component.

[0158] The RCTW system can provide visual, audible, and / or tactile notifications when an object is detected outside the rear camera range while the vehicle 700 is in reverse. Some RCTW systems include AEB to ensure that the vehicle brakes are applied to avoid a crash. The RCTW system can use one or more rear RADAR sensors 760 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback such as a display, speaker, and / or vibrating component.

[0159] Conventional ADAS systems can be prone to false positive results, which can be annoying and distracting to the driver, but are typically not catastrophic because the ADAS system alerts the driver and allows the driver to decide if a safe condition truly exists and act accordingly. However, in an autonomous vehicle 700, in the case of conflicting results, the vehicle 700 itself must decide whether to heed the results from the primary computer or the secondary computer (e.g., the first controller 736 or the second controller 736). For example, in some embodiments, the ADAS system 738 can be a secondary and / or auxiliary computer for providing perception information to a redundant computer sanity module. The redundant computer sanity monitor can run redundant and diverse software on hardware components to detect faults in perception and dynamic driving tasks. The output from the ADAS system 738 can be provided to the supervisory MCU. If the outputs from the primary computer and the secondary computer conflict, then the supervisory MCU must determine how to reconcile the conflict to ensure safe operation.

[0160] In some examples, the host computer may be configured to provide a confidence score to the supervisory MCU indicating the host computer's confidence in the selected result. If the confidence score exceeds a threshold, then the supervisory MCU may follow the direction of the host computer regardless of whether the secondary computer provides conflicting or inconsistent results. In cases where the confidence score does not meet the threshold and where the host computer and the secondary computer indicate different results (e.g., conflict), the supervisory MCU may arbitrate between these computers to determine the appropriate result.

[0161] The supervisory MCU may be configured to run a neural network that is trained and configured to determine conditions under which the secondary computer provides false alarms based on outputs from the host computer and the secondary computer. Thus, the neural network in the supervisory MCU can learn when the output of the secondary computer can be trusted and when it cannot. For example, when the secondary computer is a RADAR-based FCW system, the neural network in the supervisory MCU can learn when the FCW system is identifying metal objects that are not in fact dangerous, such as a drain grate or manhole cover that triggers an alarm. Similarly, when the secondary computer is a camera-based LDW system, the neural network in the supervisory MCU can learn to disregard the LDW when a cyclist or pedestrian is present and lane departure is actually the safest strategy. In embodiments that include a neural network running on the supervisory MCU, the supervisory MCU may include at least one of a DLA or a GPU suitable for running the neural network with associated memory. In a preferred embodiment, the supervisory MCU may include components of the SoC 704 and / or be included as a component of the SoC 704.

[0162] In other examples, the ADAS system 738 may include a secondary computer that performs ADAS functions using traditional computer vision rules. In this way, the secondary computer may use classical computer vision rules (if - then), and the presence of a neural network in the supervisory MCU can improve reliability, safety, and performance. For example, the diverse implementations and intentional non-identity make the overall system more fault-tolerant, especially for failures caused by software (or software-hardware interface) functions. For example, if there is a software vulnerability or error in the software running on the host computer and the non-identical software code running on the secondary computer provides the same overall result, then the supervisory MCU can be more confident that the overall result is correct and that the vulnerability in the software or hardware on the host computer does not cause a substantial error.

[0163] In some examples, the output of the ADAS system 738 can be fed to the perception block of the main computer and / or the dynamic driving task block of the main computer. For example, if the ADAS system 738 indicates a forward collision warning due to an object being immediately in front, the perception block can use this information when identifying the object. In other examples, the auxiliary computer can have its own neural network, which is trained and thus reduces the risk of false positives as described herein.

[0164] The vehicle 700 can further include an infotainment SoC 730 (e.g., an in-vehicle infotainment system (IVI)). Although illustrated and described as an SoC, the infotainment system can not be an SoC and can include two or more discrete components. The infotainment SoC 730 can include a combination of hardware and software that can be used to provide audio (e.g., music, personal digital assistant, navigation instructions, news, radio, etc.), video (e.g., TV, movies, streaming, etc.), telephone (e.g., hands-free calling), network connectivity (e.g., LTE, Wi-Fi, etc.), and / or information services (e.g., navigation system, rear parking assistance, radio data system, vehicle-related information such as fuel level, total distance covered, brake fuel level, oil level, door open / close, air filter information, etc.) to the vehicle 700. For example, the infotainment SoC 730 can include a radio, a disc player, a navigation system, a video player, USB and Bluetooth connectivity, an in-vehicle computer, in-vehicle entertainment, Wi-Fi, steering wheel audio controls, hands-free voice controls, a head-up display (HUD), an HMI display 734, a telematics device, a control panel (e.g., for controlling various components, features, and / or systems, and / or interacting therewith), and / or other components. The infotainment SoC 730 can further be used to provide information (e.g., visual and / or auditory) to the user of the vehicle, such as information from the ADAS system 738, autonomous driving information such as planned vehicle maneuvers, trajectories, surrounding environment information (e.g., intersection information, vehicle information, road information, etc.), and / or other information.

[0165] The infotainment SoC 730 can include GPU functionality. The infotainment SoC 730 can communicate with other devices, systems, and / or components of the vehicle 700 via a bus 702 (e.g., a CAN bus, Ethernet, etc.). In some examples, the infotainment SoC 730 can be coupled to a supervisory MCU such that in the event of a failure of the main controller 736 (e.g., the main and / or standby computer of the vehicle 700), the GPU of the infotainment system can perform some self-driving functions. In such examples, the infotainment SoC 730 can place the vehicle 700 in a driver safe parking mode as described herein.

[0166] Vehicle 700 may further include an instrument cluster 732 (such as a digital instrument panel, an electronic instrument cluster, a digital instrument panel, etc.). The instrument cluster 732 may include a controller and / or a supercomputer (such as a discrete controller or supercomputer). The instrument cluster 732 may include a set of instruments, such as a speedometer, fuel level, oil pressure, tachometer, odometer, turn indicator, shift position indicator, seat belt warning light, parking brake warning light, engine fault light, airbag (SRS) system information, lighting controls, safety system controls, navigation information, and so on. In some examples, information may be displayed and / or shared between the infotainment SoC 730 and the instrument cluster 732. Thus, the instrument cluster 732 may be included as part of the infotainment SoC 730, or vice versa.

[0167] Figure 7D A system schematic diagram of communication between a cloud-based server and Figure 7A an exemplary autonomous vehicle 700 according to some embodiments of the present disclosure. The system 776 may include a server 778, a network 790, and vehicles including the vehicle 700. The server 778 may include a plurality of GPUs 784(A)-784(H) (collectively referred to herein as GPUs 784), PCIe switches 782(A)-782(D) (collectively referred to herein as PCIe switches 782), and / or CPUs 780(A)-780(B) (collectively referred to herein as CPUs 780). The GPUs 784, CPUs 780, and PCIe switches may be interconnected by high-speed interconnects such as, for example and without limitation, the NVLink interface 788 developed by NVIDIA and / or PCIe connections 786. In some examples, the GPUs 784 are connected via NVLink and / or an NVSwitch SoC, and the GPUs 784 and the PCIe switches 782 are connected via a PCIe interconnect. Although eight GPUs 784, two CPUs 780, and two PCIe switches are illustrated, this is not intended to be limiting. Depending on the embodiment, each of the servers 778 may include any number of GPUs 784, CPUs 780, and / or PCIe switches. For example, each of the servers 778 may include eight, sixteen, thirty-two, and / or more GPUs 784.

[0168] Server 778 can receive image data via network 790 and from a vehicle, the image data representing an image showing an unexpected or changed road condition such as a recently started roadwork. Server 778 can transmit a neural network 792, an updated neural network 792, and / or map information 794 via network 790 and to the vehicle, including information about traffic and road conditions. Updates to the map information 794 can include updates to the HD map 722, such as information about construction sites, potholes, curves, floods, or other obstacles. In some examples, the neural network 792, the updated neural network 792, and / or the map information 794 can be represented and / or generated based on data received from new training and / or from any number of vehicles in the environment and / or experience from training performed at a data center (e.g., using server 778 and / or other servers).

[0169] Server 778 can be used to train a machine learning model (e.g., a neural network) based on training data. The training data can be generated using vehicles, and / or can be generated in a simulation (e.g., using a game engine). In some examples, the training data is labeled (e.g., in cases where the neural network benefits from supervised learning) and / or undergoes other preprocessing, while in other examples, the training data is not labeled and / or preprocessed (e.g., in cases where the neural network does not require supervised learning). The training can be performed according to any one or more classes of machine learning techniques, including but not limited to the following classes: supervised training, semi-supervised training, unsupervised training, self-learning, reinforcement learning, federated learning, transfer learning, feature learning (including principal component and clustering analysis), multilinear subspace learning, manifold learning, representation learning (including alternative dictionary learning), rule-based machine learning, anomaly detection, and any variants or combinations thereof. Once the machine learning model is trained, the machine learning model can be used by the vehicle (e.g., transmitted to the vehicle via network 790), and / or the machine learning model can be used by server 778 to remotely monitor the vehicle.

[0170] In some examples, server 778 can receive data from a vehicle and apply the data to the latest real-time neural network for real-time intelligent inference. Server 778 can include a deep learning supercomputer powered by GPU 784 and / or a dedicated AI computer, such as the DGX and DGX Station machines developed by NVIDIA. However, in some examples, server 778 can include a deep learning infrastructure of a data center powered only by a CPU.

[0171] The deep learning infrastructure of server 778 may be capable of fast real-time inference and can use this ability to evaluate and verify the health of the processors, software, and / or associated hardware in vehicle 700. For example, the deep learning infrastructure can receive periodic updates from vehicle 700, such as an image sequence and / or objects located in that image sequence that vehicle 700 has identified (e.g., via computer vision and / or other machine learning object classification techniques). The deep learning infrastructure can run its own neural network to identify the objects and compare them to the objects identified by vehicle 700. If the results do not match and the infrastructure concludes that the AI in vehicle 700 has malfunctioned, then server 778 can transmit a signal to vehicle 700, instructing the fail-safe computer in vehicle 700 to take control, notify the passengers, and complete a safe parking operation.

[0172] For inference, server 778 can include a GPU 784 and one or more programmable inference accelerators (e.g., NVIDIA's TensorRT 3). The combination of a GPU-powered server and inference acceleration can enable real-time response. In other examples, such as when performance is less critical, a CPU, FPGA, and other processor-powered servers can be used for inference.

[0173] Example computing device

[0174] Figure 8 A block diagram of an example computing device 800 suitable for implementing some embodiments of the present disclosure. Computing device 800 can include an interconnect system 802 that directly or indirectly couples the following devices: memory 804, one or more central processing units (CPUs) 806, one or more graphics processing units (GPUs) 808, a communication interface 810, input / output (I / O) ports 812, input / output components 814, a power supply 816, one or more presentation components 818 (e.g., a display), and one or more logic units 820. In at least one embodiment, computing device 800 can include one or more virtual machines (VMs), and / or any of its components can include virtual components (e.g., virtual hardware components). For non-limiting examples, one or more GPUs 808 can include one or more vGPUs, one or more CPUs 806 can include one or more vCPUs, and / or one or more logic units 820 can include one or more virtual logic units. Thus, computing device 800 can include discrete components (e.g., a complete GPU dedicated to computing device 800), virtual components (e.g., a portion of a GPU dedicated to computing device 800), or a combination thereof.

[0175] Although Figure 8Each of the boxes is shown as being connected via an interconnect system 802 having circuitry, but this is not intended to be limiting and is for clarity only. For example, in some embodiments, a rendering component 818 such as a display device may be considered an I / O component 814 (e.g., if the display is a touch screen). As another example, the CPU 806 and / or GPU 808 may include memory (e.g., memory 804 may represent a storage device in addition to the memory of the GPU 808, CPU 806, and / or other components). Thus, Figure 8 the computing devices are merely illustrative. No distinction is made between categories such as "workstation", "server", "laptop", "desktop", "tablet", "client device", "mobile device", "handheld device", "gaming console", "electronic control unit (ECU)", "virtual reality system", and / or other device or system types because all of these are considered within the Figure 8 scope of the computing devices.

[0176] The interconnect system 802 may represent one or more links or buses, such as an address bus, a data bus, a control bus, or a combination thereof. The interconnect system 802 may include one or more types of links or buses, such as an Industry Standard Architecture (ISA) bus, an Extended Industry Standard Architecture (EISA) bus, a Video Electronics Standards Association (VESA) bus, a Peripheral Component Interconnect (PCI) bus, a Peripheral Component Interconnect Express (PCIe) bus, and / or another type of bus or link. In some embodiments, there are direct connections between components. As an example, the CPU 806 may be directly connected to the memory 804. Additionally, the CPU 806 may be directly connected to the GPU 808. In cases where there is a direct or point-to-point connection between components, the interconnect system 802 may include a PCIe link to effect the connection. In these examples, a PCI bus need not be included in the computing device 800.

[0177] The memory 804 may include any of a variety of computer-readable media. Computer-readable media can be any available media that can be accessed by the computing device 800. Computer-readable media can include volatile and non-volatile media as well as removable and non-removable media. By way of example and not limitation, computer-readable media can include computer storage media and communication media.

[0178] A computer storage medium can include volatile and non-volatile media and / or removable and non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules, and / or other data types. For example, memory 804 can store computer-readable instructions (e.g., which represent programs and / or program elements such as an operating system). A computer storage medium can include, but is not limited to, RAM, ROM, EEPROM, flash memory or other storage technologies, CD-ROM, digital versatile disks (DVDs) or other optical disk storage devices, magnetic tape cartridges, magnetic tape, magnetic disk storage devices or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by computing device 800. As used herein, a computer storage medium does not include the signal itself.

[0179] A computer storage medium can include computer-readable instructions, data structures, program modules, and / or other data types in a modulated data signal such as a carrier wave or other transmission mechanism, and includes any information conveyance medium. The term "modulated data signal" can refer to a signal that has one or more of its characteristics set or changed in such a manner as to encode information in the signal. By way of example, and not limitation, a computer storage medium can include wired media such as a wired network or direct wired connection, and wireless media such as sound, RF, infrared, and other wireless media. Any of the foregoing combinations should also be included within the scope of computer-readable media.

[0180] CPU 806 can be configured to execute at least some of the computer-readable instructions to control one or more components of computing device 800 to perform one or more of the methods and / or processes described herein. Each of CPUs 806 can include one or more cores (e.g., one, two, four, eight, twenty-eight, seventy-two, etc.) capable of simultaneously processing a large number of software threads. CPU 806 can include any type of processor, and can include different types of processors depending on the type of computing device 800 being implemented (e.g., a processor with fewer cores for a mobile device and a processor with more cores for a server). For example, depending on the type of computing device 800, the processor can be an advanced RISC machine (ARM) processor implemented using reduced instruction set computing (RISC) or an x86 processor implemented using complex instruction set computing (CISC). In addition to one or more microprocessors or supplementary coprocessors such as a math coprocessor, computing device 800 can also include one or more CPUs 806.

[0181] In addition to or instead of the CPU 806, the GPU 808 can also be configured to execute at least some computer-readable instructions to control one or more components of the computing device 800 to perform one or more methods and / or processes described herein. One or more GPUs 808 can be an integrated GPU (e.g., having one or more CPUs 806) and / or one or more GPUs 808 can be a discrete GPU. In an embodiment, one or more GPUs 808 can be a coprocessor of one or more CPUs 806. The computing device 800 can use the GPU 808 to render graphics (e.g., 3D graphics) or perform general computing. For example, the GPU 808 can be used for general-purpose computing on the GPU (GPGPU). The GPU 808 can include hundreds or thousands of cores capable of processing hundreds or thousands of software threads simultaneously. The GPU 808 can generate pixel data for an output image in response to a rendering command (e.g., a rendering command received from the CPU 806 via a host interface). The GPU 808 can include graphics memory such as display memory for storing pixel data or any other suitable data (e.g., GPGPU data). The display memory can be included as part of the memory 804. The GPU 808 can include two or more GPUs operating in parallel (e.g., via a link). The link can directly connect the GPUs (e.g., using NVLINK) or can connect the GPUs through a switch (e.g., using NVSwitch). When combined, each GPU 808 can generate pixel data or GPGPU data for different parts of the output or for different outputs (e.g., the first GPU for the first image and the second GPU for the second image). Each GPU can include its own memory or can share memory with other GPUs.

[0182] In addition to or instead of the CPU 806 and / or the GPU 808, the logic unit 820 can be configured to execute at least some computer-readable instructions to control one or more components of the computing device 800 to perform one or more methods and / or processes described herein. In an embodiment, the CPU 806, the GPU 808, and / or the logic unit 820 can perform any combination of methods, processes, and / or portions thereof discretely or jointly. One or more logic units 820 can be part of and / or integrated in one or more CPUs 806 and / or one or more GPUs 808, and / or one or more logic units 820 can be discrete components of or otherwise external to the CPU 806 and / or the GPU 808. In an embodiment, one or more logic units 820 can be a processor of one or more CPUs 806 and / or one or more GPUs 808.

[0183] Examples of the logic unit 820 include one or more processing cores and / or their components, such as data processing units (DPUs), tensor cores (TCs), tensor processing units (TPUs), pixel vision cores (PVCs), vision processing units (VPUs), graphics processing clusters (GPCs), texture processing clusters (TPCs), streaming multiprocessors (SMs), tree traversal units (TTUs), artificial intelligence accelerators (AIAs), deep learning accelerators (DLAs), arithmetic logic units (ALUs)), application specific integrated circuits (ASICs), floating point units (FPUs), input / output (I / O) elements, peripheral component interconnect (PCI) or peripheral component interconnect express (PCIe) elements, etc.

[0184] The communication interface 810 may include one or more receivers, transmitters, and / or transceivers that enable the computing device 800 to communicate with other computing devices via an electronic communication network, including wired and / or wireless communication. The communication interface 810 may include components and functionality that enable communication over any of several different networks, such as wireless networks (e.g., Wi-Fi, Z-Wave, Bluetooth, Bluetooth LE, ZigBee, etc.), wired networks (e.g., communication via Ethernet or InfiniBand), low power wide area networks (e.g., LoRaWAN, SigFox, etc.), and / or the Internet. In one or more embodiments, the logic unit 820 and / or the communication interface 810 may include one or more data processing units (DPUs) to directly transfer data received over the network and / or via the interconnect system 802 to one or more GPUs 808 (e.g., the memory in the GPUs 808).

[0185] The I / O port 812 can enable the computing device 800 to be logically coupled to other devices including I / O components 814, presentation components 818, and / or other components, some of which can be built into (e.g., integrated into) the computing device 800. Exemplary I / O components 814 include microphones, mice, keyboards, joysticks, game pads, game controllers, dish satellite antennas, scanners, printers, wireless devices, and so on. The I / O components 814 can provide a natural user interface (NUI) for processing user-generated air gestures, voice, or other physiological inputs. In some instances, the input can be transmitted to an appropriate network element for further processing. The NUI can implement any combination of speech recognition, stylus recognition, facial recognition, biometric recognition, on-screen and near-screen gesture recognition, air gestures, head and eye tracking, and touch recognition associated with the display of the computing device 800 (described in more detail below). The computing device 800 can include depth cameras such as stereo camera systems, infrared camera systems, RGB camera systems, touchscreen technologies, and combinations thereof for gesture detection and recognition. Additionally, the computing device 800 can include an accelerometer or gyroscope enabling motion detection (e.g., as part of an inertial measurement unit (IMU)). In some examples, the output of the accelerometer or gyroscope can be used by the computing device 800 to render immersive augmented reality or virtual reality.

[0186] The power supply 816 can include a hard-wired power supply, a battery power supply, or a combination thereof. The power supply 816 can supply power to the computing device 800 to enable the components of the computing device 800 to operate.

[0187] The presentation component 818 can include a display (e.g., a monitor, a touch screen, a television screen, a head-up display (HUD), other display types, or a combination thereof), speakers, and / or other presentation components. The presentation component 818 can receive data from other components (e.g., the GPU 808, the CPU 806, the DPU, etc.) and output the data (e.g., as images, videos, sounds, etc.).

[0188] Example data center

[0189] Figure 9 An example data center 900 is shown, which can be used in at least one embodiment of the present disclosure. The data center 900 can include a data center infrastructure layer 910, a framework layer 920, a software layer 930, and an application layer 940.

[0190] As Figure 9As shown, the data center infrastructure layer 910 may include a resource coordinator 912, grouped computing resources 914, and node computing resources (“node C.R.”) 916(1)-916(N), where “N” represents any whole positive integer. In at least one embodiment, the node C.R. 916(1)-916(N) may include, but is not limited to, any number of central processing units (CPUs) or other processors (including DPUs, accelerators, field programmable gate arrays (FPGAs), graphics processors or graphics processing units (GPUs), etc.), memory devices (such as dynamic read-only memory), storage devices (such as solid state drives or disk drives), network input / output (NW I / O) devices, network switches, virtual machines (VMs), power modules, and cooling modules, etc. In some embodiments, one or more of the node C.R. 916(1)-916(N) may correspond to a server having one or more of the above computing resources. Additionally, in some embodiments, the node C.R. 916(1)-916(N) may include one or more virtual components, such as vGPUs, vCPUs, etc., and / or one or more of the node C.R. 916(1)-916(N) may correspond to a virtual machine (VM).

[0191] In at least one embodiment, the grouped computing resources 914 may include separate groupings (not shown) of node C.R. 916 housed within one or more racks, or many racks (also not shown) within data centers located in various geographical locations. The separate groupings of node C.R. 916 within the grouped computing resources 914 may include grouped computing, network, memory, or storage resources that may be configured or allocated to support one or more workloads. In at least one embodiment, several node C.R. 916 including CPUs, GPUs, DPUs, and / or other processors may be grouped within one or more racks to provide computing resources to support one or more workloads. One or more racks may also include any number of power modules, cooling modules, and / or network switches in any combination.

[0192] The resource coordinator 912 may configure or otherwise control one or more of the node C.R. 916(1)-916(N) and / or the grouped computing resources 914. In at least one embodiment, the resource coordinator 912 may include a software design infrastructure (SDI) management entity for the data center 900. The resource coordinator 912 may include hardware, software, or some combination thereof.

[0193] In at least one embodiment, as Figure 9As shown, the framework layer 920 may include a job scheduler 933, a configuration manager 934, a resource manager 936, and a distributed file system 938. The framework layer 920 may include a framework for software 932 that supports the software layer 930 and / or one or more applications 942 of the application layer 940. The software 932 or the application 942 may respectively include web-based service software or applications, such as service software or applications provided by Amazon Web Services, Google Cloud, and Microsoft Azure. The framework layer 920 may be, but is not limited to, a free and open-source software web application framework, such as Apache Spark that can utilize the distributed file system 938 for large-scale data processing (e.g., "big data"). TM (hereinafter referred to as "Spark"). In at least one embodiment, the job scheduler 933 may include a Spark driver for facilitating the scheduling of workloads supported by the various layers of the data center 900. In at least one embodiment, the configuration manager 934 may be capable of configuring different layers, such as the software layer 930 and the framework layer 920 including Spark and the distributed file system 938 for supporting large-scale data processing. The resource manager 936 is capable of managing the cluster or grouped computing resources mapped to or allocated for supporting the distributed file system 938 and the job scheduler 933. In at least one embodiment, the cluster or grouped computing resources may include grouped computing resources 914 at the data center infrastructure layer 910. The resource manager 936 may coordinate with the resource coordinator 912 to manage these mapped or allocated computing resources.

[0194] In at least one embodiment, the software 932 included in the software layer 930 may include software used by at least a portion of the nodes C.R. 916(1)-916(N), the grouped computing resources 914, and / or the distributed file system 938 of the framework layer 920. One or more types of software may include, but are not limited to, Internet web search software, email virus browsing software, database software, and streaming video content software.

[0195] In at least one embodiment, one or more applications 942 included in the application layer 940 may include one or more types of applications used by at least a portion of nodes C.R. 916(1)-916(N), the grouped computing resources 914, and / or the distributed file system 938 of the framework layer 920. The one or more types of applications may include, but are not limited to, any number of genomics applications, cognitive computing and machine learning applications, including training or inference software, machine learning framework software (such as PyTorch, TensorFlow, Caffe, etc.), and / or other machine learning applications used in conjunction with one or more embodiments.

[0196] In at least one embodiment, any one of the configuration manager 934, the resource manager 936, and the resource coordinator 912 may implement any number and type of self-modifying actions based on any amount and type of data obtained in any technically feasible manner. The self-modifying actions may relieve the data center operator of the data center 900 from making potentially bad configuration decisions and may avoid underutilized and / or poorly performing portions of the data center.

[0197] The data center 900 may include tools, services, software, or other resources for training one or more machine learning models or using one or more machine learning models to predict or infer information according to one or more embodiments described herein. For example, a machine learning model may be trained by calculating weight parameters according to a neural network architecture by using the software and computing resources described above with respect to the data center 900. In at least one embodiment, by using the weight parameters calculated by one or more training techniques, the resources described above with respect to the data center 900 may be used to infer or predict information using the trained machine learning model corresponding to one or more neural networks, such as, but not limited to, those described herein.

[0198] In at least one embodiment, the data center 900 may use a CPU, an application specific integrated circuit (ASIC), a GPU, an FPGA, and / or other hardware (or corresponding virtual computing resources) to perform training and / or inference using the above resources. In addition, the one or more software and / or hardware resources described above may be configured as a service to allow a user to train or perform information inference, such as image recognition, speech recognition, or other artificial intelligence services.

[0199] Example Network Environment

[0200] The network environment suitable for implementing the embodiments of the present disclosure may include one or more client devices, servers, network attached storage (NAS), other backend devices, and / or other device types. The client devices, servers, and / or other device types (e.g., each device) may be implemented on one or more instances of the Figure 8 computing device 800 - for example, each device may include similar components, features, and / or functions of the computing device 800. Additionally, in the case of implementing backend devices (e.g., servers, NAS, etc.), the backend devices may be included as part of a data center 900, an example of which is described in more detail herein with respect to Figure 9 more detail.

[0201] The components of the network environment may communicate with each other via a network, which may be wired, wireless, or both. The network may include multiple networks, or networks within multiple networks. By way of example, the network may include one or more wide area networks (WANs), one or more local area networks (LANs), one or more public networks (e.g., the Internet and / or the public switched telephone network (PSTN)), and / or one or more private networks. In the case where the network includes a wireless telecommunications network, components such as base stations, communication towers, or even access points (and other components) may provide a wireless connection.

[0202] A compatible network environment may include one or more peer - to - peer network environments (in which case servers may not be included in the network environment), and one or more client - server network environments (in which case one or more servers may be included in the network environment). In a peer - to - peer network environment, the functions described herein with respect to servers may be implemented on any number of client devices.

[0203] In at least one embodiment, the network environment may include one or more cloud - based network environments, distributed computing environments, combinations thereof, etc. A cloud - based network environment may include a framework layer, a job scheduler, a resource manager, and a distributed file system implemented on one or more servers, which may include one or more core network servers and / or edge servers. The framework layer may include a framework for supporting one or more applications of a software layer and / or an application layer. The software or application may respectively include network - based service software or applications. In an embodiment, one or more client devices may use network - based service software or applications (e.g., by accessing the service software and / or applications via one or more application programming interfaces (APIs)). The framework layer may be, but is not limited to, a type of free and open - source software web application framework, such as may be used for large - scale data processing (e.g., “big data”) using a distributed file system.

[0204] A cloud-based network environment can provide cloud computing and / or cloud storage that perform any combination of the computing and / or data storage functions (or one or more portions thereof) described herein. Any of these various functions can be distributed across multiple locations from a central or core server (e.g., one or more data centers that can be distributed across states, regions, countries, globally, etc.). If the connection to a user (e.g., a client device) is relatively close to an edge server, the core server can assign at least a portion of the function to the edge server. A cloud-based network environment can be private (e.g., limited to a single organization), public (e.g., available to many organizations), and / or a combination thereof (e.g., a hybrid cloud environment).

[0205] Client devices can include at least some of the components, features, and functions of the example computing device 800 described herein with respect to Figure 8 As examples and not limitations, client devices can be embodied as personal computers (PCs), laptop computers, mobile devices, smartphones, tablet computers, smartwatches, wearable computers, personal digital assistants (PDAs), MP3 players, virtual reality headsets, global positioning system (GPS) or devices, video players, cameras, surveillance devices or systems, vehicles, boats, aircraft, virtual machines, drones, robots, handheld communication devices, hospital devices, gaming devices or systems, entertainment systems, in-vehicle computer systems, embedded system controllers, remote controls, appliances, consumer electronic devices, workstations, edge devices, any combination of these described devices, or any other suitable device.

[0206] The present disclosure can be described in the general context of machine-usable instructions or computer code, including computer-executable instructions such as program modules, that are executed by a computer or other machine such as a personal digital assistant or other handheld device. Generally, program modules, including routines, programs, objects, components, data structures, etc., refer to code that performs particular tasks or implements particular abstract data types. The present disclosure can be practiced in a variety of system configurations, including handheld devices, consumer electronics, general-purpose computers, more specialized computing devices, etc. The present disclosure can also be practiced in a distributed computing environment where tasks are performed by remote processing devices linked through a communications network.

[0207] As used herein, the recitation of "and / or" with respect to two or more elements shall be construed to mean only one element or a combination of elements. For example, "element A, element B, and / or element C" may include only element A, only element B, only element C, element A and element B, element A and element C, element B and element C, or element A, B, and C. Further, "at least one of element A or element B" may include at least one of element A, at least one of element B, or at least one of element A and at least one of element B. Still further, "at least one of element A and element B" may include at least one of element A, at least one of element B, or at least one of element A and at least one of element B.

[0208] The subject matter of the present disclosure is described in detail herein to meet statutory requirements. However, the description itself is not intended to limit the scope of the present disclosure. On the contrary, the inventors have contemplated that the claimed subject matter may also be embodied in other ways, including steps different from or a combination of steps similar to those described herein in connection with other current or future technologies. Moreover, although the terms "step" and / or "block" may be used herein to imply different elements of a method employed, these terms should not be construed as implying any particular order among or between the various steps disclosed herein unless the order of the steps is expressly recited.

[0209] Example text support

[0210] The disclosure of the present application also includes the following numbered clauses:

[0211] Clause 1. One or more processors, including processing circuitry configured to: apply two-dimensional (2D) sensor data generated from a source sensor device of a machine to at least a portion of a three-dimensional (3D) perception network.

[0212] Clause 2. The one or more processors of Clause 1, wherein the processing circuitry is further configured to: apply analog 2D sensor data to the at least a portion of the 3D perception network, the analog 2D sensor data being simulated based at least on an analog source sensor device and an analog target sensor device.

[0213] Clause 3. The one or more processors of Clause 1 or 2, wherein the processing circuitry is further configured to: update the 3D perception network based at least on the 2D sensor data and the analog 2D sensor data.

[0214] Clause 4. The one or more processors of Clause 1, 2, or 3, wherein the processing circuitry is further configured to perform one or more operations on the machine using the updated 3D perception network.

[0215] Clause 5. One or more processors according to Clause 1, 2, 3, or 4, wherein the processing circuitry is further configured to: inject a representation of a style extracted from real 2D sensor data into one or more features extracted from simulated 2D sensor data.

[0216] Clause 6. One or more processors according to Clause 1, 2, 3, or 4, wherein the processing circuitry is further configured to: update the 3D perception network based at least on alternating between one or more first paths that apply real 2D sensor data to at least a portion of the 3D perception network and one or more second paths that apply simulated 2D sensor data to at least a portion of the 3D perception network.

[0217] Clause 7. One or more processors according to Clause 1, 2, 3, or 4, wherein the processing circuitry is further configured to: update one or more shared weights shared by one or more first paths and one or more second paths based at least on alternating between one or more first paths that apply real 2D sensor data to at least a portion of the 3D perception network and one or more second paths that apply simulated 2D sensor data to at least a portion of the 3D perception network.

[0218] Clause 8. One or more processors according to Clause 1, 2, 3, or 4, wherein the processing circuitry is further configured to: update the 3D perception network based at least on updating a viewpoint adjustment network that includes one or more layers of the 3D perception network shared by one or more first paths that process real 2D sensor data across the viewpoint adjustment network and one or more second paths that process simulated 2D sensor data of the viewpoint adjustment network.

[0219] Clause 9. One or more processors according to Clause 1, 2, 3, or 4, wherein the source sensor device and the simulated source sensor device represent a first sensor configuration of the machine, and the first sensor configuration is different from a second sensor configuration of the simulated target sensor device.

[0220] Clause 10. One or more processors according to Clause 1, 2, 3, or 4, wherein the processing circuitry is further configured to: generate a simulated 2D sensor data frame and simulated ground truth data for at least one sensor of one or more sensors of the simulated source sensor device and at least one corresponding sensor of the simulated target sensor device within at least one of the one or more simulated time slices.

[0221] Clause 11. One or more processors according to Clause 1, 2, 3, or 4, wherein the processing circuitry is further configured to: update a first viewpoint adjustment network based at least on applying real 2D sensor data and simulated 2D sensor data to one or more layers of a 3D perception network including a first viewpoint adjustment network, and then update a second viewpoint adjustment network based at least on applying the simulated 2D sensor data to one or more layers of a 3D perception network including a second viewpoint adjustment network.

[0222] Clause 12. One or more processors according to Clause 1, 2, 3, or 4, wherein the 3D perception network is updated to perform at least one of a 2D-to-3D transformation or 2D-3D fusion.

[0223] Clause 13. One or more processors according to Clause 1, 2, 3, or 4, wherein at least one of the 3D perception network or one or more processors is included in at least one of the following: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing simulation operations; a system for performing digital twin operations; a system for performing optical transmission simulation; a system for performing collaborative content creation of 3D assets; a system for performing deep learning operations; a system for performing remote operations; a system for performing real-time streaming; a system for generating or presenting one or more of augmented reality content, virtual reality content, or mixed reality content; a system implemented using edge devices; a system implemented using robots; a system for performing conversational AI operations; a system implementing one or more language models; a system implementing one or more large language models (LLMs); a system implementing one or more vision language models (VLMs); a system for generating synthetic data; a system for generating synthetic data using AI; a system including one or more virtual machines (VMs); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources.

[0224] Clause 14. A system including one or more processors configured to perform one or more operations on a self machine using a three-dimensional (3D) perception network, the 3D perception network being updated at least based on alternating between applying real two-dimensional (2D) sensor data generated using a source sensor device of the self machine to at least a portion of the 3D perception network and applying simulated 2D sensor data to the at least a portion of the 3D perception network, the simulated 2D sensor data being simulated based at least on a simulated source sensor device and a simulated target sensor device.

[0225] Clause 15. The system according to Clause 14, wherein one or more processors are further configured to: inject a representation of a style extracted from real 2D sensor data into features extracted from simulated 2D sensor data, at least based on applying the simulated 2D sensor data to at least a portion of the 3D perception network.

[0226] Clause 16. The system according to Clause 14, wherein one or more processors are further configured to: update the 3D perception network, at least based on alternating between one or more first paths for applying real 2D sensor data to at least a portion of the 3D perception network and one or more second paths for applying simulated 2D sensor data to at least a portion of the 3D perception network.

[0227] Clause 17. The system according to Clause 14, wherein one or more processors are further configured to: update shared weights shared by one or more first paths and one or more second paths, at least based on alternating between one or more first paths for applying real 2D sensor data to at least a portion of the 3D perception network and one or more second paths for applying simulated 2D sensor data to at least a portion of the 3D perception network.

[0228] Clause 18. The system according to Clause 14, wherein one or more processors are further configured to: update the 3D perception network, at least based on updating a viewpoint adjustment network that includes one or more layers of the 3D perception network shared by one or more first paths for processing real 2D sensor data across the viewpoint adjustment network and one or more second paths for processing simulated 2D sensor data in the viewpoint adjustment network.

[0229] Clause 19. The system according to Clause 14, wherein the source sensor device and the simulated source sensor device represent a first sensor configuration of the machine, and the first sensor configuration is different from a second sensor configuration of the simulated target sensor device.

[0230] Clause 20. The system according to Clause 14, wherein one or more processors are further configured to: generate a simulated 2D sensor data frame and simulated ground truth data for at least one sensor of one or more sensors of the simulated source sensor device and at least one corresponding sensor of the simulated target sensor device, within at least one of one or more simulated time slices.

[0231] Clause 21. The system according to Clause 14, wherein at least one of the system or the 3D perception network is included in at least one of the following: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing simulation operations; a system for performing digital twin operations; a system for performing optical transmission simulation; a system for performing collaborative content creation of 3D assets; a system for performing deep learning operations; a system for performing remote operations; a system for performing real-time streaming; a system for generating or presenting one or more of augmented reality content, virtual reality content, or mixed reality content; a system implemented using edge devices; a system implemented using robots; a system for performing conversational AI operations; a system implementing one or more language models; a system implementing one or more large language models (LLMs); a system implementing one or more vision language models (VLMs); a system for generating synthetic data; a system for generating synthetic data using AI; a system including one or more virtual machines (VMs); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources.

[0232] Clause 22. A method, comprising: updating a three-dimensional (3D) reconstruction network by alternating at least based on applying real two-dimensional (2D) sensor data generated using a first sensor device of a machine to at least a portion of the 3D reconstruction network.

[0233] Clause 23. The method according to Clause 22, further comprising: applying simulated 2D sensor data generated by simulating at least a first sensor device and a second sensor device to the at least a portion of the 3D reconstruction network.

[0234] Clause 24. The method according to Clause 22 or 23, further comprising: performing one or more operations on the machine using the 3D reconstruction network.

[0235] Clause 25. The method according to Clause 22, 23, or 24, wherein the method is performed by at least one of the following or the 3D reconstruction network is included in at least one of the following: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing simulation operations; a system for performing digital twin operations; a system for performing optical transmission simulation; a system for performing collaborative content creation of 3D assets; a system for performing deep learning operations; a system for performing remote operations; a system for performing real-time streaming; a system for generating or presenting one or more of augmented reality content, virtual reality content, or mixed reality content; a system implemented using edge devices; a system implemented using robots; a system for performing conversational AI operations; a system implementing one or more language models; a system implementing one or more large language models (LLMs); a system implementing one or more vision language models (VLMs); a system for generating synthetic data; a system for generating synthetic data using AI; a system including one or more virtual machines (VMs); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources.

[0236] Clause 26. One or more processors, including processing circuitry, the processing circuitry for: performing one or more operations on a self-machine using a three-dimensional (3D) perception network, the 3D perception network being updated at least based on applying simulated two-dimensional (2D) sensor data to at least a portion of the 3D perception network.

[0237] Clause 27. The one or more processors according to Clause 26, wherein the simulated 2D sensor data is generated at least based on a simulated target sensor device and a source sensor device of the self-machine.

[0238] Clause 28. The one or more processors according to Clause 26 or 27, wherein the processing circuitry is further for: updating the 3D perception network at least based on an updated viewpoint adjustment network, the viewpoint adjustment network including one or more layers of the 3D perception network, the one or more layers being replicated as at least a portion of a plurality of channels of the viewpoint adjustment network.

[0239] Clause 29. The one or more processors according to Clause 26 or 27, wherein the processing circuitry is further for: updating the 3D perception network at least based on one or more first paths of applying the simulated 2D sensor data of the simulated source sensor device to the at least a portion of the 3D perception network and one or more second paths of applying the simulated 2D sensor data of the simulated target sensor device to the at least a portion of the 3D perception network.

[0240] Clause 30. One or more processors according to Clause 26 or 27, wherein the 3D perception network includes a view transformation that generates transformed features, and wherein the processing circuit is further configured to: update the 3D perception network at least based on a loss that minimizes a difference between transformed features of one or more first paths and one or more second paths of at least a portion of the 3D perception network.

[0241] Clause 31. One or more processors according to Clause 26 or 27, wherein the processing circuit is further configured to: update the 3D perception network at least based on a loss that minimizes a difference between view-transformed features of one or more first paths that process analog 2D sensor data corresponding to a source sensor device and view-transformed features of one or more second paths that process analog 2D sensor data corresponding to a target sensor device.

[0242] Clause 32. One or more processors according to Clause 26 or 27, wherein the source sensor device represents a first sensor configuration that is different from a second sensor configuration of the target sensor device.

[0243] Clause 33. One or more processors according to Clause 26 or 27, wherein the processing circuit is further configured to: generate analog 2D sensor data frames and analog ground truth data for at least one sensor of one or more analog sensors of the source sensor device and at least one corresponding sensor of the target sensor device within at least one time slice of one or more simulated time slices.

[0244] Clause 34. One or more processors according to Clause 26 or 27, wherein the processing circuit is further configured to: update a first viewpoint adjustment network at least based on applying real 2D sensor data and analog 2D sensor data to the first viewpoint adjustment network including one or more layers of the 3D perception network, and then update a second viewpoint adjustment network at least based on applying the analog 2D sensor data to the second viewpoint adjustment network including one or more layers of the 3D perception network.

[0245] Clause 35. One or more processors according to Clause 26 or 27, wherein the 3D perception network is updated to perform at least one of a 2D-to-3D transformation or 2D-3D fusion.

[0246] Clause 36. One or more processors according to Clause 26 or 27, wherein at least one of the 3D perception network or the one or more processors is included in at least one of the following: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing simulation operations; a system for performing digital twin operations; a system for performing optical transmission simulation; a system for performing collaborative content creation of 3D assets; a system for performing deep learning operations; a system for performing remote operations; a system for performing real-time streaming; a system for generating or presenting one or more of augmented reality content, virtual reality content, or mixed reality content; a system implemented using edge devices; a system implemented using robots; a system for performing conversational AI operations; a system implementing one or more language models; a system implementing one or more large language models (LLMs); a system implementing one or more vision language models (VLMs); a system for generating synthetic data; a system for generating synthetic data using AI; a system applying the use of generative AI to generate synthetic data; a system including one or more virtual machines (VMs); a system implemented at least in part in a data center; or a system implemented at least in part using cloud computing resources.

[0247] Clause 37. A system comprising one or more processors for performing one or more operations on a self-machine using a three-dimensional (3D) reconstruction network, the 3D reconstruction network being updated at least based on applying simulated two-dimensional (2D) sensor data to at least a portion of the 3D reconstruction network.

[0248] Clause 38. The system according to Clause 37, wherein the simulated 2D sensor data is generated at least based on a simulated target sensor device and a source sensor device of the self-machine.

[0249] Clause 39. The system according to Clause 37 or 38, wherein the one or more processors are further configured to: update the 3D reconstruction network at least based on an updated viewpoint adjustment network, the viewpoint adjustment network including one or more layers of the 3D reconstruction network, the one or more layers being replicated as at least a portion of multiple channels of the viewpoint adjustment network.

[0250] Clause 40. The system according to Clause 37 or 38, wherein the one or more processors are further configured to: update the 3D reconstruction network at least based on one or more first paths of applying the simulated 2D sensor data of the simulated source sensor device to the at least a portion of the 3D reconstruction network and one or more second paths of applying the simulated 2D sensor data of the simulated target sensor device to the at least a portion of the 3D reconstruction network.

[0251] Clause 41. The system according to Clause 37 or 38, wherein the 3D reconstruction network is updated to perform a view transformation that generates transformed features, and wherein the one or more processors are further configured to: update the 3D reconstruction network based at least on a loss that minimizes a difference between the transformed features of one or more first paths and one or more second paths of at least a portion of the 3D reconstruction network.

[0252] Clause 42. The system according to Clause 37 or 38, wherein the one or more processors are further configured to: update the 3D reconstruction network based at least on a loss that minimizes a difference between the view-transformed features of one or more first paths that process the analog 2D sensor data corresponding to the source sensor device and the view-transformed features of one or more second paths that process the analog 2D sensor data corresponding to the target sensor device.

[0253] Clause 43. The system according to Clause 37 or 38, wherein the source sensor device represents a first sensor configuration that is different from the second sensor configuration of the target sensor device.

[0254] Clause 44. The system according to Clause 37 or 38, wherein the one or more processors are further configured to: generate an analog 2D sensor data frame and analog ground truth data for at least one sensor of one or more analog sensors of the source sensor device and at least one corresponding sensor of the target sensor device within at least one time slice of the one or more simulated time slices.

[0255] Clause 45. The system according to Clause 37 or 38, wherein at least one of the system or the 3D reconstruction network is included in at least one of the following: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing simulation operations; a system for performing digital twin operations; a system for performing optical transmission simulation; a system for performing collaborative content creation of 3D assets; a system for performing deep learning operations; a system for performing remote operations; a system for performing real-time streaming; a system for generating or presenting one or more of augmented reality content, virtual reality content, or mixed reality content; a system implemented using edge devices; a system implemented using robots; a system for performing conversational AI operations; a system implementing one or more language models; a system implementing one or more large language models (LLMs); a system implementing one or more vision language models (VLMs); a system for generating synthetic data; a system for generating synthetic data using AI; a system for generating synthetic data using generative AI; a system including one or more virtual machines (VMs); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources.

[0256] Clause 46. A method, comprising: applying analog two-dimensional (2D) sensor data to at least a portion of a three-dimensional (3D) reconstruction network.

[0257] Clause 47. The method according to Clause 46, wherein the analog 2D sensor data is generated based at least on a first sensor device of a first self-machine and a second sensor device of a second self-machine.

[0258] Clause 48. The method according to Clause 46 or 47, further comprising: updating the 3D reconstruction network based at least on the analog 2D sensor data.

[0259] Clause 49. The method according to Clause 46, 47 or 48, further comprising: performing one or more operations on the second self-machine using the 3D reconstruction network.

[0260] Clause 50. The method according to Clause 46, 47, 48 or 49, wherein the method is performed by at least one of the following or the 3D reconstruction network is included in at least one of the following: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing analog operations; a system for performing digital twin operations; a system for performing optical transmission simulation; a system for performing collaborative content creation of 3D assets; a system for performing deep learning operations; a system for performing remote operations; a system for performing real-time streaming; a system for generating or presenting one or more of augmented reality content, virtual reality content or mixed reality content; a system implemented using edge devices; a system implemented using robots; a system for performing conversational AI operations; a system implementing one or more language models; a system implementing one or more large language models (LLMs); a system implementing one or more vision language models (VLMs); a system for generating synthetic data; a system for generating synthetic data using AI; a system for generating synthetic data using generative AI; a system comprising one or more virtual machines (VMs); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources.

Claims

1. One or more processors, comprising processing circuitry for performing one or more operations on a self-machine using a three-dimensional 3D perception network, wherein the 3D perception network is updated based at least on applying simulated two-dimensional 2D sensor data to at least a portion of the 3D perception network, wherein the simulated 2D sensor data is generated based at least on a simulated target sensor device and a source sensor device of the self-machine.

2. One or more processors according to claim 1, wherein the processing circuit is further used to: update the 3D perception network based at least on updating a viewpoint adjustment network, wherein the viewpoint adjustment network includes one or more layers of the 3D perception network that are copied as at least a portion of multiple channels of the viewpoint adjustment network.

3. One or more processors according to claim 1, wherein the processing circuit is further used to: update the 3D perception network based at least on one or more first paths of applying the simulated 2D sensor data of the source sensor device to the at least a portion of the 3D perception network and one or more second paths of applying the simulated 2D sensor data of the target sensor device to the at least a portion of the 3D perception network.

4. The one or more processors of claim 1 , wherein the 3D perception network comprises a view transform that generates transformed features, wherein the processing circuit is further configured to update the 3D perception network based at least on minimizing a loss of differences between the transformed features of one or more first paths and one or more second paths of the at least a portion of the 3D perception network.

5. The one or more processors of claim 1 , wherein the processing circuit is further configured to update the 3D perception network based at least on minimizing a loss of a difference between view-transformed features of one or more first paths of processing the simulated 2D sensor data corresponding to the source sensor device and view-transformed features of one or more second paths of processing the simulated 2D sensor data corresponding to the target sensor device. 6 . The one or more processors of claim 1 , wherein the source sensor device represents a first sensor configuration that is different than a second sensor configuration of the target sensor device.

7. One or more processors according to claim 1, wherein the processing circuit is further used to: generate a simulated 2D sensor data frame and simulated true value data for at least one sensor of the one or more simulated sensors of the source sensor device and at least one corresponding sensor of the target sensor device within at least one time slice of the simulation.

8. One or more processors according to claim 1, wherein the processing circuit is further used to: update the first viewpoint adjustment network based on at least applying the real 2D sensor data and the simulated 2D sensor data to a first viewpoint adjustment network including one or more layers of the 3D perception network, and then update the second viewpoint adjustment network based on at least applying the simulated 2D sensor data to a second viewpoint adjustment network including one or more layers of the 3D perception network.

9. The one or more processors of claim 1, wherein the 3D-aware network is updated to perform at least one of 2D-to-3D transformation or 2D-3D fusion.

10. The one or more processors of claim 1, wherein at least one of the 3D-aware network or the one or more processors is included in at least one of: control systems for autonomous or semi-autonomous machines; Perception systems for autonomous or semi-autonomous machines; A system for performing simulation operations; Systems for performing digital twin operations; A system for performing light transport simulations; A system for performing collaborative content creation of 3D assets; Systems for performing deep learning operations; Systems for performing remote operations; A system for performing real-time streaming; Systems for generating or presenting one or more of augmented reality content, virtual reality content, or mixed reality content; Systems implemented using edge devices; Systems implemented using robots; Systems for performing conversational AI operations; A system implementing one or more language models; A system implementing one or more large language models (LLMs); A system implementing one or more visual language models (VLMs); Systems for generating synthetic data; Systems for generating synthetic data using AI; Systems for generating synthetic data using generative AI; A system comprising one or more virtual machines VM; A system implemented at least in part in a data center; or A system implemented at least in part using cloud computing resources.

11. A system comprising one or more processors, wherein the one or more processors are used to: perform one or more operations on a self-machine using a three-dimensional (3D) reconstruction network, wherein the 3D reconstruction network is updated based at least on applying simulated two-dimensional (2D) sensor data to at least a portion of the 3D reconstruction network, wherein the simulated 2D sensor data is generated based at least on a simulated target sensor device and a source sensor device of the self-machine.

12. A system according to claim 11, wherein the one or more processors are further used to: update the 3D reconstruction network based at least on updating a viewpoint adjustment network, wherein the viewpoint adjustment network includes one or more layers of the 3D reconstruction network that are copied as at least a part of multiple channels of the viewpoint adjustment network.

13. A system according to claim 11, wherein the one or more processors are further used to update the 3D reconstruction network based at least on one or more first paths of applying the simulated 2D sensor data of the source sensor device to the at least a portion of the 3D reconstruction network and one or more second paths of applying the simulated 2D sensor data of the target sensor device to the at least a portion of the 3D reconstruction network.

14. The system of claim 11, wherein the 3D reconstruction network is updated to perform a view transformation that generates transformed features, wherein the one or more processors are further configured to update the 3D reconstruction network based at least on minimizing a loss of differences between the transformed features of one or more first paths and one or more second paths of the at least a portion of the 3D reconstruction network.

15. The system of claim 11, wherein the one or more processors are further configured to update the 3D reconstruction network based at least on minimizing a loss of a difference between view-transformed features of one or more first paths of the simulated 2D sensor data corresponding to the source sensor device and view-transformed features of one or more second paths of the simulated 2D sensor data corresponding to the target sensor device. 16 . The system of claim 11 , wherein the source sensor device represents a first sensor configuration that is different than a second sensor configuration of the target sensor device.

17. A system according to claim 11, wherein the one or more processors are further used to: generate simulated 2D sensor data frames and simulated true value data for at least one sensor of the one or more simulated sensors of the source sensor device and at least one corresponding sensor of the target sensor device within at least one time slice of the simulation.

18. The system of claim 11, wherein at least one of the system or the 3D reconstruction network is included in at least one of the following: control systems for autonomous or semi-autonomous machines; Perception systems for autonomous or semi-autonomous machines; A system for performing simulation operations; Systems for performing digital twin operations; A system for performing light transport simulations; A system for performing collaborative content creation of 3D assets; Systems for performing deep learning operations; Systems for performing remote operations; A system for performing real-time streaming; Systems for generating or presenting one or more of augmented reality content, virtual reality content, or mixed reality content; Systems implemented using edge devices; Systems implemented using robots; Systems for performing conversational AI operations; A system implementing one or more language models; A system implementing one or more large language models (LLMs); A system implementing one or more visual language models (VLMs); Systems for generating synthetic data; Systems for generating synthetic data using AI; Systems for generating synthetic data using generative AI; A system comprising one or more virtual machines VM; A system implemented at least in part in a data center; or A system implemented at least in part using cloud computing resources.

19. A method comprising: applying simulated two-dimensional (2D) sensor data to at least a portion of a three-dimensional (3D) reconstruction network, the simulated 2D sensor data being generated based on at least a first sensor device simulating a first slave machine and a second sensor device simulating a second slave machine; updating the 3D reconstruction network based at least on the simulated 2D sensor data; and One or more operations are performed on the second self-machine using the 3D reconstruction network.

20. The method according to claim 19, wherein the method is performed by at least one of the following, or the 3D reconstruction network is included in at least one of the following: control systems for autonomous or semi-autonomous machines; Perception systems for autonomous or semi-autonomous machines; A system for performing simulation operations; Systems for performing digital twin operations; A system for performing light transport simulations; A system for performing collaborative content creation of 3D assets; Systems for performing deep learning operations; Systems for performing remote operations; A system for performing real-time streaming; Systems for generating or presenting one or more of augmented reality content, virtual reality content, or mixed reality content; Systems implemented using edge devices; Systems implemented using robots; Systems for performing conversational AI operations; A system implementing one or more language models; A system implementing one or more large language models (LLMs); A system implementing one or more visual language models (VLMs); Systems for generating synthetic data; Systems for generating synthetic data using AI; Systems for generating synthetic data using generative AI; A system comprising one or more virtual machines VM; A system implemented at least in part in a data center; or A system implemented at least in part using cloud computing resources.

Citation Information

Patent Citations

  • Method for programmable timeouts of tree traversal mechanisms in hardware

    US10885698B2