Generating alternative image views from stereoscopic parallax data
By generating intermediate image representations and performing coordinate transformations on embedded processors, the problem of generating high-quality bird's-eye view images under limited resources is solved, achieving the technical effect of efficiently generating bird's-eye view images when external memory cannot be directly accessed.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NVIDIA CORP
- Filing Date
- 2025-11-05
- Publication Date
- 2026-05-08
AI Technical Summary
On resource-constrained embedded processors, it is difficult to effectively generate bird's-eye view images of a scene because it requires a large amount of memory access and processing resources. Especially when there is no direct access to external memory, existing technologies struggle to efficiently process stereo parallax data to generate high-quality bird's-eye view images.
An intermediate image representation is generated using a two-step process. First, a two-dimensional histogram image is generated, and then a bird's-eye view image is generated through coordinate transformation. Direct memory access (DMA) is used to avoid direct access to external memory. The same filter is used to process objects at different distances, reducing complexity.
It efficiently generates high-quality bird's-eye view images on resource-constrained embedded processors, reducing reliance on external memory, improving processing efficiency and image quality, and is suitable for autonomous navigation and real-time operation.
Smart Images

Figure CN122002016A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the transformation of image data between different views or representations, and more particularly to generating intermediate image representations from a set of parallax data in one or more non-limiting embodiments, which allows for processing and transformation using limited resources. Background Technology
[0002] In various computational operations, it is necessary to determine the location of various objects within a scene or geographic area. This can include (for example, but not limited to) analyzing captured image information to support tasks such as navigation, localization, controlled interaction, and collision avoidance for robots and autonomous or semi-autonomous vehicles or machines. Performing operations involving image recognition and computer vision requires significant resource capacity, including the ability to access memory with sufficient capacity to store the entire image. Tasks such as generating a bird's-eye view (BEV) representation of a scene based on captured parallax data are difficult to perform, even if possible, using resources with limited capacity (e.g., embedded processors that do not access external memory). Furthermore, tasks such as morphological filtering and motion analysis are resource-intensive when performed on bird's-eye view images because objects at different distances may have captured information of varying quality or quantity. Attached Figure Description
[0003] Various embodiments according to this disclosure will be described with reference to the accompanying drawings, in which:
[0004] Figure 1A , Figure 1B , Figure 1C and Figure 1D An image view that can be generated from captured image data according to at least one embodiment is shown;
[0005] Figure 2A An intermediate image generated using captured image data according to at least one embodiment is shown;
[0006] Figure 2B A view of similar objects in a bird's-eye view (BEV) or a top-down image and a middle histogram image is shown according to at least one embodiment;
[0007] Figure 3 The diagram illustrates corresponding image data blocks in a parallax image and an intermediate histogram image according to at least one embodiment;
[0008] Figure 4 Corresponding image data blocks in intermediate histogram images and bird's-eye view images according to at least one embodiment are shown;
[0009] Figure 5An example process for generating a bird's-eye view image from parallax image data, which can be executed using an embedded processor according to at least one embodiment, is shown;
[0010] Figure 6 An example system including an embedded processor with direct memory access (DMA) functionality is shown according to at least one embodiment;
[0011] Figure 7A A comparison graph showing the amount of detail captured for objects at different distances from the camera, according to at least one embodiment, is shown.
[0012] Figure 7B Different sized filters are shown, according to at least one embodiment, for processing the same amount of detail information of objects at different distances in a bird's-eye view image;
[0013] Figure 8 A comparison of filter sizes, according to at least one embodiment, for processing the same amount of detail information of objects at different distances in bird's-eye view images and intermediate histogram images is shown;
[0014] Figure 9 An example process is shown to perform morphological filtering on an intermediate histogram image using a single filter size for objects at different distances, according to at least one embodiment;
[0015] Figure 10 Components of a distributed system, according to at least one embodiment, are shown for generating, processing, and providing sensor-based content.
[0016] Figure 11 An example computing environment according to at least one embodiment is shown, in which one or more devices operate to process data using a SoC;
[0017] Figure 12 An example data center system according to at least one embodiment is shown;
[0018] Figure 13 A computer system according to at least one embodiment is shown;
[0019] Figure 14 A computer system according to at least one embodiment is shown;
[0020] Figure 15 At least a portion of a graphics processor according to one or more embodiments is shown;
[0021] Figure 16 At least a portion of a graphics processor according to one or more embodiments is shown;
[0022] Figure 17AAn example of an autonomous vehicle according to at least one embodiment is shown;
[0023] Figure 17B The illustration shows an embodiment according to at least one of the embodiments. Figure 17A Examples of camera positions and fields of view for autonomous vehicles;
[0024] Figure 17C It is shown that according to at least one embodiment Figure 17A A block diagram of an example system architecture for an autonomous vehicle; and
[0025] Figure 17D It is according to at least one embodiment for Figure 17A A schematic diagram of a cloud-based server communicating with an autonomous vehicle. Detailed Implementation
[0026] In the following description, various embodiments will be described. Specific configurations and details are set forth for ease of explanation in order to provide a thorough understanding of these embodiments. However, those skilled in the art will also understand that these embodiments can be practiced without these specific details. Furthermore, well-known features have been omitted or simplified to avoid obscuring the described embodiments.
[0027] The systems and methods described herein can be used, but are not limited to, non-autonomous vehicles or machines, semi-autonomous or autonomous vehicles or machines (e.g., in one or more advanced driver assistance systems (ADAS), autonomous vehicles or machines, one or more in-vehicle infotainment systems, one or more emergency vehicle detection systems), manned and unmanned robots or robotic platforms, autonomous mobile robots (AMRs), humanoid robots, warehouse vehicles, off-road vehicles, vehicles coupled to one or more trailers, aircraft, boats, reciprocating vehicles, emergency response vehicles, motorcycles, electric or motorized bicycles, aircraft, engineering vehicles, trains, underwater vehicles, remotely controlled vehicles such as drones, and / or other vehicle types. Furthermore, the systems and methods described herein can be used for a variety of purposes, such as, but not limited to, machine control, machine motion, machine driving, synthetic data generation, generative AI, model training or updating, perception, augmented reality, virtual reality, mixed reality, robotics, security and supervision, simulation and digital twins, autonomous or semi-autonomous machine applications, deep learning, environmental simulation, data center processing, conversational AI, optical transport simulation (e.g., ray tracing, path tracing, etc.), collaborative content creation of 3D assets, generative AI, cloud computing and / or any other suitable application.
[0028] The disclosed embodiments may be included in a variety of different systems, such as automotive systems (e.g., in-vehicle infotainment systems for autonomous or semi-autonomous machines, perception systems for autonomous or semi-autonomous machines), systems implemented using robots, aviation systems, medical systems, marine systems, smart area monitoring systems, systems for performing deep learning operations, systems for performing simulation operations, systems for performing digital twin operations, systems implemented using edge devices, systems containing one or more virtual machines (VMs), systems for performing synthetic data generation operations, systems implemented at least partially in a data center, systems for performing conversational AI operations, systems implementing one or more language models (e.g., large language models (LLM), visual language models (VLM), multimodal language models, etc.), systems for performing generative AI operations (e.g., using one or more language models, converter models, etc.), systems for performing optical transmission simulation, systems for performing collaborative content creation of 3D assets, systems implemented at least partially using cloud computing resources, and / or other types of systems.
[0029] Methods according to various exemplary embodiments can provide the generation of alternative image view images (e.g., bird's-eye view images) of one or more objects, based in part on image data including (or usable for determining) distance information (e.g., stereo parallax data). In at least one embodiment, this method can generate these or other such alternative views in memory-constrained environments and / or low-power operations, for example, where embedded processors equipped with direct memory access (DMA) (e.g., NVIDIA's Programmable Vision Accelerator (PVA)) may not have direct access to external memory. Partly due to DMA-related (and other similar) limitations, intermediate representations of a scene can be generated using, for example, stereo parallax data. This intermediate representation can be used to generate specific views, such as bird's-eye view (BEV) or top-down view of the scene, each of which is DMA-friendly and does not require processor access to external memory; although DMA can still access external memory. In at least one embodiment, stereo parallax data can be used to generate an intermediate representation in the form of a (quasi-bird's-eye view) 2D histogram or histogram-type image, which is a function of the camera angle θ and the distance to the camera plane z, denoted as H(z,θ). The intermediate representation H(z,θ) can then be transformed into a bird's-eye view B(z,x) in Cartesian coordinates. This approach allows for the generation of alternative view images, such as bird's-eye view images, from stereo parallax data using, for example, embedded processors with DMA capabilities.
[0030] The methods according to various exemplary embodiments may also provide data processing (e.g., filtering) for operations such as computer vision-related operations. When an image represents multiple objects at different distances (physical or virtual) from a camera, these objects are typically represented by different numbers of pixels, and therefore have different quality levels. For example, objects closer to the camera typically appear larger in the image and are represented using more pixels; while objects farther from the camera (at least similar in size) may appear smaller and be represented using fewer pixels, such as even a single pixel. Thus, objects far from the camera can be represented with very little detail in the image. For alternative image views, such as bird's-eye views generated from stereo parallax data, objects at different distances may be processed the same way, resulting in different quality outcomes. In other methods, objects at different distances may be processed differently, which introduces additional complexity and cost, partly because of the need to account for distance differences. Limitations such as those caused by the use of DMA may prevent such processing from being performed efficiently, or even impossible at all. In at least one embodiment, stereo parallax data can be used to generate intermediate representations. This intermediate representation can then be used to generate alternative views of the scene (e.g., bird's-eye view), where each of these steps is DMA-friendly (or can run using local memory) and requires no access to external memory. This intermediate representation can take the form of a (quasi-bird's-eye view) 2D histogram, which is a function of the camera angle H(z,θ). This intermediate representation H(z,θ) can be used to perform various types of processing. Because the intermediate representation is a function of the camera angle, nearby objects will appear larger (or be represented using more pixels) in the intermediate image, and this information is not lost when scaled down in the final bird's-eye view representation. Instead of using smaller filters for objects closer to the camera in the bird's-eye view (or larger filters for objects farther away), the same filter size can be used for all objects in this intermediate representation, avoiding any additional complexity arising from considering distance from the camera. A similar advantage exists for optical flow-type operations, which also do not require consideration of distance from the camera plane.
[0031] Those skilled in the art will understand that variations of this functionality and other such functionality can be used within the scope of the various embodiments, in accordance with the teachings and suggestions included herein.
[0032] Many computer processes involve determining the position of objects in a three-dimensional environment. This can include, for example, capturing images and / or sensor data of the physical environment and generating a digital reconstruction of that environment that can be used for a variety of purposes. This can include, for example, determining how to navigate a robot through a data center, or performing collision avoidance for an autonomous (or semi-autonomous) vehicle, based in part on the position and / or motion of objects generally nearby in the determined environment. The data to be analyzed can include stereo data captured using a pair of matched cameras with a determined spacing, or point cloud data captured using a LiDAR (Light Detection and Ranging) system, and other such options. Once this data is captured, it can be used to generate one or more views of at least a portion of the environment represented in the data. In some cases, a first type of view can be captured using one or more sensors, but in order to accurately perform one or more operations, it may be necessary to generate at least a second type of view. In some cases, it may be necessary to generate data representing different views, which represents nearby objects in a particular way that is beneficial or even necessary for the intended purpose.
[0033] Figure 1A An example image 100, which can be captured by a 2D camera on a vehicle according to at least one embodiment, is shown. As described above, this can be one of a pair of images captured concurrently by a pair of matched cameras in a stereo imaging system. In this example, the vehicle is traveling along a road, and the camera is positioned such that the camera view is in front of the vehicle. The camera can then capture image data representing objects that are at least partially in front of the vehicle and within the camera view. This can include movable objects, such as pedestrians or other vehicles 102, 104 at least partially in front of the (self) vehicle; and stationary objects, such as road signs 106, trees, buildings, sidewalks, etc. A “self” vehicle generally refers to a vehicle with a set of sensors arranged around it that are capable of capturing sensor data, thereby enabling the vehicle’s control and / or operating system to perceive its surrounding environment. In at least one embodiment, such a vehicle may have cameras arranged around it to capture image data of the full 360 degrees, such as a 360-degree top-down view or “bird’s-eye view” of the physical environment around the self vehicle.
[0034] For tasks such as collision avoidance and route determination, it is crucial to accurately identify the relative positions of various objects that are generally close to the vehicle (or other such controllable systems, devices, or components) to ensure that the vehicle does not inappropriately collide with or interact with any of these objects. As mentioned above, a pair of images (e.g. Figure 1A The image 100 shown can be captured using a pair of matching cameras (e.g., cameras with similar camera and imaging parameters, the same focal length, and slightly laterally separated) and used to generate the parallax image 130, as shown. Figure 1BAs shown. Due to the lateral spacing between a pair of matched cameras, each object will appear to be located in a slightly different position in the images captured by these cameras, based on the slightly different viewpoints used for that capture. The apparent location difference will be greater for objects closer to the camera than for objects farther away. By knowing the camera parameters, the distance (or "parallax") between the object's position in each captured image (often referred to as the "left" image and the "right" image due to the lateral spacing) can be used to calculate the distance between the object and the pair of cameras. This can be calculated for individual pixel positions in the left and right images, and a distance can be calculated for each such pixel position. The calculated distances to the object represented by each pixel position can then be used to generate a disparity map or disparity image 130, as shown. Figure 1B As shown. In this example parallax image, objects closer to the camera appear brighter, closer to white color values, while objects farther from the camera appear darker. For portions of the image without detectable objects (e.g., the sky), these portions can be considered as having objects located at infinity or beyond measurable distance and can be represented using black color values. As shown, one of the vehicles 102 can be determined to be closer to the camera than another vehicle 104 based on the fact that the closer vehicle 102 has a brighter color (indicating a shorter distance from the camera). The parallax data for objects in the parallax image will only represent the visible portion of the object in the image, which in this example includes the rear and right sides of the vehicle. If complete shape information is needed, another process can be used to attempt to identify the type of object and then infer additional shape information based on the object type, such as the specific vehicle brand, model, and year. However, for tasks such as collision avoidance, knowing, for example, the visible portion of the vehicle facing itself and therefore most likely to be affected, is sufficient.
[0035] In at least one embodiment, captured image and / or parallax data can be analyzed to attempt to identify specific objects in the data. Object identification may include identifying pixels determined to correspond to a single object based on factors such as similarity in location and color, and then possibly identifying the type of object. For example, if a connected component analysis method is used... Figure 1B The parallax image 130 in the image is used to identify objects within a given distance of the camera. This method may identify three sets of pixels that may correspond to a specific object. In some embodiments, specific types of objects (e.g., roads and sidewalks) may be excluded from consideration. The identified objects can then be treated as separate objects, such as... Figure 1CThe schematic diagram is shown in image 160. In this example, the two closest vehicles 102 and 104 and the road sign are identified as objects that meet the current recognition criteria. It should be understood that, for ease of explanation, the number of objects identified in this example is relatively small, and the actual object detection or recognition process may identify more. Figure 1B Many other objects are present in the parallax image 130. In some embodiments, a connected component type approach can be used to identify groups of pixels that may correspond to individual objects, and then an object recognition approach can be used to attempt to identify the type of object, for example, to distinguish between vehicles 102, 104 and a street sign 106. This distinction is important for tasks such as object avoidance, because the street sign 106 is fixed in place and will not move in the relevant future time period, while the vehicles may be in motion or able to move in that future time period, where the motion should be taken into account when determining a suitable navigation path or other similar options. Figure 1C The example image illustrates a schematic of a typical camera view from a vehicle's perspective, where the image has an image coordinate system originating from the top left corner and pixel coordinates are expressed as (i,j), where i corresponds to the vertical axis and j corresponds to the horizontal axis in the image. The object represented can be considered to be located in camera space, as it can correspond to... Or a similar coordinate system.
[0036] For at least some applications, it is desirable to generate a top-down view or bird's-eye view (BEV) of at least detected and / or identified objects in a scene, or to compute such view data of at least detected and / or identified objects in a scene. If data captured from a single camera (or stereo camera assembly) is used, the available data will be limited to objects within the camera's view unless data from multiple cameras can be stitched together or otherwise processed to generate a larger view. As an example, Figure 1D It shows the basis Figure 1A A bird's-eye view image 180 of the camera image. In this bird's-eye view image, objects 102, 104, and 106 are represented using accurate relative size and position information. As shown, since only a portion of each object is visible to the camera, such as one or more sides facing the camera, the representation of each object will only include that portion visible to the camera, unless additional processing is performed to identify the object type and fill the object's space. As shown, such a bird's-eye view image can be located in a conventional Cartesian coordinate space and represented as B(x,y). In such a coordinate system, the position of the origin is crucial, typically corresponding to the center point of the camera lens (or sensor, etc.) that captured the image data of the scene. The coordinate system of the bird's-eye view can be at an angle relative to the origin, for example, coordinate system... This represents the pitch and yaw angles relative to the origin.
[0037] In conventional methods, one or more processors (e.g., a central processing unit (CPU) or a graphics processing unit (GPU)) can be used to generate a bird's-eye view image 180 from a parallax image. This process can also be performed using hardware acceleration (e.g., using a programmable vision accelerator (PVA), a deep learning accelerator (DLA), an optical flow accelerator (OFA), etc.), which allows for very fast bird's-eye view generation, which is necessary for real-time operations such as autonomous or semi-autonomous navigation. However, to perform such processing, the processor needs access to sufficient memory space, such as external memory, capable of storing all image data at once. For high-resolution parallax images, this can include more data than can be held by hardware with limited capacity (e.g., memory accessible to an embedded processor via DMA or other similar mechanisms).
[0038] In conventional CPU-based systems, bird's-eye view images can be generated from parallax data without much concern for data transfer. However, there are situations where the processing unit may not have direct access to the external memory containing the entire set of image data. Systems used to process captured images and / or sensor data may be limited by the amount of available memory or processing capacity. For example, a system might use an embedded processor equipped with DMA that does not have direct access to external memory. Such a device may lack the memory required to transform a stereo parallax image of a scene into an alternative type of image (e.g., a bird's-eye view of the scene). While such tasks may be relatively straightforward on devices with CPUs or GPUs, they can be daunting when they need to be performed on processing units that do not have direct access to external memory, such as embedded processors.
[0039] In at least one embodiment, a two-step process can be used to transform a parallax image of a scene into an alternative view image (e.g., a bird's-eye view image). Both steps are relatively lightweight, so they can both run on DMA-based hardware. As an example, DMA can be used to perform data transfers for each of these steps on data in the rectangular input region of the parallax image, as well as data in the intermediate and generated bird's-eye view images. It should be understood that, within the scope of at least one embodiment, additional steps may also be present, such as for preprocessing, post-processing, data transformation, and other such functions.
[0040] In one example, the disparity value at a given location (i,j) in a disparity image can be represented as D(i,j). The indices i and j are vertical angles. (representing pitch angle) and horizontal angle θ (representing yaw angle) are functions. In the case of the simplest corrected image, θ is measured from the camera axis, while i and j are the usual image coordinates, with the origin at the top right corner, as given by the following formula:
[0041]
[0042] j=cθ+d
[0043] Here, parameters a, b, c, and d depend at least in part on the inherent properties of the camera. Therefore, parallax can be expressed as... Instead of D(i,j), the relationship between the stereo parallax value D and the distance z of the object from the camera image plane can be expressed as:
[0044] z = E / D
[0045] Here, E is a parameter that depends on the camera's intrinsic parameters.
[0046] In at least one embodiment, the first step of such a process may be generating an intermediate image, for example... Figure 2A The image shown is a two-dimensional (2D) histogram or intermediate image 200. In this example, the camera is not located at a single point at the bottom center as in the bird's-eye view image, but actually spans the bottom of the histogram. The histogram values can be represented by H(z,θ), and the element H(z,θ) can be applied to each pixel. The resolution error is added or subtracted by increasing z = E / D. This transformation can create an intermediate image 200, which is similar to a top-down or bird's-eye view image 180 of the scene represented in the input image. However, in this intermediate image 200, the representation of objects in the scene may have at least some inaccuracies. These inaccuracies are partly due to the fact that objects closer to the (physical or virtual) camera tend to be represented as expanded outwards or stretched laterally, because objects closer to the camera tend to occupy a larger angular range. Figure 2B An example of this lateral stretching effect based on the distance from camera 260 is shown. Figure 2B In the bird's-eye view 250, two objects 252 and 254 of the same size and shape are shown. Specifically, in this example, the two objects have the same width. However, when analyzed in histogram space based on angles, the angular range 258 occupied by the object 254, which is closer to the camera, will be larger than the angular range 252 occupied by the object 252, which is farther away. In the example shown, the angular range 258 of the closer object 254 is two to three times larger than the angular range 256 of the object 252, which is farther away from the camera 260. Figure 2BAs shown, when generating the intermediate histogram image 280 as a function of angle, the height z of the two objects will not be stretched, but the width of the closer object 254 will be stretched by an amount in the θ direction. This amount causes the width of the closer object 254 to appear to be two to three times the width of the object 252, which is farther from the virtual camera. This corresponds to the difference in the distance-based angular range occupied by these objects, even though objects 252 and 254 actually have the same width. (Review) Figure 1D Aerial view image 180 and Figure 2A The stretching of the intermediate histogram image 200 also affects the appearance of the object, because the stretching of object 102 will also cause it to appear to have a different shape (e.g., a distorted shape) in the intermediate histogram image 200 than in the bird's-eye view image 180.
[0047] The second step in this example process can be performed to attempt to at least correct this stretching effect. It will use... Figure 3 The example shown uses a parallax image 300 to generate an intermediate image 310, which has the same example objects 102, 104, and 106 as in the previous example, for ease of explanation. One advantage of generating and using the intermediate image 310 in the H(z,θ) coordinate system is... rectangular area Any data will result in a rectangular region 304(0:z) max The increment of H(z,θ) in (θ1:θ2) allows the use of DMA. In a parallax image, the value of a pixel within a rectangular region depends on its distance from these objects, which can be from the camera lens (distance zero) to an effective infinity or the maximum detectable (or maximum permissible) distance (set to z here). max Any position of the rectangle. Therefore, when the rectangle is transformed into the quasi-bird's-eye view intermediate image 310, a portion of the corresponding data advances from the camera lens position (represented by the bottom edge of the intermediate image 310) all the way to the maximum distance from the camera (represented by the top edge of the image). It should be understood that other rectangles may exist above and / or below the example rectangular region 302 in the parallax image 300, which also map to the same rectangular region 304 in the intermediate image 310.
[0048] An example process according to at least one embodiment can transform an intermediate image 400H(z,θ) into a bird's-eye view image 410B(z,x), such as Figure 4 As shown, B is then processed. The example procedure can alternatively process H, generating a list of object centroids (or other representative locations) (or sets, etc.) and other statistics in H, and then transforming this list into a corresponding list of the bird's-eye view image. In either method, the coordinate transformation can be given by the following equation:
[0049] x=ztanθ
[0050] Such tasks can also be DMA-friendly, for example, because information from any position in row x of the intermediate image H(z,θ) only affects row x of the final bird's-eye view image B(z,x). Figure 4 As shown, the rectangular region 402 of the intermediate histogram image 400 has the same height as the corresponding rectangle 404 representing the same portion of the object in the bird's-eye view image. This is because the distance from the camera to the object is the same in both images, only lateral stretching occurs (horizontally from left to right in the image). However, due to this stretching, the rectangular region 402 in the intermediate image 400 is generally wider than the corresponding rectangle 404 in the bird's-eye view image 410, where the amount of stretching can be partly based on the distance to the camera, as objects closer to the camera are represented with greater lateral stretching in the intermediate image.
[0051] Due to the discrete nature of digital images and the shrinkage effect of multiplying by tan(θ), two or more elements in H can map to the same element in B. In this case, the values in H can be summed or their maximum values taken to try to solve the problem. This may not be a problem in the case of list transformations, as long as the method allows different objects to occupy the same position in the bird's-eye view image. As mentioned above, the advantage of this method is that these steps or tasks can be performed using limited resources, such as embedded processors with DMA.
[0052] One advantage of using this block-based approach is that rectangles (e.g.) Figure 4The image data within rectangles 402 and 404 shown can be selected in terms of size, allowing data within the rectangle to be transferred and stored using capacity-limited techniques such as DMA. A given rectangular region 402 can be selected from an image (e.g., intermediate image 400), the data is transferred via DMA and processed by, for example, an embedded processor, and then stored with respect to the corresponding rectangular region 404 or block in the resulting image (e.g., bird's-eye view image 410). The result can be transferred via DMA to a result memory location for the output representation. The size of the rectangular region can be selected based on various factors, such as the size of the image being processed and the amount of available memory transfer, among other such options. For example, the memory may be able to hold approximately 32kb of input and output data, so the rectangle size can have a selected upper limit such that the amount of data within the region falls within the 32kb available portion, while taking into account factors such as image resolution. The number of rectangular regions used can then be calculated based on the total amount of data and the amount of data that can be held in each rectangular region. In at least one embodiment, all rectangular regions can have the same size, but in other embodiments, the size or shape of the rectangular regions can vary if there is a performance advantage, as long as the size and shape remain within permissible parameters. In at least one embodiment, intermediate images can be processed using 10 horizontal rectangular regions and 10 vertical rectangular regions (a total of 100 rectangular regions or tiles). Data blocks can be processed as tiles of an image, where each pixel falls within a given tile, and these tiles can be processed individually without affecting the quality of the final output. Partly due to the highly local nature of this type of processing, processing such as sharpening or low-pass filtering can be performed on individual tiles. The advantage of the rectangular correspondence between intermediate images and bird's-eye view images, especially in terms of data flow and the way data is associated between the two images, makes it possible to implement relatively complex transformations from camera view images to bird's-eye view images using hardware with limited capacity (e.g., embedded processors with DMA, or another dedicated processing unit or core with limited memory and / or transfer capabilities).
[0053] Figure 5An example computational process 500, executable according to at least one embodiment, is illustrated to generate an alternative view image (e.g., a bird's-eye view image) from a camera view image. It should be understood that, for these and other processes presented herein, additional, fewer, or alternative steps may be performed in a similar or alternative order, or at least partially in parallel, within the scope of the various embodiments, unless otherwise explicitly stated. Furthermore, although this example will discuss camera view images and bird's-eye view images, other types of image transformations can be performed using such processes within the scope of the various embodiments. Such computational processes can be performed using any combination of hardware, firmware, and / or software. For example, various functions can be performed using one or more processors that execute instructions stored in one or more memories. Such processes can also be embodied as computer-usable instructions stored on a computer storage medium. This process can be provided by a standalone application, service, or managed service (independently or in combination with another managed service), or as a microservice via an application programming interface (API), or as a plug-in to another product, etc. Furthermore, this process will relate to... Figure 6 The systems described herein are given by way of example. However, this process may be performed additionally or alternatively by any system or combination of systems, including but not limited to the systems described herein.
[0054] In this example calculation process 500, 502 parallax image data is obtained, which includes representations of one or more objects in the scene. This may include, for example, receiving parallax image data from a stereo camera assembly (or imaging device) positioned such that one or more objects fall within the view of the camera assembly. Other and / or supplementary types of data may also be present, captured using one or more sensors, where the supplementary types of data can provide visual, shape, motion, or other such data about the objects, where this data may have specific values or values relative to the camera assembly or the vehicle / system to which the camera assembly is attached. In this example, the hardware allocated for processing the parallax image data may include hardware with limited capacity, such as an embedded processor with DMA functionality. The embedded processor may be used to generate a 504 two-dimensional histogram, which includes representations of one or more objects as a function of the angle with respect to the camera assembly used to capture the parallax data. This histogram or intermediate image may be a function of distance and angle with the camera assembly and may be used as a quasi-bird's-eye view image. The embedded processor can also be used to generate a list of centroids and statistics (or other such position indicators and / or measures) of one or more objects in a 506 2D histogram, and the values in this list can be transformed into values represented in a Cartesian coordinate system. A 508 bird's-eye view image of the scene can be generated, which includes a top-view representation of one or more objects after transformation using the list of coordinates and statistics. The transfer and analysis of image data can be performed using blocks, tiles, or rectangular regions of pixels in an intermediate image, which allows the embedded processor to transfer, process, and store the image data without having to access all image data that can be stored in external memory at any given time. A 510 bird's-eye view image can then be provided for performing at least one method on the scene and / or one or more objects. For example, this could include determining a navigation path or sequence of interactions relative to objects in the scene or the surrounding environment. In other embodiments, the data can be stored for later use and analysis, or provided for performing other types of tasks.
[0055] Figure 6An example system 600 according to at least one embodiment is illustrated, wherein an embedded processor 614 is used to perform tasks such as image transformation. In this example, a computer system 602 includes a stereo camera 604 (or at least communicates with a stereo camera 604) capable of capturing stereoscopic image data of one or more objects 608 within a field of view 606 (or at least an overlapping field of view) of a stereo camera assembly. As described above, various other types of sensors or devices may also be used to capture information about objects within the scope of the various embodiments. In this example, the captured image data may be stored in local memory 616, external memory, or other such locations. The local memory may be connected to a central processing unit (CPU) 618 or other such processor (e.g., a GPU or DPU), for example via a system bus 622, which allows the CPU to process all image data accessible from the local storage device 616. However, in this example, the computing system 602 may include an image processing module 610, or at least work in conjunction with an image processing module 610. The image processing module may include an embedded processor 614, which may not have access to local storage device 616 (located external to the image processing module) and may only be able to access portions of the image data via DMA controller 612 or other such data transfer mechanisms. As discussed herein, blocks of image data may be transferred via DMA for processing by the embedded processor 614. For image transformation, stereo image data blocks (or parallax image data) are received from stereo camera 604 by the embedded processor 614 to generate intermediate images. The embedded processor can then transform the intermediate images into alternative view images, such as bird's-eye view images, using a block-based approach (which processes portions of the image data separately). Tasks for the transformation (e.g., connected component analysis and centroid calculation) can be performed using the embedded processor. The bird's-eye view image can then be provided directly or via CPU 618 or system bus 622 to control system 620 or other such destinations or receivers for one or more tasks, such as autonomous navigation, collision avoidance, or object interaction, and other such options. In this example, the image processing module 610 may be a system-on-a-chip (SoC) that can be used by the camera circuitry, or may include at least a portion of the camera circuitry. The embedded processor 614 may be used as a coprocessor or offload processor for the CPU 618 or other similar processors, such as a digital signal processor (DSP). Using DMA in a system including a DSP allows for tight control over data movement while avoiding large amounts of data caching, memory address space management, and other such tasks.
[0056] As discussed above, one of the challenges in computer vision and image understanding is the fact that objects represented at a large distance from a (physical or virtual) camera will inevitably appear smaller in images captured or generated using that camera. Consider... Figure 7A A camera view image 700 captured in the image represents one of a pair of stereoscopic images. In this image, there are two objects 702 and 704 of approximately the same size. As shown, the object 702, which is farther from the virtual camera, appears smaller and is represented by a relatively small number (e.g., 9) of pixels 706 or an array of pixels. The object 704, which is closer to the camera, has a larger representation in the image, represented here by a larger number (e.g., 56) of pixels. The smaller number of pixels associated with the objects farther from the camera relative to the closer objects results in a lower quality representation of these images, including less information about the shape, appearance, and other attributes of these distant objects. This is especially true when the scene in the camera view is transformed and viewed from a bird's-eye view (e.g., a bird's-eye view). Figure 7B This effect can be observed when processing images (750) in the image, particularly when considering occupancy grids or occupancy maps. In this example, there are two objects 752 and 754 that are similar in size but located at different distances from the camera 758. These objects at different distances can be processed using the same analysis function, but this approach may be suboptimal, partly because the information content of these objects differs. An alternative approach is to use different analysis functions for objects at different distances, but this adds additional complexity and cost due to the need to determine and consider the different distances, including processing objects differently based on their respective distances from the camera. As mentioned earlier, in some cases, it is desirable to process such images on hardware with limited capacity, and this additional processing and complexity can pose problems for that hardware, at least in terms of meeting various performance criteria.
[0057] The method discussed earlier in this paper allows for the two-step transformation of a parallax image into an alternative view image of the scene (e.g., a bird's-eye view), which can be implemented using one or more limited-capacity resources (e.g., an embedded processor with DMA capabilities). This method can generate an intermediate representation H(z,θ) of the bird's-eye view, which inherently has a greater number of occupancy cells covering objects closer to the camera, partly due to the lateral stretching effect discussed earlier. This is consistent with... Figure 7B The generated bird's-eye view image 750 shown provides a contrast. In this example image, objects are depicted as having the same size. However, the amount of information available for each object will be limited, such as... Figure 7A As shown, this is partly due to the different number of pixels based on the distance from the camera. In an isometric bird's-eye view, the object occupies the same number of grid cells regardless of its position or distance from the camera. However, in... Figure 7A In the example, because the representation of the captured object 704 contains a larger number of pixels, the object 704, which is closer to the camera, has approximately six times more usable information. Figure 7BIn the bird's-eye view, making objects appear to be the same size means either stretching the pixels of object 752, which is farther from the camera, or compressing the pixels of object 754, which is closer to the camera. Pixels in the parallax view are accumulated into the bird's-eye view, and various types of image processing can then be performed to join these pixels into a single object of appropriate size.
[0058] In such a bird's-eye view 750, using filters of the same size would result in processing different amounts of information for different objects at different distances, which would affect quality, as discussed earlier. As mentioned above, one way to ensure that a similar amount of information is used for each filter (or algorithm, etc.) is to use filters of different sizes for objects at different distances. One approach is to try using filters that allow each filter to capture the same number of pixels or data from the captured image. As shown, this might result in using a filter 756 of a first size for objects at a greater distance from camera 758. This filter could capture approximately 9 pixels of information for the object at that distance (e.g., including some pixels located in the object region but not corresponding to the object). If a similar filter is to be used for a closer object 754 to capture approximately 9 pixels of information for the closer object 754, then the size of this filter needs to be smaller. In this approach, it is necessary to use filters of different sizes for each different distance, or at least within a range of distances that can be partially based on the resolution of the original captured image data. As mentioned above, this need to determine and use multiple filter sizes (or different algorithms, etc.) can lead to additional processing and memory requirements that may be difficult to meet using, for example, hardware with limited capacity.
[0059] To avoid using different sized filters or algorithms for objects at different distances from the camera, methods according to various embodiments can use intermediate representation images, such as those discussed earlier herein, which allow the use of filters of the same size for all objects, regardless of their distance from the camera. As discussed above, such intermediate images do not suffer from the problem of varying numbers of occupancy cells covering objects closer to the camera. In such an intermediate representation of an occupancy grid, nearby objects will appear larger due to the lateral stretching effect with respect to objects farther from the camera. The amount of stretching is inversely proportional to the distance from the camera (although in other embodiments, there may be compression that is directly proportional to the distance). Because the amount of stretching is inversely proportional to the distance, filters of the same size can be used for all objects in the intermediate image, regardless of their distance from the camera. For example, in Figure 8In the intermediate image 850, two filters 852 and 854 are shown, respectively applied to the corresponding objects 702 and 704. Partly due to the stretching of the closer object 704 in the intermediate image 850, each filter 852 and 854 will process the region corresponding to approximately the same number of pixels in the original captured image. This method also helps prevent information loss from objects closer to the camera, which could otherwise occur if these objects are shrunk or compressed in the final occupied grid.
[0060] In at least one embodiment, applying a single filter of a defined size to all objects in the intermediate image is equivalent to using a smaller (or finer) filter for objects closer to the camera, and / or a larger filter for objects farther away in the final occupancy grid. It is also possible to use a filter of a single size without considering any additional complexity arising from the distance to the camera. Further advantages are gained when using optical flow maps to estimate the motion of objects. When calculated by averaging the optical flow over the objects in the camera view, this motion can be directly applied to the (z,θ) space without considering the distance (z) to the camera plane. Therefore, processing of occupancy grid information can be performed in this (z,θ) spatial representation of the occupancy grid, rather than in the occupancy grid domain itself.
[0061] In one example of the type of analysis that can be performed using such filters or algorithms, pixel data is analyzed to attempt to identify objects in an image and determine which pixels correspond to specific objects. While such a process might be relatively simple for a person viewing the image, performing this determination in software is relatively complex and / or time- and resource-intensive. As an example, it might be necessary to use a connected component algorithm (or a similar method) to analyze the input image to connect relevant pixels as associated with a single object. The individual pixels can then be labeled or otherwise indicated as being associated with a specific object. Using filters that are too large makes it difficult to distinguish nearby objects, potentially leading to these objects being incorrectly identified as larger single objects. Similarly, using filters that are too small can potentially lead to one object being incorrectly identified as two or more smaller objects.
[0062] To provide accurate pixel grouping, additional preprocessing may be required, such as dilation and / or erosion operations, to attempt to denoise the data (since input image data in various systems may contain unacceptable amounts of noise in many cases). This is achieved in part by altering the size or shape of one or more objects in the image. Erosion typically involves removing pixels from object boundaries to reduce the overall size of the object representation in the image data and eliminate edge pixels whose values may be significantly affected by regions unrelated to the object (e.g., background objects). Dilation can be used to add pixels near object boundaries to increase the size of the object representation. This also helps to connect broken or separated parts of objects in the image, which is helpful for performing connected component analysis and other types of analysis. Erosion can be used to remove noise but results in a smaller object representation; therefore, dilation can be used to recover lost object regions.
[0063] Pixels may undergo some degree of morphological filtering before being processed using connected component analysis (or similar) methods. Image data of objects that are farther from the camera and appear smaller in the image tend to be noisier than image data of closer objects because fewer pixels represent the appearance (and other such attributes) of more distant objects. This can lead to a loss of fine detail and inaccurate pixel values, where the final pixel value at that location in the image may differ significantly from the actual color of that location, as it requires trying to select the pixel value partially based on the many different colors that may exist around that location. In this case, it is preferable to use a larger filter for noisier objects. Using a larger filter for closer, less noisy images may result in a decrease in the accuracy or quality of the captured image data representing these larger, less noisy objects.
[0064] As mentioned above, performing such morphological (and other types of) filtering using intermediate representations of the image can be highly beneficial because a single filter of size and type can be used for all regions of the image, regardless of the distance of the corresponding object from the camera. Using a filter of the same size at every location also allows for simpler algorithms and reduces processing and data transfer. Figure 8A bird's-eye view 800 of a pair of objects in a scene is shown again, where filters of different sizes are needed to process the same amount of actual captured image (or other such) data. In contrast, in an intermediate image 850 of the same set of objects, objects closer to the camera are stretched, the stretching being inversely proportional to the distance from the camera. Therefore, filters of the same size can be used for each region, and these filters will contain the same number of pixels of information from the original captured image data. When comparing the portions of each object represented by the filters in the bird's-eye view 800 with those in the intermediate image, it can be seen that the portions of the objects represented by each filter are substantially the same; for example, the portions of the closer object 704 represented by the smaller filter 708 in the bird's-eye view 800 and the same-sized filter 854 in the intermediate representation 850 are very similar. Thus, each filter processes the same amount of pixel data without requiring different filters for objects at different distances from the camera. Similar advantages can be obtained when performing motion analysis or estimation. The size and type of filters can vary depending on the use case, and in some cases, different filters may be used to determine the preferred visual quality, which can be subjective and vary depending on the use case or intended purpose. For example, in a navigation use case, getting the correct shape may be more important than making the object look as accurate as possible; while in a presentation-based use case, high quality of certain visual aspects may be more important, and the precise shape or position of a given object may not matter.
[0065] Figure 9 An example computational process 900, according to at least one embodiment, can be performed during image transformation to perform consistent and efficient filtering. In this example, disparity image data 902 is obtained, which includes representations of one or more images in the scene, such as those relating to... Figure 5The example process discussed herein. In this example, the hardware allocated for processing the parallax image data may include hardware with limited capacity, such as an embedded processor with DMA capabilities. This embedded processor may be used to generate a 904 two-dimensional histogram, which includes representations of one or more objects as a function of their angles relative to the camera component used to capture the parallax data. This histogram (or intermediate image) may be a function of distance and angle relative to the camera component and may be used as a quasi-bird's-eye view image. In this example, morphological filtering will be performed to attempt to reduce noise and otherwise improve the quality of the image data to be transformed. A 906 filter, determining its size and shape, may be selected and used for morphological filtering and / or other such processing or preprocessing. A 908 morphological filtering may then be performed using the same size and shape-determining filter for all locations in the input image (including each of one or more objects), regardless of the location or distance of these objects from the camera capturing the parallax image data. Filtering may include, for example, erosion (for noise removal) and dilation (for attempting to recover any data lost during erosion). In this example, 910 connected component analysis can be performed on the filtered image data to identify pixels associated with an individual or specific object among one or more objects. This approach can effectively identify which pixels (at least most) in the image data are associated with each object. This approach is particularly useful when performing image transformations based on things like a list of object centroids, where determining the accurate size and shape of the objects is crucial for accurately determining the centroids. Based on the transformation of the 2D histogram, 912 alternative view images, such as bird's-eye view images, can be generated for a top-down representation of a scene including one or more objects. 914 bird's-eye view images can then be provided for performing at least one method on the scene and / or one or more objects. For example, this could include determining navigation paths or sequences of interactions with objects relative to the scene or the surrounding environment. In other embodiments, the data may be stored for later use and analysis, or provided for performing other types of tasks.
[0066] The various methods proposed in this paper are lightweight enough to be executed in a variety of locations, such as in real time on client devices including personal computers or game consoles. Such processing can be performed on content generated or received on the client device, or on content received from an external source, such as streaming data or other content received from a cloud server 1020 or a third-party service 1060 via at least one network, and other such options, such as... Figure 10 As shown. In some cases, the processing, generation, synthesis, and / or determination of at least a portion of the content may be performed by one of these other devices, systems, or entities and then provided to a client device (or another such recipient) for presentation or other such purposes.
[0067] As an example, Figure 10An example network configuration 1000 is illustrated that can be used to provide, generate, modify, encode, process, and / or transmit data, requests, or other such content. In at least one embodiment, client device 1002 may use components of content application 1004 on client device 1002 to generate or receive session data and to generate or receive data locally stored on the client device. In at least one embodiment, content application 1024 executing on server 1020 (e.g., cloud server or edge server) may initiate a session associated with at least one client device 1002, such as by utilizing a session manager and user data stored in user database 1036, and may result in, for example, capturing or retrieving content (e.g., one or more images or image data) from asset repository 1034, as determined by content manager 1026. Content manager 1026 may cooperate with one or more transformation modules 1028 to transform between image views, such as from parallax image to bird's-eye view image. Content application 1026 can also work in conjunction with sensor control module 1030 and control module 1032. Sensor control module 1030 can enable the capture and / or preprocessing of sensor data, while control module 1032 can perform various operations based on sensor data transformed using transformation module 1028. Transformed image data can also be provided for processing or presentation via client device 1002. In this example, content application 1024 can receive parallax data captured by client device 1002 and can return an alternative view image transformed by transformation module 1028. In at least one embodiment, content application 1024 can work in conjunction with one or more encoders, transcoders, and / or compressors that can perform tasks such as encoding, decoding, compression, and / or decompression on content instances (e.g., image data before or after transformation), where different compressions or encodings may be beneficial for different operations, such as storage and processing. At least a portion of the generated, captured, transformed, and / or compressed content can be transmitted to client device 1002 using a suitable transmission manager 1022 for delivery via download, streaming, or other such transmission channels. An encoder can be used to encode and / or compress at least a portion of the data before transmission to client device 1002. In at least one embodiment, client device 1002 receiving such content can provide it to a corresponding content application 1004, which may also or alternatively include a graphical user interface 1010, an imaging control module 1012, and a transformation module 1014 for providing, compositing, rendering, combining, modifying, transforming, or using image- or sensor-based content for presentation (or other purposes) on or by client device 1002.The decoder can also be used to decode data received via network 1040 for presentation via client device 1002, such as displaying images or video content via display 1006, and audio (e.g., sound and music) via at least one audio playback device 1008 (e.g., speakers or headphones). In at least one embodiment, at least a portion of the content may already be stored on client device 1002, rendered on client device 1002, or accessible to client device 1002, so that at least this portion of the content does not need to be transmitted via network 1040, for example, it may have been previously downloaded or locally stored on a hard drive or optical disc. In at least one embodiment, the content may be transmitted from server 1020 or user database 1036 to client device 1002 using a transmission mechanism such as data streaming. In at least one embodiment, at least a portion of the content may be obtained, enhanced, and / or streamed from another source (e.g., third-party service 1060 or other client device 1050), which may also include content application 1062 for generating, enhancing, or providing content. In at least one embodiment, portions of this functionality may be performed using multiple computing devices or multiple processors within one or more computing devices (e.g., which may include a combination of CPU and GPU).
[0068] Figure 11 Components of an example system or operating environment capable of performing image transformations according to at least one embodiment are illustrated. The environment 1100 may include a processor 1102, a memory 1104, an instruction switch 1106, a memory 1108 (sometimes referred to as dynamic random access memory or DRAM), and functional blocks 1110a and 1110b (unless otherwise stated, individually referred to as functional block 1110, collectively referred to as functional block 1110). In some embodiments, the processor 1102, memory 1104, instruction switch 1106, memory 1108, and functional block 1110 may be interconnected via wired and / or wireless connections (e.g., establishing connections for communication, etc.). In some embodiments, the components of environment 1100 may be included in a system-on-a-chip (SoC). For example, the components of environment 1100 may be included in one or more SoCs that form an integrated circuit by combining some or all of the components of environment 1100.
[0069] Processor 1102 may include one or more processors, such as one or more central processing units (CPUs), graphics processing units (GPUs), microprocessors, microcontrollers, etc. Processor 1102 may be interconnected with an instruction cache (not explicitly shown) that stores instructions for execution by processor 1102. In some embodiments, processor 1102 may be configured to output to and from... Figure 11Data associated with the configuration and / or control of one or more devices. For example, processor 1102 may be configured to output data associated with the configuration of direct memory access (DMA) hardware sequencer 1114a and / or DMA hardware sequencer 1114b to control DMA transfers to and from vector memory (VMEM) 1112a and / or VMEM 1112b of function blocks 1110a and 1110b, respectively.
[0070] Memory 1104 (sometimes referred to as an L2 buffer or L2 cache) may include a storage device interconnected with DMA hardware sequencers 1114a and / or DMA hardware sequencer 1114b of functional block 1110. In some embodiments, memory 1104 may be configured to receive and store data from DMA hardware sequencers 1114a and / or DMA hardware sequencer 1114b of functional block 1110, as described herein. In some embodiments, memory 1104 may have one or more (e.g., two) banks for enabling simultaneous read or write requests. For example, memory 1104 may have a first bank associated with DMA hardware sequencer 1114a and a second bank associated with DMA hardware sequencer 1114b.
[0071] Instruction switch 1106 may include one or more processors configured to scan memory 1108, receive data from memory 1108, cause data stored in memory 1108 and / or the local memory of instruction switch 1106 to be loaded into VMEM 1112, etc. For example, instruction switch 1106 may be coupled to memory 1108 and / or include internal memory storing instructions relating to operating one or more devices of the corresponding function block 1110. In one example, instruction switch 1106 may be configured to fetch and provide data associated with instructions to perform one or more DMA transfers as described herein. In another example, instruction switch 1106 may be configured to fetch and provide data associated with instructions to perform one or more operations specific to one or more devices of function block 1110. In an illustrative example, instruction switch 1106 may be configured to fetch and provide data associated with instructions to perform one or more filtering operations, and instruction switch 1106 may transfer data to cache 1120 of the corresponding function block 1110. In this illustrative example, the corresponding cache 1120 can be configured to transfer (e.g., load) instruction-related data into the VPU 1116 or PPE 1118 to enable the corresponding device to perform one or more filtering operations.
[0072] Memory 1108 may include a storage device interconnected with DMA hardware sequencers 1114a and / or DMA hardware sequencer 1114b of functional block 1110. In some embodiments, memory 1108 may receive and store sensor data generated by one or more sensors of the robot. For example, during robot operation, memory 1108 may be configured to receive data at least in part based on direct interconnection with one or more sensors or indirect interconnection with one or more sensors (e.g., via communication via a CAN bus, etc.). In these examples, sensor data may include image data associated with one or more images generated by one or more cameras, one or more LiDAR data associated with one or more point clouds generated by one or more LiDAR sensors, radar data associated with one or more radar images generated by one or more radar sensors, etc. In some embodiments, memory 1108 may be configured to provide (e.g., transmit) the sensor data stored therein to one or more components of functional block 1110. For example, during the processing of one or more images generated by one or more cameras of the robot, DMA hardware sequencers 1114a and / or DMA hardware sequencer 1114b can retrieve image data from memory 1108, causing the image data to be stored in VMEM 1112a and / or VMEM 1112b, respectively. In some embodiments, memory 1108 can receive and store data from DMA hardware sequencers 1114a and / or DMA hardware sequencer 1114b of function block 1110. For example, DMA hardware sequencers 1114a and / or DMA hardware sequencer 1114b can provide memory 1108 with image data updated at least in part based on the processing of the image data, and memory 1108 can store the updated image data in memory 1108.
[0073] Functional block 1110 may include VMEM 1112a, 1112b, DMA hardware sequencers 1114a, 1114b, vector processing units (VPU) 1116a, 1116b, pixel processing engines (PPE) 1118a, 1118b, caches 1120a, 1120b, 1120c, 1120d, and decoupled lookup tables (DLUT) 1122a, 1122b. For clarity, unless otherwise specified, each component will be referred to as VMEM 1112, DMA hardware sequencer 1114, VPU 1116, PPE 1118, cache 1120, and DLUT 1122, respectively; and collectively as VMEM 1112, DMA hardware sequencer 1114, VPU 1116, PPE 1118, cache 1120, and DLUT 1122, unless otherwise specified. Although some interconnections are shown in the figure, it should be understood that the connections shown are for simplicity only, and unless otherwise explicitly stated, one or more devices in function block 1110 may be interconnected with one or more other devices in function block 1110.
[0074] VMEM 1112 may include a storage device interconnected with the processor 1102 and the corresponding DMA hardware sequencer 1114, VPU 1116, PPE 1118, and cache 1120 of function block 1110. In some embodiments, VMEM 1112 may receive and store sensor data acquired from memory 1108. For example, VMEM 1112 may receive and store sensor data acquired from memory 1108 by DMA hardware sequencer 1114. Additionally or alternatively, VMEM 1112 may receive and store sensor data acquired from memory 1108 via instruction switch 1106. In some embodiments, VMEM 1112 may be interconnected with PPE 1118 via a decoupled load / store unit (DLSU) 1124. As described herein, the DLSU 1124 can be configured to buffer data communicated between the VMEM 1112 and the PPE 1118 to reduce the latency associated with communication between the VMEM 1112 and the PPE 1118.
[0075] The DMA hardware sequencer 1114 may include one or more processors for controlling the execution of one or more instructions. For example, the DMA hardware sequencer 1114 may receive instructions from processor 1102, a corresponding VPU 1116 or PPE 1118, and / or storage devices (e.g., devices associated with the DMA hardware sequencer 1114, such as internal or external memory, not explicitly shown), and the DMA hardware sequencer 1114 may cooperate with the corresponding VPU 1116 and / or PPE 1118 to perform one or more operations during instruction execution. In an illustrative example, the DMA hardware sequencer 1114 may receive instructions that cause it to retrieve data (e.g., sensor data, etc.) from memory 1108 and store that data in a corresponding VMEM 1112. In some embodiments, the DMA hardware sequencer 1114 may perform one or more operations at least in part based on the data retrieved from memory 1108. For example, the DMA hardware sequencer 1114 can fill frames (e.g., image frames), manipulate addresses, manage overlapping data, manage different traversal orders, consider different frame sizes, and so on. In some embodiments, the DMA hardware sequencer 1114 can receive signals (e.g., from VPU 1116 or PPE 1118) indicating that one or more operations have been performed on data stored in VMEM 1112, updating one or more descriptors at least in part based on updates to the data, and performing operations on the data again.
[0076] VPU 1116 may include one or more processors that execute one or more instructions. For example, VPU 1116 may receive instructions from processor 1102, and the corresponding VPU 1116 may cooperate with DMA hardware sequencer 1114 and / or PPE 1118 to perform one or more operations during instruction execution. In an illustrative example, VPU 1116 may receive instructions from processor 1102 that cause VPU 1116 to trigger the corresponding DMA hardware sequencer 1114 to retrieve sensor data from memory 1108 and store the sensor data in the corresponding VMEM 1112. In the example, VPU 1116 may process the data stored in the corresponding VMEM 1112 and write the data back to VMEM 1112. In these examples, the data written by VPU 1116 to the corresponding VMEM 1112 may include updated sensor data and / or data generated at least in part based on the analysis performed by VPU 1116 on the sensor data, including the location of objects or features within a frame, classification indicating the type of object or agent, etc. In some embodiments, VPU 1116 may provide (e.g., send, transmit, transfer, etc.) signals to the corresponding DMA hardware sequencer 1114 to cause DMA hardware sequencer 1114 to update one or more descriptors (described herein). For example, VPU 1116 may send a signal to the corresponding DMA hardware sequencer 1114 to cause DMA hardware sequencer 1114 to update one or more descriptors at least in part based on the data written by VPU 1116 to the corresponding VMEM 1112.
[0077] PPE 1118 may include one or more processors that execute one or more instructions. For example, PPE 1118 may receive instructions from processor 1102, and the corresponding PPE 1118 may cooperate with DMA hardware sequencer 1114 and / or VPU 1116 to perform one or more operations during instruction execution. In an illustrative example, PPE 1118 may receive instructions from processor 1102 that cause PPE 1118 to trigger the corresponding DMA hardware sequencer 1114 to retrieve (e.g., receive, acquire, capture, etc.) sensor data from memory 1108 and store the sensor data in the corresponding VMEM 1112. In the example, PPE 1118 may process the data stored in the corresponding VMEM 1112 and write the data back to VMEM 1112. In these examples, the data written by PPE 1118 to the corresponding VMEM 1112 may include updated sensor data and / or data generated at least in part based on the analysis performed by PPE 1118 on the sensor data, including the location of objects or features within a frame, classification indicating the type of object or agent, etc. In some embodiments, PPE 1118 may signal to the corresponding DMA hardware sequencer 1114 to update one or more descriptors (described herein). For example, PPE 1118 may signal to the corresponding DMA hardware sequencer 1114 to update one or more descriptors at least in part based on the data written by PPE 1118 to the corresponding VMEM 1112.
[0078] Cache 1120 may include a storage device interconnected with VMEM 1112 and / or instruction switch 1106. As described above, cache 1120 may receive instruction-associated data from instruction switch 1106 and load instructions into one or more devices of function block 1110 to cause the one or more devices to operate according to the instructions. DLUT 1122 may include a processor and / or memory configured to store one or more lookup tables. In some embodiments, DLUT 1122 may be configured to implement communication between processor 1102 and one or more components of function block 1110. For example, DLUT 1122 may be configured to communicate with processor 1102 and / or Figure 11 The processor 1102 communicates with one or more memory devices (e.g., memory 1108 and / or memory 1104). The DLUT 1122 can then manage the communication between the processor 1102 and... Figure 11The DLSU 1124 may include storage devices interconnected with VMEM 1112 and PPE 1118 of a given functional block 1110. For example, the DLSU 1124 may receive and store sensor data acquired by VMEM 1112 from memory 1108. Additionally or alternatively, the DLSU 1124 may receive and store data provided as output by PPE 1118.
[0079] In some embodiments, the systems and methods described herein can be performed in a simulated environment (e.g., NVIDIA's DriveSIM) using simulated data (e.g., simulated sensor data from simulated sensors of a virtual machine or simulated machine). For example, simulated sensor data and / or map data can be used to identify regions of interest (e.g., parking spaces) and sub-regions of interest (e.g., sub-regions of parking spaces including curbs, wheel stops, etc.) in the simulated environment, and this information can be used to perform operations associated with the virtual machine in the environment (e.g., parking). These simulated operations can be used to test the performance of the underlying algorithms, systems, and / or processes before deploying them to the real world. In some cases, simulation can be used to generate synthetic training data, for example, training data including regions of interest and / or sub-regions of interest from the simulation. The synthetic training data can then be processed (as a supplement or alternative to real-world data) to determine the geometry and / or other information associated with the region of interest, such as, for example, the location of a parking space or pallet delivery within a warehouse. In any example, such as when using a simulated environment for testing, validation, training, etc., one or more optical transport algorithms (e.g., ray tracing and / or path tracing algorithms) can be used to render or otherwise generate the simulated environment and / or associated training data. In some embodiments, simulated environments and / or one or more of their objects, features, or components can be generated or managed within a 3D content collaboration platform (e.g., NVIDIA's OMNIVERSE) for use in industrial digitization, generative physics AI, and / or other use cases, applications, or services. For example, the content collaboration platform or system may include a system that uses or develops generic scene descriptors (USD) (e.g., OpenUSD) data to manage objects, features, scenes, etc., in simulated environments, digital environments, etc. The platform may include realistic physical simulations, such as using NVIDIA's PhysX SDK, to simulate real physical phenomena and physical interactions with simulations hosted on the platform. This platform can integrate OpenUSD and ray tracing / path tracing / light transport simulations (such as NVIDIA's RTX rendering technology) into software tools and simulation workflows for building, training, deploying, or testing AI systems (such as systems for testing, validating, training (e.g., machine learning models, neural networks, etc.) and / or other tasks related to automobiles, robots, machines, or other applications).
[0080] In at least one embodiment, a small language model optimized for at least one target language can be hosted in a cloud environment and made available to various people, entities, systems, operations, etc. In at least one embodiment, such a model can also be provided for various entities to deploy and use on their resources, whether local resources or allocated portions of multi-tenant physical or virtual resources, and other such options.
[0081] In some embodiments, the model may be deployed as part of a software container (such as NVIDIA's NIM), which may include the code and support required to run inference using the model. Such containers may include a set of easy-to-use inference microservices to accelerate the deployment of the underlying model in the cloud or data center and to help manage the security of request and generated response data. The container may be pre-configured for easy deployment and may include one or more optimized inference engines. The container may also include management functions for handling tasks such as identity management, metric generation, health checks, and status monitoring. Therefore, in some examples, a machine learning model (a small language model) may be packaged as a microservice (e.g., an inference microservice) that may include a container (e.g., an operating system (OS) level virtualization package) that may include an application programming interface (API) layer, a server layer, a runtime layer, and / or a model "engine". For example, the inference microservice may include the container itself and the model (e.g., weights and biases). In some cases, such as when the machine learning model is small enough (e.g., with a small number of parameters), the model may be included within the container itself. In other examples (e.g., when the model is large), the model may be hosted / stored in the cloud (e.g., a data center) and / or hosted on-premises and / or at the edge (e.g., on a local server or computing device, but outside of a container). In these embodiments, the model may be accessed via one or more APIs (e.g., a REST API). Therefore, in some embodiments, the machine learning models described herein may be deployed as inference microservices to accelerate model deployment on any cloud, data center, or edge computing system while ensuring data security. For example, an inference microservice may include one or more APIs, pre-configured containers for simplified deployment, an optimized inference engine (e.g., execution software built using standardized AI model deployments, such as NVIDIA's Triton Inference Server, and / or one or more APIs for high-performance deep learning inference, which may include inference runtime and model optimizations that provide low latency and high throughput for production applications—e.g., NVIDIA's TensorRT), and / or enterprise management data for telemetry (e.g., including identity, metrics, health checks, and / or monitoring). The machine learning models described in this paper can be included as part of a microservice along with an acceleration infrastructure that can be deployed with a single command and / or orchestrated and automatically scaled using a container orchestration system on the acceleration infrastructure (e.g., reaching data center scale on a single device).Therefore, the inference microservice may include a machine learning model (e.g., a machine learning model optimized for high-performance inference), inference runtime software for executing the machine learning model and providing output / response to inputs (e.g., user queries, prompts, etc.), and enterprise management software for providing health checks, identification, and / or other monitoring functions. In some embodiments, the inference microservice may include software for in-situ replacement and / or updating of the machine learning model. During replacement or updating, the software performing the replacement / update may maintain the user configurations of the inference runtime software and the enterprise management software.
[0082] The systems and methods described herein can be used for a variety of purposes, such as, but not limited to, machine (e.g., robots, vehicles, construction machinery, warehouse vehicles / machines, autonomous, semi-autonomous and / or other machine types) control, machine motion, machine driving, synthetic data generation, model training (e.g., using real data, augmented data and / or synthetic data, such as synthetic data generated using simulation platforms or systems, synthetic data generation techniques (e.g., but not limited to those described herein), perception, augmented reality (AR), virtual reality (VR), mixed reality (MR), robotics, security and supervision (e.g., in smart city implementations), autonomous or semi-autonomous machine applications, deep learning, environmental simulation, object or actor simulation and / or digital twins, data center processing, conversational AI, optical transport simulation (e.g., ray tracing, path tracing, etc.), distributed or collaborative content creation of 3D assets (e.g., using generic scene descriptor (USD) data, such as OpenUSD and / or other data types), cloud computing, generative artificial intelligence (e.g., using one or more diffusion models, converter models, etc.) and / or any other suitable application.
[0083] The disclosed embodiments can be included in a variety of different systems, such as automotive systems (e.g., control systems for autonomous or semi-autonomous machines, perception systems for autonomous or semi-autonomous machines), systems implemented using robots or robotic platforms, aviation systems, medical systems, marine systems, smart area monitoring systems, systems for performing deep learning operations, systems for performing simulation operations (e.g., in driving or vehicle simulations, in robot simulations, in smart city or surveillance simulations, etc.), systems for performing digital twin operations (e.g., in conjunction with a collaborative content creation platform or system, such as, but not limited to, NVIDIA's OMNIVERSE and / or another platform, system, or service using USD or OpenUSD data types), systems implemented using edge devices, systems containing one or more virtual machines (VMs), and systems for... Systems that perform synthetic data generation operations (e.g., using one or more neural rendering fields (NERF), Gaussian sputtering techniques, diffusion models, converter models, etc.), systems located at least partially in a data center, systems for performing conversational AI operations, systems that implement one or more language models (e.g., one or more large language models (LLM), one or more visual language models (VLM), one or more multimodal language models, etc.), systems for performing optical transport simulations, systems for performing collaborative content creation of 3D assets (e.g., using generic scene descriptor (USD) data, such as OpenUSD, computer-aided design (CAD) data, 2D and / or 3D graphics or design data and / or other data types), systems implemented using at least partially cloud computing resources, and / or other types of systems.
[0084] Data Center
[0085] Figure 12 An example data center 1200 that can be used with at least one embodiment is shown. In at least one embodiment, the data center 1200 includes a data center infrastructure layer 1210, a framework layer 1220, a software layer 1230, and an application layer 1240.
[0086] In at least one embodiment, such as Figure 12As shown, the data center infrastructure layer 1210 may include a resource coordinator 1212, grouped computing resources 1214, and node computing resources (“nodes CR”) 1216(1)-1216(N), where “N” represents any positive integer. In at least one embodiment, nodes CR 1216(1)-1216(N) may include, but are not limited to, any number of central processing units (“CPUs”) or other processors (including accelerators, field-programmable gate arrays (FPGAs), graphics processors, etc.), memory devices (e.g., dynamic read-only memory), storage devices (e.g., solid-state drives or disk drives), network input / output (“NWI / O”) devices, network switches, virtual machines (“VMs”), power modules, and cooling modules, etc. In at least one embodiment, one or more nodes CR 1216(1)-1216(N) may be servers having one or more of the aforementioned computing resources.
[0087] In at least one embodiment, the grouped computing resources 1214 may include individual groups (not shown) of node CRs housed in one or more racks, or a plurality of racks (also not shown) housed in data centers in various geographic locations. The individual groups of node CRs within the grouped computing resources 1214 may include computing, networking, memory, or storage resources that can be configured or allocated to support groups of one or more workloads. In at least one embodiment, several node CRs, including CPUs or processors, may be grouped within one or more racks to provide computing resources to support one or more workloads. In at least one embodiment, the one or more racks may also include any number of power modules, cooling modules, and network switches, in any combination.
[0088] In at least one embodiment, resource coordinator 1212 may be configured or otherwise control one or more nodes CR1216(1)-1216(N) and / or grouped computing resources 1214. In at least one embodiment, resource coordinator 1212 may include a software design infrastructure (“SDI”) management entity for data center 1200. In at least one embodiment, resource coordinator 1212 may include hardware, software, or some combination thereof.
[0089] In at least one embodiment, such as Figure 12As shown, framework layer 1220 includes job scheduler 1222, configuration manager 1224, resource manager 1226, and distributed file system 1228. In at least one embodiment, framework layer 1220 may include a framework of software 1232 supporting software layer 1230 and / or one or more applications 1242 supporting application layer 1240. In at least one embodiment, software 1232 or application 1242 may respectively include web-based service software or applications, such as services or applications provided by Amazon Web Services, Google Cloud, and Microsoft Azure. In at least one embodiment, framework layer 1220 may be, but is not limited to, a free and open-source software web application framework, such as Apache Spark™ (hereinafter referred to as "Spark") which can utilize distributed file system 1228 for large-scale data processing (e.g., "big data"). In at least one embodiment, job scheduler 1232 may include Spark drivers to facilitate the scheduling of workloads supported by the various layers of data center 1200. In at least one embodiment, configuration manager 1224 may be able to configure different layers, such as software layer 1230 and framework layer 1220 including Spark and distributed file system 1228 for supporting large-scale data processing. In at least one embodiment, resource manager 1226 is able to manage cluster or group computing resources mapped to or allocated to support distributed file system 1228 and job scheduler 1222. In at least one embodiment, cluster or group computing resources may include group computing resources 1214 on data center infrastructure layer 1210. In at least one embodiment, resource manager 1226 may coordinate with resource coordinator 1212 to manage these mapped or allocated computing resources.
[0090] In at least one embodiment, the software 1232 included in the software layer 1230 may include software used by at least a portion of the nodes CR1216(1)-1216(N), the grouped computing resources 1214, and / or the distributed file system 1228 of the framework layer 1220. One or more types of software may include, but are not limited to, Internet web page search software, email virus scanning software, database software, and streaming video content software.
[0091] In at least one embodiment, the one or more applications 1242 included in the application layer 1240 may include one or more types of applications used by at least a portion of nodes CR1216(1)-1216(N), grouped computing resources 1214, and / or the distributed file system 1228 of the framework layer 1220. The one or more types of applications may include, but are not limited to, any number of genomics applications, cognitive computing and machine learning applications, including training or inference software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), or other machine learning applications used in conjunction with one or more embodiments.
[0092] In at least one embodiment, any of the configuration manager 1224, resource manager 1226, and resource coordinator 1212 can implement any number and type of self-modification actions based on any amount and type of data acquired in any technically feasible manner. In at least one embodiment, self-modification actions can mitigate potentially poor configuration decisions by data center operators of data center 1200 and can prevent underutilization and / or poor performance of the data center.
[0093] In at least one embodiment, data center 1200 may include tools, services, software, or other resources to train one or more machine learning models or to use one or more machine learning models to predict or infer information according to one or more embodiments described herein. For example, in at least one embodiment, a machine learning model can be trained by calculating weight parameters based on a neural network architecture using the software and computing resources described above with respect to data center 1200. In at least one embodiment, information can be inferred or predicted using trained machine learning models corresponding to one or more neural networks using the resources described above with respect to data center 1200, by using weight parameters calculated through one or more training techniques described herein.
[0094] In at least one embodiment, the data center may use a CPU, application-specific integrated circuit (ASIC), GPU, FPGA, or other hardware to utilize the aforementioned resources to perform training and / or inference. Furthermore, one or more of the aforementioned software and / or hardware resources may be configured as a service to allow a user to train or perform information inference, such as image recognition, speech recognition, or other artificial intelligence services.
[0095] Such components can be used to generate alternative view images, such as bird's-eye view images, from parallax data using hardware with limited capacity (such as embedded processors that do not access external memory).
[0096] Computer System
[0097] Figure 13 This is a block diagram illustrating an exemplary computer system according to at least one embodiment. The exemplary computer system may be a system of interconnected devices and components, a system-on-a-chip (SOC), or some combination thereof formed with a processor, which may include an execution unit to execute instructions. In at least one embodiment, according to this disclosure, such as the embodiments described herein, computer system 1300 may include, but is not limited to, components such as processor 1302, whose execution unit includes logic to execute algorithms for process data. In at least one embodiment, computer system 1300 may include a processor, such as those available from Intel Corporation of Santa Clara, California. Processor family, Xeon TM , XScale TM and / or StrongARM TM , Core TM or Nervana TM A microprocessor may be used, although other systems (including PCs, engineering workstations, set-top boxes, etc.) with other microprocessors may also be used. In at least one embodiment, computer system 1300 may execute a version of the Windows operating system available from Microsoft Corporation of Redmond, Washington, although other operating systems (such as UNIX and Linux), embedded software, and / or graphical user interfaces may also be used.
[0098] The embodiments can be used in other devices, such as handheld devices and embedded applications. Some examples of handheld devices include cellular phones, Internet Protocol (IP) devices, digital cameras, personal digital assistants (“PDAs”), and handheld PCs. In at least one embodiment, the embedded application may include a microcontroller, a digital signal processor (“DSP”), a system-on-a-chip (SoC), a network computer (“NetPC”), a set-top box, a network hub, a wide area network (“WAN”) switch, or any other system that can execute one or more instructions according to at least one embodiment.
[0099] In at least one embodiment, the computer system 1300 may include, but is not limited to, a processor 1302, which may include, but is not limited to, one or more execution units 1308, to perform machine learning model training and / or inference according to the techniques described herein. In at least one embodiment, the computer system 1300 is a single-processor desktop or server system, but in another embodiment, the computer system 1300 may be a multiprocessor system. In at least one embodiment, the processor 1302 may include, but is not limited to, a Complex Instruction Set Computing (“CISC”) microprocessor, a Reduced Instruction Set Computing (“RISC”) microprocessor, a Very Long Instruction Word (“VLIW”) microprocessor, a processor implementing instruction set combination, or any other processor device, such as a digital signal processor. In at least one embodiment, the processor 1302 may be coupled to a processor bus 1310, which can transmit data signals between the processor 1302 and other components in the computer system 1300.
[0100] In at least one embodiment, processor 1302 may include, but is not limited to, a Level 1 (“L1”) internal cache memory (“cache”) 1304. In at least one embodiment, processor 1302 may have a single internal cache or multiple levels of internal cache. In at least one embodiment, the cache memory may reside external to processor 1302. Depending on specific implementation and requirements, other embodiments may also include a combination of internal and external caches. In at least one embodiment, register file 1306 may store different types of data in various registers, including but not limited to integer registers, floating-point registers, status registers, and instruction pointer registers.
[0101] In at least one embodiment, a logic execution unit 1308, which performs integer and floating-point operations, is also located within the processor 1302. In at least one embodiment, the processor 1302 may also include a microcode (“ucode”) read-only memory (“ROM”) for storing microcode of certain macro instructions. In at least one embodiment, the execution unit 1308 may include logic for processing a packaged instruction set 1309. In at least one embodiment, by including the packaged instruction set 1309 in the instruction set of a general-purpose processor, along with the associated circuitry for executing the instructions, the packaged data in the processor 1302 can be used to perform operations used by numerous multimedia applications. In one or more embodiments, many multimedia applications can be executed more quickly and efficiently by using the full width of the processor’s data bus to perform operations on the packaged data, which may eliminate the need to transfer smaller data units on the processor’s data bus to perform one or more operations on one data element at a time.
[0102] In at least one embodiment, execution unit 1308 may also be used in a microcontroller, embedded processor, graphics device, DSP, and other types of logic circuitry. In at least one embodiment, computer system 1300 may include, but is not limited to, memory 1320. In at least one embodiment, memory 1320 may be implemented as a dynamic random access memory (“DRAM”) device, a static random access memory (“SRAM”) device, a flash memory device, or other storage device. In at least one embodiment, memory 1320 may store instructions 1319 and / or data 1321 represented by data signals that can be executed by processor 1302.
[0103] In at least one embodiment, the system logic chip may be coupled to the processor bus 1310 and the memory 1320. In at least one embodiment, the system logic chip may include, but is not limited to, a memory controller hub (“MCH”) 1316, and the processor 1302 may communicate with the MCH 1316 via the processor bus 1310. In at least one embodiment, the MCH 1316 may provide a high-bandwidth memory path 1318 to the memory 1320 for instruction and data storage, as well as for storage of graphics commands, data, and textures. In at least one embodiment, the MCH 1316 may initiate data signals between the processor 1302, the memory 1320, and other components in the computer system 1300, and bridge data signals between the processor bus 1310, the memory 1320, and the system I / O 1322. In at least one embodiment, the system logic chip may provide a graphics port for coupling to a graphics controller. In at least one embodiment, MCH 1316 may be coupled to memory 1320 via high-bandwidth memory path 1318, and graphics / video card 1312 may be coupled to MCH 1316 via Accelerated Graphics Port (“AGP”) interconnect 1314.
[0104] In at least one embodiment, computer system 1300 may use system I / O 1322, which is a proprietary hub interface bus, to couple MCH 1316 to I / O controller hub (“ICH”) 1330. In at least one embodiment, ICH 1330 may provide direct connectivity to certain I / O devices via a local I / O bus. In at least one embodiment, the local I / O bus may include, but is not limited to, a high-speed I / O bus for connecting peripheral devices to memory 1320, chipset, and processor 1302. Examples may include, but are not limited to, an audio controller 1329, a firmware hub (“FlashBIOS”) 1328, a wireless transceiver 1326, a data storage 1324, a conventional I / O controller 1323 including a user input and keyboard interface 1325, a serial expansion port 1327 (e.g., a Universal Serial Bus (USB) port), and a network controller 1334. Data storage 1324 may include a hard disk drive, floppy disk drive, CD-ROM device, flash memory device, or other mass storage device.
[0105] In at least one embodiment, Figure 13 A system including interconnected hardware devices or "chips" is shown, while in other embodiments, Figure 13 An exemplary system-on-a-chip (SoC) may be illustrated. In at least one embodiment, the device may be interconnected with a proprietary interconnect, a standardized interconnect (e.g., PCIe), or some combination thereof. In at least one embodiment, one or more components of the computer system 1300 are interconnected using a compute fast link (CXL) interconnect.
[0106] Such components can be used to generate alternative view images, such as bird's-eye view images, from parallax data using hardware with limited capacity (such as embedded processors that do not access external memory).
[0107] Figure 14 This is a block diagram illustrating an electronic device 1400 for utilizing a processor 1410 according to at least one embodiment. In at least one embodiment, the electronic device 1400 may be, for example, but not limited to, a laptop computer, tower server, rack server, blade server, laptop computer, desktop computer, tablet computer, mobile device, telephone, embedded computer, or any other suitable electronic device.
[0108] In at least one embodiment, the electronic device 1400 may, but is not limited to, a processor 1410 communicatively coupled to any suitable number or type of components, peripherals, modules, or devices. In at least one embodiment, the processor 1410 is coupled using a bus or interface, such as an I2C bus, a system management bus (“SMBus”), a low pin count (LPC) bus, a serial peripheral interface (“SPI”), a high-definition audio (“HDA”) bus, a serial advanced technology accessory (“SATA”) bus, a universal serial bus (“USB”) (versions 1, 2, and 3), or a universal asynchronous receiver / transmitter (“UART”) bus.
[0109] In at least one embodiment, Figure 14 The system shown includes interconnected hardware devices or "chips," while in other embodiments, Figure 14 An exemplary system-on-a-chip (SoC) can be illustrated. In at least one embodiment, Figure 14 The device shown can be interconnected with proprietary interconnects, standardized interconnects (e.g., PCIe), or some combination thereof. In at least one embodiment, Figure 14 One or more components are interconnected using Computational Fast Link (CXL) interconnects.
[0110] In at least one embodiment, Figure 14 It may include a display 1424, a touch screen 1425, a touchpad 1430, a near field communication unit (“NFC”) 1445, a sensor hub 1440, a thermal sensor 1446, a fast chipset (“EC”) 1435, a trusted platform module (“TPM”) 1438, a BIOS / firmware / flash (“BIOS, FW Flash”) 1422, a DSP 1460, a drive 1420 (e.g., a solid-state drive (“SSD”) or a hard disk drive (“HDD”)), a wireless local area network unit (“WLAN”) 1450, a Bluetooth unit 1452, a wireless wide area network unit (“WWAN”) 1456, a global positioning system (GPS) 1455, a camera (“USB 3.0 camera”) 1454 (e.g., a USB 3.0 camera), and / or a low-power double data rate (“LPDDR”) memory unit (“LPDDR3”) 1415 implemented in, for example, the LPDDR3 standard. These components can each be implemented in any suitable way.
[0111] In at least one embodiment, other components may be communicatively coupled to processor 1410 via the components described above. In at least one embodiment, accelerometer 1441, ambient light sensor (“ALS”) 1442, compass 1443, and gyroscope 1444 may be communicatively coupled to sensor hub 1440. In at least one embodiment, thermal sensor 1439, fan 1437, keyboard 1436, and touchpad 1430 may be communicatively coupled to EC 1435. In at least one embodiment, speaker 1463, earphone 1464, and microphone (“mic”) 1465 may be communicatively coupled to audio unit (“audio codec and Class D amplifier”) 1462, which in turn may be communicatively coupled to DSP 1460. In at least one embodiment, audio unit 1462 may include, for example, but not limited to, audio encoder / decoder (“codec”) and Class D amplifier. In at least one embodiment, SIM card (“SIM”) 1457 may be communicatively coupled to WWAN unit 1456. In at least one embodiment, components such as WLAN unit 1450, Bluetooth unit 1452, and WWAN unit 1456 can be implemented as next-generation form factor (NGFF).
[0112] Such components can be used to generate alternative view images, such as bird's-eye view images, from parallax data using hardware with limited capacity (such as embedded processors that do not access external memory).
[0113] Figure 15 This is a block diagram of a processing system according to at least one embodiment. In at least one embodiment, system 1500 includes one or more processors 1502 and one or more graphics processors 1508, and may be a single-processor desktop system, a multi-processor workstation system, or a server system having a large number of processors 1502 or processor cores 1507. In at least one embodiment, system 1500 is a processing platform incorporated within a system-on-a-chip (SoC) integrated circuit for use in mobile, handheld, or embedded devices.
[0114] In at least one embodiment, system 1500 may include or be integrated into a server-based gaming platform, including a game console, mobile game console, handheld game console, or online game console, which are game and media consoles. In at least one embodiment, system 1500 is a mobile phone, smartphone, tablet computing device, or mobile internet device. In at least one embodiment, processing system 1500 may also include components coupled to or integrated into a wearable device, such as a smartwatch, smart glasses, augmented reality, or virtual reality device. In at least one embodiment, processing system 1500 is a television or set-top box device having one or more processors 1502 and a graphical interface generated by one or more graphics processors 1508.
[0115] In at least one embodiment, each of the one or more processors 1502 includes one or more processor cores 1507 for processing instructions that, when executed, perform operations against the system and user software. In at least one embodiment, each of the one or more processor cores 1507 is configured to process a specific instruction set 1509. In at least one embodiment, the instruction set 1509 may facilitate Complex Instruction Set Computing (CISC), Reduced Instruction Set Computing (RISC), or computation via Very Long Instruction Word (VLIW). In at least one embodiment, the one or more processor cores 1507 may each process a different instruction set 1509, which may include instructions that facilitate the emulation of other instruction sets. In at least one embodiment, the one or more processor cores 1507 may also include other processing devices, such as digital signal processors (DSPs).
[0116] In at least one embodiment, one or more processors 1502 include cache memory 1504. In at least one embodiment, one or more processors 1502 may have a single internal cache or multiple levels of internal caches. In at least one embodiment, the cache memory is shared among the various components of one or more processors 1502. In at least one embodiment, one or more processors 1502 also use an external cache (e.g., a Level 3 (L3) cache or a Last Level Cache (LLC)) (not shown), which can be shared among one or more processor cores 1507 using known cache coherence techniques. In at least one embodiment, one or more processors 1502 further include a register file 1506, and the processor may include different types of registers (e.g., integer registers, floating-point registers, status registers, and instruction pointer registers) for storing different types of data. In at least one embodiment, register file 1506 may include general-purpose registers or other registers.
[0117] In at least one embodiment, one or more processors 1502 are coupled to one or more interface buses 1510 to transmit communication signals, such as address, data, or control signals, between the processors 1502 and other components in the system 1500. In at least one embodiment, the interface buses 1510 may be processor buses, such as a version of the Direct Media Interface (DMI) bus. In at least one embodiment, the interface buses 1510 are not limited to DMI buses and may include one or more peripheral component interconnect buses (e.g., PCI, PCI Express), memory buses, or other types of interface buses. In at least one embodiment, the processor 1502 includes an integrated memory controller 1516 and a platform controller hub 1530. In at least one embodiment, the memory controller 1516 facilitates communication between memory devices and other components of the processing system 1500, while the platform controller hub (PCH) 1530 provides connectivity to I / O devices via a local I / O bus.
[0118] In at least one embodiment, memory device 1520 may be a dynamic random access memory (DRAM) device, a static random access memory (SRAM) device, a flash memory device, a phase-change memory device, or a device with suitable performance for use as processor memory. In at least one embodiment, memory device 1520 may be used as system memory of processing system 1500 to store data 1522 and instructions 1521 for use when one or more processors 1502 execute an application or process. In at least one embodiment, memory controller 1516 is also coupled to an optional external graphics processor 1512, which may communicate with one or more graphics processors 1508 of one or more processors 1502 to perform graphics and media operations. In at least one embodiment, display device 1511 may be connected to processor 1502. In at least one embodiment, display device 1511 may include one or more internal display devices, such as in mobile electronic devices or laptop devices, or external display devices connected via a display interface (e.g., DisplayPort). In at least one embodiment, the display device 1511 may include a head-mounted display (HMD), such as a stereoscopic display device for virtual reality (VR) or augmented reality (AR) applications.
[0119] In at least one embodiment, the platform controller hub 1530 enables peripheral devices to connect to the memory device 1520 and one or more processors 1502 via a high-speed I / O bus. In at least one embodiment, the I / O peripheral devices include, but are not limited to, an audio controller 1546, a network controller 1534, a firmware interface 1528, a wireless transceiver 1526, a touch sensor 1525, and a data storage device 1524 (e.g., a hard disk drive, flash memory, etc.). In at least one embodiment, the data storage device 1524 may be connected via a storage interface (e.g., SATA) or via a peripheral bus, such as a peripheral component interconnect bus (e.g., PCI, PCIe). In at least one embodiment, the touch sensor 1525 may include a touchscreen sensor, a pressure sensor, or a fingerprint sensor. In at least one embodiment, the wireless transceiver 1526 may be a Wi-Fi transceiver, a Bluetooth transceiver, or a mobile network transceiver, such as a 3G, 4G, or LTE transceiver. In at least one embodiment, the firmware interface 1528 enables communication with the system firmware and may be, for example, a Unified Extensible Firmware Interface (UEFI). In at least one embodiment, network controller 1534 may enable network connectivity to a wired network. In at least one embodiment, a high-performance network controller (not shown) is coupled to one or more interface buses 1510. In at least one embodiment, audio controller 1546 is a multi-channel high-definition audio controller. In at least one embodiment, processing system 1500 includes an optional legacy I / O controller 1540 for coupling legacy (e.g., Personal System 2 (PS / 2)) devices to system 1500. In at least one embodiment, platform controller hub 1530 may also be connected to one or more Universal Serial Bus (USB) controllers 1542 that connect input devices, such as a keyboard and mouse combination 1543, a camera 1544, or other USB input devices.
[0120] In at least one embodiment, instances of the memory controller 1516 and platform controller hub 1530 may be integrated into a discrete external graphics processor, such as external graphics processor 1512. In at least one embodiment, the platform controller hub 1530 and / or the memory controller 1516 may be external to one or more processors 1502. For example, in at least one embodiment, system 1500 may include an external memory controller 1516 and a platform controller hub 1530, which may be configured as a memory controller hub and a peripheral controller hub in a system chipset communicating with processor 1502.
[0121] Such components can be used to generate alternative view images, such as bird's-eye view images, from parallax data using hardware with limited capacity (such as embedded processors that do not access external memory).
[0122] Figure 16 This is a block diagram of a processor 1600 having one or more processor cores 1602A-1602N, an integrated memory controller 1614, and an integrated graphics processor 1608 according to at least one embodiment. In at least one embodiment, the processor 1600 may include additional cores, up to and including additional cores 1602N, indicated by dashed boxes. In at least one embodiment, each of the one or more processor cores 1602A-1602N includes one or more internal cache units 1604A-1604N. In at least one embodiment, each processor core may also access one or more shared cache units 1606.
[0123] In at least one embodiment, one or more internal cache units 1604A-1604N and one or more shared cache units 1606 represent a cache memory hierarchy within the processor 1600. In at least one embodiment, one or more cache memory units 1604A-1604N may include at least one level of instruction and data cache within each processor core and one or more levels of cache in a shared intermediate cache, such as Level 2 (L2), Level 3 (L3), Level 4 (L4), or other levels of cache, wherein the highest level of cache preceding external memory is classified as LLC. In at least one embodiment, cache coherence logic maintains coherence between the various cache units 1606 and 1604A-1604N.
[0124] In at least one embodiment, the processor 1600 may further include a set of one or more bus controller units 1616 and a system agent core 1610. In at least one embodiment, the one or more bus controller units 1616 manage a set of peripheral buses, such as one or more PCI or PCIe buses. In at least one embodiment, the system agent core 1610 provides management functions for various processor components. In at least one embodiment, the system agent core 1610 includes one or more integrated memory controllers 1614 to manage access to various external memory devices (not shown).
[0125] In at least one embodiment, one or more processor cores 1602A-1602N include support for concurrent multithreading. In at least one embodiment, system agent core 1610 includes components for coordinating and operating one or more processor cores 1602A-1602N during multithreaded processing. In at least one embodiment, system agent core 1610 may additionally include a power control unit (PCU) including logic and components for regulating one or more power states of one or more processor cores 1602A-1602N and graphics processor 1608.
[0126] In at least one embodiment, processor 1600 further includes a graphics processor 1608 for performing graph processing operations. In at least one embodiment, graphics processor 1608 is coupled to one or more shared cache units 1606 and a system proxy core 1610 including one or more integrated memory controllers 1614. In at least one embodiment, system proxy core 1610 further includes a display controller 1611 for driving graphics processor output to one or more coupled displays. In at least one embodiment, display controller 1611 may also be a separate module coupled to graphics processor 1608 via at least one interconnect, or it may be integrated within graphics processor 1608.
[0127] In at least one embodiment, the ring-based interconnect unit 1612 is used to couple internal components of the processor 1600. In at least one embodiment, alternative interconnect units, such as point-to-point interconnects, switched interconnects, or other technologies, may be used. In at least one embodiment, the graphics processor 1608 is coupled to the ring-based interconnect unit 1612 via I / O link 1613.
[0128] In at least one embodiment, I / O link 1613 represents at least one of a variety of I / O interconnects, including packaged I / O interconnects that facilitate communication between various processor components and high-performance embedded memory module 1618 (e.g., eDRAM module). In at least one embodiment, each of one or more processor cores 1602A-1602N and graphics processor 1608 uses embedded memory module 1618 as a shared last-level cache.
[0129] In at least one embodiment, one or more processor cores 1602A-1602N are homogeneous cores executing a common instruction set architecture. In at least one embodiment, one or more processor cores 1602A-1602N are heterogeneous in terms of instruction set architecture (ISA), with one or more processor cores 1602A-1602N executing a common instruction set, while one or more other cores of one or more processor cores 1602A-1602N execute a subset of the common instruction set or a different instruction set. In at least one embodiment, in terms of microarchitecture, one or more processor cores 1602A-1602N are heterogeneous, with one or more cores having relatively high power consumption coupled to one or more power cores having lower power consumption. In at least one embodiment, the processor 1600 may be implemented on one or more chips or implemented as a SoC integrated circuit.
[0130] Such components can be used to generate alternative view images, such as bird's-eye view images, from parallax data using hardware with limited capacity (such as embedded processors that do not access external memory).
[0131] Autonomous vehicles
[0132] Figure 17A An example of an autonomous vehicle 1700 according to at least one embodiment is shown. In at least one embodiment, the autonomous vehicle 1700 (which may alternatively be referred to herein as "vehicle 1700") may be, but is not limited to, a passenger vehicle, such as a car, truck, bus, and / or another type of vehicle capable of accommodating one or more passengers. In at least one embodiment, vehicle 1700 may be a semi-tractor-trailer for hauling cargo. In at least one embodiment, vehicle 1700 may be an aircraft, a robotic vehicle, or other type of vehicle.
[0133] Autonomous vehicles can be described according to the levels of automation defined by the National Highway Traffic Safety Administration (“NHTSA”) and the Society of Automotive Engineers (“SAE”) of the U.S. Department of Transportation in their standard “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles” (e.g., standard number J3016-201806, published June 15, 2018; standard number J3016-201609, published September 30, 2016; and previous and future versions of this standard). In at least one embodiment, vehicle 1700 may be able to function according to one or more of the levels of automation, from Level 1 to Level 5. For example, in at least one embodiment, vehicle 1700 may be able to perform conditional automation (Level 3), high automation (Level 4), and / or full automation (Level 5).
[0134] In at least one embodiment, vehicle 1700 may include, but is not limited to, components such as chassis, body, wheels (e.g., 2, 4, 6, 8, 18, etc.), tires, axles, and other vehicle components. In at least one embodiment, vehicle 1700 may include, but is not limited to, propulsion system 1750, such as an internal combustion engine, a hybrid powertrain, an all-electric motor, and / or another type of propulsion system. In at least one embodiment, propulsion system 1750 may be connected to the drivetrain of vehicle 1700, which may include, but is not limited to, a transmission, to enable propulsion of vehicle 1700. In at least one embodiment, propulsion system 1750 may be controlled in response to receiving a signal from throttle / accelerator 1752.
[0135] In at least one embodiment, when the propulsion system 1750 is operating (e.g., when the vehicle is in motion), the steering system 1754 (which may include, but is not limited to, a steering wheel) is used to steer the vehicle 1700 (e.g., along a desired path or route). In at least one embodiment, the steering system 1754 may receive signals from the steering actuator 1756. In at least one embodiment, the steering wheel may be optional for fully automated (Level 5) functionality. In at least one embodiment, the brake sensor system 1746 may be used to operate the vehicle brakes in response to signals received from the brake actuator 1748 and / or brake sensors.
[0136] In at least one embodiment, the controller 1736 may include, but is not limited to, one or more system-on-chips (“SoCs”). Figure 17AA controller 1736 (not shown) and / or a graphics processing unit (“GPU”) provides signals (e.g., representing commands) to one or more components and / or systems of vehicle 1700. For example, in at least one embodiment, controller 1736 may send signals to operate vehicle braking via brake actuator 1748, to operate steering system 1754 via one or more steering actuators 1756, and to operate propulsion system 1750 via one or more throttles / accelerators 1752. In at least one embodiment, one or more controllers 1736 may include one or more onboard (e.g., integrated) computing devices that process sensor signals and output operating commands (e.g., signals representing commands) to enable autonomous driving and / or assist a driver in driving vehicle 1700. In at least one embodiment, one or more controllers 1736 may include a first controller 1736 for autonomous driving functions, a second controller 1736 for functional safety functions, a third controller 1736 for artificial intelligence functions (e.g., computer vision), a fourth controller 1736 for infotainment functions, a fifth controller 1736 for redundancy in emergency situations, and / or other controllers. In at least one embodiment, a single controller 1736 may handle two or more of the functions described above, and two or more controllers 1736 may handle a single function and / or any combination thereof.
[0137] In at least one embodiment, one or more controllers 1736 provide signals for controlling one or more components and / or systems of vehicle 1700 in response to sensor data received from one or more sensors (e.g., sensor inputs). In at least one embodiment, sensor data can be received from sensors, including but not limited to one or more Global Navigation Satellite System (“GNSS”) sensors 1758 (e.g., one or more Global Positioning System sensors), one or more RADAR sensors 1760, one or more ultrasonic sensors 1762, one or more LIDAR sensors 1764, one or more inertial measurement unit (IMU) sensors 1766 (e.g., one or more accelerometers, one or more gyroscopes, one or more magnetic compasses, one or more magnetometers, etc.), one or more microphones 1796, one or more stereo cameras 1768, one or more wide-angle cameras 1770 (e.g., fisheye cameras), one or more infrared cameras 1772, one or more surround cameras 1774 (e.g., 360-degree cameras), and remote cameras (…). Figure 17A (not shown in the image), medium-range camera ( Figure 17A(Not shown in the image) One or more speed sensors 1744 (e.g., for measuring the speed of vehicle 1500), one or more vibration sensors 1742, one or more steering sensors 1740, one or more brake sensors (e.g., as part of brake sensor system 1746) and / or other sensor types are received.
[0138] In at least one embodiment, one or more controllers 1736 may receive input (e.g., represented by input data) from the instrument panel 1732 of the vehicle 1700 and provide output (e.g., represented by output data, display data, etc.) via a human-machine interface (“HMI”) display 1734, a voice signaler, a speaker, and / or other components of the vehicle 1700. In at least one embodiment, the output may include information such as vehicle speed, velocity, time, map data (e.g., high-definition map). Figure 17A The HMI display 1734 may display information such as (not shown in the image), location data (e.g., the location of vehicle 1700, for example, on a map), direction, the location of other vehicles (e.g., occupancy raster), information about objects, and the state of objects perceived by one or more controllers 1736. For example, in at least one embodiment, the HMI display 1734 may display information about the presence of one or more objects (e.g., road signs, warning signs, traffic light changes, etc.) and / or information about driving operations that have been, are being, or will be made (e.g., changing lanes now, exiting exit 34B within two miles, etc.).
[0139] In at least one embodiment, vehicle 1700 further includes a network interface 1724 that can communicate over one or more networks using one or more wireless antennas 1726 and / or one or more modems. For example, in at least one embodiment, network interface 1724 may be able to communicate over Long Term Evolution (“LTE”), Wideband Code Division Multiple Access (“WCDMA”), Universal Mobile Telecommunications System (“UMTS”), Global System for Mobile Communications (“GSM”), IMT-CDMA Multicarrier (“CDMA2000”) networks, etc. In at least one embodiment, one or more wireless antennas 1726 may also enable communication between objects in the environment (e.g., vehicles, mobile devices) using one or more local area networks (e.g., Bluetooth, Bluetooth Low Energy (LE), Z-Wave, ZigBee, etc.) and / or one or more low-power wide area networks (hereinafter “LPWAN”) (e.g., LoRaWAN, SigFox, etc. protocols).
[0140] Such components can be used to generate alternative view images, such as bird's-eye view images, from parallax data using hardware with limited capacity (such as embedded processors that do not access external memory).
[0141] Figure 17B The illustration shows an embodiment according to at least one of the embodiments. Figure 17A Examples of camera positions and fields of view for the autonomous vehicle 1700. In at least one embodiment, the camera and its respective field of view are an example embodiment and are not intended to be limiting. For example, in at least one embodiment, additional and / or alternative cameras may be included and / or the cameras may be located at different positions on the vehicle 1700.
[0142] In at least one embodiment, the camera type used for the camera may include, but is not limited to, a digital camera suitable for use with components and / or systems of vehicle 1700. In at least one embodiment, the camera may operate at Automotive Safety Integrity Level (“ASIL”) B and / or other ASILs. In at least one embodiment, the camera type may have any image capture rate, such as 60 frames per second (fps), 1220 fps, 240 fps, etc. In at least one embodiment, the camera may be able to use a rolling shutter, a global shutter, another type of shutter, or a combination thereof. In at least one embodiment, the color filter array may include a red-to-clear (“RCCC”) color filter array, a red-to-clear-blue (“RCCB”) color filter array, a red-blue-green (“RBGC”) color filter array, a Foveon X3 color filter array, a Bayer sensor (“RGGB”) color filter array, a monochrome sensor color filter array, and / or other types of color filter arrays. In at least one embodiment, a transparent pixel camera, such as a camera with RCCC, RCCB, and / or RBGC color filter arrays, may be used to improve photosensitivity.
[0143] In at least one embodiment, one or more cameras may be used to perform advanced driver assistance system (“ADAS”) functions (e.g., as part of a redundancy or fail-safe design). For example, in at least one embodiment, a multi-function mono camera may be installed to provide functions including lane departure warning, traffic sign assist, and intelligent headlight control. In at least one embodiment, one or more cameras (e.g., all cameras) may simultaneously record and provide image data (e.g., video).
[0144] In at least one embodiment, one or more cameras may be mounted in a mounting assembly, such as a custom-designed (3D-printed) assembly, to cut out stray light and reflections within the vehicle 1700 (e.g., reflections from the dashboard in the windshield mirror), which may interfere with the camera's image data capture capabilities. Regarding the rearview mirror mounting assembly, in at least one embodiment, the rearview mirror assembly may be 3D-printed custom-made such that the camera mounting plate matches the shape of the rearview mirror. In at least one embodiment, one or more cameras may be integrated into the rearview mirror.
[0145] In at least one embodiment, for side-view cameras, one or more cameras may also be integrated into four pillars in each corner of the cabin.
[0146] In at least one embodiment, a camera (e.g., a forward-facing camera) having a field of view including a portion of the environment in front of the vehicle 1700 can be used for surround view and, with the assistance of one or more controllers 1736 and / or control SoCs, to help identify the forward path and obstacles, thereby providing information crucial for generating an occupancy grid and / or determining a preferred vehicle path. In at least one embodiment, the forward-facing camera can be used to perform many ADAS functions similar to LIDAR, including but not limited to emergency braking, pedestrian detection, and collision avoidance. In at least one embodiment, the forward-facing camera can also be used for ADAS functions and systems, including but not limited to lane departure warning (“LDW”), adaptive cruise control (“ACC”), and / or other functions (e.g., traffic sign recognition).
[0147] In at least one embodiment, various cameras can be used in a forward-facing configuration, including, for example, a monocular camera platform including a CMOS (“complementary metal-oxide-semiconductor”) color imager. In at least one embodiment, a wide-angle camera 1770 can be used to sense objects entering from the periphery (e.g., pedestrians, people crossing the street, or bicycles). Although in Figure 17B Only one wide-angle camera 1770 is shown; however, in other embodiments, the vehicle 1700 may have any number (including zero) of wide-angle cameras. In at least one embodiment, any number of remote cameras 1798 (e.g., a pair of remote stereo cameras) can be used for depth-based object detection, especially for objects for which a neural network has not yet been trained. In at least one embodiment, the remote camera 1798 can also be used for object detection and classification, as well as basic object tracking.
[0148] In at least one embodiment, any number of stereo cameras 1768 may also be included in a forward configuration. In at least one embodiment, one or more stereo cameras 1768 may include an integrated control unit comprising a scalable processing unit that may provide programmable logic (“FPGA”) and a multi-core microprocessor with a controller area network (“CAN”) or Ethernet interface integrated on a single chip. In at least one embodiment, such a unit may be used to generate a 3D map of the environment of the vehicle 1700, including distance estimates for all points in the image. In at least one embodiment, one or more stereo cameras 1768 may include, but are not limited to, a compact stereo vision sensor, which may include, but is not limited to, two camera lenses (one on the left and one on the right) and an image processing chip that can measure the distance from the vehicle 1700 to a target object and use the generated information (e.g., metadata) to activate autonomous emergency braking and lane departure warning functions. In at least one embodiment, other types of stereo cameras 1768 may also be used in addition to those described herein.
[0149] In at least one embodiment, a camera (e.g., a side-view camera) having a field of view including a portion of the environment on the side of the vehicle 1700 can be used for surround viewing, thereby providing information for creating and updating the occupied grid, and generating a side collision warning. For example, in at least one embodiment, one or more surround cameras 1774 (e.g., such as...) Figure 17B The four surround cameras 1774 shown can be positioned on the vehicle 1700. In at least one embodiment, one or more surround cameras 1774 can be, but are not limited to, any number and combination of one or more wide-angle cameras, one or more fisheye lenses, one or more 360-degree cameras, and / or similar cameras. For example, in at least one embodiment, four fisheye lens cameras can be located at the front, rear, and sides of the vehicle 1700. In at least one embodiment, the vehicle 1700 can use three surround cameras 1774 (e.g., left, right, and rear) and can utilize one or more other cameras (e.g., forward-facing cameras) as a fourth surround-view camera.
[0150] In at least one embodiment, a camera (e.g., a rear-view camera) having a field of view including a portion of the environment behind the vehicle 1700 can be used for parking assistance, surround view, rear collision warning, and creating and updating occupancy raster. In at least one embodiment, a wide variety of cameras can be used, including but not limited to cameras that are also suitable as one or more forward-facing cameras (e.g., long-range camera 1798 and / or one or more mid-range cameras 1776, one or more stereo cameras 1768, one or more infrared cameras 1772, etc.), as described herein.
[0151] Such components can be used to generate alternative view images, such as bird's-eye view images, from parallax data using hardware with limited capacity (such as embedded processors that do not access external memory).
[0152] Figure 17C It is shown that, according to at least one embodiment, it is used for Figure 17A A block diagram of an example system architecture for an autonomous vehicle 1700. In at least one embodiment, Figure 17C Each component, feature, and system of vehicle 1700 is connected via bus 1702. In at least one embodiment, bus 1702 may include, but is not limited to, a CAN data interface (also referred to herein as a “CAN bus”). In at least one embodiment, CAN may be a network within vehicle 1700 used to help control various features and functions of vehicle 1700, such as brake actuation, acceleration, braking, steering, windshield wipers, etc. In at least one embodiment, bus 1702 may be configured to have dozens or even hundreds of nodes, each node having its own unique identifier (e.g., a CAN ID). In at least one embodiment, bus 1702 can be read to locate steering wheel angle, ground speed, engine revolutions per minute (“RPM”), button positions, and / or other vehicle status indicators. In at least one embodiment, bus 1702 may be an ASIL B compliant CAN bus.
[0153] In at least one embodiment, FlexRay and / or Ethernet protocols may be used in addition to or as an alternative to CAN. In at least one embodiment, any number of buses forming bus 1702 may be present, including but not limited to zero or more CAN buses, zero or more FlexRay buses, zero or more Ethernet buses, and / or zero or more other types of buses using different protocols. In at least one embodiment, two or more buses may be used to perform different functions and / or for redundancy. For example, a first bus may be used for a collision avoidance function, while a second bus may be used for actuation control. In at least one embodiment, each bus in bus 1702 may communicate with any component of vehicle 1700, and two or more buses in bus 1702 may communicate with corresponding components. In at least one embodiment, any number of System-on-Chip (“SoC”) 1704 (such as SoC 1704(A) and SoC 1704(B)), each controller 1736, and / or each computer within the vehicle may access the same input data (e.g., input from sensors of vehicle 1700) and may be connected to a common bus, such as a CAN bus.
[0154] In at least one embodiment, vehicle 1700 may include one or more controllers 1736, such as those described herein. Figure 17A As described above. In at least one embodiment, controller 1736 can be used for a variety of functions. In at least one embodiment, controller 1736 can be coupled to any of various other components and systems of vehicle 1700 and can be used for control of vehicle 1700, artificial intelligence of vehicle 1700, infotainment and / or other functions of vehicle 1700.
[0155] In at least one embodiment, vehicle 1700 may include any number of SoCs 1704. In at least one embodiment, each SoC 1704 may include, but is not limited to, a central processing unit (“CPU”) 1706, a graphics processing unit (“GPU”) 1708, a processor 1710, a cache 1712, an accelerator 1714, a data storage 1716, and / or other components and features not shown. In at least one embodiment, the SoC 1704 can be used to control vehicle 1700 on various platforms and systems. For example, in at least one embodiment, the SoC 1704 may interface with one or more servers via a network (...). Figure 17C (Not shown in the image) A high-definition (“HD”) map 1722 is obtained by refreshing and / or updating the map and combined in the system (e.g., the system of vehicle 1700).
[0156] In at least one embodiment, the CPU 1706 may include a CPU cluster or CPU complex (or "CCPLEX" herein). In at least one embodiment, the CPU 1706 may include multiple cores and / or a secondary ("L2") cache. For example, in at least one embodiment, the CPU 1706 may include eight cores in a consistent multiprocessor configuration. In at least one embodiment, the CPU 1706 may include four dual-core clusters, each cluster having a dedicated L2 cache (e.g., a 2-megabyte (MB) L2 cache). In at least one embodiment, the CPU 1706 (e.g., CCPLEX) may be configured to support simultaneous cluster operation, such that any combination of clusters of the CPU 1706 can be active at any given time.
[0157] In at least one embodiment, one or more CPU 1706s may implement power management capabilities, including but not limited to one or more of the following features: automatic clock gating of individual hardware blocks to conserve dynamic power when idle; each core clock may be gated when such cores are not actively executing instructions due to executing Wait for Interrupt (“WFI”) / Wait for Event (“WFE”) instructions; each core may be independently power-gated; each core cluster may be independently clock-gated when all cores are clock-gated or power-gated; and / or each core cluster may be independently power-gated when all cores are power-gated. In at least one embodiment, the CPU 1706 may further implement an enhanced algorithm for managing power states, wherein allowed power states and expected wake-up times are specified, and the hardware / microcode determines which optimal power state a core, cluster, and CCPLEX should enter. In at least one embodiment, the processing core may support a simplified power state entry sequence in software, wherein work is offloaded to the microcode.
[0158] In at least one embodiment, GPU 1708 may include an integrated GPU (also referred to herein as an "iGPU"). In at least one embodiment, GPU 1708 may be programmable and efficient for parallel workloads. In at least one embodiment, GPU 1708 may use an enhanced tensor instruction set. In at least one embodiment, GPU 1708 may include one or more streaming microprocessors, wherein each streaming microprocessor may include a Level 1 ("L1") cache (e.g., an L1 cache with at least 96KB of storage capacity), and two or more streaming microprocessors may share an L2 cache (e.g., an L2 cache with 512KB of storage capacity). In at least one embodiment, GPU 1708 may include at least eight streaming microprocessors. In at least one embodiment, GPU 1708 may use a computation application programming interface (API). In at least one embodiment, GPU 1708 may use one or more parallel computing platforms and / or programming models (e.g., NVIDIA's CUDA model).
[0159] In at least one embodiment, one or more of the GPU 1708 can be power-optimized for optimal performance in automotive and embedded use cases. For example, in at least one embodiment, the GPU 1708 can be fabricated on FinFET (“FinFET”) circuitry. In at least one embodiment, each streaming microprocessor can incorporate multiple mixed-precision processing cores, which are divided into multiple, for example, but not limited to, 64 PF32 cores and 32 PF64 cores, which can be divided into four processing blocks. In at least one embodiment, each processing block can be allocated 16 FP32 cores, 8 FP64 cores, 16 INT32 cores, two mixed-precision NVIDIA Tensor cores for deep learning matrix arithmetic, a level-zero (“L0”) instruction cache, a scheduler (e.g., a warp scheduler) or sequencer, dispatch units, and / or a 64KB register file. In at least one embodiment, the streaming microprocessor can include independent parallel integer and floating-point data paths to provide efficient execution of workloads through a mixture of computation and addressing computation. In at least one embodiment, the streaming microprocessor may include independent thread scheduling capabilities to enable finer-grained synchronization and cooperation between parallel threads. In at least one embodiment, the streaming microprocessor may include a combined L1 data cache and shared memory unit to improve performance while simplifying programming.
[0160] In at least one embodiment, one or more GPUs 1708 may include high-bandwidth memory (HBM) and / or a 16GB HBM2 memory subsystem to provide a peak memory bandwidth of approximately 900GB / s in some examples. In at least one embodiment, in addition to or as an alternative to HBM memory, synchronous graphics random access memory (“SGRAM”), such as graphics double data rate type 5 synchronous random access memory (“GDDR5”), may be used.
[0161] In at least one embodiment, GPU 1708 may include unified memory technology. In at least one embodiment, an address translation service (“ATS”) is supported to allow GPU 1708 to directly access CPU 1706 page tables. In at least one embodiment, when the GPU experiences a miss at GPU 1708 Memory Management Unit (“MMU”), an address translation request can be sent to CPU 1706. In response, in at least one embodiment, CPU 1706 can look up the virtual-to-physical mapping for the address in the page table and translate it back to GPU 1708. In at least one embodiment, unified memory technology can allow a single unified virtual address space to be used for both CPU 1706 and GPU 1708 memory, thereby simplifying GPU 1708 programming and porting applications to GPU 1708.
[0162] In at least one embodiment, the GPU 1708 may include any number of access counters that track the frequency with which the GPU 1708 accesses the memory of other processors. In at least one embodiment, the access counters can help ensure that memory pages are moved to the physical memory of the processor that accesses the pages most frequently, thereby improving the efficiency of shared memory ranges between processors.
[0163] In at least one embodiment, one or more SoCs 1704 may include any number of caches 1712, including those described herein. For example, in at least one embodiment, cache 1712 may include a Level 3 (“L3”) cache available to both CPU 1706 and GPU 1708 (e.g., it is connected to both CPU 1706 and GPU 1708). In at least one embodiment, cache 1712 may include a write-back cache, which may track the state of the line, for example, by using a cache coherence protocol (e.g., MEI, MESI, MSI, etc.). In at least one embodiment, the L3 cache may include 4 MB of memory or more, depending on the embodiment, although a smaller cache size may be used.
[0164] In at least one embodiment, one or more SoCs 1704 may include one or more accelerators 1714 (e.g., hardware accelerators, software accelerators, or a combination thereof). In at least one embodiment, the SoC 1704 may include a hardware acceleration cluster that may include optimized hardware accelerators and / or large on-chip memory. In at least one embodiment, large on-chip memory (e.g., 4MB SRAM) may enable the hardware acceleration cluster to accelerate neural networks and other computations. In at least one embodiment, the hardware acceleration cluster may be used to supplement the GPU 1708 and offload some tasks from the GPU 1708 (e.g., freeing up more cycles of the GPU 1708 to perform other tasks). In at least one embodiment, the accelerator 1714 may be used for a target workload (e.g., perceptual, convolutional neural network (“CNN”), recurrent neural network (“RNN”), etc.) that is stable enough to be changed to acceleration. In at least one embodiment, the CNN may include region-based or region convolutional neural networks (“RCNN”) and fast RCNN (e.g., for object detection) or other types of CNNs.
[0165] In at least one embodiment, accelerator 1714 (e.g., a hardware acceleration cluster) may include one or more deep learning accelerators (“DLAs”). In at least one embodiment, the DLA may include, but is not limited to, one or more tensor processing units (“TPUs”) configured to provide an additional 10 trillion operations per second for deep learning applications and inference. In at least one embodiment, the TPU may be an accelerator configured to perform and optimize image processing functions (e.g., for CNNs, RCNNs, etc.). In at least one embodiment, the DLA may be further optimized for a specific set of neural network types and floating-point operations and inference. In at least one embodiment, the DLA is designed to provide higher performance per millimeter than a typical general-purpose GPU and typically significantly outperforms the performance of a CPU. In at least one embodiment, the TPU may perform a variety of functions, including single-instance convolution functions, such as supporting INT8, INT16, and FP16 data types for features and weights, and post-processor functions. In at least one embodiment, the DLA can execute neural networks, particularly CNNs, quickly and efficiently on processed or unprocessed data for any of a variety of functions, including but not limited to: CNNs for object recognition and detection using data from camera sensors; CNNs for distance estimation using data from camera sensors; CNNs for emergency vehicle detection and recognition using data from microphones; CNNs for facial recognition and vehicle owner recognition using data from camera sensors; and / or CNNs for safety and / or safety-related events.
[0166] In at least one embodiment, the DLA can perform any function of the GPU 1708, and by using an inference accelerator, for example, the designer can target either the DLA or the GPU 1708 for any function. For example, in at least one embodiment, the designer can centralize the CNN processing and floating-point operations on the DLA, leaving other functions to the GPU 1708 and / or accelerator 1714.
[0167] In at least one embodiment, accelerator 1714 may include a programmable vision accelerator (“PVA”), which may alternatively be referred to herein as a computer vision accelerator. In at least one embodiment, the PVA may be designed and configured to accelerate computer vision algorithms for advanced driver assistance systems (“ADAS”) 1738, autonomous driving, augmented reality (“AR”) applications, and / or virtual reality (“VR”) applications. In at least one embodiment, the PVA may provide a balance between performance and flexibility. For example, in at least one embodiment, each PVA may include, for example, but not limited to, any number of Reduced Instruction Set Computer (“RISC”) cores, Direct Memory Access (“DMA”), and / or any number of vector processors.
[0168] In at least one embodiment, the RISC core can interact with an image sensor (e.g., the image sensor of any camera described herein), an image signal processor, etc. In at least one embodiment, each RISC core may include any number of memories. In at least one embodiment, the RISC core can use any of a variety of protocols, depending on the embodiment. In at least one embodiment, the RISC core can execute a real-time operating system (“RTOS”). In at least one embodiment, the RISC core can be implemented using one or more integrated circuit devices, application-specific integrated circuits (“ASICs”), and / or memory devices. For example, in at least one embodiment, the RISC core may include an instruction cache and / or tightly coupled RAM.
[0169] In at least one embodiment, DMA enables PVA components to access system memory independently of the CPU 1706. In at least one embodiment, DMA can support any number of features for providing optimization to the PVA, including, but not limited to, support for multidimensional addressing and / or circular addressing. In at least one embodiment, DMA can support up to six or more addressing dimensions, which may include, but are not limited to, block width, block height, block depth, horizontal block step, vertical block step, and / or depth step.
[0170] In at least one embodiment, the vector processor may be a programmable processor designed to efficiently and flexibly execute the programming of computer vision algorithms and provide signal processing capabilities. In at least one embodiment, the PVA may include a PVA core and two vector processing subsystem partitions. In at least one embodiment, the PVA core may include a processor subsystem, a DMA engine (e.g., two DMA engines), and / or other peripheral devices. In at least one embodiment, the vector processing subsystem may operate as the main processing engine of the PVA and may include a vector processing unit (“VPU”), an instruction cache, and / or a vector memory (e.g., “VMEM”). In at least one embodiment, the VPU core may include a digital signal processor, such as a single-instruction, multiple-data (“SIMD”), very long instruction word (“VLIW”) digital signal processor. In at least one embodiment, the combination of SIMD and VLIW can improve throughput and speed.
[0171] In at least one embodiment, each vector processor may include an instruction cache and may be coupled to dedicated memory. As a result, in at least one embodiment, each vector processor may be configured to execute independently of other vector processors. In at least one embodiment, vector processors included in a particular PVA may be configured to employ data parallelism. For example, in at least one embodiment, multiple vector processors included in a single PVA may execute a common computer vision algorithm, but on different regions of an image. In at least one embodiment, vector processors included in a particular PVA may simultaneously execute different computer vision algorithms on an image, or even execute different algorithms on consecutive images or portions of an image. In at least one embodiment, among others, any number of PVAs may include a hardware-accelerated cluster, and any number of vector processors may include in each PVA. In at least one embodiment, the PVA may include additional error-correcting code (“ECC”) memory to enhance the security of the entire system.
[0172] In at least one embodiment, accelerator 1714 may include an on-chip computer vision network and static random access memory (“SRAM”) to provide high-bandwidth, low-latency SRAM for accelerator 1714. In at least one embodiment, the on-chip memory may include at least 4 MB of SRAM, including, for example, but not limited to, eight field-configurable memory blocks accessible by both the PVA and DLA. In at least one embodiment, each pair of memory blocks may include an advanced peripheral bus (“APB”) interface, configuration circuitry, a controller, and a multiplexer. In at least one embodiment, any type of memory may be used. In at least one embodiment, the PVA and DLA may access the memory via a backbone that provides high-speed access to the memory for the PVA and DLA. In at least one embodiment, the backbone may include an on-chip computer vision network that interconnects the PVA and DLA to the memory (e.g., using an APB).
[0173] In at least one embodiment, the on-chip computer vision network may include an interface that determines whether both the PVA and DLA provide ready and valid signals before transmitting any control signals / addresses / data. In at least one embodiment, the interface may provide separate stages and separate channels for transmitting control signals / addresses / data, as well as bursty communication for continuous data transmission. In at least one embodiment, the interface may conform to the International Organization for Standardization (“ISO”) 26262 or the International Electrotechnical Commission (“IEC”) 61508 standard, although other standards and protocols may be used.
[0174] In at least one embodiment, one or more SoCs 1704 may include a real-time ray tracing hardware accelerator. In at least one embodiment, the real-time ray tracing hardware accelerator may be used to rapidly and efficiently determine the location and extent of an object (e.g., within a world model) to generate real-time visualization simulations, for radar signal interpretation, for sound propagation synthesis and / or analysis, for simulation of SONAR systems, for general wave propagation simulation, for comparison with LiDAR data for localization and / or other functions, and / or for other uses.
[0175] In at least one embodiment, accelerator 1714 can have a wide range of applications in autonomous driving. In at least one embodiment, PVA can be used in critical processing stages in ADAS and autonomous vehicles. In at least one embodiment, the capabilities of PVA are well-suited for algorithmic domains requiring predictable processing, low power consumption, and low latency. In other words, PVA performs well on semi-intensive or intensive conventional computations, even on small datasets that may require predictable runtimes with low latency and low power consumption. In at least one embodiment, in an autonomous vehicle, such as vehicle 1700, PVA may be designed to run classical computer vision algorithms, as they would be efficient in object detection and integer mathematical operations.
[0176] For example, according to at least one embodiment of the technology, PVA is used to perform computer stereo vision. In at least one embodiment, a semi-global matching-based algorithm may be used in some examples, although this is not intended to be limiting. In at least one embodiment, applications for Level 3-5 autonomous driving use dynamic motion estimation / stereo matching (e.g., from moving structures, pedestrian recognition, lane detection, etc.). In at least one embodiment, PVA can perform computer stereo vision functions on input from two monocular cameras.
[0177] In at least one embodiment, the PVA can be used to perform dense optical flow. For example, in at least one embodiment, the PVA can process raw RADAR data (e.g., using 4D Fast Fourier Transform) to provide processed RADAR data. In at least one embodiment, the PVA is used for time-of-flight depth processing, for example, by processing raw time-of-flight data to provide processed time-of-flight data.
[0178] In at least one embodiment, the DLA can be used to run any type of network to enhance control and driving safety, including, for example, but not limited to, neural networks that output a confidence measurement for each object detection. In at least one embodiment, the confidence level can be represented or interpreted as a probability, or provide a relative “weight” for each detection compared to other detections. In at least one embodiment, the confidence metric enables the system to make further determinations about which detections should be considered true positives rather than false positives. In at least one embodiment, the system can set a threshold for the confidence level and only consider detections exceeding the threshold as true positives. In embodiments using an Automatic Emergency Braking (“AEB”) system, false positives would cause the vehicle to automatically perform emergency braking, which is obviously undesirable. In at least one embodiment, highly confident detections can be considered triggers for AEB. In at least one embodiment, the DLA can run a neural network for regressing confidence values. In at least one embodiment, the neural network may take at least some subset of parameters as its input, such as bounding box size, obtained ground plane estimate (e.g. from another subsystem), and output from MU sensor 1766, which are related to the vehicle 1700 orientation, distance, and 3D position estimate of the object obtained from the neural network and / or other sensors (e.g., LIDAR sensor 1764 or RADAR sensor 1760).
[0179] In at least one embodiment, one or more SoC 1704s may include one or more data storage 1716s (e.g., memory). In at least one embodiment, the data storage 1716 may be on-chip memory of the SoC 1704, which may store neural networks to be executed on the GPU 1708 and / or DLA. In at least one embodiment, the capacity of the data storage 1716 may be large enough to store multiple instances of the neural network for redundancy and security. In at least one embodiment, the data storage 1716 may include an L2 or L3 cache.
[0180] In at least one embodiment, one or more SoCs 1704 may include any number of processors 1710 (e.g., embedded processors). In at least one embodiment, processor 1710 may include a boot and power management processor, which may be a dedicated processor and subsystem to handle boot power and management functions and associated security implementations. In at least one embodiment, the boot and power management processor may be part of the boot sequence of the SoC 1704 and may provide runtime power management services. In at least one embodiment, the boot power and management processor may provide clock and voltage programming, assist with system low-power state transitions, manage SoC 1704 thermal and temperature sensors, and / or manage SoC 1704 power states. In at least one embodiment, each temperature sensor may be implemented as a ring oscillator with an output frequency proportional to temperature, and the SoC 1704 may use the ring oscillator to detect the temperature of the CPU 1706, GPU 1708, and / or accelerator 1714. In at least one embodiment, if it is determined that the temperature exceeds a threshold, the startup and power management processor may enter a temperature fault routine and place the SoC 1704 into a lower power state and / or place the vehicle 1700 into a driver safety stop mode (e.g., bring the vehicle 1700 to a safe stop).
[0181] In at least one embodiment, the processor 1710 may further include a set of embedded processors that can be used as an audio processing engine. This can be an audio subsystem capable of providing full hardware support for multi-channel audio through multiple interfaces, as well as a wide and flexible range of audio I / O interfaces. In at least one embodiment, the audio processing engine is a dedicated processor core with a digital signal processor and dedicated RAM.
[0182] In at least one embodiment, the processor 1710 may further include an always-on processor engine that can provide the necessary hardware features to support low-power sensor management and wake-up use cases. In at least one embodiment, the always-on processor engine may include, but is not limited to, a processor core, tightly coupled RAM, support for peripherals (e.g., timers and interrupt controllers), various I / O controller peripherals, and routing logic.
[0183] In at least one embodiment, processor 1710 may further include a secure clustering engine, which includes, but is not limited to, a dedicated processor subsystem for handling automotive application security management. In at least one embodiment, the secure clustering engine may include, but is not limited to, two or more processor cores, tightly coupled RAM, support for peripheral devices (e.g., timers, interrupt controllers, etc.), and / or routing logic. In a secure mode, in at least one embodiment, two or more cores may operate in lockstep mode and function as a single core with comparison logic to detect any differences between their operations. In at least one embodiment, processor 1710 may further include a real-time camera engine, which may include, but is not limited to, a dedicated processor subsystem for handling real-time camera management. In at least one embodiment, processor 1710 may further include a high dynamic range signal processor, which may include, but is not limited to, an image signal processor, which is a hardware engine as part of the camera processing pipeline.
[0184] In at least one embodiment, processor 1710 may include a video image synthesizer, which may be a processing block (e.g., implemented on a microprocessor) that implements video post-processing functions required for video playback applications to produce an image of the final player window. In at least one embodiment, the video image synthesizer may perform lens distortion correction on the wide-angle camera 1770, the surround camera 1774, and / or the in-vehicle monitoring camera sensor. In at least one embodiment, the in-vehicle monitoring camera sensor is preferably monitored by a neural network running on another instance of SoC 1704, configured to identify in-vehicle events and respond accordingly. In at least one embodiment, the in-vehicle system may perform, but is not limited to, lip reading to activate cellular service and make phone calls, dictate emails, change the vehicle's destination, activate or change the vehicle's infotainment system and settings, or provide voice-activated web browsing. In at least one embodiment, certain functions are available to the driver when the vehicle is operating in autonomous mode, and are otherwise disabled.
[0185] In at least one embodiment, the video image synthesizer may include enhanced temporal denoising for both spatial and temporal noise reduction. For example, in at least one embodiment where motion occurs in the video, the denoising appropriately weights spatial information, reducing the weight of information provided by adjacent frames. In at least one embodiment, where the image or a portion of the image does not contain motion, the temporal denoising performed by the video image synthesizer may use information from previous images to reduce noise in the current image.
[0186] In at least one embodiment, the video image compositor can also be configured to perform stereoscopic correction on the input stereoscopic shot frames. In at least one embodiment, the video image compositor can also be used for user interface compositing when the operating system desktop is in use, and the GPU 1708 is not required to continuously render new surfaces. In at least one embodiment, when the GPU 1708 is powered on and actively performing 3D rendering, the video image compositor can be used to offload the GPU 1708 to improve performance and responsiveness.
[0187] In at least one embodiment, one or more SoCs 1704 may further include a Mobile Industrial Processor Interface (“MIPI”) camera serial interface, a high-speed interface, and / or a video input block that can be used for receiving video and input from a camera and associated pixel input functions. In at least one embodiment, one or more SoCs 1704 may further include one or more input / output controllers that may be software-controlled and can be used to receive I / O signals not assigned to a specific role.
[0188] In at least one embodiment, one or more SoCs 1704 may also include a wide range of peripheral interfaces to enable communication with peripheral devices, audio encoders / decoders (“codecs”), power management and / or other devices. In at least one embodiment, SoC 1704 may be used to process data from cameras (e.g., connected via gigabit multimedia serial links and Ethernet channels), sensors (e.g., LiDAR sensor 1764, RADAR sensor 1760, etc., which may be connected via Ethernet channels), from bus 1702 (e.g., speed, steering wheel position, etc. of vehicle 1700), from GNSS sensor 1758 (e.g., connected via Ethernet bus or CAN bus), etc. In at least one embodiment, one or more SoCs 1704 may also include dedicated high-performance, high-capacity memory controllers, which may include their own DMA engines and may be used to free CPU 1706 from routine data management tasks.
[0189] In at least one embodiment, the SoC 1704 can be an end-to-end platform with a flexible architecture spanning automation levels 3-5, providing full and efficient utilization of the diversity and redundancy of a comprehensive functional safety architecture for computing, vision, and ADAS technologies, and providing a platform for a flexible, reliable driving software stack and deep learning tools. In at least one embodiment, the SoC 1704 can be faster, more reliable, and even more energy-efficient and space-saving than traditional systems. For example, in at least one embodiment, the accelerator 1714, when combined with the CPU 1706, GPU 1708, and data storage 1716, can provide a fast and efficient platform for Level 3-5 autonomous vehicles.
[0190] In at least one embodiment, the computer vision algorithms can be executed on a CPU, which can be configured using a high-level programming language such as C to execute multiple processing algorithms across a variety of visual data. However, in at least one embodiment, the CPU often cannot meet the performance requirements of many computer vision applications, such as those related to execution time and power consumption. In at least one embodiment, many CPUs cannot execute complex object detection algorithms in real time for automotive ADAS applications and practical Level 3-5 autonomous vehicles.
[0191] The embodiments described herein allow multiple neural networks to be executed simultaneously and / or sequentially, and allow the results to be combined to achieve Level 3-5 autonomous driving capabilities. For example, in at least one embodiment, the CNN executed on a DLA or a discrete GPU (e.g., GPU 1720) may include text and character recognition, thereby allowing the reading and understanding of traffic signs, including signs for which the neural network has not yet been specifically trained. In at least one embodiment, the DLA may also include a neural network capable of recognizing, interpreting, and providing semantic understanding of symbols and passing that semantic understanding to a path planning module running on a CPU Complex.
[0192] In at least one embodiment, multiple neural networks, such as Level 3, 4, or 5 driving systems, can run simultaneously. For example, in at least one embodiment, a warning sign stating "Warning: Flashing lights indicate icy conditions" and a light can be interpreted independently or jointly by multiple neural networks. In at least one embodiment, such a warning sign itself can be recognized as a traffic sign by a first deployed neural network (e.g., a trained neural network), and the text "Flashing lights indicate icy conditions" can be interpreted by a second deployed neural network, which notifies the vehicle's path planning software (preferably executed on a CPU Complex) that icing is present when the flashing lights are detected. In at least one embodiment, the flashing lights can be identified by running a third deployed neural network over multiple frames, notifying the vehicle's path planning software of the presence (or absence) of the flashing lights. In at least one embodiment, all three neural networks can run simultaneously, for example, within a DLA and / or on a GPU1708.
[0193] In at least one embodiment, the CNN for facial recognition and vehicle owner identification can use data from camera sensors to identify the presence of an authorized driver and / or owner of vehicle 1700. In at least one embodiment, a sensor processing engine is always activated to unlock the vehicle when the owner approaches the driver's door and turns on the lights, and in a secure mode, the vehicle can be disabled when the owner leaves. In this way, SoC 1704 provides security against theft and / or carjacking.
[0194] In at least one embodiment, the CNN for emergency vehicle detection and identification can use data from microphone 1796 to detect and identify emergency vehicle sirens. In at least one embodiment, SoC 1704 uses the CNN to classify environmental and urban sounds, as well as visual data. In at least one embodiment, the CNN running on DLA is trained to identify the relative shut-off speed of emergency vehicles (e.g., by using the Doppler effect). In at least one embodiment, the CNN can also be trained to identify emergency vehicles specific to the local area where the vehicle is operating, such as those identified by GNSS sensor 1758. In at least one embodiment, when operating in Europe, the CNN will seek to detect European sirens, and when in North America, the CNN will seek to identify only North American sirens. In at least one embodiment, once an emergency vehicle is detected, a control program can be used, with the assistance of ultrasonic sensor 1762, to execute emergency vehicle safety routines, slow the vehicle, pull over to the side of the road, stop, and / or idle the vehicle until the emergency vehicle passes.
[0195] In at least one embodiment, vehicle 1700 may include CPU 1718 (e.g., a discrete CPU or dCPU) which may be coupled to SoC 1704 (e.g., PCIe) via a high-speed interconnect. For example, in at least one embodiment, CPU 1718 may include an x86 processor. CPU 1718 may be used to perform any of a variety of functions, including arbitrating potentially inconsistent results between ADAS sensors and SoC 1704, and / or monitoring the status and health of controller 1736 and / or, for example, an on-chip infotainment system (“Infotainment SoC”) 1730. In at least one embodiment, SoC 1704 includes one or more interconnects, and the interconnects may include Peripheral Component Interconnect Fast (PCIe).
[0196] In at least one embodiment, vehicle 1700 may include GPU 1720 (e.g., a discrete GPU or dGPU) which may be coupled to SoC 1704 (e.g., NVIDIA's NVLINK channel) via a high-speed interconnect. In at least one embodiment, GPU 1720 may provide additional artificial intelligence capabilities, such as by executing redundant and / or different neural networks, and may be used to train and / or update neural networks (e.g., sensor data) from sensors of vehicle 1700, at least in part, based on input.
[0197] In at least one embodiment, vehicle 1700 may further include network interface 1724, which may include, but is not limited to, wireless antenna 1726 (e.g., one or more wireless antennas 1726 for different communication protocols, such as cellular antennas, Bluetooth antennas, etc.). In at least one embodiment, network interface 1724 can be used to enable wireless connectivity with Internet cloud services (e.g., with servers and / or other network devices), with other vehicles, and / or with computing devices (e.g., passenger client devices). In at least one embodiment, for communication with other vehicles, a direct link and / or an indirect link (e.g., across networks and via the Internet) may be established between vehicle 1700 and other vehicles. In at least one embodiment, a vehicle-to-vehicle communication link may be used to provide a direct link. In at least one embodiment, an inter-vehicle communication link may provide vehicle 1700 with information about vehicles near vehicle 1700 (e.g., vehicles in front, to the side, and / or behind vehicle 1700). In at least one embodiment, such foregoing functionality may be part of a cooperative adaptive cruise control function of vehicle 1700.
[0198] In at least one embodiment, network interface 1724 may include a SoC that provides modulation and demodulation functions and enables controller 1736 to communicate over a wireless network. In at least one embodiment, network interface 1724 may include a radio frequency (RF) front-end for up-converting from baseband to RF and down-converting from RF to baseband. In at least one embodiment, frequency conversion may be performed in any technically feasible manner. For example, frequency conversion may be performed using well-known processes and / or using superheterodyne processes. In at least one embodiment, the RF front-end functionality may be provided by a separate chip. In at least one embodiment, the network interface may include wireless functions for communicating over LTE, WCDMA, UMTS, GSM, CDMA2000, Bluetooth, Bluetooth LE, Wi-Fi, Z-Wave, ZigBee, LoRaWAN, and / or other wireless protocols.
[0199] In at least one embodiment, vehicle 1700 may further include data storage 1728, which may include, but is not limited to, off-chip (e.g., off-chip SoC 1704) storage. In at least one embodiment, data storage 1728 may include, but is not limited to, one or more storage elements, including RAM, SRAM, dynamic random access memory (“DRAM”), video random access memory (“VRAM”), flash memory, hard disk, and / or other components and / or devices capable of storing at least one bit of data.
[0200] In at least one embodiment, vehicle 1700 may further include GNSS sensors 1758 (e.g., GPS and / or auxiliary GPS sensors) to assist mapping, sensing, occupancy grid generation, and / or path planning functions. In at least one embodiment, any number of GNSS sensors 1758 may be used, including, for example, but not limited to, GPS sensors using a USB connector with an Ethernet-to-serial (e.g., RS-232) bridge.
[0201] In at least one embodiment, vehicle 1700 may further include a RADAR sensor 1760. In at least one embodiment, the RADAR sensor 1760 may be used by vehicle 1700 for remote vehicle detection, even in dark and / or inclement weather conditions. In at least one embodiment, the RADAR functional safety level may be ASIL B. In at least one embodiment, the RADAR sensor 1760 may use a CAN bus and / or bus 1702 (e.g., for transmitting data generated by the RADAR sensor 1760) for control and access to object tracking data, and in some examples may access an Ethernet channel to access raw data. In at least one embodiment, various RADAR sensor types may be used. For example, but not limited to, the RADAR sensor 1760 may be suitable for front, rear, and side RADAR use. In at least one embodiment, one or more of the RADAR sensors 1760 are pulse Doppler RADAR sensors.
[0202] In at least one embodiment, the RADAR sensor 1760 may include different configurations, such as long-range with a narrow field of view, short-range with a wide field of view, short-range side coverage, etc. In at least one embodiment, the long-range radar can be used for adaptive cruise control functions. In at least one embodiment, the long-range radar system can provide a wide field of view, for example, within a range of 250 m, achieved through two or more independent scans. In at least one embodiment, the RADAR sensor 1760 can help distinguish between static and moving objects and can be used by the ADAS system 1738 for emergency braking assistance and forward collision warning. In at least one embodiment, the sensor 1760 included in the long-range radar system may include, but is not limited to, a monostatic multi-mode radar with multiple (e.g., six or more) fixed radar antennas and high-speed CAN and FlexRay interfaces. In at least one embodiment, using six antennas, the central four antennas can create a focused beam pattern designed to record the surrounding environment of the vehicle 1700 at a higher speed while minimizing interference from traffic from adjacent lanes. In at least one embodiment, the other two antennas can expand the field of view, thereby enabling rapid detection of vehicles entering or leaving the vehicle 1700's lane.
[0203] In at least one embodiment, as an example, the mid-range radar system may include a range of up to 160m (front) or 80m (rear), and a range of up to 42 degrees (front) or 150 degrees (rear). In at least one embodiment, the short-range radar system may include, but is not limited to, any number of radar sensors 1760 designed to be mounted at both ends of the rear bumper. When mounted at both ends of the rear bumper, in at least one embodiment, the RADAR sensor system may generate two beams to continuously monitor the blind spots in the rear direction and beside the vehicle. In at least one embodiment, the short-range radar system may be used in the ADAS system 1738 for blind spot detection and / or lane change assistance.
[0204] In at least one embodiment, the vehicle 1700 may further include an ultrasonic sensor 1762. In at least one embodiment, the ultrasonic sensor 1762 may be positioned at the front, rear, and / or side of the vehicle 1700, and may be used for parking assistance and / or creating and updating occupancy grids. In at least one embodiment, multiple ultrasonic sensors 1762 may be used, and different ultrasonic sensors 1762 may be used for different detection ranges (e.g., 2.5m, 4m). In at least one embodiment, the ultrasonic sensor 1762 may operate at the ASIL B functional safety level.
[0205] In at least one embodiment, vehicle 1700 may include a LiDAR sensor 1764. In at least one embodiment, LiDAR sensor 1764 may be used for object and pedestrian detection, emergency braking, collision avoidance, and / or other functions. In at least one embodiment, LiDAR sensor 1764 may operate at functional safety level ASIL B. In at least one embodiment, vehicle 1700 may include multiple LiDAR sensors 1764 (e.g., two, four, six, etc.) that can use an Ethernet channel (e.g., providing data to a Gigabit Ethernet switch).
[0206] In at least one embodiment, the LIDAR sensor 1764 may be able to provide a list of objects and their distances for a 360-degree field of view. In at least one embodiment, for example, the commercially available LIDAR sensor 1764 may have an advertising range of approximately 100m, an accuracy of 2cm-3cm, and support, for example, a 100Mbps Ethernet connection. In at least one embodiment, one or more non-protruding LIDAR sensors may be used. In such an embodiment, the LIDAR sensor 1764 may include a small device that can be embedded in the front, rear, side, and / or corner locations of the vehicle 1700. In at least one embodiment, the LIDAR sensor 1764, in such an embodiment, can provide a horizontal field of view of up to 120 degrees and a vertical field of view of 35 degrees, with a range of 200m even for low-reflectivity objects. In at least one embodiment, the front-facing LIDAR sensor 1764 may be configured for a horizontal field of view between 45 degrees and 135 degrees.
[0207] In at least one embodiment, lidar technology, such as 3D flash lidar, may also be used. In at least one embodiment, the 3D flash lidar uses a laser flash as a transmission source to illuminate the environment surrounding the vehicle 1700, up to approximately 200m away. In at least one embodiment, the flash lidar unit includes, but is not limited to, a receiver that records the laser pulse transmission time and the reflected light on each pixel, which in turn corresponds to the distance from the vehicle 1700 to the object. In at least one embodiment, the flash lidar can allow the generation of highly accurate and distortion-free images of the surrounding environment with each laser flash. In at least one embodiment, four flash lidar sensors may be deployed, one on each side of the vehicle 1700. In at least one embodiment, the 3D flash lidar system includes, but is not limited to, a solid-state 3D staring array lidar camera, which has no moving parts other than a fan (e.g., a non-scanning lidar device). In at least one embodiment, the flash lidar device can use Class I (eye-safe) laser pulses at 5 nanoseconds per frame and can capture reflected laser light as intensity data for 3D distance point clouds and co-registration.
[0208] In at least one embodiment, vehicle 1700 may further include IMU sensor 1766. In at least one embodiment, IMU sensor 1766 may be located at the center of the rear axle of vehicle 1700. In at least one embodiment, IMU sensor 1766 may include, for example, but not limited to, accelerometers, magnetometers, gyroscopes, one or more magnetic compasses, and / or other sensor types. In at least one embodiment, for example in a six-axis application, IMU sensor 1766 may include, but is not limited to, accelerometers and gyroscopes. In at least one embodiment, for example in a nine-axis application, IMU sensor 1766 may include, but is not limited to, accelerometers, gyroscopes, and magnetometers.
[0209] In at least one embodiment, the IMU sensor 1766 can be implemented as a miniature, high-performance GPS-assisted inertial navigation system (“GPS / INS”) that combines a microelectromechanical system (“MEMS”) inertial sensor, a high-sensitivity GPS receiver, and an advanced Kalman filtering algorithm to provide estimates of position, velocity, and attitude. In at least one embodiment, the IMU sensor 1766 enables the vehicle 1700 to estimate its heading by directly observing and correlating velocity changes from GPS to the IMU sensor 1766 without requiring input from a magnetic sensor. In at least one embodiment, the IMU sensor 1766 and the GNSS sensor 1758 can be combined in a single integrated unit.
[0210] In at least one embodiment, vehicle 1700 may include a microphone 1796 placed in and / or around vehicle 1700. In at least one embodiment, microphone 1796 may be used for things such as emergency vehicle detection and identification.
[0211] In at least one embodiment, vehicle 1700 may also include any number of camera types, including stereo camera 1768, wide-angle camera 1770, infrared camera 1772, surround camera 1774, long-range camera 1798, mid-range camera 1776, and / or other camera types. In at least one embodiment, the cameras may be used to capture image data around the entire perimeter of vehicle 1700. In at least one embodiment, the type of camera used depends on vehicle 1700. In at least one embodiment, any combination of camera types may be used to provide the necessary coverage around vehicle 1700. In at least one embodiment, the number of cameras deployed may vary depending on the embodiment. For example, in at least one embodiment, vehicle 1700 may include six cameras, seven cameras, ten cameras, twelve cameras, or other numbers of cameras. In at least one embodiment, by way of example but not limitation, the cameras may support Gigabit Multimedia Serial Link (“GMSL”) and / or Gigabit Ethernet communication. In at least one embodiment, each camera may be as described earlier herein. Figure 17A and Figure 17B A more detailed description was provided.
[0212] In at least one embodiment, vehicle 1700 may further include vibration sensor 1742. In at least one embodiment, vibration sensor 1742 may measure vibrations of components of vehicle 1700, such as axles. For example, in at least one embodiment, changes in vibration may indicate changes in road surface. In at least one embodiment, when two or more vibration sensors 1742 are used, differences between vibrations may be used to determine friction or slippage of the road surface (e.g., when the vibration difference is between a powered drive shaft and a freely rotating shaft).
[0213] In at least one embodiment, vehicle 1700 may include ADAS system 1738. In some examples, in at least one embodiment, ADAS system 1738 may include, but is not limited to, SoC. In at least one embodiment, ADAS system 1738 may include, but is not limited to, any number and combination of autonomous / adaptive / automatic cruise control (“ACC”) system, cooperative adaptive cruise control (“CACC”) system, forward collision warning (“FCW”) system, automatic emergency braking (“AEB”) system, lane departure warning (“LDW”) system, lane keeping assist (“LKA”) system, blind spot warning (“BSW”) system, rear cross traffic warning (“RCTW”) system, collision warning (“CW”) system, lane centering (“LC”) system, and / or other systems, features, and / or functions.
[0214] In at least one embodiment, the ACC system may use a RADAR sensor 1760, a LIDAR sensor 1764, and / or any number of cameras. In at least one embodiment, the ACC system may include a longitudinal ACC system and / or a lateral ACC system. In at least one embodiment, the longitudinal ACC system monitors and controls the distance to another vehicle directly in front of vehicle 1700 and automatically adjusts the speed of vehicle 1700 to maintain a safe distance from the vehicle ahead. In at least one embodiment, the lateral ACC system performs distance holding and, if necessary, suggests that vehicle 1700 change lanes. In at least one embodiment, lateral ACC is associated with other ADAS applications such as LC and CW.
[0215] In at least one embodiment, the CACC system uses information from other vehicles, which may be received from other vehicles via a wireless link or indirectly via a network connection (e.g., via the Internet) through network interface 1724 and / or wireless antenna 1726. In at least one embodiment, the direct link may be provided by a vehicle-to-vehicle (“V2V”) communication link, while the indirect link may be provided by an infrastructure-to-vehicle (“I2V”) communication link. Generally, V2V communication provides information about vehicles immediately ahead (e.g., vehicles directly in front and in the same lane as vehicle 1700), while I2V communication provides information about traffic further ahead. In at least one embodiment, the CACC system may include one or both of the I2V and V2V information sources. In at least one embodiment, given information about vehicles ahead of vehicle 1700, the CACC system may be more reliable, and it has the potential to improve traffic flow smoothness and reduce road congestion.
[0216] In at least one embodiment, the FCW system is designed to alert a driver to a hazard so that the driver can take corrective action. In at least one embodiment, the FCW system uses a front-facing camera and / or a RADAR sensor 1760, coupled to a dedicated processor, DSP, FPGA, and / or ASIC, electrically coupled to provide driver feedback, such as a display, speaker, and / or vibration components. In at least one embodiment, the FCW system can provide warnings, for example, in the form of audible, visual, haptic, and / or rapid braking pulses.
[0217] In at least one embodiment, the AEB system detects an impending forward collision with another vehicle or other object and can automatically apply braking if the driver does not take corrective action within a specified time or distance parameter. In at least one embodiment, the AEB system may use a front-facing camera and / or RADAR sensor 1760 coupled to a dedicated processor, DSP, FPGA, and / or ASIC. In at least one embodiment, when the AEB system detects a hazard, it will typically first alert the driver to take corrective action to avoid a collision, and if the driver does not take corrective action, the AEB system may automatically apply braking to attempt to prevent or minimize the effects of the predicted collision. In at least one embodiment, the AEB system may include techniques such as dynamic brake support and / or collision proximity braking.
[0218] In at least one embodiment, the LDW system provides visual, auditory, and / or tactile warnings, such as steering wheel or seat vibrations, to alert the driver when vehicle 1700 crosses lane markings. In at least one embodiment, the LDW system is not activated when the driver, for example, indicates intentional lane departure by activating a turn signal. In at least one embodiment, the LDW system may use a front-facing camera coupled to a dedicated processor, DSP, FPGA, and / or ASIC, whose electrical coupling provides driver feedback, such as a display, speaker, and / or vibrator components. In at least one embodiment, the LKA system is a variant of the LDW system. In at least one embodiment, if vehicle 1700 begins to leave its lane, the LKA system provides steering input or braking to correct vehicle 1700.
[0219] In at least one embodiment, the BSW system detects and warns the driver that the vehicle is in a blind spot. In at least one embodiment, the BSW system can provide visual, auditory, and / or tactile alerts to indicate that merging or changing lanes is unsafe. In at least one embodiment, the BSW system can provide additional warnings when the driver uses turn signals. In at least one embodiment, the BSW system can use a rear-side camera and / or radar sensor 1760, coupled to a dedicated processor, DSP, FPGA, and / or ASIC, electrically coupled to driver feedback, such as a display, speaker, and / or vibration assembly.
[0220] In at least one embodiment, when the vehicle 1700 is reversing, the RCTW system can provide visual, auditory, and / or tactile notifications when an object is detected outside the range of the rear camera. In at least one embodiment, the RCTW system includes an AEB system to ensure the application of the vehicle brakes to avoid a collision. In at least one embodiment, the RCTW system may use one or more rear-facing radar sensors 1760 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, electrically coupled to provide driver feedback, such as a display, speaker, and / or vibration assembly.
[0221] In at least one embodiment, conventional ADAS systems may be prone to false positives, which can be annoying and distracting to the driver, but are generally not catastrophic because conventional ADAS systems alert the driver and allow the driver to determine whether a safe condition truly exists and take appropriate action. In at least one embodiment, in the event of conflicting results, vehicle 1700 itself decides whether to pay attention to the results from the main computer or the auxiliary computer (e.g., the first or second controller in controller 1736). For example, in at least one embodiment, ADAS system 1738 may be a backup and / or auxiliary computer for providing perception information to a backup computer rationality module. In at least one embodiment, the backup computer rationality monitor may run redundant different software on hardware components to detect faults in perception and dynamic driving tasks. In at least one embodiment, the output from ADAS system 1738 may be provided to a monitoring MCU. In at least one embodiment, if the output from the main computer and the output from the auxiliary computer conflict, the monitoring MCU determines how to reconcile the conflict to ensure safe operation.
[0222] In at least one embodiment, the master computer can be configured to provide a confidence score to the supervisory MCU, indicating the master computer's confidence in the selected result. In at least one embodiment, if the confidence score exceeds a threshold, the supervisory MCU can follow the master computer's instructions regardless of whether the auxiliary computer provides conflicting or inconsistent results. In at least one embodiment, if the confidence score does not meet the threshold and the master computer and the auxiliary computer indicate different results (e.g., conflict), the supervisory MCU can arbitrate between the computers to determine the appropriate result.
[0223] In at least one embodiment, the supervisory MCU can be configured to run one or more neural networks trained and configured to determine the conditions for a false alarm provided by the auxiliary computer based at least in part on outputs from a host computer and an auxiliary computer. In at least one embodiment, the neural networks in the supervisory MCU can learn when the outputs of the auxiliary computer can be trusted and when they cannot. For example, in at least one embodiment, when the auxiliary computer is a radar-based FCW system, one or more neural networks in the supervisory MCU can learn when the FCW system identifies a metallic object that is not actually dangerous, such as a drain grille or manhole cover that triggers an alarm. In at least one embodiment, when the auxiliary computer is a camera-based LDW system, the neural networks in the supervisory MCU can learn to cover the LDW when a cyclist or pedestrian is present and lane departure is actually the safest maneuver. In at least one embodiment, the supervisory MCU can include at least one of a DLA or a GPU suitable for running neural networks with associated memory. In at least one embodiment, the supervisory MCU can include and / or be included as a component of the SoC 1704.
[0224] In at least one embodiment, the ADAS system 1738 may include an auxiliary computer that performs ADAS functions using conventional computer vision rules. In at least one embodiment, the auxiliary computer may use classic computer vision rules (if-then), and the presence of a neural network in the monitoring MCU can improve reliability, security, and performance. For example, in at least one embodiment, diverse implementations and intentional non-identity make the entire system more fault-tolerant, especially to failures caused by software (or software-hardware interface) functionality. For example, in at least one embodiment, if a software error or bug exists in the software running on the host computer, and different software code running on the auxiliary computer provides consistent overall results, the monitoring MCU may have greater confidence that the overall results are correct and that the error in the software or hardware on the host computer will not lead to a major error.
[0225] In at least one embodiment, the output of the ADAS system 1738 may be fed to the perception block and / or the dynamic driving task block of the host computer. For example, in at least one embodiment, if the ADAS system 1738 indicates a forward collision warning due to an object immediately ahead, the perception block may use this information when identifying the object. In at least one embodiment, the assistance computer may have its own neural network trained to reduce the risk of false positives, as described herein.
[0226] In at least one embodiment, vehicle 1700 may further include an infotainment SoC 1730 (e.g., an in-vehicle infotainment system (IVI)). Although illustrated and described as an SoC, in at least one embodiment, the infotainment system SoC 1730 may not be an SoC and may include, but is not limited to, two or more discrete components. In at least one embodiment, the infotainment SoC 1730 may include, but is not limited to, a combination of hardware and software that can be used to provide the vehicle 1700 with audio (e.g., music, personal digital assistant, navigation instructions, news, radio, etc.), video (e.g., television, movies, streaming media, etc.), telephone (e.g., hands-free calling), network connectivity (e.g., LTE, WiFi, etc.) and / or information services (e.g., navigation system, rear parking assist, radio data system, vehicle-related information (e.g., fuel level, total distance traveled, brake fuel level, fuel level, door opening / closing, air filter information, etc.)). For example, the infotainment SoC 1730 may include a radio, disk player, navigation system, video player, USB and Bluetooth connectivity, automotive, in-vehicle entertainment, WiFi, steering wheel audio controls, hands-free voice control, head-up display (“HUD”), HMI display 1734, telematics device, control panel (e.g., for controlling and / or interacting with various components, features and / or systems) and / or other components. In at least one embodiment, the infotainment SoC 1730 may also be used to provide information (e.g., visual and / or auditory) to a user of vehicle 1700, such as information from ADAS system 1738, autonomous driving information (e.g., planned vehicle operations, trajectory, surrounding environment information (e.g., intersection information, vehicle information, road information, etc.) and / or other information).
[0227] In at least one embodiment, the infotainment SoC 1730 may include any number and type of GPU functionality. In at least one embodiment, the infotainment SoC 1730 may communicate with other devices, systems, and / or components of the vehicle 1700 via bus 1702. In at least one embodiment, the infotainment SoC 1730 may be coupled to a monitoring MCU, enabling the GPU of the infotainment system to perform some autonomous driving functions in the event of a failure of the main controller 1736 (e.g., the primary and / or backup computer of the vehicle 1700). In at least one embodiment, the infotainment SoC 1730 may place the vehicle 1700 into a driver-safe parking mode, as described herein.
[0228] In at least one embodiment, vehicle 1700 may further include instrument cluster 1732 (e.g., digital instrument cluster, electronic instrument cluster, digital instrument panel, etc.). In at least one embodiment, instrument cluster 1732 may include, but is not limited to, controllers and / or supercomputers (e.g., discrete controllers or supercomputers). In at least one embodiment, instrument cluster 1732 may include, but is not limited to, any number and combination of instruments, such as speedometer, fuel level, fuel pressure, tachometer, odometer, turn signal indicator, shift position indicator, seatbelt warning light, parking brake warning light, engine malfunction indicator, auxiliary restraint system (e.g., airbag) information, lighting control, safety system control, navigation information, etc. In some examples, information may be displayed and / or shared between infotainment SoC 1730 and instrument cluster 1732. In at least one embodiment, instrument cluster 1732 may be included as part of infotainment SoC 1730, or vice versa.
[0229] Such components can be used to generate alternative view images, such as bird's-eye view images, from parallax data using hardware with limited capacity (such as embedded processors that do not access external memory).
[0230] Figure 17D It is according to at least one embodiment for use on a cloud-based server and Figure 17AA diagram of a system for communication between autonomous vehicles 1700. In at least one embodiment, the system may include, but is not limited to, a server 1778, a network 1790, and any number and type of vehicles, including vehicle 1700. In at least one embodiment, the server 1778 may include, but is not limited to, multiple GPUs 1784(A)-1784(H) (collectively referred to herein as GPU 1784), PCIe switches 1782(A)-1782(D) (collectively referred to herein as PCIe switch 1782), and / or CPUs 1780(A)-1780(B) (collectively referred to herein as CPU 1780). In at least one embodiment, the GPU 1784, CPU 1780, and PCIe switch 1782 may be interconnected with high-speed interconnects, such as, but not limited to, NVLink interface 1788 and / or PCIe connection 1786 developed by NVIDIA. In at least one embodiment, the GPU 1784 is connected via NVLink and / or NVSwitch SoC and the GPU 1784 and PCIe switches 1782 via PCIe interconnect. In at least one embodiment, although eight GPUs 1784, two CPUs 1780, and four PCIe switches 1782 are described, this is not intended to be limiting. In at least one embodiment, each server 1778 may include, but is not limited to, any combination of any number of GPUs 1784, CPUs 1780, and / or PCIe switches 1782. For example, in at least one embodiment, each server 1778 may include eight, sixteen, thirty-two, and / or more GPUs 1784.
[0231] In at least one embodiment, server 1778 may receive image data representing an image from a vehicle via network 1790, the image showing unexpected or changed road conditions, such as recently started road construction. In at least one embodiment, server 1778 may transmit neural network 1792, updated or otherwise, and / or map information 1794, including but not limited to information about traffic and road conditions, to the vehicle via network 1790. In at least one embodiment, updates to map information 1794 may include, but are not limited to, updates to high-definition map 1722, such as information about construction sites, potholes, detours, floods, and / or other obstacles. In at least one embodiment, neural network 1792 and / or map information 1794 may have been generated from new training and / or experience represented by data received from any number of vehicles in the environment, and / or at least based in part on training performed in a data center (e.g., using server 1778 and / or other servers).
[0232] In at least one embodiment, server 1778 can be used to train a machine learning model (e.g., a neural network) at least in part based on training data. In at least one embodiment, the training data can be generated by a vehicle, and / or can be generated in a simulation (e.g., using a game engine). In at least one embodiment, any amount of training data is labeled (e.g., the associated neural network benefits from supervised learning) and / or undergoes other preprocessing. In at least one embodiment, any amount of training data is not labeled and / or preprocessed (e.g., where the associated neural network does not require supervised learning). In at least one embodiment, once the machine learning model is trained, the machine learning model can be used by the vehicle (e.g., transmitted to the vehicle via network 1790), and / or the machine learning model can be used by server 1778 to remotely monitor the vehicle.
[0233] In at least one embodiment, server 1778 can receive data from a vehicle and apply the data to a state-of-the-art real-time neural network for real-time intelligent inference. In at least one embodiment, server 1778 may include a deep learning supercomputer and / or a dedicated AI computer powered by GPU 1784, such as the DGX and DGX Station machines developed by NVIDIA. However, in at least one embodiment, server 1778 may include a deep learning infrastructure in a CPU-driven data center.
[0234] In at least one embodiment, the deep learning infrastructure of server 1778 may be capable of fast, real-time inference and may use this capability to assess and verify the health status of processors, software, and / or related hardware. For example, in at least one embodiment, the deep learning infrastructure may receive periodic updates from vehicle 1700, such as image sequences and / or objects that vehicle 1700 has located in the image sequence (e.g., via computer vision and / or other machine learning object classification techniques). In at least one embodiment, the deep learning infrastructure may run its own neural network to identify objects and compare them with objects identified by vehicle 1700, and if the results do not match and the deep learning infrastructure concludes that the AI in vehicle 1700 has malfunctioned, server 1778 may send a signal to vehicle 1700 instructing the fail-safe computer of vehicle 1700 to take control, notify passengers, and complete a safe stopping operation.
[0235] Such components can be used to generate alternative view images, such as bird's-eye view images, from parallax data using hardware with limited capacity (such as embedded processors that do not access external memory).
[0236] The various embodiments can be described by the following terms:
[0237] 1. A system comprising:
[0238] At least one embedded processor with direct memory access (DMA) functionality is used for:
[0239] A two-dimensional 2D histogram view of one or more objects in the environment is generated, partially based on disparity data of those objects, the histogram view being a function of angle and distance relative to a plane of at least one camera used to generate the disparity data; and
[0240] A bird's-eye view image of the one or more objects is generated by transforming the 2D histogram view.
[0241] 2. The system according to Clause 1, wherein the at least one embedded processor cannot access the external memory used when generating the bird's-eye view image.
[0242] 3. The system according to Clause 1, wherein the system is further configured to: determine the parallax data using image data captured by the camera.
[0243] 4. The system according to Clause 1, wherein the camera is a stereo camera assembly or a pair of matched camera sensors.
[0244] 5. The system according to Clause 1, wherein the bird's-eye view image is generated in part by generating a list of object centroids and statistics using the 2D histogram view and transforming the list into a corresponding list in the coordinate system of the bird's-eye view image.
[0245] 6. The system according to Clause 1, wherein the data transfer of the embedded processor is performed using the DMA on a rectangular region of the image data.
[0246] 7. The system according to Clause 1, wherein the at least one embedded processor is further configured to:
[0247] Receive stereo image data from the at least one camera; and
[0248] The stereo image data is used to generate a disparity map that includes the disparity data of the one or more objects.
[0249] 8. The system according to Clause 1, wherein the parallax data includes data obtained from at least one additional sensor.
[0250] 9. The system according to Clause 1, wherein said system comprises at least one of the following:
[0251] A system used to perform simulation operations;
[0252] A system used to perform simulations to test or validate autonomous machine applications;
[0253] Systems used to perform digital twin operations;
[0254] A system for performing optical transmission simulation;
[0255] A system used for rendering graphics output;
[0256] A system used to perform deep learning operations;
[0257] A system for performing generative AI operations using large language model LLM;
[0258] A system for performing generative AI operations using a visual language model (VLM);
[0259] A system for performing generative AI operations using the multimodal language model MMLM;
[0260] A system for deploying one or more language models using an operating system-level virtualization container, wherein the operating system-level virtualization container communicates with the one or more language models using one or more application programming interface (API) methods.
[0261] Systems implemented using edge devices;
[0262] Systems used to generate or present virtual reality (VR) content;
[0263] A system for generating or presenting augmented reality (AR) content;
[0264] A system for generating or presenting mixed reality (MR) content;
[0265] A system containing one or more virtual machines (VMs);
[0266] A system that is at least partially implemented in a data center;
[0267] A system for performing hardware tests using simulation;
[0268] Systems for generating synthetic data;
[0269] A collaborative content creation platform for 3D assets; or
[0270] A system that utilizes cloud computing resources at least in part.
[0271] 10. At least one embedded processor having a direct memory access (DMA) function for generating a bird's-eye view image of a scene by generating an intermediate histogram from parallax data of the scene as a function of angle and distance relative to a camera plane, and transforming the intermediate histogram into the bird's-eye view image.
[0272] 11. The at least one embedded processor according to Clause 10, wherein the intermediate histogram includes representations of one or more objects in the scene, and wherein the at least one embedded processor is further configured to: perform connected component analysis on the intermediate histogram to identify pixel locations associated with the one or more objects.
[0273] 12. The at least one embedded processor according to Clause 11, wherein the at least one embedded processor is further configured to: generate a list of object centroids and statistics of the one or more objects using the intermediate histogram, and transform the list into a corresponding list in the coordinate system of the bird's-eye view image.
[0274] 13. At least one embedded processor as described in Clause 10, wherein the at least one embedded processor has no access to the complete set of image data stored in external memory for generating the intermediate histogram or the bird's-eye view image.
[0275] 14. At least one embedded processor according to Clause 10, wherein data transfer of the at least one embedded processor is performed using the DMA on a rectangular region of image data.
[0276] 15. The at least one embedded processor according to Clause 10, wherein the at least one embedded processor is further configured to: determine the parallax data using image data captured at least using a stereo camera assembly.
[0277] 16. At least one embedded processor according to Clause 10, wherein said at least one embedded processor is included in at least one of the following:
[0278] A system used to perform simulation operations;
[0279] A system used to perform simulations to test or validate autonomous machine applications;
[0280] Systems used to perform digital twin operations;
[0281] A system for performing optical transmission simulation;
[0282] A system used for rendering graphics output;
[0283] A system used to perform deep learning operations;
[0284] Systems implemented using edge devices;
[0285] Systems used to generate or present virtual reality (VR) content;
[0286] A system for generating or presenting augmented reality (AR) content;
[0287] A system for generating or presenting mixed reality (MR) content;
[0288] A system containing one or more virtual machines (VMs);
[0289] A system that is at least partially implemented in a data center;
[0290] A system for performing hardware tests using simulation;
[0291] Systems for generating synthetic data;
[0292] A system for performing generative AI operations using large language model LLM;
[0293] A system for performing generative AI operations using a visual language model (VLM);
[0294] A system for performing generative AI operations using the multimodal language model MMLM;
[0295] A system for deploying one or more language models using an operating system-level virtualization container, wherein the operating system-level virtualization container communicates with the one or more language models using one or more application programming interface (API) methods.
[0296] A collaborative content creation platform for 3D assets; or
[0297] A system that utilizes cloud computing resources at least in part.
[0298] 17. A computer-implemented method, comprising:
[0299] An intermediate histogram representation of the parallax image is generated using an embedded processor with DMA memory access.
[0300] Identify the position in the intermediate histogram representation associated with one or more objects; and
[0301] Using the embedded processor and in part based on the location, the intermediate histogram representation is transformed into a bird's-eye view image that includes representations of the one or more objects.
[0302] 18. The computer-implemented method according to Clause 17, wherein the embedded processor has no access to external memory used in generating the intermediate histogram representation or the bird's-eye view image.
[0303] 19. The computer-implemented method according to Clause 17 further includes:
[0304] The embedded processor is used to perform connected component analysis on the intermediate histogram representation to identify the locations associated with the one or more objects.
[0305] 20. The computer-implemented method according to Clause 18 further includes:
[0306] The embedded processor is used to generate a list of object centroids and statistics for the one or more objects from the intermediate histogram representation; and
[0307] The embedded processor is used to transform the list into a corresponding list in the coordinate system of the bird's-eye view image.
[0308] 21. The computer-implemented method according to Clause 18, wherein the bird's-eye view image is provided for use in a simulation of the one or more objects generated at least in part using a 3D content collaboration platform for 3D assets.
[0309] 22. The computer-implemented method according to Clause 21, wherein the 3D content collaboration platform for 3D assets uses OpenUSD.
[0310] Other variations are within the spirit of this disclosure. Therefore, although the disclosed technology is readily adaptable to various modifications and alternative constructions, certain embodiments thereof are illustrated in the accompanying drawings and have been described in detail above. However, it should be understood that the disclosure is not intended to be limited to one or more specific forms disclosed, but rather, it is intended to cover all modifications, alternative constructions, and equivalents falling within the spirit and scope of this disclosure as defined in the appended claims.
[0311] Unless otherwise stated or obviously contradicted by the context, the terms “a,” “an,” and “the,” and similar references, used in the context of describing the disclosed embodiments (particularly in the context of the appended claims), should be interpreted as encompassing both singular and plural forms, rather than as definitions of the terms. Unless otherwise stated, the terms “comprising,” “having,” “including,” and “containing” should be interpreted as open-ended terms (meaning “including, but not limited to”). The term “connection” (referring to a physical connection where not modified) should be interpreted as partially or wholly contained, attached to, or joined together, even with some intervention. Unless otherwise indicated herein, references to numerical ranges herein are intended only as a way of abbreviating each individual value falling within that range, and each individual value is incorporated into the specification as if it were separately described herein. Unless otherwise indicated or contradicted by the context, the use of the terms “set” (e.g., “item set”) or “subset” should be interpreted as a non-empty set comprising one or more members. Furthermore, unless otherwise indicated or contradicted by the context, the term "subset" of a corresponding set does not necessarily refer to an appropriate subset of the corresponding set, but rather the subset and the corresponding set can be equal.
[0312] Unless otherwise explicitly stated or clearly contradicted by the context, connective phrases such as “at least one of A, B, and C” or “at least one of A, B, and C” are understood in the context to generally refer to items, terms, etc., which can be A or B or C, or any non-empty subset of the set A, B, and C. For example, in an illustrative example of a set with three members, the connective phrases “at least one of A, B, and C” and “at least one of A, B, and C” refer to any of the following sets: {A}, {B}, {C}, {A, B}, {A, C}, {B, C}, {A, B, C}. Therefore, such connective language is generally not intended to imply that some embodiments require the presence of at least one of A, at least one of B, and at least one of C. Additionally, unless otherwise stated or contradicted by the context, the term “multiple” indicates a plural state (e.g., “multiple items” means multiple items). The number of items in multiple items is at least two, but may be more if explicitly indicated or indicated by the context. Furthermore, unless otherwise stated or clearly understood from the context, the phrase “based on” means “at least partially based on” rather than “based on only”.
[0313] Unless otherwise indicated herein or clearly contradicted by the context, the operations of the processes described herein may be performed in any suitable order. In at least one embodiment, processes such as those described herein (or variations thereof and / or combinations thereof) are executed under the control of one or more computer systems configured with executable instructions and are implemented as code (e.g., executable instructions, one or more computer programs, or one or more application programs) that is executed jointly on one or more processors via hardware or a combination thereof. In at least one embodiment, the code is stored on a computer-readable storage medium, for example, in the form of a computer program comprising a plurality of instructions executable by one or more processors. In at least one embodiment, the computer-readable storage medium is a non-transitory computer-readable storage medium that excludes transient signals (e.g., propagating transient electrical or electromagnetic transmissions) but includes non-transitory data storage circuitry (e.g., buffers, caches, and queues). In at least one embodiment, code (e.g., executable code or source code) is stored on one or more non-transitory computer-readable storage media (or other memory for storing executable instructions) on which executable instructions are stored, which, when executed by one or more processors of a computer system (e.g., as a result of execution), cause the computer system to perform the operations described herein. In at least one embodiment, the set of non-transitory computer-readable storage media comprises multiple non-transitory computer-readable storage media, and one or more of the individual non-transitory storage media lack all the code, but the multiple non-transitory computer-readable storage media collectively store all the code. In at least one embodiment, the executable instructions are executed such that different instructions are executed by different processors; for example, the non-transitory computer-readable storage media store the instructions, and the main central processing unit (“CPU”) executes some instructions while the graphics processing unit (“GPU”) executes other instructions. In at least one embodiment, different components of the computer system have separate processors, and the different processors execute different subsets of the instructions.
[0314] Therefore, in at least one embodiment, the computer system is configured to implement one or more services that perform the operations of the processes described herein, either individually or collectively, and such a computer system is configured with suitable hardware and / or software to enable the implementation of the operations. Furthermore, the computer system implementing at least one embodiment of this disclosure is a single device, and in another embodiment it is a distributed computer system comprising multiple devices operating in different ways, such that the distributed computer system performs the operations described herein, and that a single device does not perform all the operations.
[0315] The use of any and all examples or exemplary language (e.g., “such as”) provided herein is intended only to better illustrate embodiments of this disclosure and does not constitute a limitation on the scope of the disclosure unless otherwise required. No language in the specification should be construed as indicating that any unclaimed element is essential to the practice of the disclosure.
[0316] All references cited in this article, including publications, patent applications and patents, are incorporated herein by reference as if each reference were individually and specifically indicated to be incorporated herein by reference and the entire contents of which are described herein.
[0317] The terms “coupled” and “connected”, and their derivatives, may be used in the specification and claims. It should be understood that these terms may not be intended to be synonyms with each other. Rather, in certain examples, “connected” or “coupled” may be used to indicate that two or more elements are in direct or indirect physical or electrical contact with each other. “Coupled” may also mean that two or more elements are not in direct contact with each other, but still cooperate or interact with each other.
[0318] Unless otherwise expressly stated, it will be understood that throughout this specification, terms such as “processing,” “computing,” “determining,” etc., refer to the actions and / or processes of a computer or computing system or similar electronic computing device that process and / or convert data represented as physical quantities (e.g., electrons) in the registers and / or memory of the computing system into other data represented as physical quantities in the memory, registers, or other such information storage, transmission, or display devices of the computing system.
[0319] In a similar manner, the term "processor" can refer to any device or part of memory that processes electronic data from registers and / or memory and converts that electronic data into other electronic data that can be stored in registers and / or memory. As a non-limiting example, a "processor" can be a CPU or a GPU. A "computing platform" can include one or more processors. As used herein, a "software" process can include, for example, software and / or hardware entities that perform work over time, such as tasks, threads, and intelligent agents. Similarly, each process can refer to multiple processes that execute instructions sequentially or intermittently, sequentially, or in parallel. The terms "system" and "method" are used interchangeably herein, provided that a system can embody one or more methods, and a method can be considered a system.
[0320] This document refers to the process of acquiring, obtaining, receiving, or inputting analog or digital data into a subsystem, computer system, or computer-implemented machine. Analog and digital data can be acquired, obtained, received, or input in various ways, such as by receiving data as a parameter to a function call or a call to an application programming interface (API). In some implementations, the process of acquiring, obtaining, receiving, or inputting analog or digital data can be accomplished by transmitting data via a serial or parallel interface. In another implementation, the process of acquiring, obtaining, receiving, or inputting analog or digital data can be accomplished by transmitting data from a providing entity to an acquiring entity via a computer network. Reference can also be made to providing, outputting, transmitting, sending, or presenting analog or digital data. In various examples, the process of providing, outputting, transmitting, sending, or presenting analog or digital data can be implemented by transmitting data as an input or output parameter to a function call, an API, or an inter-process communication mechanism.
[0321] While the discussion above illustrates example implementations of the described technologies, other architectures can be used to implement the described functionality and are intended to fall within the scope of this disclosure. Furthermore, although specific assignments of responsibilities have been defined above for discussion purposes, various functions and responsibilities can be assigned and divided in different ways depending on the circumstances.
[0322] Furthermore, although the subject matter has been described in language specific to structural features and / or methodological actions, it should be understood that the subject matter claimed in the appended claims is not necessarily limited to the specific features or actions described. Rather, specific features and actions are disclosed as exemplary forms for implementing the claims.
Claims
1. A system comprising: At least one embedded processor with direct memory access (DMA) functionality is used for: A two-dimensional 2D histogram view of one or more objects is generated, in part based on disparity data of one or more objects in the environment, the two-dimensional histogram view being a function of angle and distance relative to a plane of at least one camera used to generate the disparity data; as well as A bird's-eye view image of the one or more objects is generated by transforming the 2D histogram view.
2. The system of claim 1, wherein the at least one embedded processor cannot access the external memory used when generating the bird's-eye view image.
3. The system of claim 1, wherein the system is further configured to: determine the parallax data using image data captured by the camera.
4. The system of claim 1, wherein the camera is a stereo camera assembly or a pair of matched camera sensors.
5. The system of claim 1, wherein the bird's-eye view image is generated in part by generating a list of object centroids and statistics using the 2D histogram view and transforming the list into a corresponding list in the coordinate system of the bird's-eye view image.
6. The system of claim 1, wherein the data transfer of the embedded processor is performed using the DMA on a rectangular region of the image data.
7. The system of claim 1, wherein the at least one embedded processor is further configured to: Receive stereo image data from the at least one camera; and The stereo image data is used to generate a disparity map that includes the disparity data of the one or more objects.
8. The system of claim 1, wherein the parallax data includes data obtained from at least one additional sensor.
9. The system of claim 1, wherein the system comprises at least one of the following: A system used to perform simulation operations; A system used to perform simulations to test or validate autonomous machine applications; Systems used to perform digital twin operations; A system for performing optical transmission simulation; A system used for rendering graphics output; A system used to perform deep learning operations; A system for performing generative AI operations using large language model LLM; A system for performing generative AI operations using a visual language model (VLM); A system for performing generative AI operations using the multimodal language model MMLM; A system for deploying one or more language models using an operating system-level virtualization container, wherein the operating system-level virtualization container communicates with the one or more language models using one or more application programming interface (API) methods. Systems implemented using edge devices; Systems used to generate or present virtual reality (VR) content; A system for generating or presenting augmented reality (AR) content; A system for generating or presenting mixed reality (MR) content; A system containing one or more virtual machines (VMs); A system that is at least partially implemented in a data center; A system for performing hardware tests using simulation; Systems for generating synthetic data; A collaborative content creation platform for 3D assets; or A system that utilizes cloud computing resources at least in part.
10. At least one embedded processor having a direct memory access (DMA) function for generating a bird's-eye view image of a scene by generating an intermediate histogram from parallax data of the scene as a function of angle and distance relative to a camera plane, and transforming the intermediate histogram into the bird's-eye view image.
11. The at least one embedded processor of claim 10, wherein the intermediate histogram includes representations of one or more objects in the scene, and wherein the at least one embedded processor is further configured to: perform connected component analysis on the intermediate histogram to identify pixel locations associated with the one or more objects.
12. The at least one embedded processor according to claim 11, wherein the at least one embedded processor is further configured to: generate a list of object centroids and statistics of the one or more objects using the intermediate histogram, and transform the list into a corresponding list in the coordinate system of the bird's-eye view image.
13. The at least one embedded processor of claim 10, wherein the at least one embedded processor cannot access the complete set of image data stored in external memory for generating the intermediate histogram or the bird's-eye view image.
14. The at least one embedded processor of claim 10, wherein the data transfer of the at least one embedded processor is performed using the DMA on a rectangular region of the image data.
15. The at least one embedded processor of claim 10, wherein the at least one embedded processor is further configured to: determine the parallax data using image data captured at least using a stereo camera assembly.
16. The at least one embedded processor of claim 10, wherein the at least one embedded processor is comprised of at least one of the following: A system used to perform simulation operations; A system used to perform simulations to test or validate autonomous machine applications; Systems used to perform digital twin operations; A system for performing optical transmission simulation; A system used for rendering graphics output; A system used to perform deep learning operations; Systems implemented using edge devices; Systems used to generate or present virtual reality (VR) content; A system for generating or presenting augmented reality (AR) content; A system for generating or presenting mixed reality (MR) content; A system containing one or more virtual machines (VMs); A system that is at least partially implemented in a data center; A system for performing hardware tests using simulation; Systems for generating synthetic data; A system for performing generative AI operations using large language model LLM; A system for performing generative AI operations using a visual language model (VLM); A system for performing generative AI operations using the multimodal language model MMLM; A system for deploying one or more language models using an operating system-level virtualization container, wherein the operating system-level virtualization container communicates with the one or more language models using one or more application programming interface (API) methods. A collaborative content creation platform for 3D assets; or A system that utilizes cloud computing resources at least in part.
17. A computer-implemented method, comprising: An intermediate histogram representation of the parallax image is generated using an embedded processor with DMA memory access. Identify the position in the intermediate histogram representation associated with one or more objects; and Using the embedded processor and in part based on the location, the intermediate histogram representation is transformed into a bird's-eye view image that includes representations of the one or more objects.
18. The computer-implemented method of claim 17, wherein the embedded processor has no access to external memory used in generating the intermediate histogram representation or the bird's-eye view image.
19. The computer-implemented method according to claim 17, further comprising: The embedded processor is used to perform connected component analysis on the intermediate histogram representation to identify the locations associated with the one or more objects.
20. The computer-implemented method according to claim 18, further comprising: The embedded processor is used to generate a list of object centroids and statistics for one or more objects from the intermediate histogram representation; as well as The embedded processor is used to transform the list into a corresponding list in the coordinate system of the bird's-eye view image.
21. The computer-implemented method of claim 18, wherein the bird's-eye view image is provided for use in a simulation of the one or more objects generated at least in part using a 3D content collaboration platform for 3D assets.
22. The computer-implemented method of claim 21, wherein the 3D content collaboration platform for 3D assets uses OpenUSD.