Three-dimensional multi-camera perception system and application
By directly processing image data through a multi-camera perception system to generate bird's-eye view features, the problem of accuracy in three-dimensional position determination in complex environments is solved, and higher-precision three-dimensional position determination is achieved.
Patent Information
- Application Number
- CN202510151554.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-09-26
- Filing Date
- 2025-02-11
- Publication Date
- 2025-09-19
AI Technical Summary
Existing 3D position determination systems suffer from projection errors caused by occlusion, inaccurate camera calibration, and overlapping fields of view in complex environments, which affect the accuracy of the 3D position of objects, especially in retail and warehouse environments.
By using a multi-camera perception system, image data is directly processed to generate multi-view image features, and a spatiotemporal transformer and spatial encoder are used to generate bird's-eye view features. Combined with the temporal encoder and decoder, the three-dimensional position of the object is determined, avoiding the projection process of the two-dimensional position.
The accuracy of three-dimensional position determination is improved, especially in complex environments, errors caused by occlusion and inaccurate calibration are eliminated, and the precision of object position is enhanced.
Smart Images

Figure CN120676130A_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims the benefit of U.S. Provisional Application No. 63 / 566,549, filed March 18, 2024, and Italian Patent Application No. 102024000020065, filed September 9, 2024. Each application is incorporated herein by reference in its entirety. Background Art
[0003] Determining the three-dimensional (3D) position of an object in a particular environment is important for many tasks, such as tracking objects in a retail and / or warehouse environment. Conventional systems for determining 3D positions within an environment may receive image data generated using multiple cameras positioned throughout the environment, where each camera includes a respective field of view that captures a portion of the environment. The conventional system may then process the image data from each camera individually to determine the two-dimensional (2D) position of an object in the image represented by the image data. Next, to determine the 3D position, the conventional system may project the 2D position of the object in the image into a 3D coordinate space associated with the environment using calibration information associated with the cameras.
[0004] However, many problems may arise when projecting a 2D position into a 3D coordinate space associated with an environment. For example, the projection of the 2D position may be affected by various factors, such as occlusions in the image (e.g., objects are occluded by other objects), inaccurate calibration of the camera, and / or misaligned object detection between cameras that include overlapping fields of view. As a result, the accuracy of these conventional systems for determining the 3D position of an object may be reduced, which may further cause problems for downstream tasks (e.g., using the 3D position to track objects in the environment). In addition, these problems with conventional systems may be more prevalent in certain environments, such as complex environments (e.g., retail environments, warehouse environments, etc.) that include a large number of cameras throughout the environment and / or a large amount of space that is occluded by one or more cameras. Summary of the Invention
[0005] Embodiments of the present disclosure relate to three-dimensional multi-camera perception systems and applications. Disclosed herein are systems and methods for processing image data generated using multiple cameras positioned throughout an environment to directly determine three-dimensional (3D) information associated with objects located within the environment. For example, one or more feature extractors (e.g., one or more backbones) can be used to process image data to determine multi-view image features associated with the image data. These multi-view image features and calibration data associated with the cameras can then be processed using one or more spatiotemporal transformers to determine the 3D position of objects within the environment. For example, as described in more detail herein, a spatial encoder can process multi-view image features and calibration data to generate bird's eye view (BEV) features. A temporal encoder can then fuse current BEV features with instances of previous BEV features associated with previous time periods. A decoder can then process these fused BEV features to determine the 3D position of an object.
[0006] Compared to conventional systems (such as the systems described above), the systems of the present disclosure can directly determine the 3D position associated with an object without initially determining the 2D position corresponding to a different image and / or projecting the 2D position from the image space to the 3D coordinate space associated with the environment. Therefore, the systems of the present disclosure can eliminate projection errors that may be caused by occlusions within the image (e.g., an object is occluded by other objects), inaccurate calibration of the camera, and / or misalignment of object detection between cameras with overlapping fields of view. By eliminating these projection errors, the systems of the present disclosure can improve the overall accuracy of determining the 3D position, particularly in complex environments (e.g., retail environments, warehouse environments, etc.) that include a large number of cameras positioned throughout the environment and / or a large amount of space occluded by one or more cameras. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] The present system and method for a three-dimensional multi-camera perception system and application will be described in detail below with reference to the accompanying drawings, wherein:
[0008] Figure 1A An example of a first process of performing three-dimensional multi-camera perception to determine three-dimensional information related to an object according to some embodiments of the present disclosure is shown;
[0009] Figure 1B An example of a second process of performing three-dimensional multi-camera perception to determine three-dimensional information related to an object according to some embodiments of the present disclosure is shown;
[0010] Figure 2 shows an example of an environment including cameras located at various locations according to some embodiments of the present disclosure;
[0011] Figures 3A to 3Bshows an example of a camera generating image data representing an environment according to some embodiments of the present disclosure;
[0012] Figure 4 shows examples of three-dimensional information that may be output according to some embodiments of the present disclosure;
[0013] Figure 5 An example of a process for classifying an object using three-dimensional information associated with the object according to some embodiments of the present disclosure is shown;
[0014] Figure 6 shows a data flow diagram illustrating a process for training one or more networks to perform three-dimensional multi-camera perception according to some embodiments of the present disclosure;
[0015] Figures 7 and 8 A flowchart illustrating a method of performing three-dimensional multi-camera perception associated with an environment according to some embodiments of the present disclosure is shown;
[0016] Figure 9 A flowchart illustrating a method for determining bird's-eye view features based at least on multi-view image features according to some embodiments of the present disclosure is shown;
[0017] Figure 10 shows an example architecture that may perform one or more of the processes described herein, according to some embodiments of the present disclosure;
[0018] Figure 11 is a block diagram of an example computing device suitable for implementing some embodiments of the present disclosure; and
[0019] Figure 12 is a block diagram of an example data center suitable for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION
[0020] Systems and methods related to three-dimensional multi-camera perception systems and applications are disclosed. For example, the system may receive instances of image data generated using cameras positioned throughout an environment. As described herein, the environment may include indoor environments, such as retail environments, warehouse environments, office environments, educational environments, and / or any other indoor environments, and / or the environment may include outdoor environments. In addition, the environment may include static objects that are stationary in the environment, such as shelves, tables, racks, walls, doors, fixtures, furniture, appliances, and / or any other type of static objects, as well as dynamic objects that move throughout the environment, such as people, animals, machines (e.g., robots, etc.), and / or any other type of dynamic objects. In some examples, the system may receive image data continuously from the cameras. In some examples, the system may receive image data at a given time instance (e.g., every second, every minute, every hour, etc.). However, in some examples, the image data may have been previously generated using the cameras and stored in one or more databases for later processing.
[0021] As described herein, cameras can be positioned throughout an environment such that the cameras include fields of view (FOVs) that capture different portions of the environment (e.g., within the environment). For example, a first camera may include a first FOV that includes a first portion of the environment, a second camera may include a second FOV that includes a second portion of the environment, a third camera may include a third FOV that includes a third portion of the environment, and so on. Additionally, in some examples, at least some of the cameras may include overlapping FOVs within the environment. For example, a first camera's first FOV may at least partially overlap with a second camera's second FOV, such that the first and second cameras capture similar portions of the environment. Furthermore, in some examples, such as when cameras include overlapping FOVs, a portion of the environment may be obscured by one camera while still visible to the other camera. For example, a static object located within the environment may obscure a portion of a first camera's first FOV, rendering a portion of the environment invisible to the first camera while still visible to the second camera.
[0022] The system can then process the image data to determine information associated with objects located in the environment. As described herein, this information can include, but is not limited to, 3D information of objects within the environment (e.g., 3D positions), classifications associated with objects within the environment, identifiers associated with objects within the environment, and / or any other information. Furthermore, in some examples, the 3D information of an object can include, but is not limited to, a 3D point within the environment, a 3D bounding shape within the environment (e.g., a bounding box, a bounding cuboid, a bounding cylinder, etc.), a 3D pose of the object (e.g., a skeleton, etc.), a 3D shape representing the object (e.g., a rod, a cylinder, etc.), a 3D point representing the position of the object (e.g., a point on the ground, etc.), a relative position relative to a reference object, and / or any other type of 3D information indicating the position of an object in the environment. In addition, in some examples, one or more parameters can be used to represent a 3D bounding shape, for example, three parameters for representing the scale of the bounding shape (e.g., length, width, and height), three parameters for representing the center position of the bounding shape (e.g., x-coordinate position, y-coordinate position, and z-coordinate position), two parameters for representing the yaw of the object (e.g., cosine angle and sine angle), and / or two parameters for representing the speed of the object (e.g., speed in the x direction and speed in the y direction).
[0023] To determine 3D information, the system may first process the image data using one or more feature extractors configured to generate feature data representing multi-view image features. For example, the system may process the image data using one or more backbones configured to generate first feature data associated with first image data generated using a first camera, second feature data associated with second image data generated using a second camera, third feature data associated with third image data generated using a third camera, and so on. In such an example, the first feature data may represent a first feature associated with a first image depicting a first portion of the environment, the second feature data may represent a second feature associated with a second image depicting a second portion of the environment, the third feature data may represent a third feature associated with a third image depicting a third portion of the environment, and so on. This is why the feature data may be referred to as representing "multi-view" image features of the environment.
[0024] The system may then process the feature data and calibration data associated with a camera within the environment using one or more spatiotemporal transformers configured to determine 3D information associated with the object. As described herein, the calibration data for the camera may associate 3D coordinates (e.g., 3D points) within the environment with 2D coordinates (e.g., 2D points) associated with an image generated using the camera. For example, as described in greater detail herein, the calibration data for the camera may include a matrix (e.g., a 3x4 projection matrix) that associates 3D points within the environment with 2D points associated with the image. In some examples, the system (and / or another system) may generate the calibration data using one or more inputs indicating information associated with the camera, such as intrinsic parameters (e.g., focal length, principal point, scale factor, etc.) and / or extrinsic parameters (e.g., position, orientation, etc.) associated with the camera. In some examples, the system (and / or another system) may automatically generate the calibration data based at least on processed data generated using the camera.
[0025] For determining more details of the 3D information, the system may use one or more spatial encoders to process feature data, calibration data, and / or query data representing one or more queries. Based at least on the processing, the spatial encoder may generate aggregated feature data representing features associated with the environment, which features are also referred to as "BEV features." For example, the spatial encoder may generate aggregated feature data by aggregating features represented by individual feature data associated with different images of the image data. In some examples, to perform the aggregation, the spatial encoder may use the calibration data to project 3D points associated with the aggregated feature data (e.g., a BEV image) to 2D points associated with the feature data (e.g., an image associated with the feature data). The spatial encoder may then use the projection to determine features corresponding to the 2D points associated with the feature data and map these features to 3D points associated with the aggregated feature data.
[0026] The system can then use one or more temporal encoders to process the aggregated feature data with respect to one or more additional instances of aggregated feature data associated with one or more previous time instances. For example, the temporal encoder can be configured to concatenate the aggregated feature data with previous instances of aggregated feature data to generate fused feature data. In some examples, when performing the concatenation, the temporal encoder can use one or more temporal self-attention layers to model the temporal connections between BEV features of different instances of aggregated feature data so as to construct accurate associations between similar objects represented by the aggregated feature data at different time instances. In addition, the present invention describes in more detail the details of how the temporal encoder generates fused feature data.
[0027] The system can then process the fused feature data using one or more decoders, such as one or more DETR decoders (and / or any other type of decoder) that are configured to generate data representing 3D information associated with the object. For example, in some examples, the decoder can include one or more self-attention layers and / or one or more cross-attention layers that can be stacked together alternatively, where the layers are used to process the fused feature data to determine the 3D information. However, in other examples, the decoder can include any other type of layer to perform one or more processes described herein. Additionally, in some examples, the decoder can use a set of learned embeddings as object queries, where the object queries can indicate where a target object may be located in the environment to determine 3D information associated with the object.
[0028] In some examples, the system can continue to perform these processes to continue generating data representing 3D information associated with the object at different time instances. For example, the system can generate data for every frame generated using the camera, every other frame generated using the camera, every fourth frame generated using the camera, and / or using any other type of interval. Furthermore, in some examples, the system (and / or another system) can subsequently perform one or more operations using the data representing the 3D information. For example, the system can use the data to track one or more objects within the environment, determine 2D information associated with the one or more objects (e.g., determine a 2D bounding shape associated with an object within the image), determine additional 3D information associated with the one or more objects (e.g., determine a 3D bounding shape associated with an object as represented by the image), determine one or more classifications associated with the one or more objects, and / or perform any other operations. While these are just a few examples of additional processes that the system can perform using the data, in other examples, the system can perform additional and / or alternative processes using the data.
[0029] As described herein, by performing these processes to determine a 3D position associated with an object, the system may not need to first determine a 2D position associated with the object within the image and / or may not need to project the 2D position in the image into a 3D space associated with the environment. Thus, the systems described herein can more accurately determine the 3D position associated with an object by eliminating detection and / or projection errors. Furthermore, these improvements may be more prevalent in certain environments, such as complex environments that include numerous cameras monitoring the interior of the environment and / or environments that include large areas that are occluded by at least some of the cameras. For example, by first determining multi-view image features associated with the image and then using the multi-view image features to determine BEV features associated with the environment, the system is able to process image data from numerous cameras while also capturing features associated with objects that may be occluded by some cameras but visible to other cameras.
[0030] The systems and methods described herein may be used by, but are not limited to, non-autonomous vehicles or machines, semi-autonomous vehicles or machines (e.g., in one or more adaptive driver assistance systems (ADAS)), autonomous vehicles or machines, manned and unmanned robots or robotic platforms, warehouse vehicles, off-road vehicles, vehicles coupled to one or more trailers, airships, boats, space shuttles, emergency response vehicles, motorcycles, electric or motorized bicycles, aircraft, engineering vehicles, underwater vehicles, drones, and / or other vehicle types. Furthermore, the systems and methods described herein may be used for a variety of purposes, such as, but not limited to, machine control, machine motion, machine driving, synthetic data generation, model training, perception, augmented reality, virtual reality, mixed reality, robotics, safety and surveillance, simulation and digital twins, autonomous or semi-autonomous machine applications, deep learning, environmental simulation, object or participant simulation and / or digital twins, data center processing, conversational AI, light transport simulation (e.g., ray tracing, path tracing, etc.), collaborative content creation of 3D assets, cloud computing, and / or any other suitable application.
[0031] The disclosed embodiments can be included in a variety of different systems, such as automotive systems (e.g., control systems for autonomous or semi-autonomous machines, perception systems for autonomous or semi-autonomous machines), systems implemented using robots, aviation systems, medical systems, boating systems, smart area monitoring systems, systems for performing deep learning operations, systems for performing simulation operations, systems for performing digital twin operations, systems implemented using edge devices, systems implementing large language models (LLMs), systems implementing small language models (SLMs), systems implementing visual language models (VLMs), systems including one or more virtual machines (VMs), systems for performing synthetic data generation operations, systems implemented at least in part in a data center, systems for performing conversational AI operations, systems for performing light transport simulations, systems for performing collaborative content creation of 3D assets, systems for performing generative AI operations, systems implemented at least in part using cloud computing resources, and / or other types of systems.
[0032] refer to Figure 1A , Figure 1AAn example of a first process for performing three-dimensional multi-camera perception to determine 3D information related to an object according to some embodiments of the present disclosure is shown. It should be understood that this arrangement and other arrangements described herein are presented as examples only. In addition to the arrangements and elements shown, other arrangements and elements (e.g., machines, interfaces, functions, sequences, functional groupings, etc.) can also be used, and some elements can be omitted entirely. In addition, many of the elements described herein are functional entities that can be implemented as discrete or distributed components or in combination with other components, and can be implemented in any suitable combination and position. The various functions performed by the entities described herein can be performed by hardware, firmware, and / or software. For example, the various functions can be implemented by a processor executing instructions stored in a memory.
[0033] Process 100 may include cameras 102(1)-(M) (also referred to as "camera 102" or "cameras 102") generating image data 104(1)-(M) (also referred to as "image data 104"). As described herein, in some examples, cameras 102 may be located throughout an environment, such as a retail environment, a warehouse environment, an office environment, an educational environment, an outdoor environment, and / or any other type of environment. Furthermore, cameras 102 may be located throughout the environment such that cameras 102 include fields of view (FOVs) that capture different portions of the environment (e.g., an interior of the environment). For example, first camera 102(1) may include a first FOV that includes a first portion of the environment, second camera 102(2) may include a second FOV that includes a second portion of the environment, third camera 102(3) may include a third FOV that includes a third portion of the environment, and so on, until finally camera 102(M) includes a final FOV that includes a final portion of the environment.
[0034] Furthermore, in some examples, at least some of the cameras 102 may include overlapping FOVs within the environment. For example, a first FOV of a first camera 102(1) may at least partially overlap with a second FOV of a second camera 102(2), such that the first camera 102(1) and the second camera 102(2) capture similar portions of the environment. Furthermore, in some examples, such as when the cameras 102 include overlapping FOVs, a portion of the environment may be occluded by one of the cameras 102 but still be visible by the other camera 102. For example, a static object located within the environment may occlude a portion of the first FOV of the first camera 102(1), such that a portion of the environment is not visible to the first camera 102(1), but the portion of the environment is still visible to the second FOV of the second camera 102(2).
[0035] For example, Figure 2An example of an environment 202 according to some embodiments of the present disclosure is shown that includes cameras 204(1)-(4) (also referred to singly as "camera 204" or plurally as "cameras 204") located at various locations. As shown, a first camera 204(1) (which may represent the first camera 102(1)) may be located at a first location and include a first FOV 206(1) of the environment 202, a second camera 204(2) (which may represent the second camera 102(2)) may be located at a second location and include a second FOV 206(2) of the environment 202, a third camera 204(3) (which may represent the third camera 102(3)) may be located at a third location and include a third FOV 206(3) of the environment 202, and a fourth camera 204(4) (which may represent the final camera 102(M)) may be located at a fourth location and include a fourth FOV 206(4) of the environment 202.
[0036] In some examples, one or more of the FOVs 206(1)-(4) may at least partially overlap with each other. For example, each of the FOVs 206(1)-(4) may include at least a portion of the environment 202, such as a central portion of the environment 202. However, the FOVs 206(1)-(2) of the cameras 204(1)-(2) may be partially obscured by the obstacle 208 (e.g., a static object), such that the cameras 204(1)-(2) may not be able to capture a portion of the environment 202 that is to the right of the obstacle 208. Additionally, the FOVs 206(3)-(4) of the cameras 204(3)-(4) may be partially obscured by the obstacle 208, such that the cameras 204(3)-(4) may not be able to capture a portion of the environment 202 that is to the left of the obstacle 208.
[0037] In some examples, environment 202 may include a complex environment, such as a warehouse environment, a retail environment, an educational environment, and / or any other type of environment. Figure 2 The environment 202 represented in the example may include an interior within four sides of the environment 202, wherein the FOV 206 (1)-(4) of the camera 204 includes different portions of the interior of the environment 202. Furthermore, in some examples, the camera 204 may be stationary within the environment 202, such that the position of the camera 204 does not change and / or does not substantially change. Thus, after calibrating the camera 204 (which will be described in greater detail herein), the calibration information may associate the same 3D point within the environment 202 with a 2D point associated with images generated using the camera 204 over a period of time.
[0038] Figures 3A to 3B 2 shows an example of camera 204 generating image data representing environment 202 according to some embodiments of the present disclosure. Figure 3AAs shown in the example of FIG, at least a first object 302(1) and a second object 302(2) can be located within the environment 202. Figure 3A The example of shows objects 302(1)-(2) as including people, but in other examples, the objects may include any other type of object, such as a machine, a robot, an animal, etc. Then, the first camera 204(1) may generate first image data (e.g., first image data 104(1)), the second camera 204(2) may generate second image data (e.g., second image data 104(2)), the third camera 204(3) may generate third image data (e.g., third image data 104(3)), and the fourth camera 204(4) may generate fourth image data (e.g., final image data 104(M)).
[0039] For example, Figure 3B As shown in the example of , first image data generated using first camera 204(1) may represent first image 304(1), wherein first image 304(1) depicts at least first object 302(1). However, second object 302(2) may be obscured by obstacle 208 in first image 304(1). Furthermore, second image data generated using second camera 204(2) may represent second image 304(2), wherein second image 304(2) also depicts first object 302(1). However, similar to first image 304(1), second object 302(2) may be obscured by obstacle 208 in second image 304(2). Furthermore, third image data generated using third camera 204(3) may represent third image 304(3), wherein third image 304(3) depicts at least second object 302(2). However, first object 302(1) may be obscured by obstacle 208 in third image 304(3). Additionally, fourth image data generated using fourth camera 204(4) may represent fourth image 304(4), where fourth image 304(4) also depicts second object 302(2). However, similar to third image 304(3), first object 302(1) may be obscured by obstacle 208 in fourth image 304(4).
[0040] review Figure 1AAs an example, process 100 may include processing image data 104 using feature extractors 106(1)-(M) (also referred to singly as "feature extractor 106" or plurally as "feature extractors 106"). As described herein, feature extractor 106 may include, but is not limited to, one or more backbones, one or more neural network layers, one or more neural networks, one or more encoders, one or more decoders, and / or any other type of processing component configured to perform one or more processes described herein. As shown, based at least on the processing of image data 104, process 100 may include feature extractor 106 generating feature data 108(1)-(M) (also referred to as "feature data 108"). For example, the first feature extractor 106(1) may generate first feature data 108(1) representing a first feature associated with a first image represented by the first image data 104(1), the second feature extractor 106(2) may generate second feature data 108(2) representing a second feature associated with a second image represented by the second image data 104(2), the third feature extractor 106(3) may generate third feature data 108(3) representing a third feature associated with a third image represented by the third image data 104(3), and the last feature extractor 106(M) may generate last feature data 108(M) representing a last feature associated with a last image represented by the last image data 104(M).
[0041] The process 100 may then include one or more spatiotemporal transformers 110 processing the feature data 108 and / or calibration data 112. As described in more detail herein, and as Figure 1B As shown in the example of , the spatiotemporal transformer 110 may include and / or use one or more spatial encoders, one or more temporal encoders, one or more decoders, one or more layers of one or more neural networks, one or more neural networks, and / or any other type of processing component configured to perform at least a portion of the processes described herein. Furthermore, the calibration data 112 for the camera 102 may associate 3D coordinates (e.g., 3D points) within the environment with 2D coordinates (e.g., 2D points) associated with an image generated using the camera 102. For example, the calibration data 112 for the camera 102 may include a matrix, such as a 3x4 projection matrix, that associates 3D points within the environment with 2D points associated with the image.
[0042] As an example of a matrix, camera perspective projection can be written as a linear mapping between homogenous coordinates using:
[0043]
[0044] In equation (1), the 3x4 projection matrix represents the mapping from 3D coordinates to 2D coordinates.
[0045] Additionally, for images, intrinsic and / or extrinsic parameters can be added to the projection matrix. For example, a camera calibration matrix might include:
[0046]
[0047] In equation (2), α u =fk u And α v =-fk v , where k is the pixel, α u is the scaling ratio of the image in the x-coordinate direction, α v is the scaling ratio of the image in the y-coordinate direction, and (u0, v0) is the principal point where the optical axis intersects the image plane.
[0048] Therefore, the Euclidean transformation between camera and world coordinates may include X c =RX w +T, where:
[0049]
[0050] Furthermore, when these three matrices are concatenated, the following is formed:
[0051]
[0052] Then, equation (4) defines the 3x4 projection matrix from Euclidean 3-space to the image as follows:
[0053]
[0054] While this is just one example of a projection matrix that may be represented by calibration data 112 , in other examples, calibration data 112 may represent one or more additional and / or alternative matrices that relate 3D points within the environment to 2D points associated with images generated using camera 102 .
[0055] Furthermore, in some examples, calibration data 112 may represent additional information associated with camera 102. For example, as described herein, at least a portion of one or more cameras 102 may be occluded by one or more objects (e.g., one or more static objects) located within an environment. Thus, in some examples, if a portion of camera 102 is occluded, calibration data 112 may represent the occluded portion of the environment. For example, calibration data 112 may represent a 2D location within an image generated using the occluded camera 102. In such examples, the 2D location may include, but is not limited to, a 2D point (e.g., a 2D pixel) location, a boundary shape, and / or any other type of 2D location.
[0056] Then, process 100 may include a spatiotemporal transformer 110 generating and / or outputting object data 114 that represents 3D information associated with one or more objects located in the environment. For example, as described herein, object data 114 may represent at least one or more 3D positions associated with one or more objects located in the environment. In some examples, the 3D position may be represented using one or more parameters, such as three parameters representing the scale of a bounding shape (e.g., length, width, and height), three parameters representing the center position of a bounding shape (e.g., x-coordinate position, y-coordinate position, and z-coordinate position), two parameters representing the yaw of an object (e.g., cosine angle and sine angle), and / or two parameters representing the velocity of an object (e.g., velocity in the x-direction and velocity in the y-direction).
[0057] For example, Figure 4 1 shows an example of 3D information that can be output according to some embodiments of the present disclosure. Figure 4 As shown in the example of , the output data (e.g., object data 114) can represent a top-down image 402 (e.g., a BEV image) of the environment 202, which includes at least a first 3D position associated with a first object 302(1), wherein the first 3D position includes a first bounding shape 404(1) (e.g., a 3D bounding shape in which the third dimension is oriented toward the ground plane), and a second 3D position associated with a second object 302(2), wherein the second 3D position includes a second bounding shape 404(2) (e.g., a 3D bounding shape in which the third dimension is oriented toward the ground plane). Although Figure 4 The example shows that the 3D position includes a bounding shape, in other examples, the output data may represent any other type of 3D position associated with one or more of the objects 302(1)-(2). In addition, although Figure 4 The example shows that the output data represents a top-down image 402 with an indicated 3D position, but in other examples, the output data may only represent the 3D position (eg, 3D coordinates associated with the 3D position).
[0058] review Figure 1A In some examples, process 100 can continue to repeat as camera 102 continues to generate new image data 104 representing the environment. For example, process 100 can repeat such that object data 114 is generated for every frame generated using camera 102, every other frame generated using camera 102, every fourth frame generated using camera 102, and / or using any other interval associated with frames generated using camera 102.
[0059] like Figure 1A As further illustrated by the example of , process 100 may include one or more processing components 116 performing one or more operations using object data 114. For example, as described herein, the operations may include, but are not limited to, tracking one or more objects within an environment, determining 2D information associated with one or more objects (e.g., determining a 2D bounding shape associated with an object represented by an image), determining additional 3D information associated with one or more objects (e.g., determining a 3D bounding shape associated with an image), determining one or more classifications associated with one or more objects, and / or performing any other operations.
[0060] In further detail, processing component 116 may be configured to track one or more objects located within an environment. For example, processing component 116 may use object data 114 (e.g., 3D information) and at least a portion of calibration data 112 (e.g., camera calibration information) to project a 3D location (e.g., a 3D bounding shape) associated with an object to a 2D location (e.g., a 2D bounding shape, such as a 2D bounding box) within an image represented by image data 104. In some examples, to perform the projection, processing component 116 may first determine image pixels associated with the 3D location. Processing component 114 may then use the pixels to generate a 2D bounding shape within the image. Furthermore, processing component 114 may use one or more techniques to associate the 3D bounding shape with the 2D bounding shape, such as by using one or more algorithms (e.g., a Hungarian algorithm, etc.).
[0061] The processing component 116 can then use the 2D position to track the object. For example, the processing component 116 can use one or more machine learning models to track the 2D position of the object between images over a period of time (e.g., 10 seconds, 30 seconds, 1 minute, 5 minutes, etc.). In some examples, the machine learning model can track the 2D position based at least on extracting appearance features associated with the 2D position (e.g., 2D bounding shape) and then using the extracted appearance features to track the object (e.g., determining that the embeddings are similar and / or related between the images). In addition, in some examples, the processing component 116 can track the object in a 3D environment, for example by again associating the 3D position (e.g., 3D bounding shape) with the 2D position (e.g., 2D bounding shape).
[0062] As described herein, the feature extractor 106 and / or the spatiotemporal transformer 110 may use any type of processing components to perform one or more of the processes described herein. For example, Figure 1B An example of a second process 118 for performing three-dimensional multi-camera perception to determine 3D information associated with an object according to some embodiments of the present disclosure is shown. It should be understood that this arrangement and other arrangements described herein are presented as examples only. In addition to or in place of the arrangements and elements shown, other arrangements and elements (e.g., machines, interfaces, functions, sequences, functional groupings, etc.) may be used, and some elements may be omitted entirely. In addition, many of the elements described herein are functional entities that may be implemented as discrete or distributed components or with other components and in any suitable combination and location. The various functions performed by the entities described herein may be performed by hardware, firmware, and / or software. For example, the various functions may be implemented by a processor executing instructions stored in a memory.
[0063] As shown, process 118 may include one or more backbones 120 processing image data representing at least images 122(1)-(N) (also referred to as "image 122" in the singular or "images 122" in the plural). In some examples, image 122 may be generated from Figure 1A 4. For example, first image 122(1) may be represented by first image data 104(1), second image 122(2) may be represented by second image data 104(2), third image 122(3) may be represented by third image data 104(3), and / or final image 122(N) may be represented by final image data 104(M). In other words, each image 122 may be generated using a respective camera 102 located within the environment. Furthermore, in some examples, images 122 may be generated at substantially the same time such that images 122 depict the environment in a similar state (e.g., objects in the environment are located at similar locations within the environment and / or are positioned in similar orientations).
[0064] exist Figure 1B In the example of , the backbone 120 can include any type of backbone network, such as a residual neural network (ResNet), VoVNet, ImageNet, Deep Layer Attack (DLA), a transformer-based base backbone (e.g., DINO, Swin, etc.), a convolutional neural network, and / or any other type of backbone network. Thus, the process 118 can include the backbone 120 processing the image 122 and generating feature data representing multi-view image features 124 based at least on the processing. For example, the backbone 120 can generate first feature data representing a first feature 124 associated with the first image 122(1), second feature data representing a second feature 124 associated with the second image 122(2), third feature data representing a third feature 124 associated with the third image 122(3), and last feature data representing a last feature 124 associated with the last image 122(N).
[0065] The process 118 may then include one or more spatial encoders 126 processing the multi-view image features, the calibration data 112, and / or the query data 128. In some examples, the query data 128 may represent one or more queries (e.g., one or more BEV queries) associated with one or more features (e.g., one or more features for which 3D information is being generated). The process 118 may then include, based at least on the processing, the spatial encoder 126 generating and / or outputting current BEV features 130(1) associated with the multi-view image features 124. As described herein, in some examples, the spatial encoder 126 may project the multi-view image features 124 using the calibration data 112 to generate the current BEV features 130(1).
[0066] For example, the spatial encoder 126 can be configured to sample 3D reference points associated with the environment and then project the 3D reference points onto the 2D views of the image 122 using the calibration data 112. In some examples, for a query represented by the query data 128, the projected 2D points may fall at one or more 2D points on one or more 2D views, which may be referred to as reference points. The spatial encoder 126 can then sample features 124 from the reference points and / or one or more points at least partially surrounding the reference points and output the sampled features 124 as spatial cross-attentions. In some examples, when outputting the sampled features 124 as spatial cross-attentions, the spatial encoder 126 can perform a weighted summation on the sampled features.
[0067] In some examples, when sampling features, the spatial encoder 126 can perform various techniques to obtain reference points in the image 122. For example, the spatial encoder 126 can lift each query Q on the BEV plane to a pillar-like query, sampling N from the pillars. ref 3D reference points are then projected onto 2D views. For a BEV query, the projected 2D point may fall on a certain view but not on other views. Therefore, the hit view can be called V hit Afterwards, the spatial encoder 126 can treat the 2D point as a query Q p reference points and hit views around these reference points V hit Finally, the spatial encoder 126 may perform weighted summation on the sampled feature points as the output of spatial cross attention, where the process of spatial cross attention (SCA) can be expressed as follows:
[0068]
[0069] In Equation (7), i can index the camera view, j can index the reference point, and N ref Can be the overall reference point for BEV enquiries. In addition, can be the feature in the i-th camera view. Therefore, for each query, the projection function P(p,i,j) can be used to obtain the j-th reference point on the i-th view image.
[0070] In 3D space, an object at (xv, y′) may appear at height z′ on the z-axis. Therefore, in some examples, a set of anchor heights can be predefined. To capture objects that appear at different heights. Therefore, for each query, a 3D reference point can be obtained Finally, the spatial encoder 126 can project the 3D points to different image views through the camera's projection matrix, which can be written as:
[0071] P(p,i,j)=(x ij ,y ij )
[0072] where z ij ·[x ij y ij 1] T =T i [x′y i z′ j 1] T (8)
[0073] In equation (8), P(p,i,j) is the sum of the j-th 3D point (x′,y′,zj ) is projected onto the 2D point on the i-th view, and is the known projection matrix of the ith camera (e.g., from above).
[0074] The process 118 may then include one or more temporal encoders 132 processing the current BEV feature 130(1) and one or more previous BEV features 130(0) associated with one or more previous time instances (e.g., time instances associated with images previously generated using a camera). For example, in some examples, the temporal encoder 132 may be configured to associate the BEV features 130(1)-(0) between different time instances. In some examples, the temporal encoder 132 may perform these associations using one or more techniques, such as modeling the temporal connections between the BEV features 130(1)-(0) via one or more self-attention layers.
[0075] Additionally or alternatively, in some examples, the temporal encoder 132 may use a warping and concatenation strategy to perform temporal encoding. For example, given BEV features at a different frame (e.g., previous BEV features 130(0)), the temporal encoder 132 may warp the BEV features to the current frame (e.g., current BEV features 130(1)) based on a reference frame transformation matrix between the previous frame and the current frame. Next, the temporal encoder 132 may concatenate the previous BEV features 130(0) with the current BEV features 130(1) along the channel dimension to perform dimensionality reduction using a residual block. While these are just a few example techniques for how the temporal encoder 132 may perform temporal encoding using the BEV features 130(1)-(0), in other examples, the temporal encoder 132 may use any other techniques.
[0076] Then, the process 118 may include one or more decoders 134 processing the output from the temporal encoder 132 and / or the query data 136. As described herein, in some examples, the query data 136 may represent one or more queries indicating where a target object may be located, wherein the one or more queries may be learned through training. In addition, the decoder 134 may include any type of decoder, such as a detector transformer (DETR) decoder, a binary decoder, an image decoder, and / or any other type of decoder.
[0077] Process 118 may then include, based at least on processing the output from temporal encoder 132 and / or query data 136, decoder 134 generating and / or outputting object data 138 (which may be similar to and / or representative of object data 114), object data 138 representing 3D information associated with one or more objects. For example, as described herein, object data 114 may represent at least one or more 3D positions associated with objects located in an environment.
[0078] With further detail regarding the decoder 134, in some examples, the decoder 134 can include one or more self-attention layers and one or more criss-cross attention layers, which can be alternately stacked on top of each other. The criss-cross attention layer can then take as input (1) the query features to generate sampling offsets and attention weights, (2) 2D points on the value features as sampling references for each query, and (3) the BEV features output by the temporal encoder 132. The decoder 134 can then process the inputs and, based at least on the processing, determine the center of the projected box on the BEV plane to use as a per-image reference point, where the per-image reference point can indicate a possible location of an object in the BEV plane. Thus, the decoder 134 can use these per-image reference points to determine 3D information associated with the object.
[0079] As described herein, in some examples, additional processing can be performed on the 3D information, such as processing to classify objects. Figure 5 An example of a process for classifying an object using 3D information associated with the object is shown in accordance with some embodiments of the present disclosure. As shown, process 500 may initially include performing at least a portion of process 118 to generate object data 138 representing a 3D location associated with the object. However, process 500 may subsequently include one or more projection components 502 further processing object data 138 to generate 2D data 504 representing 2D information associated with the object. For example, in some examples, projection component 502 may project the 3D location associated with the object to a 2D location within image 122, where 2D data 504 represents the 2D location. As described herein, in some examples, the 2D location may include a bounding shape (e.g., a bounding box, etc.) associated with the object.
[0080] Process 500 may also include one or more detection components 506 processing at least a portion of the image data representing image 122. As described herein, detection component 506 may include and use one or more machine learning models, one or more neural networks, one or more algorithms, and / or any other type of processing component configured to perform one or more processes described herein. Based at least on the processing, process 500 may include detection component 506 generating and / or outputting detection data 508 associated with the object. For example, detection data 508 may represent at least a location associated with the object within image 122, an identifier associated with the object, a classification associated with the object, and / or any other information.
[0081] As described herein, in some examples, an identifier can include, but is not limited to, a numeric identifier, an alphabetic identifier, an alphanumeric identifier, and / or any other type of identifier that can be used to identify an object. Additionally, in some examples, detection component 506 can initially assign an identifier to an object when the object is first detected (e.g., in image 122). Detection component 506 can then continue to assign the same identifier to an object when the object is detected in additional images 122 (e.g., images 122 representing different views of the object and / or images 122 generated later using a camera).
[0082] The process 500 may then include one or more association components 510 associating the identifier with the 2D location of the object within the image using the 2D data 504 and the detection data 508. Figure 1B Process 118 (and / or Figure 1A The process 100 of FIG. 1 may first be used to determine the precise location of an object within the image 122 , and then post-processing may be used to track identifiers associated with the object within the image 122 .
[0083] Figure 6 A data flow diagram of a process for training one or more networks to perform three-dimensional multi-camera perception according to some embodiments of the present disclosure is shown. As shown in process 600, training can include inputting training images 602 into backbone 120. In some examples, training images 602 can be generated using multiple cameras located at various locations within one or more environments and / or can be generated over a period of time. For example, training images 602 can correspond to video generated by a camera over one second, five seconds, ten seconds, twenty seconds, one minute, and / or any other time period. In some examples, training images 602 can be synthetically generated (e.g., generated from a computer model or rendering), realistically generated (e.g., designed and generated from real-world data (e.g., image data generated using a camera)), machine-automated, and / or a combination thereof.
[0084] Training may also include using ground truth data 604 corresponding to the training image 602. Ground truth data 604 may include annotations, labels, masks, and the like. For example, as shown, ground truth data 604 may include at least 3D information 606 associated with the object represented by the training image 602. Ground truth data 604 may be synthetically generated (e.g., generated from a computer model or rendering), realistically generated (e.g., designed and generated from real-world data), machine-automated (e.g., using feature analysis and learning to extract features from the data and then generate labels), manually annotated (e.g., annotator or annotation expert defines the location of the labels), and / or a combination thereof. In some examples, for each instance of the training image 602, there may be corresponding ground truth data 604.
[0085] like Figure 6 As further shown in the example of , one or more training engines 608 may use one or more loss functions to measure the loss (e.g., error) of output data 610 compared to ground truth data 604. In some examples, output data 610 may be similar to and / or include object data 114 and / or object data 138. For example, output data 610 may represent at least 3D information associated with an object, where the 3D information is determined using one or more processes described herein. In some examples, training engine 608 may use any type of loss function, such as cross entropy loss, mean squared error, mean absolute error, mean deviation error, and / or other loss function types. In some examples, different outputs may have different loss functions. In such examples, the loss functions may be combined to form an overall loss (where one or more losses may be weighted), and the overall loss may be used to train backbone network 120, spatial encoder 126, temporal encoder 132, and / or decoder 134. Furthermore, in some examples, training engine 608 can update query data 128 and / or query data 136 based on at least the total loss. In any example, a backward pass can be performed to recursively calculate the gradient of the loss function with respect to the training parameters. In some examples, these gradients can be calculated using weights and / or biases.
[0086] Now refer to Figures 7 to 9, each block of methods 700, 800, and 900 described herein comprises a computational process that may be performed using any combination of hardware, firmware, and / or software. For example, the various functions may be implemented by a processor executing instructions stored in a memory. Methods 700, 800, and 900 may also be embodied as computer-usable instructions stored on a computer storage medium. Methods 700, 800, and 900 may be provided by a standalone application, a service, or a hosted service (standalone or in combination with another hosted service), or a plug-in to another product, to name a few. Furthermore, methods 700, 800, and 900 may be combined, by way of example, with Figure 1A to Figure 1B However, these methods 700, 800, and 900 may additionally or alternatively be performed by any one system or any combination of systems, including but not limited to the systems described herein.
[0087] Figure 7 A flow chart of a method 700 for performing three-dimensional multi-camera perception associated with an environment, according to some embodiments of the present disclosure, is shown. Method 700, at block B702, may include acquiring image data generated using cameras located within the environment, the image data representing an image. For example, cameras 102 located throughout the environment may generate image data 104 representing image 122. As described herein, in some examples, cameras 102 may be located at different locations within the environment and / or may include different orientations within the environment, such that cameras 102 include different fields of view (FOVs) of the environment. Furthermore, in some examples, at least some of the FOVs of cameras 102 may overlap with each other.
[0088] At block B704, method 700 may include processing the image data based at least on one or more backbones to determine a first feature associated with the image. For example, backbone 120 (and / or feature extractor 106) may process image data 104 representing image 122. Based at least on this processing, backbone 120 may generate feature data 108 representing multi-view image features 124 associated with image 122. For example, in some examples, backbone 120 may generate corresponding feature data 108 representing one or more image features for each image 122.
[0089] The method 700 may include, at block B706, determining a second feature associated with the environment based at least on the first feature and calibration data that associates three-dimensional (3D) coordinates associated with the environment with two-dimensional (2D) coordinates associated with the image. For example, the spatial encoder 126 (and / or the spatiotemporal transformer 110) may process the feature data 108 representing the multi-view image feature 124 and the calibration data 112 to determine the current BEV feature 130(1) associated with the image 122. As described herein, in some examples, the spatial encoder 126 may process additional data (e.g., query data 128) to determine the current BEV feature 130(1).
[0090] The method 700 may include, at block B708, determining one or more 3D locations associated with one or more objects located within the environment based at least on the second feature. For example, the decoder 134 (and / or the spatiotemporal transformer 110) may determine the 3D location associated with the object based at least on the current BEV feature 130(1). In some examples, and as described herein, the temporal encoder 132 (and / or the spatiotemporal transformer 110) may initially associate (e.g., concatenate, combine, etc.) the current BEV feature 130(1) with one or more previous BEV features 130(0). In such examples, the decoder 134 may then process the associated BEV features 130(1)-(0) to determine the 3D location associated with the object.
[0091] Method 700 may include, at block B710, performing one or more operations based at least on the one or more 3D positions. For example, processing component 116 may perform operations using object data 114 representing 3D positions associated with objects. As described herein, the operations may include, but are not limited to, tracking one or more objects within an environment, determining 2D information associated with one or more objects (e.g., determining a 2D bounding shape associated with an object within an image), determining additional 3D information associated with one or more objects (e.g., determining a 3D bounding shape associated with an object represented by an image), determining one or more classifications associated with one or more objects, and / or any other operations.
[0092] Figure 8A flowchart of another method 800 for performing three-dimensional multi-camera perception associated with an environment, according to some embodiments of the present disclosure, is shown. Method 800 may include, at block B802, determining multi-view image features based at least on one or more feature extractors processing image data generated using cameras located in the environment. For example, backbone 120 (e.g., feature extractor 106) may process image data 104 representing image 122, where image data 104 was generated using camera 102 located in the environment. Based at least on this processing, backbone 120 may generate feature data 108 representing multi-view image features 124 associated with image 122. For example, in some examples, backbone 120 may generate corresponding feature data 108 representing one or more image features for each image 122.
[0093] The method 800 may include, at block B804, determining a current bird's eye view (BEV) feature based at least on processing the multi-view image features and calibration data associated with the camera by one or more spatial encoders. For example, the spatial encoder 126 (and / or the spatiotemporal transformer 110) may process the feature data 108 representing the multi-view image features 124 and the calibration data 112 to determine the current BEV feature 130(1) associated with the image 122. As described herein, in some examples, the spatial encoder 126 may process additional data (e.g., query data 128) to determine the current BEV feature 130(1).
[0094] The method 800 may include, at block B806, determining a fused feature based at least on processing the current BEV feature and the one or more previous BEV features by the one or more temporal encoders. For example, the temporal encoder 132 (and / or the spatiotemporal transformer 110) may process the current BEV feature 130(1) and the one or more previous BEV features 130(0) associated with the one or more previous time instances. Based at least on this processing, the temporal encoder 132 may generate a fused feature. As described herein, in some examples, the temporal encoder 132 may perform any technique to determine the fused feature, such as by associating the BEV features 130(1)-(0) with each other.
[0095] At block B808, method 800 may include determining three-dimensional information associated with one or more objects located within the environment based at least on one or more decoder-processed fused features. For example, decoder 134 (and / or spatiotemporal transformer 110) may determine 3D information associated with the objects based at least on the fused features. As described herein, the 3D information may include at least one or more 3D locations associated with the objects, such as one or more 3D boundary shapes.
[0096] Figure 9A flowchart of a method 900 for determining a bird's-eye view feature based at least on multi-view image features is shown in accordance with some embodiments of the present disclosure. Method 900 may include, at block B902, obtaining first feature data representing multi-view image features associated with an image generated using a camera. For example, a feature extractor 106 (e.g., backbone 120) may process image data 104 representing an image 122, where the image data 104 was generated using a camera 102 located in an environment. Based at least on the processing, the feature extractor 106 may generate feature data 108 representing multi-view image features 124 associated with the image 122. For example, in some examples, the feature extractor 106 may generate corresponding feature data 108 representing one or more image features for each image 122.
[0097] Method 900 may include, at block B904, determining that a three-dimensional (3D) point associated with the environment corresponds to a two-dimensional (2D) point associated with the image based at least on calibration data associated with the camera. For example, calibration data 112 for camera 102 may associate a 3D coordinate within the environment with a 2D coordinate associated with image 122. In this manner, calibration data 112 may be used to project a 3D point within the environment to a 2D point associated with image 122. In some examples, because camera 102 may be stationary within the environment, calibration data 112 may associate the same 3D point within the environment with the same 2D point within images generated using camera 102 over a period of time.
[0098] Method 900 may include, at block B906, determining that the 2D point is associated with one or more features of the multi-view image features. For example, spatiotemporal transformer 110 (eg, spatial encoder 126) may determine that a feature is associated with the 2D point within image 122 based at least on the projection.
[0099] Method 900 may include, at block B 908, generating second feature data representing one or more features associated with environment 908. For example, spatiotemporal transformer 110 (e.g., spatial encoder 126) may generate second feature data representing features, where the features correspond to BEV features 130(1)-(0). The second feature data may then be used to determine 3D information associated with one or more objects located within the environment, as described herein.
[0100] Figure 10An example architecture is shown in which one or more processes described herein can be performed according to some embodiments of the present disclosure. As shown, the architecture can include at least one or more systems 1002 (which can be similar to and / or represent example computing device 1100 and / or example data center 1200) and environment 1004 (which can represent and / or include environment 202), which includes cameras 102 located at various locations. In addition, system 1002 can include at least one or more processors 1006 (which can be similar to, and / or include CPU 1106 and / or GPU 1108), one or more network interfaces 1008 (which can be similar to, and / or include communication interface 1110), and memory 1010 (which can be similar to, and / or include memory 1104). In some examples, system 1002 can be remote from environment 1004, for example, by including one or more edge devices and communicating with camera 102.
[0101] As further shown, the memory 1010 may include at least the feature extractor 106, the spatiotemporal transformer 110, the processing component 116, the backbone 120, the spatial encoder 126, the temporal encoder 132, and / or the decoder 134. Thus, the system 1002 may be configured to perform at least a portion of the processes described herein, such as the process 100 and / or the process 118, to perform three-dimensional multi-camera perception. Figure 10 The example shows that the feature extractor 106, the spatiotemporal transformer 110, the processing component 116, the backbone 120, the spatial encoder 126, the temporal encoder 132 and / or the decoder 134 are stored in the memory 1010, in other examples, one or more of the feature extractor 106, the spatiotemporal transformer 110, the processing component 116, the backbone 120, the spatial encoder 126, the temporal encoder 132 and / or the decoder 134 may include hardware components located outside the memory 1010.
[0102] Example computing device
[0103] Figure 111 is a block diagram of an example computing device 1100 suitable for implementing some embodiments of the present disclosure. Computing device 1100 may include an interconnect system 1102 that directly or indirectly couples the following devices: memory 1104, one or more central processing units (CPUs) 1106, one or more graphics processing units (GPUs) 1108, a communication interface 1110, input / output (I / O) ports 1112, I / O components 1114, a power supply 1116, one or more presentation components 1118 (e.g., display(s)), and one or more logic units 1120. In at least one embodiment, computing device(s) 1100 may include one or more virtual machines (VMs), and / or any of its components may include virtual components (e.g., virtual hardware components). For non-limiting examples, one or more of GPUs 1108 may include one or more vGPUs, one or more of CPUs 1106 may include one or more vCPUs, and / or one or more of logic units 1120 may include one or more virtual logic units. As such, computing device(s) 1100 may include discrete components (eg, a full GPU dedicated to computing device 1100 ), virtual components (eg, a portion of a GPU dedicated to computing device 1100 ), or a combination thereof.
[0104] although Figure 11 The various blocks of are shown as being connected with lines via interconnect system 1102, but this is not intended to be limiting and is provided merely for clarity. For example, in some embodiments, presentation component 1118 (such as a display device) may be considered to be I / O component 1114 (e.g., if the display is a touch screen). As another example, CPU 1106 and / or GPU 1108 may include memory (e.g., memory 1104 may represent a storage device in addition to the memory of GPU 1108, CPU 1106, and / or other components). In other words, Figure 11 The computing devices referred to herein are illustrative only. No distinction is made between categories such as "workstation," "server," "laptop," "desktop," "tablet," "client device," "mobile device," "handheld device," "game console," "electronic control unit (ECU)," "virtual reality system," and / or other device or system types, as all are contemplated. Figure 11 within the range of computing devices.
[0105] Interconnect system 1102 can represent one or more links or buses, such as an address bus, a data bus, a control bus, or a combination thereof. Interconnect system 1102 can include one or more bus or link types, such as an Industry Standard Architecture (ISA) bus, an Extended Industry Standard Architecture (EISA) bus, a Video Electronics Standards Association (VESA) bus, a Peripheral Component Interconnect (PCI) bus, a Peripheral Component Interconnect Express (PCIe) bus, and / or another type of bus or link. In some embodiments, there is a direct connection between components. For example, CPU 1106 can be directly connected to memory 1104. Further, CPU 1106 can be directly connected to GPU 1108. In the case where there is a direct connection or a point-to-point connection between components, interconnect system 1102 can include a PCIe link to perform the connection. In these examples, it is not necessary to include a PCI bus in computing device 1100.
[0106] Memory 1104 may include any of a variety of computer-readable media. Computer-readable media can be any available media that can be accessed by computing device 1100. Computer-readable media can include volatile and non-volatile media, as well as removable and non-removable media. By way of example and not limitation, computer-readable media can include computer storage media and communication media.
[0107] Computer storage media may include volatile and nonvolatile media and / or removable and non-removable media implemented in any method or technology for storage of information such as computer-readable instructions, data structures, program modules, and / or other data types. For example, memory 1104 may store computer-readable instructions (e.g., representing programs and / or program elements, such as an operating system). Computer storage media may include, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and that can be accessed by the computing device 1100. As used herein, computer storage media does not include the signals themselves.
[0108] Computer storage media can embody computer-readable instructions, data structures, program modules, and / or other data types in a modulated data signal (such as a carrier wave or other transport mechanism), and include any information delivery media. The term "modulated data signal" may refer to a signal that has one or more of its characteristics set or changed in a manner that encodes information in the signal. By way of example, and not limitation, computer storage media may include wired media (such as a wired network or a direct wired connection) and wireless media (such as acoustic, RF, infrared, and other wireless media). Combinations of any of the above should also be included within the scope of computer-readable media.
[0109] The CPU 1106 may be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 1100 to perform one or more of the methods and / or processes described herein. Each of the CPUs 1106 may include one or more cores (e.g., 1, 2, 4, 8, 28, 72, etc.) capable of processing multiple software threads simultaneously. The CPU 1106 may include any type of processor and may include different types of processors (e.g., a processor with fewer cores for mobile devices and a processor with more cores for servers) depending on the type of computing device 1100 implemented. For example, depending on the type of computing device 1100, the processor may be an Advanced RISC Machine (ARM) processor implemented using Reduced Instruction Set Computing (RISC) or an x86 processor implemented using Complex Instruction Set Computing (CISC). In addition to one or more microprocessors or supplementary coprocessors such as math coprocessors, the computing device 1100 may also include one or more CPUs 1106.
[0110] In addition to or in place of the CPU 1106, the GPU(s) 1108 may be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 1100 to perform one or more of the methods and / or processes described herein. One or more of the GPUs 1108 may be integrated GPUs (e.g., with one or more of the CPUs 1106) and / or one or more of the GPUs 1108 may be discrete GPUs. In embodiments, one or more of the GPU(s) 1108 may be a coprocessor to one or more of the CPU(s) 1106. The GPU 1108 may be used by the computing device 1100 to render graphics (e.g., 3D graphics) or perform general-purpose computations. For example, the GPU 1108 may be used for general-purpose computing on a GPU (GPGPU). The GPU 1108 may include hundreds or thousands of cores capable of processing hundreds or thousands of software threads simultaneously. The GPU 1108 may generate pixel data for an output image in response to a rendering command (e.g., a rendering command received from the CPU 1106 via a host interface). The GPU 1108 may include graphics memory, such as display memory, for storing pixel data or any other suitable data, such as GPGPU data. The display memory may be included as part of the memory 1104. The GPU 1108 may include two or more GPUs operating in parallel (e.g., via a link). The link may connect the GPUs directly (e.g., using NVLINK) or may connect the GPUs via a switch (e.g., using NVSwitch). When combined, each GPU 1108 may generate pixel data or GPGPU data for different portions of the output or for different outputs (e.g., a first GPU for a first image and a second GPU for a second image). Each GPU may include its own memory or may share memory with other GPUs.
[0111] In addition to or in lieu of the CPU 1106 and / or GPU 1108, the logic unit(s) 1120 may be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 1100 to perform one or more of the methods and / or processes described herein. In embodiments, the CPU(s) 1106, the GPU(s) 1108, and / or the logic unit(s) 1120 may independently or jointly perform any combination of methods, processes, and / or portions thereof. One or more of the logic units 1120 may be part of and / or integrated into one or more of the CPU 1106 and / or GPU 1108, and / or one or more of the logic units 1120 may be discrete components or otherwise external to the CPU 1106 and / or GPU 1108. In embodiments, one or more of logic units 1120 may be co-processors for one or more of CPU(s) 1106 and / or one or more of GPU(s) 1108 .
[0112] Examples of logic unit 1120 include one or more processing cores and / or components thereof, such as a data processing unit (DPU), a tensor core (TC), a tensor processing unit (TPU), a pixel vision core (PVC), a vision processing unit (VPU), a graphics processing cluster (GPC), a texture processing cluster (TPC), a streaming multiprocessor (SM), a tree traversal unit (TTU), an artificial intelligence accelerator (AIA), a deep learning accelerator (DLA), an arithmetic logic unit (ALU), an application-specific integrated circuit (ASIC), a floating point unit (FPU), an input / output (I / O) element, a peripheral component interconnect (PCI) or a peripheral component interconnect express (PCIe) element, etc.
[0113] The communication interface 1110 may include one or more receivers, transmitters, and / or transceivers that enable the computing device 1100 to communicate with other computing devices via an electronic communication network (including wired and / or wireless communications). The communication interface 1110 may include components and functionality that enable communication over any of a number of different networks, such as wireless networks (e.g., Wi-Fi, Z-Wave, Bluetooth, Bluetooth LE, ZigBee, etc.), wired networks (e.g., via Ethernet or InfiniBand communication), low-power wide-area networks (e.g., LoRaWAN, SigFox, etc.), and / or the Internet. In one or more embodiments, the logic unit 1120 and / or the communication interface 1110 may include one or more data processing units (DPUs) to transmit data received over the network and / or through the interconnect system 1102 directly to one or more GPUs 1108 (e.g., memory of one or more GPUs 1108).
[0114] The I / O ports 1112 can enable the computing device 1100 to be logically coupled to other devices including I / O components 1114, presentation components 1118, and / or other components, some of which can be built into (e.g., integrated into) the computing device 1100. Illustrative I / O components 1114 include a microphone, a mouse, a keyboard, a joystick, a gamepad, a game controller, a satellite dish, a scanner, a printer, a wireless device, and the like. The I / O components 1114 can provide a natural user interface (NUI) that processes air gestures, voice, or other physiological input generated by the user. In some instances, the input can be transmitted to an appropriate network element for further processing. The NUI can implement any combination of voice recognition, stylus recognition, facial recognition, biometric recognition, gesture recognition on and near the screen, air gestures, head and eye tracking, and touch recognition associated with the display of the computing device 1100 (as described in more detail below). Computing device 1100 may include a depth camera, such as a stereo camera system, an infrared camera system, an RGB camera system, touch screen technology, or a combination thereof, for gesture detection and recognition. Additionally, computing device 1100 may include an accelerometer or gyroscope (e.g., as part of an inertial measurement unit (IMU)) to enable motion detection. In some examples, computing device 1100 may use the output of the accelerometer or gyroscope to render immersive augmented reality or virtual reality.
[0115] The power supply 1116 may include a hardwired power supply, a battery power supply, or a combination thereof. The power supply 1116 may provide power to the computing device 1100 to enable the components of the computing device 1100 to operate.
[0116] One or more presentation components 1118 may include a display (e.g., a monitor, a touch screen, a television screen, a head-up display (HUD), other display types, or a combination thereof), speakers, and / or other presentation components. Presentation component 1118 may receive data from other components (e.g., GPU 1108, CPU 1106, DPU, etc.) and output data (e.g., as images, video, sound, etc.).
[0117] Sample Data Center
[0118] Figure 12 An example data center 1200 is shown that can be used in at least one embodiment of the present disclosure. The data center 1200 can include a data center infrastructure layer 1210, a framework layer 1220, a software layer 1230, and / or an application layer 1240.
[0119] like Figure 12 As shown, the data center infrastructure layer 1210 may include a resource coordinator 1212, grouped computing resources 1214, and node computing resources ("node CRs") 1216(1)-1216(N), where "N" represents any integer, positive integer. In at least one embodiment, the node CRs 1216(1)-1216(N) may include, but are not limited to, any number of central processing units ("CPUs") or other processors (including DPUs, accelerators, field programmable gate arrays (FPGAs), graphics processors or graphics processing units (GPUs), etc.), memory devices (e.g., dynamic read-only memories), storage devices (e.g., solid-state or disk drives), network input / output ("NW I / O") devices, network switches, virtual machines ("VMs"), power modules and / or cooling modules, etc. In some embodiments, one or more of the node CRs 1216(1)-1216(N) may correspond to a server having one or more of the above-mentioned computing resources. Furthermore, in some embodiments, node CRs 1216(1)-12161(N) may include one or more virtual components, such as vGPUs, vCPUs, etc., and / or one or more of node CRs 1216(1)-1216(N) may correspond to a virtual machine (VM).
[0120] In at least one embodiment, the grouped computing resources 1214 may include separate groups of node CRs 1216 housed in one or more racks (not shown) or in many racks in data centers at different geographical locations (also not shown). The separate groups of node CRs 1216 within the grouped computing resources 1214 may include grouped computing, network, memory, or storage resources that can be configured or allocated to support one or more workloads. In at least one embodiment, several node CRs 1216 including CPUs, GPUs, DPUs, and / or other processors may be grouped in one or more racks to provide computing resources to support one or more workloads. One or more racks may also include any number of power modules, cooling modules, and / or network switches in any combination.
[0121] Resource coordinator 1212 may configure or otherwise control one or more node CRs 1216(1)-1216(N) and / or grouped computing resources 1214. In at least one embodiment, resource coordinator 1212 may comprise a software design infrastructure ("SDI") management entity for data center 1200. Resource coordinator 1212 may comprise hardware, software, or some combination thereof.
[0122] In at least one embodiment, Figure 12 As shown, the framework layer 1220 may include a job scheduler 1228, a configuration manager 1234, a resource manager 1236, and / or a distributed file system 1238. The framework layer 1220 may include a framework that supports the software 1232 of the software layer 1230 and / or one or more applications 1242 of the application layer 1240. The software 1232 or the application 1242 may include network-based service software or applications, such as those provided by Amazon Web Services, Google Cloud, and Microsoft Azure. The framework layer 1220 may be, but is not limited to, a type of free and open source software web application framework that can utilize the distributed file system 1238 for large-scale data processing (e.g., "big data"), such as Apache Spark. TM(hereinafter referred to as "Spark"). In at least one embodiment, the job scheduler 1228 may include a Spark driver to facilitate the scheduling of workloads supported by the various layers of the data center 1200. The configuration manager 1234 may be capable of configuring different layers, such as the software layer 1230 and the framework layer 1220 including Spark and a distributed file system 1238 for supporting large-scale data processing. The resource manager 1236 may be capable of managing the mapping or allocation of clustered or grouped computing resources used to support the distributed file system 1238 and the job scheduler 1228. In at least one embodiment, the clustered or grouped computing resources may include the grouped computing resources 1214 at the data center infrastructure layer 1210. The resource manager 1236 may coordinate with the resource coordinator 1212 to manage these mapped or allocated computing resources.
[0123] In at least one embodiment, the software 1232 included in the software layer 1230 may include software used by at least portions of the node CRs 1216(1)-1216(N), the grouped computing resources 1214, and / or the distributed file system 1238 of the framework layer 1220. The one or more types of software may include, but are not limited to, Internet web search software, email virus scanning software, database software, and streaming video content software.
[0124] In at least one embodiment, the applications 1242 included in the application layer 1240 may include one or more types of applications used by the node CRs 1216(1)-1216(N), the grouped computing resources 1214, and / or at least portions of the distributed file system 1238 of the framework layer 1220. The one or more types of applications may include, but are not limited to, any number of genomics applications, cognitive computing, and machine learning applications (including training or inference software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), and / or other machine learning applications used in conjunction with one or more embodiments).
[0125] In at least one embodiment, any of configuration manager 1234, resource manager 1236, and resource coordinator 1212 can implement any number and type of self-modification actions based on any number and type of data obtained in any technically feasible manner. The self-modification actions can save the data center operator of data center 1200 from making potentially poor configuration decisions and potentially avoiding underutilized and / or underperforming portions of the data center.
[0126] According to one or more embodiments described herein, data center 1200 can include tools, services, software, or other resources for training one or more machine learning models or using one or more machine learning models to predict or infer information. For example, the (one or more) machine learning models can be trained by computing weight parameters according to a neural network architecture using the software and / or computing resources described above with respect to data center 1200. In at least one embodiment, a trained or deployed machine learning model corresponding to one or more neural networks can be used to infer or predict information using the resources described above with respect to data center 1200 using weight parameters computed by one or more training techniques, such as, but not limited to, those described herein.
[0127] In at least one embodiment, data center 1200 may use a CPU, an application-specific integrated circuit (ASIC), a GPU, an FPGA, and / or other hardware (or virtual computing resources corresponding thereto) to perform training and / or reasoning using the aforementioned resources. In addition, the aforementioned one or more software and / or hardware resources may be configured to allow users to train or perform information reasoning services, such as image recognition, speech recognition, or other artificial intelligence services.
[0128] Sample network environment
[0129] A network environment suitable for implementing embodiments of the present disclosure may include one or more client devices, servers, network attached storage (NAS), other backend devices, and / or other device types. The client devices, servers, and / or other device types (e.g., each device) may be configured to: Figure 11 1100, for example, each device may include similar components, features, and / or functions of the computing device 1100. In addition, in the case of implementing a backend device (e.g., a server, NAS, etc.), the backend device may be included as part of the data center 1200, and the example of the data center 1200 is referred to herein with respect to Figure 12 Describe in more detail.
[0130] The components of the network environment can communicate with each other via one or more networks that can be wired, wireless, or both. The network can include multiple networks or networks of networks. As an example, the network can include one or more wide area networks (WANs), one or more local area networks (LANs), one or more public networks (such as the Internet and / or the Public Switched Telephone Network (PSTN)), and / or one or more private networks. In the case where the network includes a wireless telecommunications network, components such as base stations, communication towers, or even access points (and other components) can provide wireless connections.
[0131] Compatible network environments may include one or more peer-to-peer network environments, in which case the network environment may not include a server, and one or more client-server network environments, in which case the network environment may include one or more servers. In a peer-to-peer network environment, the functionality described herein with respect to the server(s) may be implemented on any number of client devices.
[0132] In at least one embodiment, the network environment may include one or more cloud-based network environments, distributed computing environments, combinations thereof, and the like. The cloud-based network environment may include a framework layer, a job scheduler, a resource manager, and a distributed file system implemented on one or more of the servers, which may include one or more core network servers and / or edge servers. The framework layer may include software supporting the software layer and / or a framework for one or more applications of the application layer. The software or application may include network-based service software or applications, respectively. In an embodiment, one or more of the client devices may use web-based service software or applications (e.g., by accessing the service software and / or application via one or more application programming interfaces (APIs)). The framework layer may be, but is not limited to, a type of free and open source software web application framework, such as one that can use a distributed file system for large-scale data processing (e.g., "big data").
[0133] A cloud-based network environment can provide cloud computing and / or cloud storage that performs any combination of the computing and / or data storage functions described herein (or one or more portions thereof). Any of these different functions can be distributed across multiple locations from a central or core server (e.g., one or more data centers that can be distributed across states, regions, countries, the world, etc.). If the connection to the user (e.g., a client device) is relatively close to an edge server, the core server can assign at least a portion of the functionality to the edge server. A cloud-based network environment can be private (e.g., limited to a single organization), public (e.g., available to many organizations), and / or a combination thereof (e.g., a hybrid cloud environment).
[0134] One or more client devices may include herein Figure 11At least some of the components, features, and functionality of one or more example computing devices 1100 are described. By way of example and not limitation, a client device may be embodied as a personal computer (PC), a laptop computer, a mobile device, a smartphone, a tablet computer, a smartwatch, a wearable computer, a personal digital assistant (PDA), an MP3 player, a virtual reality headset, a global positioning system (GPS) or device, a video player, a camera, a surveillance device or system, a vehicle, a vessel, a spacecraft, a virtual machine, a drone, a robot, a handheld communication device, a hospital device, a gaming device or system, an entertainment system, a vehicle computer system, an embedded system controller, a remote control, an appliance, a consumer electronic device, a workstation, an edge device, any combination of the depicted devices, or any other suitable device.
[0135] The present disclosure may be described in the general context of computer code or machine-usable instructions (including computer-executable instructions, such as program modules) executed by a computer or other machine, such as a personal data assistant or other handheld device. Generally, program modules, including routines, programs, objects, components, data structures, and the like, refer to code that performs a particular task or implements a particular abstract data type. The present disclosure may be implemented in a variety of system configurations, including handheld devices, consumer electronics, general-purpose computers, more specialized computing devices, and the like. The present disclosure may also be implemented in distributed computing environments where tasks are performed by remote processing devices linked through a communications network.
[0136] As used herein, the statement "and / or" with respect to two or more elements should be interpreted as meaning only one element, or a combination of elements. For example, "element A, element B, and / or element C" may include only element A, only element B, only element C, element A and element B, element A and element C, element B and element C, or elements A, B, and C. Furthermore, "at least one of element A or element B" may include at least one of element A, at least one of element B, or at least one of element A and at least one of element B. Furthermore, "at least one of element A and element B" may include at least one of element A, at least one of element B, or at least one of element A and at least one of element B.
[0137] The subject matter of the present disclosure is described herein with specificity to satisfy statutory requirements. However, the description itself is not intended to limit the scope of the present disclosure. Rather, the inventors have contemplated that the claimed subject matter may also be embodied in other ways, in conjunction with other current or future technologies, to include different steps or combinations of steps similar to the steps described in this document. Furthermore, although the terms "step" and / or "box" may be used herein to refer to different elements of the method employed, such terms should not be construed to imply any particular order among or between the various steps disclosed herein unless and except where the order of the individual steps is explicitly described.
[0138] Example paragraph
[0139] A: A method comprising: determining one or more first features associated with a plurality of images represented by the image data based at least on processing image data generated using a plurality of cameras located in an environment; determining one or more second features associated with the environment based at least on one or more spatial encoders processing the one or more first features and calibration data relating one or more three-dimensional (3D) coordinates associated with the environment to one or more two-dimensional (2D) coordinates associated with the plurality of images; determining one or more 3D positions associated with one or more objects located in the environment based at least on one or more decoders processing the one or more second features; and performing one or more operations based at least on the one or more 3D positions.
[0140] B: The method as described in paragraph A also includes: generating one or more fused features based at least on one or more temporal encoders processing the one or more second features and one or more previous features associated with the environment, wherein determining the one or more 3D positions is based at least on the one or more decoders processing the one or more fused features.
[0141] C: A method as described in paragraph B, wherein: the one or more second features are associated with the image data generated using the multiple cameras during a first time period; and the one or more previous features are associated with second image data generated using the multiple cameras during a second time period before the first time period.
[0142] D: A method as described in any of paragraphs AC, wherein: a first portion of the one or more first features is associated with a first image in the multiple images, and a second portion of the one or more first features is associated with a second image in the multiple images; and determining the one or more second features is achieved by at least aggregating the first portion of the one or more first features with the second portion of the one or more first features.
[0143] E: A method as described in any of paragraphs AD, wherein the calibration data represents at least a matrix for projecting the 3D coordinates within the environment to the 2D coordinates associated with the multiple images.
[0144] F: A method as described in any of paragraphs AE, wherein determining the one or more second features includes: processing the calibration data based at least on the one or more spatial encoders to determine that a 3D point within the environment is associated with a 2D point within the multiple images; determining that the 2D point is associated with the one or more second features based at least on the one or more first features; and associating the one or more second features with the 3D point based at least on the association of the 2D point with the one or more second features.
[0145] G: A method as described in any of paragraphs AF, wherein: the multiple cameras are located within the environment and are oriented so that the multiple cameras include a field of view representing at least a portion of the interior of the environment; and the one or more objects include one or more dynamic objects located within the interior of the environment.
[0146] H: A method as described in any of paragraphs AG, wherein the one or more operations include at least one of: determining one or more trajectories associated with the one or more objects within the environment; determining one or more classifications associated with the one or more objects; determining one or more 2D positions associated with the one or more objects within the multiple images; or causing presentation of information associated with the one or more 3D positions.
[0147] I: A system comprising: one or more processors for: determining, based at least on image data generated using a plurality of cameras located within an environment, one or more first features associated with a plurality of images represented by the image data; determining, based at least on the one or more first features and calibration data relating three-dimensional (3D) points within the environment to two-dimensional (2D) points associated with the plurality of images, one or more second features associated with the environment; determining, based at least on the one or more second features, one or more 3D locations associated with one or more objects located within the environment; and performing one or more operations based at least on the one or more 3D locations.
[0148] J: A system as described in paragraph I, wherein the one or more processors are further used to: determine one or more third features associated with a second image represented by second image data generated at least based on second image data generated using the multiple cameras located in the environment; and determine one or more fourth features associated with the environment based at least on the one or more third features and the calibration data, wherein the one or more 3D positions are also determined based at least on the one or more fourth features.
[0149] K: A system as described in paragraph J, wherein: the image data is generated using the multiple cameras during a first time period; and the second image data is generated using the multiple cameras during a second time period different from the first time period.
[0150] L: A system as described in paragraph J, wherein the one or more processors are further used to: process the one or more second features and the one or more fourth features based at least on one or more time encoders to generate one or more fused features; wherein the one or more 3D positions are determined based at least on the one or more fused features.
[0151] M: A system as described in any of paragraphs IL, wherein: a first portion of the one or more first features is associated with a first image in the plurality of images, and a second portion of the one or more first features is associated with a second image in the plurality of images; and determining the one or more second features includes: processing the one or more first features and the calibration data based at least on one or more spatial encoders, and determining the one or more second features by aggregating the first portion of the one or more first features with the second portion of the one or more first features.
[0152] N: A system as described in any of paragraphs IM, wherein the determination of the one or more second features includes: determining that a 3D point within the environment is associated with a 2D point within the multiple images based at least on one or more spatial encoders processing the calibration data; determining that the 2D point is associated with the one or more second features based at least on the one or more first features; and associating the one or more second features with the 3D point based at least on the association of the 2D point with the one or more second features.
[0153] O: A system as described in any of paragraphs IN, wherein the calibration data represents a matrix for projecting the 3D points associated with the environment to the 2D points associated with the multiple cameras.
[0154] P: A system as described in any of paragraphs IO, wherein the determination of the one or more 3D positions associated with the one or more objects is also based at least on data representing one or more object queries, wherein the one or more object queries indicate that the one or more objects may be located at one or more locations in the environment.
[0155] Q: A system as described in any of paragraphs IP, wherein: the multiple cameras are located within the environment and oriented so that the multiple cameras include a field of view representing at least a portion of the interior of the environment; and the one or more objects include one or more dynamic objects located within the interior of the environment.
[0156] R: A system as described in any of paragraphs IQ, wherein the system is included in at least one of the following items: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing one or more simulation operations; a system for performing one or more digital twin operations; a system for performing light transport simulation; a system for performing collaborative content creation of 3D assets; a system for performing one or more deep learning operations; a system implemented using an edge device; a system implemented using a robot; a system for performing one or more generative AI operations; a system for performing operations using one or more small language models; a system for performing operations using one or more large language models; a system for performing operations using one or more visual language models (VLMs); a system for performing one or more conversational AI operations; a system for generating synthetic data; a system for presenting at least one of virtual reality content, augmented reality content, or mixed reality content; a system comprising one or more virtual machines (VMs); a system implemented at least in part in a data center; or a system implemented at least in part using cloud computing resources.
[0157] S: One or more processors, comprising: processing circuitry for: generating second feature data associated with the environment based at least on first feature data associated with image data generated using multiple cameras in the environment and calibration data associating three-dimensional (3D) points in the environment with two-dimensional (2D) points associated with the multiple cameras; determining one or more 3D positions associated with one or more objects located in the environment based at least on the second feature data; and performing one or more operations based at least on the one or more 3D positions.
[0158] T: One or more processors as described in paragraph S, wherein the one or more processors are included in at least one of the following items: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing one or more simulation operations; a system for performing one or more digital twin operations; a system for performing light transport simulation; a system for performing collaborative content creation of 3D assets; a system for performing one or more deep learning operations; a system implemented using an edge device; a system implemented using a robot; a system for performing one or more generative AI operations; a system for performing operations using one or more small language models; a system for performing operations using one or more large language models; a system for performing operations using one or more visual language models (VLMs); a system for performing one or more conversational AI operations; a system for generating synthetic data; a system for presenting at least one of virtual reality content, augmented reality content, or mixed reality content; a system comprising one or more virtual machines (VMs); a system implemented at least in part in a data center; or a system implemented at least in part using cloud computing resources.
Claims
1. A method comprising: determining one or more first features associated with a plurality of images represented by the image data based at least on processing image data generated using a plurality of cameras located in the environment; determining one or more second features associated with the environment based at least on processing the one or more first features with one or more spatial encoders and calibration data correlating one or more three-dimensional (3D) coordinates associated with the environment with one or more two-dimensional (2D) coordinates associated with the plurality of images; determining one or more 3D positions associated with one or more objects located in the environment based at least on processing the one or more second features by the one or more decoders; and One or more operations are performed based at least on the one or more 3D positions.
2. The method according to claim 1, further comprising: generating one or more fused features based at least on one or more temporal encoders processing the one or more second features and one or more previous features associated with the environment, Wherein determining the one or more 3D positions is based at least on processing the one or more fused features by the one or more decoders.
3. The method according to claim 2, wherein: The one or more second features are associated with the image data generated using the plurality of cameras during a first time period; as well as The one or more previous features are associated with second image data generated using the plurality of cameras during a second time period prior to the first time period.
4. The method according to claim 1, wherein: A first portion of the one or more first features is associated with a first image in the plurality of images, and a second portion of the one or more first features is associated with a second image in the plurality of images; as well as Determining the one or more second features is accomplished by aggregating at least the first portion of the one or more first features with the second portion of the one or more first features. 5 . The method of claim 1 , wherein the calibration data represents at least a matrix for projecting the 3D coordinates within the environment to the 2D coordinates associated with the plurality of images.
6. The method of claim 1 , wherein determining the one or more second characteristics comprises: processing the calibration data based at least on the one or more spatial encoders to determine associations between 3D points within the environment and 2D points within the plurality of images; determining, based at least on the one or more first features, that the 2D point is associated with the one or more second features; as well as The one or more second features are associated with the 3D point based at least on the 2D point being associated with the one or more second features.
7. The method according to claim 1, wherein: The plurality of cameras are located within the environment and oriented such that the plurality of cameras include a field of view representing at least a portion of an interior of the environment; as well as The one or more objects include one or more dynamic objects located within the interior of the environment.
8. The method of claim 1 , wherein the one or more operations include at least one of: determining one or more trajectories associated with the one or more objects within the environment; determining one or more classifications associated with the one or more objects; determining one or more 2D locations associated with the one or more objects within the plurality of images; or Information associated with the one or more 3D locations is caused to be presented.
9. A system comprising: One or more processors for: determining, based at least on image data generated using a plurality of cameras located within an environment, one or more first features associated with a plurality of images represented by the image data; determining one or more second characteristics associated with the environment based at least on the one or more first characteristics and calibration data correlating three-dimensional (3D) points within the environment with two-dimensional (2D) points associated with the plurality of images; determining one or more 3D positions associated with one or more objects located within the environment based at least on the one or more second features; and One or more operations are performed based at least on the one or more 3D positions.
10. The system of claim 9, wherein the one or more processors are further configured to: determining, based at least on second image data generated using the plurality of cameras located in the environment, one or more third features associated with a second image represented by the second image data; and determining one or more fourth characteristics associated with the environment based at least on the one or more third characteristics and the calibration data, Wherein the one or more 3D positions are further determined based at least on the one or more fourth features.
11. The system of claim 10, wherein: The image data is generated using the plurality of cameras during a first time period; and The second image data is generated using the plurality of cameras during a second time period different from the first time period.
12. The system of claim 10, wherein the one or more processors are further configured to: processing the one or more second features and the one or more fourth features based on at least one or more temporal encoders to generate one or more fused features; The one or more 3D positions are determined based at least on the one or more fused features.
13. The system of claim 9, wherein: A first portion of the one or more first features is associated with a first image in the plurality of images, and a second portion of the one or more first features is associated with a second image in the plurality of images; as well as The determining of the one or more second features includes: processing the one or more first features and the calibration data based at least on one or more spatial encoders, and determining the one or more second features by aggregating the first part of the one or more first features with the second part of the one or more first features.
14. The system of claim 9, wherein determining the one or more second characteristics comprises: processing the calibration data based at least on one or more spatial encoders to determine that 3D points within the environment are associated with 2D points within the plurality of images; determining, based at least on the one or more first features, that the 2D point is associated with the one or more second features; as well as The one or more second features are associated with the 3D point based at least on the 2D point being associated with the one or more second features.
15. The system of claim 9, wherein the calibration data represents a matrix for projecting the 3D points associated with the environment to the 2D points associated with the plurality of cameras.
16. The system of claim 9, wherein the determination of the one or more 3D positions associated with the one or more objects is further based at least on data representing one or more object queries indicating one or more locations in the environment where the one or more objects may be located.
17. The system of claim 9, wherein: The plurality of cameras are located within the environment and oriented such that the plurality of cameras include a field of view representing at least a portion of an interior of the environment; as well as The one or more objects include one or more dynamic objects located within the interior of the environment.
18. The system of claim 9, wherein the system is included in at least one of the following: control systems for autonomous or semi-autonomous machines; Perception systems for autonomous or semi-autonomous machines; a system for performing one or more simulation operations; a system for performing one or more digital twin operations; a system for performing light transport simulations; A system for performing collaborative content creation of 3D assets; a system for performing one or more deep learning operations; Systems implemented using edge devices; Systems implemented using robots; A system for performing one or more generative AI operations; a system for performing operations using one or more small language models; A system for performing operations using one or more large language models; A system for performing operations using one or more visual language models (VLMs); A system for performing one or more conversational AI operations; Systems for generating synthetic data; a system for presenting at least one of virtual reality content, augmented reality content, or mixed reality content; A system comprising one or more virtual machines VM; A system implemented at least in part in a data center; or A system implemented at least in part using cloud computing resources.
19. One or more processors comprising: Processing circuitry for: generating second feature data associated with the environment based on at least first feature data associated with image data generated using a plurality of cameras in the environment and calibration data associating three-dimensional (3D) points in the environment with two-dimensional (2D) points associated with the plurality of cameras; determining one or more 3D positions associated with one or more objects located in the environment based at least on the second feature data; and One or more operations are performed based at least on the one or more 3D positions.
20. The one or more processors of claim 19, wherein the one or more processors are included in at least one of the following: control systems for autonomous or semi-autonomous machines; Perception systems for autonomous or semi-autonomous machines; a system for performing one or more simulation operations; a system for performing one or more digital twin operations; a system for performing light transport simulations; A system for performing collaborative content creation of 3D assets; a system for performing one or more deep learning operations; Systems implemented using edge devices; Systems implemented using robots; A system for performing one or more generative AI operations; a system for performing operations using one or more small language models; A system for performing operations using one or more large language models; A system for performing operations using one or more visual language models (VLMs); A system for performing one or more conversational AI operations; Systems for generating synthetic data; a system for presenting at least one of virtual reality content, augmented reality content, or mixed reality content; A system comprising one or more virtual machines VM; A system implemented at least in part in a data center; or a system implemented at least in part using cloud computing resources.