Detecting occluded objects in images for autonomous systems and applications

By combining image processing technology and point cloud processing technology, the obscured traffic objects in the image are detected, and the problem of inaccurate truth labels in the prior art is solved, achieving more efficient truth data generation and reducing manual participation.

CN120070881APending Publication Date: 2025-05-30NVIDIA CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411727495.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-30
Filing Date
2024-11-28
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In the process of generating images, it is difficult to accurately identify traffic objects blocked in the image, resulting in inaccurate truth tags generated, which increases the workload and time of the manual marker.

Method used

By detecting occluded objects in image or other sensor data representations, using a combination of image processing technology and point cloud processing technology to determine whether the object or feature is occluded. The specific method includes using map data and sensor data to determine the three-dimensional position of the object and projecting it to the two-dimensional position of the image, and combining distance information to determine whether the object is blocked.

Benefits of technology

More accurate and accurate true value data generation is achieved, reducing the need for manual participation and improving the performance of autonomous or semi-autonomous systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070881A_ABST
    Figure CN120070881A_ABST
Patent Text Reader

Abstract

The disclosure relates to detecting occluded objects in images for autonomous systems and applications. In various examples, occluded objects in detection images or other sensor data representations for autonomous or semi-autonomous systems and applications are described herein. The systems and methods described herein may use various techniques to determine when an object is occluded at certain portions of an image. For example, an image may be processed to determine a classification associated with an object depicted by the image, and then whether one or more objects in the image are occluded may be determined using the classification and tags projected on the image using a map. Still for example, a first distance to a point within the environment may be determined using a map, and a second distance to a point within the environment may be determined using the point cloud. These distances may then be used to determine whether one or more objects within the image are occluded.
Need to check novelty before this filing date? Find Prior Art

Description

BACKGROUND OF THE INVENTION

[0001] Marked images (or other sensor data representations) generated by machines that navigate in an environment can be important for many purposes. For example, images can be marked to generate ground truth data, which can then be used to train machine learning models to perform various tasks. Thus, traditional systems can use map-marked images, for example, by using labels of various objects (e.g., traffic signs, traffic poles, driving surfaces, etc.) represented by the map to generate labels of the corresponding objects depicted in the image. However, in some cases, an object may be occluded in the image, such as when a dynamic object and / or a static object is within the environment and between the machine that generates the image and the object to be marked. Thus, if the map indicates that a traffic object (e.g., a road boundary, a road line, etc.) should be located in a particular part (e.g., a pixel, etc.) of the image, but the image depicts a vehicle in that part of the image, the traditional system may incorrectly generate a ground truth label for the image. Then, this mislabeled ground truth data may be used to train the model to make inaccurate predictions, or a human annotator may be required to perform a quality check on the label, update the label, thereby increasing the manual effort and overall time required to generate high-quality ground truth data. SUMMARY OF THE INVENTION

[0002] Embodiments of the present disclosure relate to detecting occluded objects in a detected image or other sensor data representation (e.g., a LiDAR (Light Detection and Ranging) point cloud, RADAR (Radio Detection and Ranging) data, a range or projection image, etc.) for autonomous or semi-autonomous systems and applications. For example, the systems and methods described herein can use one or more techniques to determine when an object or feature (e.g., a traffic object or feature, a dynamic or static object, etc.) is occluded at a part (e.g., a pixel, a point, etc.) of an image or other sensor data representation. For a first example, and for an image, the image can be processed to determine a classification associated with the object depicted in the image. Then this classification can be used together with a label projected onto the image using a map to determine whether one or more objects are occluded. For a second example, also for an image, a map can be used to determine a first distance to a point within the environment, and a point cloud can be used to determine a second distance to a point within the environment. Then these distances can be used to determine whether one or more objects in the image are occluded. In some examples, the systems and methods can combine these techniques to determine a final result indicating whether an object is occluded in the image.

[0003] In some embodiments, compared to traditional systems, the current system is capable of determining when an object that should be depicted in an image or other sensor data representation is actually occluded by one or more other objects. For example, traditional systems may only use the labels associated with objects from a map to generate labels for the corresponding objects depicted in an image. However, if an object is occluded in the image, e.g., by one or more other dynamic and / or static objects, the labels of the image may be incorrect. Thus, by performing one or more of the techniques described herein, the current system is capable of generating labels for the objects depicted in an image, as well as generating additional labels indicating whether various parts of the objects are occluded at various parts of the image. Because of this, the current system can provide many improvements, such as generating more accurate and precise ground truth data, generating ground truth data that requires little or no human intervention, generating training data that can be relied upon to train machine learning models to perform various tasks (e.g., detecting objects, determining paths, identifying road features (e.g., lane lines, road boundary lines, crosswalks, signs, waiting conditions, etc.) and / or the like). BRIEF DESCRIPTION OF THE DRAWINGS

[0004] The present system and method for detecting occluded objects in an image or other sensor data representation for autonomous or semi-autonomous systems and applications are described in detail below with reference to the drawings, in which:

[0005] Figures 1A to 1C FIG. 1 shows an example data flow diagram of a process for detecting occluded objects within an image, in accordance with some embodiments of the present disclosure;

[0006] Figure 2 FIG. 2 shows an example of an image that has been segmented to determine the classifications associated with the objects depicted by the image, in accordance with some embodiments of the present disclosure;

[0007] Figure 3 FIG. 3 shows an example of projecting the three-dimensional position of an object from a map to a two-dimensional position in an image, in accordance with some embodiments of the present disclosure;

[0008] Figure 4A FIG. 4 shows an example of determining whether a traffic object is occluded at a portion of an image, in accordance with some embodiments of the present disclosure;

[0009] Figure 4B FIG. 5 shows an example of generating occlusion information associated with an image, in accordance with some embodiments of the present disclosure;

[0010] Figure 5 FIG. 6 shows an example of projecting three-dimensional points within an environment to a two-dimensional portion of an image and then using the projection to determine information associated with that portion of the image, in accordance with some embodiments of the present disclosure;

[0011] Figure 6Shows an example of an occupancy map that can be generated using point cloud data and LiDAR data according to one or more embodiments of the present disclosure;

[0012] Figure 7 Shows an example of determining a threshold distance associated with determining whether a portion of an image is occluded according to some embodiments of the present disclosure;

[0013] Figure 8 Shows an example of determining whether a traffic object is occluded in certain portions of an image according to some embodiments of the present disclosure;

[0014] Figure 9 Is a flowchart showing a method of detecting an occluded object in an image using one or more image processing techniques according to some embodiments of the present disclosure;

[0015] Figure 10 Is a flowchart showing a method of detecting an occluded object in an image using one or more point cloud techniques according to some embodiments of the present disclosure;

[0016] Figure 11 Is a flowchart showing a method of detecting an occluded object in an image using one or more image processing techniques and one or more point cloud techniques according to some embodiments of the present disclosure;

[0017] Figure 12A Is an illustration of an example autonomous vehicle according to some embodiments of the present disclosure;

[0018] Figure 12B Is according to some embodiments of the present disclosure Figure 12A An example of the camera position and field of view of an example autonomous vehicle;

[0019] Figure 12C Is according to some embodiments of the present disclosure Figure 12A A block diagram of an example system architecture of an example autonomous vehicle;

[0020] Figure 12D Is for use in a cloud-based server and Figure 12A An example system diagram for communication between an example autonomous vehicle;

[0021] Figure 13 Is a block diagram of an example computing device suitable for implementing some embodiments of the present disclosure; and

[0022] Figure 14 Is a block diagram of an example data center suitable for implementing some embodiments of the present disclosure. Detailed Description

[0023] The disclosed systems and methods relate to detecting occluded objects in detected image or other sensor data representations for autonomous or semi-autonomous systems and applications. Although the present disclosure may be described with respect to an example autonomous or semi-autonomous vehicle or machine 1200 (alternatively referred to herein as "vehicle 1200", "self-vehicle 1200", "self-machine 1200", or "machine 1200"), this is not meant to be limiting. For example, the systems and methods described herein may be used, without limitation, for non-autonomous vehicles or machines, semi-autonomous vehicles or machines (e.g., in one or more adaptive driving assistance systems (ADAS)), autonomous vehicles or machines, manned and unmanned robots or robotic platforms, warehouse vehicles, off-road vehicles, vehicles connected to one or more trailers, airships, vessels, shuttles, emergency vehicles, motorcycles, electric or motorized bicycles, airplanes, construction vehicles, underwater vehicles, drones, and / or other vehicle types. Additionally, although the present disclosure may be described with respect to occlusion detection for ground truth data generation in autonomous or semi-autonomous systems and applications, this is not intended to be limiting, and the systems and methods described herein may be used in augmented reality, virtual reality, mixed reality, robotics, security and surveillance, autonomous or semi-autonomous machine applications, and / or any other technical field where object detection and / or map creation may be used.

[0024] For example, the system may receive image data and / or other sensor data generated using one or more sensors of one or more machines navigating in an environment. In an example using image data, the image data may represent or correspond to one or more images depicting the environment. In an example using other sensor modalities (e.g., LiDAR, RADAR, ultrasonic, etc.), the sensor data may represent other sensor data representations, such as range images, projection images, point clouds, etc. For simplicity, the examples presented herein mainly discuss images, but the techniques described herein may also be applied to other sensor modalities and sensor data representations. Then, the system may be configured to process the image to generate labels for the objects depicted in the image. For example, in some examples, for an image, the system may be configured to generate labels for traffic objects depicted in the image, such as but not limited to roads, road boundaries, road lines, traffic lights, traffic signs, traffic polling stations, parking spaces, waiting conditions, static objects or features, warehouses, offices, retail, commercial, aerial, amphibious, and / or other features or objects in the environment, and / or any other type of traffic, road, environment, and / or perceivable object or feature. In some examples, the system may use a map associated with the environment to label the traffic objects depicted by the image. Additionally, in some examples, the system may be configured to generate labels for other types of objects, such as structures, pedestrians, vehicles, vegetation, and / or any other type of object. As described herein, in some examples, at least a portion of at least one traffic object or feature (and / or other types of objects or features) may be occluded, e.g., by another object. Thus, the system may be configured to perform one or more techniques to indicate the portions of the image associated with the occluded portions of the traffic object or feature.

[0025] For example, in some examples, the system can use image processing techniques to determine an image portion associated with an occluded traffic object located in the environment. For example, for an image, the system can use one or more models to process the image data representing the image, such as one or more machine learning models, one or more neural networks, one or more segmentation models, one or more classification models, one or more perception models, and / or any other type of model trained to perform object segmentation and / or classification. For example, at least based on the processing, the model can generate various classifications and / or segmentation masks associated with the objects depicted by the image. As described herein, classifications can include, but are not limited to, roads, sidewalks, buildings, walls, fences, human heads, traffic lights, traffic signs, vegetation, terrain, sky, people, riders, cars, trucks, buses, trains, motorcycles, bicycles, and / or any other type of object classification. Additionally, the system can use a map to generate labels for one or more objects depicted by the image (such as traffic objects (e.g., road boundaries, road markings, etc.)). As described in more detail herein, the system can project the labels from the map to the image using localization and / or three-dimensional (3D) to two-dimensional (2D) projection.

[0026] Then, the system can use the classifications and labels to determine whether portions of the image are associated with an occluded traffic object. For example, the system can use the labels to determine that a portion of the image (e.g., a pixel) is associated with a traffic object, such as a road boundary or a lane marking. Then, the system can use the classification associated with that portion of the image to determine whether the traffic object is occluded in that portion of the image. In some examples, the system can determine that the traffic object is not occluded in that portion of the image at least based on the classification corresponding to (e.g., being similar to) the label of the traffic object, or determine that the traffic object is occluded in that portion of the image at least based on the classification not corresponding to (e.g., being different from) the label.

[0027] For a first example, if the traffic object includes a road boundary, the system can determine that the road boundary is not occluded in that portion of the image at least based on the classification including one or more first classifications (e.g., road or sidewalk), or determine that the road boundary is occluded in that portion of the image at least based on the classification including one or more second classifications (e.g., utility pole, vegetation, truck, car, train, person, etc.) (e.g., any classification different from the first classification). For a second example, if the traffic object includes a traffic sign, the system can determine that the traffic sign is not occluded in that portion of the image at least based on the classification including one or more first classifications (e.g., traffic sign), or determine that the traffic sign is occluded in that portion of the image at least based on the classification including one or more second classifications (e.g., vegetation, truck, car, train, person, etc.).

[0028] Then, the system can continue to perform these processes for one or more additional portions of the image, such as additional portions labeled as associated with one or more additional traffic objects (e.g., pixels labeled as road markings, road boundaries, etc.). Additionally, the system can generate data (referred to as "first occlusion data" in some examples) indicating which portions of the image are associated with occluded traffic objects and / or which portions of the image are associated with unoccluded traffic objects. In some examples, for portions of the image associated with occluded traffic objects, the first occlusion data can further indicate whether the occlusion is caused by a dynamic object (e.g., a car, truck, bus, train, motorcycle, bicycle, person, etc.) or a static object (e.g., a traffic light, traffic sign, vegetation, terrain, fence, wall, building, etc.). Additionally, the system can continue to perform these processes on one or more additional images represented by the image data.

[0029] In some examples, in addition to or as an alternative to using image processing techniques, the system can use point cloud processing techniques to determine portions of the image associated with occluded traffic objects located in the environment. For example, also for the image, the system can determine portions of the image (e.g., pixels) associated with traffic objects, such as by using classification, labeling, and / or any other techniques. Then, the system can use a map to determine 3D coordinates associated with points in the environment associated with the portions of the image. Additionally, the system can use the 3D coordinates to determine a first distance associated with the portions of the image (e.g., the distance between the machine that generated the image data and the point located in the environment).

[0030] The system can also generate point cloud data (e.g., a 3D occupancy grid) representing points in the environment. In some examples, the system uses a point cloud map representing initial points in the environment and / or LiDAR data generated by one or more LiDAR sensors of the machine to generate the point cloud data. For example, the system can generate the point cloud data by combining at least a portion of the points represented by the point cloud map with at least a portion of the points represented by the LiDAR data. Then, the system can use the point cloud data to determine a second distance associated with the portions of the image (e.g., another distance between the machine and the point located in the environment). For example, the system can project multiple rays from the portions of the image towards the points in the environment to determine multiple distances. Then, the system can use the multiple distances to determine the second distance, such as by using the minimum distance, median distance, average distance, maximum distance, and / or any other distance.

[0031] Then, the system can use the first distance and the second distance to determine whether the traffic object is occluded in that portion of the image. For example, the system can determine that the traffic object is occluded in that portion of the image at least based on the second distance being within a threshold distance from the first distance, or determine that the traffic object is not occluded in that portion of the image at least based on the second distance being outside the threshold distance from the first distance. In such an example, the system can use one or more techniques to determine the threshold distance. For example, as described in more detail herein, the system can use at least the incident angle associated with the ray, the depth of the ground point associated with the point in the environment, and / or any other factors to determine the threshold distance.

[0032] Then, the system can continue to perform these processes for one or more additional portions of the image, such as additional portions labeled as associated with one or more additional traffic objects (e.g., pixels labeled as road markings, road boundaries, etc.). Additionally, the system can generate data (referred to as "second occlusion data" in some examples) indicating which portions of the image are associated with occluded traffic objects and / or which portions of the image are associated with non-occluded traffic objects. Additionally, the system can continue to perform these processes for one or more additional images represented by the image data.

[0033] In some examples, the system can combine image processing techniques and point cloud processing techniques to determine the portions of the image associated with occluded traffic objects located in the environment. For example, for a portion of the image, the system can use at least first occlusion data (indicating whether the traffic object is occluded in that portion of the image) and second occlusion data (also indicating whether the traffic object is occluded in that portion of the image) to ultimately determine whether the traffic object is occluded in that portion of the image. In some examples, the system can use additional data when making the final determination, such as data representing depth values associated with 3D points corresponding to the image portion and / or data representing uncertainty values associated with one or more of the labels and / or classifications associated with the image portion. Techniques for making the final determination using such data are described in more detail herein.

[0034] Then, the system can perform a similar process to generate data (referred to as "final occlusion data" in some examples) indicating which portions of the image are associated with occluded traffic objects and / or which portions of the image are associated with non-occluded traffic objects. Similar to the first occlusion data, in some examples, the final occlusion data can further indicate whether the occlusion is caused by a dynamic object (e.g., a car, truck, bus, train, motorcycle, bicycle, person, etc.) or a static object (e.g., a traffic light, traffic sign, vegetation, terrain, fence, wall, building, etc.). Then, the system can continue to perform these processes to generate final occlusion data for one or more additional images represented by the image data.

[0035] Although mainly described with respect to traffic or other road or vehicle-related features and objects, the processes described herein can be used to identify occluded features for any type of feature or object from map data, blueprint data, schematic data, and / or other data representing an environment, space, region, area, building, etc., in order to more effectively and accurately identify quality images or other sensor data representations and corresponding ground truth data.

[0036] The systems and methods described herein can be used in non-autonomous vehicles or machines, semi-autonomous vehicles or machines (e.g., in one or more adaptive driving assistance systems (ADAS)), autonomous vehicles or machines, manned and unmanned robots or robotic platforms, warehouse vehicles, off-road vehicles, vehicles coupled to one or more trailers, aircraft, boats, shuttles, emergency response vehicles, motorcycles, electric or motorized bicycles, airplanes, construction vehicles, underwater vehicles, drones, and / or other vehicle types, but are not limited thereto. Additionally, the systems and methods described herein can be used for various purposes, such as but not limited to, for machine control, machine motion, machine driving, synthetic data generation, model training, perception, augmented reality, virtual reality, mixed reality, robotics, security and surveillance, simulation and digital twins, autonomous or semi-autonomous machine applications, deep learning, environment simulation, object or participant simulation, and / or digital twins, data center processing, conversational artificial intelligence (AI), light transport simulation (e.g., ray tracing, path tracing, etc.), collaborative content creation of 3D assets, cloud computing, and / or any other suitable application.

[0037] The disclosed embodiments can be included in various different systems, such as automotive systems (e.g., control systems for autonomous or semi-autonomous machines, perception systems for autonomous or semi-autonomous machines), systems implemented using robots, aviation systems, medical systems, boating systems, intelligent area monitoring systems, systems for performing deep learning operations, systems for performing simulation operations, systems for performing digital twin operations, systems implemented using edge devices, systems implementing large language models (LLMs), systems containing one or more virtual machines (VMs), systems for performing synthetic data generation operations, systems implemented at least partially in a data center, systems for performing conversational AI operations, systems for performing light transport simulation, systems for performing collaborative content creation of 3D assets, systems for performing generative AI operations, systems implemented at least partially using cloud computing resources, and / or other types of systems.

[0038] Reference Figures 1A to 1C , Figures 1A to 1CShows an example data flow diagram of a process for detecting occluded objects within an image according to some embodiments of the present disclosure. It should be understood that such arrangements and other arrangements described herein are presented only as examples. Other arrangements and elements (e.g., machines, interfaces, functions, sequences, function groupings, etc.) may be used in addition to or in place of the shown arrangements and elements, and some elements may be omitted entirely. Moreover, many of the elements described herein are functional entities that may be implemented as discrete or distributed components or in combination with other components and may be implemented in any suitable combination and location. The various functions performed by the entities described herein may be executed by hardware, firmware, and / or software. For example, the various functions may be executed by a processor executing instructions stored in a memory. In some embodiments, the systems, methods, and processes described herein may use components, features, and / or functions similar to those of the exemplary autonomous vehicle 1200 in Figures 12A to 12D and the example computing device 1300 of Figure 13 and / or the example data center 1400 of Figure 14 to perform.

[0039] More specifically, Figure 1A shows an example data flow diagram of a process 100 for detecting occluded objects in an image using one or more image processing techniques. As shown, process 100 may include a classification component 102 receiving image data 104 representing one or more images depicting an environment. As described herein, in some examples, the classification component 102 may receive the image data 104 from one or more machines (e.g., vehicle 1200) navigating in the environment, while in other examples, the classification component 102 may receive the image data 104 from any other source. Then, process 100 may include the classification component 102 processing the image data 104 using one or more models associated with object segmentation and / or object classification. For example, the models may include one or more machine learning models, one or more neural networks, one or more segmentation models, one or more classification models, one or more perception models, and / or any other type of model trained to perform object segmentation and / or classification.

[0040] In some examples, a model can be trained to label objects using one or more classifications. As described herein, classifications can include, but are not limited to, roads, sidewalks, buildings, walls, fences, poles, traffic lights, traffic signs, vegetation, terrain, sky, people, riders, cars, trucks, buses, trains, motorcycles, bicycles, and / or any other type of object classification. Then, process 100 can include classification component 102 generating and / or outputting classification data 106 representative of the classifications associated with the object. For example, for an image represented by image data 104, classification data 106 can represent a classification mask associated with respective portions of the image (e.g., pixels, regions, tiles, chunks, etc.). For example, classification data 106 can represent a first pixel associated with a first classification (e.g., road), a second pixel associated with a second classification (e.g., sidewalk), a third pixel associated with a third classification (e.g., car), and so on.

[0041] In some examples, when describing object classifications and / or labels, objects can be grouped into one or more groups. For a first example, traffic objects can include, but are not limited to, roads, road boundaries, sidewalks, road markings, traffic poles, traffic signs, traffic signals, and / or any other type of object associated with navigation machines within an environment. For a second example, travel surfaces can include, but are not limited to, roads, road boundaries, road markings, parking surfaces, and / or any other object associated with a surface on which a machine can navigate. For a third example, dynamic objects can include, but are not limited to, people, riders, cars, trucks, buses, trains, motorcycles, bicycles, and / or any other object that may be associated with movement. For a fourth example, static objects can include, but are not limited to, roads, sidewalks, buildings, walls, fences, poles, traffic lights, traffic signs, terrain, and / or any other object not associated with movement.

[0042] For example, Figure 2 An example of an image 202 in accordance with some embodiments of the present disclosure is shown, which has been segmented to determine the classifications associated with the objects depicted by the image 202. As shown, classification component 102 may have analyzed image data representative of the image 202 (e.g., image data 104) to classify the objects as pedestrians 204 (although only one is labeled for clarity), bicycles 206, vehicles 208 (although only one is labeled for clarity), roads 210, road lines 212 (although only one is labeled for clarity), poles 214, traffic signs 216, vegetation 218 (although only one is labeled for clarity), sky 220, trash cans 222 (although only one is labeled for clarity), and buildings 224 (although only one is labeled for clarity).

[0043] Then, the classification component 102 can generate classification data (e.g., classification data 106) representing the classification associated with the object. For example, the classification data can represent a mask indicating the classification associated with the object. For example, the classification data can represent pixel positions associated with the pedestrian 204, pixel positions associated with the bicycle 206, pixel positions associated with the vehicle 208, pixel positions associated with the road 210, pixel positions associated with the road line 212, pixel positions associated with the pole 214, pixel positions associated with the traffic sign 216, pixel positions associated with the vegetation 218, pixel positions associated with the sky 220, pixel positions associated with the trash can 222, and / or pixel positions associated with the building 224.

[0044] Return reference Figure 1A For an example of, process 100 can include a projection component 108 receiving map data 110 representing at least a portion of the environment associated with the image data 104. For example, the map data 110 can represent labels and three-dimensional (3D) positions of objects located within the environment, such as roads, road lines, road boundaries, traffic poles, traffic signs, traffic signals, and / or any other type of traffic object located within the environment. As described herein, the 3D position can include coordinate positions, such as an x coordinate position, a y coordinate position, and a z coordinate position. Then, the projection component 108 can project the 3D positions associated with the objects to two-dimensional (2D) positions associated with the image represented by the image data 104. As described herein, the 2D position can include pixel positions, coordinates (e.g., an x coordinate position and a y coordinate position), and / or any other type of position associated with a portion of the image. At least based on the projection, the projection component 108 can generate labels for at least a portion of the objects depicted in the image using the labels from the map data 110, where the generated labels can be represented by the projected label data 112.

[0045] For example, for an image, the projection component 108 can position the machine that generates the image data 104 representing the image relative to the map represented by the map data 110. In some examples, the projection component 108 can use any technique to position the machine, such as by using sensor data (e.g., position data, image data, LiDAR data, RADAR data, etc.) generated using the machine. For example, the projection component 108 can position the machine by determining an initial pose of the machine using first sensor data generated using one or more position sensors of the machine (e.g., Global Positioning System (GPS)). Then, the projection component 108 can use second sensor data generated using one or more other sensors of the machine (e.g., one or more image sensors, one or more LiDAR sensors, one or more RADAR sensors, etc.) to refine and / or update the initial pose of the machine. For example, to refine the initial pose, the projection component 108 can compare one or more features represented by the second sensor data with one or more features represented by the map. At least based on the comparison, the projection component 108 can update the initial pose of the machine to an estimated pose (e.g., a positioning pose) of the machine in the environment. As described herein, the pose can represent the position of the machine (e.g., x-coordinate position, y-coordinate position, and / or z-coordinate position), the orientation of the machine (e.g., yaw, pitch, and / or roll), and / or any other position, pose, or orientation information.

[0046] At least based on the positioning, the projection component 108 can then use one or more techniques to project the 3D position associated with the object to a 2D position associated with the image. Additionally, the projection component 108 can use the label associated with the object being projected to generate a corresponding label of the object as depicted by the image. For a first example, if the projection component 108 projects the 3D coordinates of a road to the 2D coordinates of the image, the projection component 108 can generate a label of "road" indicating the 2D coordinates (e.g., pixels) of the image. For a second example, if the projection component 108 projects the 3D coordinates of a traffic line to the 2D coordinates of the image, the projection component 108 can generate a label of "traffic line" indicating the 2D coordinates (e.g., pixels) of the image.

[0047] For example, Figure 3Shows an example of projecting the 3D position of an object in a map (which may be represented by map data 110) to the 2D position of an image 202. As shown, at least based on performing the projection 302, the projection component 108 can generate labels for at least a road 304 (although only one is labeled for clarity), a road line 306 (although only one is labeled for clarity), a road boundary (although only one is labeled for clarity), a traffic sign 310, and a traffic pole 312. However, in other examples, the projection component 108 can use the map to generate labels for additional and / or alternative objects depicted by the image 202.

[0048] Return reference Figure 1A In an example, the process 100 can include the occlusion component 114 using the classification data 106 and the projected label data 112 to determine whether traffic objects (and / or other types of objects) are occluded in various parts of the image. For example, for an image, the occlusion component 114 can use the projected label data 112 to determine that a portion of the image (e.g., a pixel) is associated with a traffic object (e.g., a road line or a road boundary). Then, the occlusion component 114 can use the classification data 106 to determine the classification associated with that portion of the image. Additionally, the occlusion component 114 can use the classification to determine whether the traffic object is occluded in that portion of the image. In some examples, the occlusion component 114 can determine that the traffic object is not occluded in that portion of the image at least based on the classification corresponding (e.g., being similar) to the label of the traffic object, or determine that the traffic object is occluded in that portion of the image at least based on the classification not corresponding (e.g., being different) to the label.

[0049] For a first example, if the label of the traffic object indicates a road boundary, the occlusion component 114 may determine that the traffic object is not occluded at that portion of the image when the classification includes one or more first classifications corresponding to the road boundary (e.g., road or sidewalk), or determine that the traffic object is occluded at that portion of the image when the classification includes one or more second classifications not corresponding to the road boundary (e.g., car, vegetation, person, etc.). In some embodiments, when using a machine learning model (e.g., a machine learning model trained to detect dynamic objects such as cars, pedestrians, etc.) to determine that there is no classification of the image, and the label from the map data indicates a lane line or a road boundary or other feature or object of interest, then that feature or object may be determined to be unoccluded. As another example, if the label of the traffic object indicates a traffic sign, the occlusion component 114 may determine that the traffic object is not occluded at that portion of the image when the classification includes one or more first classifications corresponding to the traffic sign (e.g., traffic sign), or determine that the traffic object is occluded at that portion of the image when the classification includes one or more second classifications not corresponding to the traffic sign (e.g., car, vegetation, person, etc.).

[0050] In some examples, the occlusion component 114 may perform a similar process on one or more additional portions of the image. For example, the occlusion component 114 may perform a similar process on the image portions (e.g., pixels) associated with the traffic object. Additionally, process 100 may include the occlusion component 114 generating occlusion data 116 that indicates the image portions where the traffic object is occluded and / or the image portions where the traffic object is not occluded. For example, for a portion of the image, the occlusion data 116 may represent one of the following labels: a first label indicating that the traffic object is not occluded, a second label indicating that the traffic object is occluded, a third label indicating that the traffic object is occluded by a dynamic object, a fourth label indicating that the traffic object is occluded by a static object, and / or any other label. In some examples, process 100 may then continue to repeat for one or more additional images represented by the image data 104.

[0051] For example, Figure 4A illustrates an example of determining whether a traffic object is occluded at a portion of an image according to some embodiments of the present disclosure. In Figure 4AIn the example, the occlusion component 114 may analyze a portion 402 of the projection 302 relative to a portion 404 of the image 202. For a first example, the occlusion component 114 may determine that a first point 406(1) (e.g., a first pixel) associated with the image 202 is associated with a first point label 408(1) of the road boundary 308. Then, the occlusion component 114 may determine that the road boundary 308 is occluded at the first point 406(1) at least based on the first point 406(1) being classified as a vehicle 208 in the image 202. For a second example, the occlusion component 114 may determine that a second point 406(2) (e.g., a second pixel) associated with the image 202 is also associated with a second point label 408(2) of the road boundary 308. Then, the occlusion component 114 may determine that the road boundary 308 is not occluded at the second point 406(2) at least based on the second point 406(2) being classified as a road 210 in the image 202. Then, the occlusion component 114 may perform a similar process on one or more additional points associated with the image 202.

[0052] Next, Figure 4B An example of generating occlusion information 410 (which may be represented by occlusion data 116) associated with the image 202 in accordance with some embodiments of the present disclosure is shown. As shown, the occlusion information 410 may include a first label 412 of traffic objects (e.g., road boundaries and road lines) that are not occluded within the image 202 (although only one is labeled for clarity), where the first label 412 includes a dark line within the occlusion information 410. Additionally, the occlusion information 410 may include a second label 414 of traffic objects (e.g., road boundaries and road lines) that are occluded within the image 202 (although only one is labeled for clarity), where the second label 414 includes a gray line within the occlusion information 410.

[0053] Figure 1B An example data flow diagram of a process 118 for detecting occluded objects within an image using one or more point cloud processing techniques is shown. The process 118 may include a location component 120 receiving label data 122 and map data 110 (although, in some examples, the map data 110 may include the label data 122). In some examples, the label data 122 may represent labels projected onto the image (e.g., using the techniques described herein with respect to Figure 1Alabels of 3D points within the environment represented by the map data 110 and / or in one or more processes described herein. For an image, the location component 120 can then use one or more processes described herein (e.g., using machine-generated sensor data) to locate the machine that generates the image data 104 representing the image relative to the map. Then, the location component 120 can project the 3D position (e.g., 3D coordinates) of a point in the map associated with a traffic object to the 2D position of a portion (e.g., a pixel) of the image.

[0054] Then, the location component 120 can use the projection to determine information associated with this portion of the image. For example, the location component 120 can determine at least one ray that projects from this portion of the image (e.g., the 2D position of a pixel) to the 3D position of a point within the environment. Additionally, the location component 120 can generate ray data 124 representing information associated with the ray, such as 2D image position, direction, label (e.g., traffic object label), and / or any other information associated with the ray. The location component 120 can also determine the distance between the 2D image position and the 3D position of the point within the environment. Additionally, the location component 120 can generate distance data 126 representing the distance. Then, the location component 120 can perform similar processes on one or more other portions of the image.

[0055] For example, Figure 5 illustrates an example of projecting a 3D position associated with a point within an environment to a 2D position associated with a portion of an image 202 and then using the projection to determine information associated with this portion of the image 202. As shown, the location component 120 can use a map 502 (where, Figure 5 represents a simplified map for illustrative purposes) to project the 3D position 504 associated with a point in the environment to the 2D position 506 associated with a portion (e.g., a pixel) of the image 202, which is represented by the projection 508. As described herein, the 3D position 504 can at least include an x-coordinate position, a y-coordinate position, and a z-coordinate position associated with the point. Additionally, the 2D position can include an x-coordinate position and a y-coordinate position associated with this portion of the image 202. As described herein, in some examples, the location component 120 can use any technique to perform the projection 508.

[0056] Then, the location component 120 can determine information associated with this portion of the image 202. For example, the location component 120 can determine at least one ray projecting from a 2D location 506 associated with the image 202 to a 3D location 504 associated with a point in the environment, where the ray can also be represented by the projection 508. Additionally, the location component 120 can generate data representing information associated with the ray (e.g., ray data 124), such as the 2D location 506, direction, label (e.g., traffic object label), and / or any other information associated with the ray. The location component 120 can also use the 2D location 506 associated with the image 202 and the 3D location 504 associated with a point in the environment to determine a distance, which can be represented by the length of the projection. Additionally, the location component 120 can generate data representing the distance (e.g., distance data 126). Then, the location component 120 can perform a similar process on one or more other portions of the image 202.

[0057] Return reference Figure 1B Returning to the example of , the process 118 can include the occupancy component 128 receiving point cloud data 130 and / or LiDAR data 132. As described herein, the point cloud data 130 can represent a point cloud associated with the environment (e.g., a point cloud map, an occupancy map, etc.). For example, the point cloud data 130 can be generated using LiDAR data (and / or other types of distance data, such as RADAR data), which is generated by one or more machines as the machines navigate within the environment. For example, LiDAR data from the machines can be combined to generate the point cloud data 130 representing points within the environment. By combining LiDAR data from multiple sensors associated with machines within the environment, the point cloud data 130 can represent a dense number of points within the environment.

[0058] Then, the LiDAR data 132 can be generated by the machine as the machine navigates within the environment, where the machine also generates the image data 104 that is being processed for traffic object occlusion. For example, the machine can use one or more LiDAR sensors to generate the LiDAR data 132 while also using one or more image sensors to generate the image data 104. Thus, the LiDAR data 132 can better represent the current occupancy associated with the environment because the LiDAR data 132 can represent the current positions of objects in the environment, such as dynamic objects that may change with different time instances.

[0059] As shown, process 118 may subsequently include occupancy component 128 using point cloud data 130 and / or LiDAR data 132 to generate occupancy data 134 associated with the environment. For example, as described herein, occupancy data 134 may represent occupancy associated with the environment at a time approximate to the time when the image being processed is generated. For example, occupancy data 134 may represent points within the environment associated with objects located in the environment at an approximate time of image generation. In some examples, occupancy component 128 may generate occupancy data 134 at least by updating point cloud data 130 using LiDAR data 132.

[0060] For example, occupancy component 128 may insert at least a portion of the points represented by LiDAR data 132 into point cloud data 130. In some examples, if the points inserted into point cloud data 130 are farther away from the initial points represented by point cloud data 130 and in the same direction as the initial points, occupancy component 128 may remove the initial points from point cloud data 130. This is because the object associated with the initial points (e.g., reflecting away) may no longer be located within the environment. Thus, by removing the initial points, occupancy data 134 may indicate that the environmental region associated with the initial points is no longer occupied by an object. In other words, occupancy component 128 may generate occupancy data 134 by updating point cloud data 130 to indicate the current positions of the objects currently located within the environment at the time of image generation.

[0061] For example, Figure 6 An example of an occupancy map 602 that may be generated using point cloud data (e.g., point cloud data 130) and LiDAR data (e.g., LiDAR data 132) in accordance with one or more embodiments of the present disclosure is shown. As shown, occupancy map 602 may include a plurality of points 604 (although only one point is labeled for clarity), where at least a portion of the points are from the point cloud data and at least a portion of the points are from the LiDAR data. For example, the points 604 associated with static objects (e.g., road 210, road lines 212, traffic pole 214, traffic sign 216, and building 224) may be from the point cloud data and / or the LiDAR data. However, the points 604 associated with dynamic objects (e.g., vehicle 208, pedestrian 204, bicycle 206, and trash can 222) may be from the LiDAR data. Although Figure 6 the example shows that occupancy map 602 includes a specific number of points 604, in other examples, occupancy map 602 may include a smaller number of points or a larger number of points.

[0062] Returning to Figure 1BIn an example, process 118 may include distance component 136 using at least ray data 124 and occupancy data 134 to determine distances to points within the environment. For example, for a portion of the image (e.g., a pixel), distance component 136 may analyze ray data 124 to determine a 2D position associated with that portion of the image (e.g., the 2D position of the pixel within the image). Then, distance component 136 may project one or more rays from the 2D position into the occupancy map represented by occupancy data 134. As described herein, the number of rays may include, but is not limited to, one ray, two rays, five rays, ten rays, fifty rays, and / or any other number of rays. Then, distance component 136 may use the rays to determine the distance from the 2D position to a point in the occupancy map (e.g., the ray projection depth).

[0063] For example, distance component 136 may determine one or more distances associated with one or more rays, such as a respective distance for each ray. Then, distance component 136 may use the distances to determine a final distance associated with the 2D position. In some examples, occupancy component 136 may determine the final distance as the maximum distance among the distances. However, in other examples, distance component 136 may determine the final distance as the minimum distance, the average of the distances, the median of the distances, and / or using any other technique. Then, distance component 136 may generate distance data 138 representing the final distance associated with the 2D position within the image. Additionally, distance component 136 may perform a similar process to generate distance data 138 associated with one or more additional portions of the image.

[0064] Then, process 118 may include occlusion component 140 using distance data 126 and distance data 138 to determine whether traffic objects (and / or other types of objects) are occluded at various portions of the image. For example, for the image, occlusion component 140 may use distance data 126 to determine a first distance associated with a portion of the image (e.g., a pixel) where the portion of the image is associated with a traffic object. Then, occlusion component 140 may use distance data 138 to determine a second distance associated with that portion of the image (e.g., the distance determined using occupancy data 134). Then, occlusion component 140 may use at least the first distance and the second distance to determine whether the traffic object is occluded at that portion of the image.

[0065] For example, the occlusion component 140 can determine that a traffic object is occluded at this portion of the image at a threshold distance outside the first distance based at least on the second distance, or determine that the traffic object is not occluded at this portion of the image within the threshold distance from the first distance, where the threshold distance can be represented by threshold data 142. In some examples, the occlusion component 140 can use a set threshold distance for the points, such as but not limited to 0.1 meter, 0.5 meter, 1 meter, and / or any other distance. However, in other examples, the occlusion component 140 can dynamically determine the threshold distance for this portion of the image. In such examples, the occlusion component 140 can use one or more factors to determine the threshold distance, such as the incident angle associated with the projection ray and the ground thickness (e.g., average ground point cloud thickness).

[0066] For example, Figure 7 illustrates an example of determining a threshold distance associated with determining whether a portion of an image is occluded according to some embodiments of the present disclosure. As shown, the occlusion component 140 can determine an angle 702 associated with a ray 704 projected from a portion of the image and a thickness 706 associated with a ground point 708 from an occupancy map (e.g., which can be represented by occupancy data 134). Then, the occlusion component 140 can determine the threshold distance using the angle 702 and the thickness 706. For example, in Figure 7 the example, the occlusion component 140 can determine the threshold distance to include the length of a segment 710 of the ray 704 that is within the ground point 708. However, in other examples, the occlusion component 140 can use any other technique to determine the threshold distance.

[0067] In addition, Figure 8 illustrates an example of determining whether a traffic object is occluded in certain portions of an image 202 according to some embodiments of the present disclosure. As shown, the occlusion component 140 can determine a first distance 802 associated with a first portion of the image 202 using first distance data (e.g., distance data 126) and determine a second distance 804 associated with the first portion of the image 202 using second distance data (e.g., distance data 138). The occlusion component 140 can also perform one or more processes described herein to determine a first threshold distance associated with the first portion of the image 202. In addition, the occlusion component 140 can then determine that a traffic object (e.g., a road boundary) is occluded at the first portion of the image 202 based at least on the second distance 804 being outside the first threshold distance from the first distance 802. As shown, the traffic object can be occluded based on an object 806 (which can correspond to a vehicle 208) located within an environment 808 associated with the image 202.

[0068] Additionally, the occlusion component 140 may determine a first distance 810 associated with a second portion of the image 202 using first distance data (e.g., distance data 126), and determine a second distance 812 associated with the second portion of the image 202 using second distance data (e.g., distance data 138). The occlusion component 140 may also perform one or more processes described herein to determine a threshold distance associated with the second portion of the image 202. Additionally, the occlusion component 140 may then determine that a traffic object (e.g., a road boundary) is not occluded at the second portion of the image 202 based at least on the second distance 812 being within a second threshold distance from the first distance 810.

[0069] Although Figure 8 the example of only shows the object 806 that may correspond to the vehicle 208 depicted by the image 202 for clarity, in other examples, the environment 808 may include, but is not limited to, each object depicted by the image 202.

[0070] Returning to Figure 1B the example of, the process 118 may include the occlusion component 140 generating occlusion data 144 that indicates the portions of the image where the traffic object is occluded and / or the portions of the image where the traffic object is not occluded. For example, for a portion of the image, the occlusion data 144 may represent one of a first label, a second label, a third label, and / or any other label, where the first label indicates that the traffic object is not occluded, the second label indicates that the traffic object is occluded, and the third label indicates that occlusion cannot be determined. In some examples, the process 100 may then continue to repeat for one or more additional images represented by the image data 104. Additionally, in some examples, the occlusion data 144 may represent occlusion information similar to the occlusion data 116, such as the occlusion information 410 from Figure 4B the example of.

[0071] In some examples, one or more techniques may be used to perform the process 118 in order to optimize the processing performed for the process 118. For example, one or more criteria may be used to divide the image data 104 being processed into groups (e.g., chunks). For a first example, the image data 104 may be divided based at least on the distance traveled by the machine that generated the image data 104 such that each group includes the images generated by the machine as the machine travels a set distance. For example, a first group may include one or more first images generated as the machine travels a first set distance, a second group may include one or more second images generated as the machine travels a second set distance, a third group may include one or more third images generated as the machine travels a third set distance, and so on. In such an example, the set distance may include, but is not limited to, 1 meter, 5 meters, 10 meters, 20 meters, 50 meters, and / or any other distance.

[0072] For a second example, the image data 104 can be divided based at least on the time elapsed by the machine that generated the image data 104 such that each group includes the images generated by the machine when the machine has elapsed a set time. For example, the first group can include one or more first images generated when the machine has elapsed a first set time, the second group can include one or more second images generated when the machine has elapsed a second set time, the third group can include one or more third images generated when the machine has elapsed a third set time, and so on. In such an example, the set time can include, but is not limited to, 1 second, 10 seconds, 1 minute, 5 minutes, 10 minutes, and / or any other time period. Although these are just two example techniques for dividing the images represented by the image data 104 into groups, in other examples, any other technique can be used to divide the images into groups.

[0073] In some examples, the process 118 can then include processing these groups at different time instances. For example, the first group can be processed, then the second group, then the third group, and so on. By processing one group at a single time instance, the process 118 can also reduce the amount of map data 110 and / or point cloud data 130 that is processed. For example, if the group is associated with a set distance within the environment, a portion of the map data 110 and / or a portion of the point cloud data 130 associated with the set distance can be used during processing without using other portions of the map data 110 and / or other portions of the point cloud data 130. By reducing the amount of map data 110 and / or point cloud data 130 required for processing, the process 118 can again be optimized by reducing the amount of time required to process the image groups.

[0074] As described herein, in some examples, more than one processing technique can be combined to detect occluded objects in an image. For example, Figure 1C An example data flow diagram of a process 146 for detecting occluded objects in an image using image processing techniques and point cloud processing techniques is shown. As shown, the arbitration component 148 can receive at least occlusion data 116 generated using Figure 1A image processing techniques and occlusion data 144 generated using Figure 1B point cloud processing techniques. As described herein, for an image, the occlusion data 116 can indicate whether a traffic object is occluded at certain portions of the image, and the occlusion data 144 can also indicate whether a traffic object is occluded at certain portions of the image. In some examples, the occlusion data 116 can indicate the same traffic object occlusion as the occlusion data 144. However, in other examples, the occlusion data 116 can indicate one or more different traffic object occlusions compared to the occlusion data 144.

[0075] AsFigure 1C As shown in the example of, process 146 may also include arbitration component 148 receiving additional data, such as uncertainty data 150 and / or location data 152. As described herein, uncertainty data 150 may represent an uncertainty value associated with classifying an image using classification component 102. For example, for a portion of an image (e.g., a pixel), a low uncertainty value may indicate a high probability that the classification of that portion of the image is correct, while a high uncertainty value may indicate a low probability that the classification of that portion of the image is correct. In some examples, one or more techniques may be used to determine the uncertainty value.

[0076] For a first example, one or more feature tracking techniques may be used to track feature points between the images represented by image data 104. The uncertainty value may then be determined using the classifications associated with the tracked feature points between the images. For example, if the classifications between the images are similar for the tracked feature points, the uncertainty values for those points in the images may be low, and if the classifications between the images are different for the tracked points, the uncertainty values for those points in the images may be high. For a second example, a model (e.g., a model associated with classification component 102) may determine the uncertainty value associated with the classification. Although these are just a few example techniques for determining the uncertainty value associated with a classification, in other examples, additional and / or alternative techniques may be used to determine the uncertainty value.

[0077] Location data 152 may indicate one or more coordinates associated with the portion of the image (e.g., a pixel) that is labeled and / or classified. As described herein, in some examples, location data 152 may indicate the z - coordinate position associated with the portion of the image within the environment. However, in other examples, location data 152 may also represent the x - coordinate position and / or the y - coordinate position associated with the portion of the image.

[0078] Then, process 146 may include arbitration component 148 processing occlusion data 116, occlusion data 144, uncertainty data 150, and / or location data 152, and generating final occlusion data 154 associated with the image based at least on that processing. In some examples, arbitration component 148 may use one or more machine - learning models, one or more neural networks, one or more algorithms, one or more rules, etc. to process the data, which are configured to determine the final portions of the image where the traffic object is occluded and / or the final portions of the image where the traffic object is not occluded.

[0079] For a first example, and for a portion of an image, if both occlusion data 116 and occlusion data 144 indicate that a traffic object at that portion of the image is occluded, then arbitration component 148 may generate final occlusion data 154 to also indicate that the traffic object at that portion of the image is occluded. For a second example, and for a portion of an image, if both occlusion data 116 and occlusion data 144 indicate that a traffic object at that portion of the image is not occluded, then arbitration component 148 may generate final occlusion data 154 to also indicate that the traffic object at that portion of the image is not occluded.

[0080] For a third example, again for a portion of an image, if there is a difference between occlusion data 116 and occlusion data 144 as to whether a traffic object at that portion of the image is occluded, then arbitration component 148 may generate final occlusion data 154 to indicate that the traffic object at that portion of the image is not occluded. For a fourth example, again for a portion of an image, if there is a difference between occlusion data 116 and occlusion data 144 as to whether a traffic object at that portion of the image is occluded, then arbitration component 148 may generate final occlusion data 154 to indicate that the traffic object at that portion of the image is occluded.

[0081] For a fifth example, and again for a portion of an image, if there is again a difference between occlusion data 116 and occlusion data 144 as to whether a traffic object at that portion of the image is occluded, then arbitration component 148 may generate final occlusion data 154 to be similar to one of occlusion data 116 or occlusion data 144. For example, if occlusion data 116 indicates that a traffic object is occluded at that portion of the image and occlusion data 144 indicates that the traffic object is not occluded at that portion of the image, then arbitration component 148 may generate final occlusion data 154 to indicate that the traffic object is occluded at that portion of the image. In some examples, for this fifth example, arbitration component 148 may use uncertainty data 150 and / or location data 152 when determining whether to use occlusion data 116 or occlusion data 144.

[0082] For example, in some examples, when the uncertainty value associated with a portion of the image meets (e.g., is less than or equal to) a threshold (e.g., 1%, 5%, 10%, etc.), the arbitration component 148 may determine to use the occlusion data 116, or when the uncertainty value does not meet (e.g., is greater than) the threshold, the arbitration component 148 may determine not to use the occlusion data 144. The arbitration component 148 may make such a decision because the occlusion data 116 may be more reliable in cases of lower uncertainty values than in cases of higher uncertainty values. In some examples, when the z - coordinate position associated with a portion of the image meets (e.g., is equal to or greater than) a threshold distance (e.g., 1 meter, 5 meters, 10 meters, etc.), the arbitration component 148 may determine to use the occlusion data 116, or when the z - coordinate value does not meet (e.g., is less than) the threshold distance, the arbitration component 148 may determine to use the occlusion data 144. The arbitration component 148 may make such a determination because a larger z - coordinate value may indicate that the labeled point associated with that portion of the image is not contained on the same surface (e.g., road) as the machine, and thus, the occlusion data 116 may be less reliable.

[0083] Although these are just a few example techniques of how the arbitration component 148 may use the occlusion data 116, the occlusion data 144, the uncertainty data 150, and / or the position data 152 to generate the final occlusion data 154, in other examples, the arbitration component 148 may use additional and / or alternative techniques.

[0084] The process 146 may include the arbitration component 148 outputting the final occlusion data 154 that indicates the portions of the image where the traffic object is occluded and / or the portions of the image where the traffic object is not occluded. For example, for a portion of the image, the final occlusion data 154 may represent one of a first label, a second label, a third label, a fourth label, a fifth label, and / or any other label, where the first label indicates that the traffic object is not occluded, the second label indicates that the traffic object is occluded, the third label indicates that the traffic object is occluded by a dynamic object, the fourth label indicates that the traffic object is occluded by a static object, and the fifth label indicates that the occlusion at that portion of the image may not be determined. Additionally, in some examples, the occlusion data 154 may represent occlusion information similar to the occlusion data 116, such as the occlusion information 410 from Figure 4B the example of

[0085] Now refer to Figures 9 to 11, each block of methods 900, 1000, and 1100 described herein includes a computational process that can be performed using any combination of hardware, firmware, and / or software. For example, various functions can be performed by a processor executing instructions stored in a memory. Methods 900, 1000, and 1100 can also be embodied as computer-usable instructions stored on a computer storage medium. Methods 900, 1000, and 1100 can be provided by a stand-alone application, a service, or a hosted service (stand-alone or in combination with another hosted service), or a plug-in of another product, to name a few. Additionally, methods 900, 1000, and 1100 are described Figures 1A to 1C by way of example. However, these methods 900, 1000, and 1100 can alternatively or additionally be performed by any one system or any combination of systems, including but not limited to the systems described herein.

[0086] Figure 9 is a flowchart showing method 900 for detecting occluded objects within an image using one or more image processing techniques according to some embodiments of the present disclosure. Method 900 may include, at block B902, determining a classification corresponding to a portion of an image using one or more machine learning models and at least based on image data representative of the image. For example, classification component 102 may process image data 104 representative of the image using a machine learning model, where the machine learning model is trained to generate various classifications and / or segmentation masks associated with the objects depicted by the image. For example, at least based on the processing, the machine learning model may at least determine a classification associated with a portion of the image. As described herein, a portion of an image may include pixels, regions, zones, tiles, blocks, and / or any other portion associated with the image.

[0087] Method 900 may include, at block B904, determining that a point within the environment corresponding to a portion of the image is associated with a traffic object at least based on map data associated with the environment. For example, projection component 108 may process at least map data 110 representing at least a portion of the environment depicted by the image. At least based on the processing, projection component 108 may project a 3D position associated with a point within the environment to a 2D position associated with that portion of the image. Additionally, projection component 108 may subsequently use a label associated with the point in the map to generate a label associated with that portion of the image, where the label includes a traffic object. As described herein, traffic objects may include, but are not limited to, roads, sidewalks, road markings, traffic poles, traffic signs, traffic signals, and / or any other type of object associated with a navigation machine within the environment.

[0088] Method 900 may include, at block B906, determining whether a traffic object is occluded at that portion of the image based at least on classification. For example, the occlusion component 114 may use classification data 106 representing the classification associated with that portion of the image and projected label data 112 representing the label generated for that portion of the image to determine whether the traffic object is occluded at that portion of the image. As described herein, in some examples, the occlusion component 114 may determine that the traffic object is not occluded at that portion of the image based at least on a classification corresponding to (e.g., similar to) the label of the traffic object, or determine that the traffic object is occluded at that portion of the image based at least on a classification not corresponding to (e.g., different from) the label. For example, if the traffic object includes a road boundary or a road line, the occlusion component 114 may determine that the traffic object is not occluded at that portion of the image when the classification includes a road or a sidewalk, or determine that the traffic object is occluded at that portion of the image when the classification includes a car, vegetation, or a pedestrian.

[0089] Method 900 may include, at block B908, generating data indicating whether the traffic object is occluded at that portion of the image. For example, the occlusion component 114 may generate occlusion data 116 indicating whether the traffic object is occluded at that portion of the image. As described herein, the occlusion data 116 may represent one of a first label indicating that the traffic object is not occluded, a second label indicating that the traffic object is occluded, a third label indicating that the traffic object is occluded by a dynamic object, a fourth label indicating that the traffic object is occluded by a static object, and / or any other label.

[0090] Figure 10 is a flowchart showing a method 1000 for detecting occluded objects in an image using one or more point cloud techniques according to some embodiments of the present disclosure. Method 1000 may include, at block B1002, determining a first distance associated with a point corresponding to a portion of the image within the environment based at least on map data associated with the environment. For example, the location component 120 may use the map data 110 to project a 3D location associated with a point within the environment to a 2D location associated with a portion of the image (e.g., a pixel). Then, the location component 120 may use the 2D location and the 3D location to determine the first distance associated with the point. As described herein, the point may be associated with a traffic object located within the environment.

[0091] Method 1000 may include, at block B1004, determining a second distance associated with a point in the environment based at least on point cloud data. For example, distance component 136 may use occupancy data 134 representing the point cloud to determine the second distance associated with a point in the environment. As described herein, in some examples, to determine the second distance, distance component 136 may project a plurality of rays from a 2D position associated with a portion of the image using the point cloud. Distance component 136 may then use the distances associated with the plurality of rays to determine the second distance.

[0092] Method 1000 may include, at block B1006, determining whether a traffic object is occluded at a portion of the image based at least on the first distance and the second distance. For example, occlusion component 140 may determine whether a traffic object is occluded at a portion of the image based at least on the first distance and the second distance. As described herein, in some examples, occlusion component 140 may determine that the traffic object is occluded at a portion of the image based at least on the second distance being outside a threshold distance from the first distance, or determine that the traffic object is not occluded at a portion of the image based at least on the second distance being within a threshold distance from the first distance. Additionally, in such examples, occlusion component 140 may use a set threshold distance and / or may dynamically determine the threshold distance.

[0093] Method 1000 may include, at block B1008, generating data indicating whether a traffic object is occluded at a portion of the image. For example, occlusion component 140 may generate occlusion data 144 indicating whether a traffic object is occluded at a portion of the image. As described herein, occlusion data 144 may represent one of a first label indicating that the traffic object is not occluded, a second label indicating that the object is occluded, a third label indicating that the traffic object is occluded by a dynamic object, a fourth label indicating that the traffic object is occluded by a static object, and / or any other label.

[0094] Figure 11 is a flowchart of a method 1100 for detecting occluded objects in an image using one or more image processing techniques and one or more point cloud techniques according to some embodiments of the present disclosure. For example, at block B1102, process 1100 may include receiving first data indicating whether a traffic object is occluded at a portion of the image, the first data being generated using image processing. For example, arbitration component 148 may receive occlusion data 116 indicating whether a traffic object is occluded at a portion of the image, where occlusion data 116 is generated using Figure 1A process 100.

[0095] Method 1100 may include, at block B1104, receiving second data that indicates whether a traffic object is occluded at that portion of the image, the second data being generated using point cloud data. For example, arbitration component 148 may receive occlusion data 144, which also indicates whether a traffic object is occluded at that portion of the image, where occlusion data 144 is generated using Figure 1B process 118.

[0096] Method 1100 may include, at block B1106, determining whether a traffic object is occluded at that portion of the image based at least on the first data and the second data. For example, arbitration component 148 may determine whether a traffic object is occluded at that portion of the image based at least on occlusion data 116 and occlusion data 144. In some examples, arbitration component 148 may make the determination using additional data (such as uncertainty data 150 and / or location data 152). Additionally, arbitration component 148 may use one or more techniques to make the determination, which will be described in more detail herein.

[0097] Method 1100 may include, at block B1108, generating third data that indicates whether a traffic object is occluded at that portion of the image. For example, arbitration component 148 may generate occlusion data 154, which indicates whether a traffic object is occluded at that portion of the image. As described herein, occlusion data 154 may represent one of a first label indicating that the traffic object is not occluded, a second label indicating that the traffic object is occluded, a third label indicating that the traffic object is occluded by a dynamic object, a fourth label indicating that the traffic object is occluded by a static object, and / or any other label.

[0098] Example Autonomous Vehicle

[0099] Figure 12AFIG. is an illustration of an exemplary autonomous vehicle 1200 in accordance with some embodiments of the present disclosure. The autonomous vehicle 1200 (alternatively, referred to herein as “vehicle 1200”) may include, but is not limited to, passenger vehicles such as cars, trucks, buses, first responder vehicles, shuttle vehicles, electric or motorized bicycles, motorcycles, fire trucks, police vehicles, ambulances, boats, construction vehicles, underwater vessels, robotic vehicles, drones, airplanes, vehicles coupled to trailers (e.g., semi-trailer trucks for hauling cargo) and / or another type of vehicle (e.g., driverless and / or accommodating one or more passengers). Autonomous vehicles are generally described according to the levels of automation defined by the National Highway Traffic Safety Administration (NHTSA), a division of the U.S. Department of Transportation, and the Society of Automotive Engineers (SAE) in “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles” (Standard No. J3016-201806, issued June 15, 2018, Standard No. J3016-201609, issued September 30, 2016, and previous and future versions of the standard). The vehicle 1200 may be capable of implementing functions corresponding to one or more of Levels 3-5 of the autonomous driving level. For example, depending on the embodiment, the vehicle 1200 may be capable of driver assistance (Level 1), partial automation (Level 2), conditional automation (Level 3), highly automated (Level 4), and / or fully automated (Level 5). The term “autonomous” as used herein may include any and / or all types of autonomy of the vehicle 1200 or other machines, such as fully autonomous, highly autonomous, conditionally autonomous, partially autonomous, providing assisted autonomy, semi-autonomous, primarily autonomous, or other designations.

[0100] The vehicle 1200 may include components such as a chassis, a body, wheels (e.g., 2, 4, 6, 8, 18, etc.), tires, axles, and other components of the vehicle. The vehicle 1200 may include a propulsion system 1250, such as an internal combustion engine, a hybrid power plant, a fully electric motor, and / or another type of propulsion system. The propulsion system 1250 may be connected to the driveline of the vehicle 1200, which may include a transmission, to effect the propulsion of the vehicle 1200. The propulsion system 1250 may be controlled in response to signals received from the throttle / accelerator 1252.

[0101] A steering system 1254 that may include a steering wheel can be used to steer a vehicle 1200 (e.g., along a desired path or route) while the propulsion system 1250 is operating (e.g., while the vehicle is in motion). The steering system 1254 can receive a signal from a steering actuator 1256. For fully autonomous (Level 5) functionality, the steering wheel can be optional.

[0102] A brake sensor system 1246 can be used to operate vehicle brakes in response to receiving a signal from a brake actuator 1248 and / or a brake sensor.

[0103] One or more controllers 1236 that may include one or more system-on-chips (SoCs) 1204 ( Figure 12C ) and / or one or more GPUs can provide signals (e.g., representing commands) to one or more components and / or systems of the vehicle 1200. For example, one or more controllers can send signals to operate vehicle brakes via one or more brake actuators 1248, operate the steering system 1254 via one or more steering actuators 1156, and operate the propulsion system 1250 via one or more throttles / accelerators 1252. One or more controllers 1236 can include one or more on-board (e.g., integrated) computing devices (e.g., supercomputers) that process sensor signals and output operation commands (e.g., signals representing commands) to enable autonomous driving and / or assist a human driver in driving the vehicle 1200. One or more controllers 1236 can include a first controller 1236 for autonomous driving functionality, a second controller 1236 for functional safety functionality, a third controller 1236 for artificial intelligence functionality (e.g., computer vision), a fourth controller 1236 for infotainment functionality, a fifth controller 1236 for redundancy in emergency situations, and / or other controllers. In some examples, a single controller 1236 can handle two or more of the above functions, two or more controllers 1236 can handle a single function, and / or any combination thereof.

[0104] One or more controllers 1236 may provide signals for controlling one or more components and / or systems of vehicle 1200 in response to sensor data (e.g., sensor inputs) received from one or more sensors. Sensor data may be received from, for example and without limitation, a Global Navigation Satellite System (“GNSS”) sensor 1258 (e.g., a Global Positioning System sensor), a RADAR sensor 1260, an ultrasonic sensor 1262, a LIDAR sensor 1264, an Inertial Measurement Unit (IMU) sensor 1266 (e.g., an accelerometer, a gyroscope, a magnetic compass, a magnetometer, etc.), a microphone 1296, a stereo camera 1268, a wide-angle camera 1270 (e.g., a fisheye camera), an infrared camera 1272, a surround camera 1274 (e.g., a 360-degree camera), a long-range and / or mid-range camera 1298, a speed sensor 1244 (e.g., for measuring the speed of vehicle 1200), a vibration sensor 1242, a steering sensor 1240, a brake sensor (e.g., as part of a brake sensor system 1246), and / or other sensor types.

[0105] One or more of the controllers 1236 may receive inputs (e.g., represented by input data) from the instrument cluster 1232 of vehicle 1200 and provide outputs (e.g., represented by output data, display data, etc.) via a Human Machine Interface (HMI) display 1234, an audible annunciator, a speaker, and / or via other components of vehicle 1200. These outputs may include information such as vehicle speed, rate, time, map data (e.g., Figure 12C a high-definition (“HD”) map 1222), location data (e.g., the location of vehicle 1200 on a map, for example), direction, the locations of other vehicles (e.g., occupancy grids), and information about objects and object states as perceived by the controller 1236, and so on. For example, the HMI display 1234 may display information about the presence of one or more objects (e.g., street signs, warning signs, traffic light changes, etc.) and / or information about driving maneuvers that the vehicle has made, is making, or will make (e.g., changing lanes now, exiting 34B in two miles, etc.).

[0106] Vehicle 1200 also includes a network interface 1224, which may communicate over one or more networks using one or more wireless antennas 1226 and / or a modem. For example, network interface 1224 may be capable of communicating via Long Term Evolution (“LTE”), Wideband Code Division Multiple Access (“WCDMA”), Universal Mobile Telecommunications System (“UMTS”), Global System for Mobile Communications (“GSM”), IMT-CDMA Multi-Carrier (“CDMA2000”), etc. One or more wireless antennas 1226 may also enable communication between objects (such as vehicles, mobile devices, etc.) in an environment using one or more local area networks such as Bluetooth, Bluetooth Low Energy (“LE”), Z-Wave, ZigBee, etc. and / or one or more low power wide area networks (“LPWAN”) such as LoRaWAN, SigFox, etc.

[0107] Figure 12B For an example autonomous vehicle 1200 in accordance with some embodiments of the present disclosure for Figure 12A Example camera positions and fields of view for the example autonomous vehicle 1200. The cameras and respective fields of view are one example embodiment and are not intended to be limiting. For example, additional and / or alternative cameras may be included, and / or these cameras may be located at different positions on vehicle 1200.

[0108] The camera type for the cameras may include, but is not limited to, digital cameras that may be adapted to be used with components and / or systems of vehicle 1200. The cameras may operate under an Automotive Safety Integrity Level (ASIL) B and / or under another ASIL. The camera type may have any image capture rate, such as 60 frames per second (fps), 120 fps, 240 fps, etc., depending on the embodiment. The cameras may be capable of using a rolling shutter, a global shutter, another type of shutter, or a combination thereof. In some examples, the color filter array may include a Red Clear Clear Clear (RCCC) color filter array, a Red Clear Clear Blue (RCCB) color filter array, a Red Blue Green Clear (RBGC) color filter array, a Foveon X3 color filter array, a Bayer sensor (RGGB) color filter array, a monochrome sensor color filter array, and / or another type of color filter array. In some embodiments, clear pixel cameras such as cameras with RCCC, RCCB, and / or RBGC color filter arrays may be used in an effort to increase light sensitivity.

[0109] In some examples, one or more of the cameras can be used to perform Advanced Driver Assistance System (ADAS) functions (e.g., as part of a redundant or fail-safe design). For example, a multi-functional monocular camera can be installed to provide functions including lane departure warning, traffic sign assistance, and intelligent headlight control. One or more of the cameras (e.g., all cameras) can record and provide image data (e.g., video) simultaneously.

[0110] One or more of the cameras can be installed in mounting components such as custom-designed (three-dimensional (“3D”) printed) components to cut off stray light and reflections from inside the vehicle (e.g., reflections from the dashboard reflected in the windshield mirror) that may interfere with the image data capture ability of the cameras. Regarding the wing mirror mounting component, the wing mirror assembly can be custom 3D printed such that the camera mounting plate matches the shape of the wing mirror. In some examples, one or more cameras can be integrated into the wing mirror. For side-view cameras, one or more cameras can also be integrated into the four pillars at each corner of the cab.

[0111] A camera having a field of view that includes an environmental portion in front of the vehicle 1200 (e.g., a front camera) can be used for surround view to help identify forward paths and obstacles and, with the help of one or more controllers 1236 and / or a control SoC, assist in providing information crucial for generating an occupancy grid and / or determining a preferred vehicle path. The front camera can be used to perform many of the same ADAS functions as LIDAR, including emergency braking, pedestrian detection, and collision avoidance. The front camera can also be used for ADAS functions and systems, including lane departure warning (“LDW”), adaptive cruise control (“ACC”), and / or other functions such as traffic sign recognition.

[0112] A variety of cameras can be used in a front-mounted configuration, including, for example, a monocular camera platform including a complementary metal oxide semiconductor (“CMOS”) color imager. Another example can be a wide-angle camera 1270, which can be used to sense objects (e.g., pedestrians, intersection traffic, or bicycles) entering the field of view from the periphery. Although Figure 12B only one wide-angle camera is illustrated, any number (including zero) of wide-angle cameras 1270 can be present on the vehicle 1200. Additionally, a long-range camera 1298 (e.g., a long-range stereo camera pair) can be used for depth-based object detection, especially for objects for which a neural network has not been trained. The long-range camera 1298 can also be used for object detection and classification and basic object tracking.

[0113] Any number of stereo cameras 1268 may also be included in a front-facing configuration. In at least one embodiment, one or more stereo cameras 1268 may include an integrated control unit that includes a scalable processing unit that may provide a multi-core microprocessor and programmable logic ("FPGA") with an integrated controller area network ("CAN") or Ethernet interface on a single chip. Such a unit may be used to generate a 3D map of the vehicle environment, including distance estimates for all points in the image. Alternatively, the stereo camera 1268 may include a compact stereo vision sensor that may include two camera lenses (one on the left and one on the right) and an image processing chip that may measure the distance from the vehicle to a target object and activate autonomous emergency braking and lane departure warning functions using the generated information (e.g., metadata). Other types of stereo cameras 1268 may be used in addition to or in place of those described herein.

[0114] Cameras having a field of view that includes an environmental portion of the side of the vehicle 1200 (e.g., side-view cameras) may be used for surround view, providing information used to create and update an occupancy grid and generate side-impact collision warnings. For example, surround cameras 1274 (e.g., four surround cameras 1274 as shown in Figure 12B may be disposed on the vehicle 1200. The surround cameras 1274 may include wide-angle cameras 1270, fish-eye cameras, 360-degree cameras, and / or the like. By way of example, four fish-eye cameras may be disposed on the front, rear, and sides of the vehicle. In an alternative arrangement, the vehicle may use three surround cameras 1274 (e.g., left, right, and rear) and may utilize one or more other cameras (e.g., a forward-facing camera) as a fourth surround camera.

[0115] Cameras having a field of view that includes an environmental portion of the rear of the vehicle 1200 (e.g., rear-view cameras) may be used for assisting with parking, surround view, rear collision warnings, and creating and updating an occupancy grid. A variety of cameras may be used, including but not limited to cameras that are also suitable as front-facing cameras as described herein (e.g., long-range and / or mid-range cameras 1298, stereo cameras 1268, infrared cameras 1272, etc.).

[0116] Figure 12C For use in accordance with some embodiments of the present disclosure Figure 12ABlock diagram of an example system architecture of an example autonomous vehicle 1200. It should be understood that this and other arrangements described herein are set forth only as examples. Other arrangements and elements (e.g., machines, interfaces, functions, orders, function groupings, etc.) may be used in addition to or instead of those shown, and some elements may be entirely omitted. Further, many of the elements described herein are functional entities that may be implemented as discrete or distributed components or in combination with other components, and in any suitable combination and location. The various functions described herein as being performed by entities may be implemented by hardware, firmware, and / or software. For example, the various functions may be implemented by a processor executing instructions stored in memory.

[0117] Figure 12C Each of the components, features, and systems in vehicle 1200 is illustrated as being connected via bus 1202. Bus 1202 may include a Controller Area Network (CAN) data interface (alternatively referred to herein as the "CAN bus"). CAN may be a network within vehicle 1200 used to assist in controlling various features and functions of vehicle 1200, such as driving of brakes, acceleration, braking, steering, windshield wipers, and the like. The CAN bus may be configured to have dozens or even hundreds of nodes, each with its own unique identifier (e.g., CAN ID). The CAN bus may be read to find steering wheel angle, ground speed, engine revolutions per minute (RPM), button positions, and / or other vehicle status indicators. The CAN bus may be ASIL B compliant.

[0118] Although bus 1202 is described herein as a CAN bus, this is not intended to be limiting. For example, in addition to or instead of a CAN bus, FlexRay and / or Ethernet may be used. Further, although bus 1202 is shown as a single line, this is not intended to be limiting. For example, any number of buses 1202 may exist, which may include one or more CAN buses, one or more FlexRay buses, one or more Ethernet buses, and / or one or more other types of buses using different protocols. In some examples, two or more buses 1202 may be used to perform different functions, and / or may be used for redundancy. For example, a first bus 1202 may be used for collision avoidance functions, and a second bus 1202 may be used for drive control. In any example, each bus 1202 may communicate with any component of vehicle 1200, and two or more buses 1202 may communicate with the same component. In some examples, each SoC 1204, each controller 1236, and / or each computer within the vehicle may have access to the same input data (e.g., input from sensors of vehicle 1200), and may be connected to a common bus such as a CAN bus.

[0119] Vehicle 1200 may include one or more controllers 1236, such as those described herein with respect to Figure 12A the controllers described. The controller 1236 may be used for a variety of functions. The controller 1236 may be coupled to any other different components and systems of the vehicle 1200 and may be used for the control of the vehicle 1200, the artificial intelligence of the vehicle 1200, the infotainment for the vehicle 1200, and / or the like.

[0120] Vehicle 1200 may include one or more system-on-chips (SoCs) 1204. The SoC 1204 may include a CPU 1206, a GPU 1208, a processor 1210, a cache 1212, an accelerator 1214, a data store 1216, and / or other components and features not shown. In a variety of platforms and systems, the SoC 1204 may be used to control the vehicle 1200. For example, one or more SoCs 1204 may be combined with an HD map 1222 in a system (such as a system of the vehicle 1200), and the HD map may obtain map refreshes and / or updates from one or more servers (such as Figure 12D one or more servers 1278) via a network interface 1224.

[0121] The CPU 1206 may include a CPU cluster or a CPU complex (alternatively, referred to herein as "CCPLEX"). The CPU 1206 may include multiple cores and / or an L2 cache. For example, in some embodiments, the CPU 1206 may include eight cores in a coherent multi-processor configuration. In some embodiments, the CPU 1206 may include four dual-core clusters, each of which has a dedicated L2 cache (such as a 2MB L2 cache). The CPU 1206 (such as CCPLEX) may be configured to support simultaneous cluster operation such that any combination of the clusters of the CPU 1206 can be active at any given time.

[0122] The CPU 1206 may implement power management capabilities including one or more of the following features: Each hardware block may automatically perform clock gating when idle to save dynamic power; Due to the execution of WFI / WFE instructions, each core clock may be gated when the core is not actively executing instructions; Each core may independently perform power gating; When all cores perform clock gating or power gating, each core cluster may be independently clock gated; and / or When all cores perform power gating, each core cluster may be independently power gated. The CPU 1206 may further implement an enhanced algorithm for managing power states, where allowed power states and desired wake-up times are specified, and the hardware / microcode determines the optimal power state for the cores, clusters, and CCPLEX to enter. The processing cores may support a simplified power state entry sequence in software, and this work is offloaded to the microcode.

[0123] The GPU 1208 may include an integrated GPU (alternatively, referred to herein as "iGPU"). The GPU 1208 may be programmable and efficient for parallel workloads. In some examples, the GPU 1208 may use an enhanced tensor instruction set. The GPU 1208 may include one or more streaming microprocessors, where each streaming microprocessor may include an L1 cache (e.g., an L1 cache with at least 96KB of storage capacity), and two or more of these streaming microprocessors may share an L2 cache (e.g., an L2 cache with 512KB of storage capacity). In some embodiments, the GPU 1208 may include at least eight streaming microprocessors. The GPU 1208 may use a compute application programming interface (API). Additionally, the GPU 1208 may use one or more parallel computing platforms and / or programming models (e.g., NVIDIA's CUDA).

[0124] In automotive and embedded use cases, the GPU 1208 can be power optimized for best performance. For example, the GPU 1208 can be fabricated on fin field-effect transistors (FinFETs). However, this is not intended to be limiting, and the GPU 1208 can be fabricated using other semiconductor manufacturing processes. Each streaming microprocessor can incorporate a number of mixed-precision processing cores divided into multiple blocks. By way of example and not limitation, 64 PF32 cores and 32 PF64 cores can be divided into four processing blocks. In such an example, each processing block can be allocated 16 FP32 cores, 8 FP64 cores, 16 INT32 cores, two mixed-precision NVIDIA tensor cores for deep learning matrix arithmetic, an L0 instruction cache, a warp scheduler, a dispatch unit, and / or a 64KB register file. Additionally, the streaming microprocessor can include separate parallel integer and floating-point data paths to enable efficient execution of workloads leveraging a mix of compute and addressing computations. The streaming microprocessor can include separate thread scheduling capabilities to allow for finer-grained synchronization and cooperation between parallel threads. The streaming microprocessor can include a combined L1 data cache and shared memory unit to improve performance while simplifying programming.

[0125] The GPU 1208 can include, in some examples, high-bandwidth memory (HBM) that provides a peak memory bandwidth of approximately 900GB / s and / or a 16GB HBM2 memory subsystem. In some examples, in addition to or alternatively to HBM memory, synchronous graphics random access memory (SGRAM), such as fifth-generation graphics double data rate synchronous random access memory (GDDR5), can be used.

[0126] The GPU 1208 can include unified memory technology that includes access counters to allow memory pages to be more precisely migrated to the processors that most frequently access them, thereby improving the efficiency of the memory ranges shared between processors. In some examples, address translation service (ATS) support can be used to allow the GPU 1208 to directly access the CPU 1206 page tables. In such an example, when the GPU 1208 memory management unit (MMU) experiences a miss, an address translation request can be transmitted to the CPU 1206. In response, the CPU 1206 can look up the virtual-physical mapping for the address in its page table and transmit the translation back to the GPU 1208. In this way, the unified memory technology can allow for a single unified virtual address space for the memory of both the CPU 1206 and the GPU 1208, thus simplifying GPU 1208 programming and porting applications to the GPU 1208.

[0127] In addition, the GPU 1208 may include an access counter that can track how frequently the GPU 1208 accesses the memory of other processors. The access counter can help ensure that memory pages are moved to the physical memory of the processor that most frequently accesses those pages.

[0128] The SoC 1204 may include any number of caches 1212, including those described herein. For example, the cache 1212 may include an L3 cache that is available to both the CPU 1206 and the GPU 1208 (e.g., that is connected to both the CPU 1206 and the GPU 1208). The cache 1212 may include a write-back cache that can track the state of lines, for example, by using a cache coherence protocol (such as MEI, MESI, MSI, etc.). Depending on the embodiment, the L3 cache may include 4MB or more, but smaller cache sizes may also be used.

[0129] The SoC 1204 may include an arithmetic logic unit (ALU) that can be utilized in the processing of performing any of a variety of tasks or operations regarding the vehicle 1200, such as processing a DNN. In addition, the SoC 1204 may include a floating point unit (FPU) (or other math co-processor or digital co-processor type) for performing mathematical operations within the system. For example, the SoC 104 may include one or more FPUs integrated as execution units within the CPU 1206 and / or the GPU 1208.

[0130] The SoC 1204 may include one or more accelerators 1214 (such as hardware accelerators, software accelerators, or a combination thereof). For example, the SoC 1204 may include a hardware accelerator cluster that can include optimized hardware accelerators and / or large on-chip memory. This large on-chip memory (e.g., 4MB SRAM) can enable the hardware accelerator cluster to accelerate neural networks and other computations. The hardware accelerator cluster can be used to supplement the GPU 1208 and offload some of the tasks of the GPU 1208 (e.g., freeing up more cycles of the GPU 1208 for performing other tasks). As an example, the accelerator 1214 can be used for targeted workloads that are stable enough to be easily accelerated (such as perception, convolutional neural networks (CNNs), etc.). When used herein, the term "CNN" may include all types of CNNs, including region-based or region convolutional neural networks (RCNNs) and fast RCNNs (e.g., for object detection).

[0131] The accelerator 1214 (e.g., a hardware accelerator cluster) may include a Deep Learning Accelerator (DLA). The DLA may include one or more Tensor Processing Units (TPUs) that can be configured to provide an additional one trillion operations per second for deep learning applications and inference. The TPU may be an accelerator configured to perform image processing functions (e.g., for CNN, RCNN, etc.) and optimized for executing image processing functions. The DLA may be further optimized for a specific set of neural network types and floating-point operations, as well as inference. The design of the DLA may provide higher performance per millimeter than a general-purpose GPU and far exceed the performance of a CPU. The TPU may perform several functions, including single-instance convolution functions, supporting INT8, INT16, and FP16 data types for both features and weights, as well as post-processor functions.

[0132] The DLA may execute neural networks, especially CNNs, quickly and efficiently for any of a variety of functions on processed or unprocessed data, such as, for example and without limitation: CNNs for object recognition and detection using data from a camera sensor; CNNs for distance estimation using data from a camera sensor; CNNs for emergency vehicle detection and identification and detection using data from a microphone; CNNs for face recognition and vehicle owner recognition using data from a camera sensor; and / or CNNs for security and / or safety-related events.

[0133] The DLA may perform any function of the GPU 1208, and by using inference accelerators, for example, a designer may configure the DLA or the GPU 1208 for any function. For example, a designer may focus the processing and floating-point operations of a CNN on the DLA and leave other functions to the GPU 1208 and / or other accelerators 1214.

[0134] The accelerator 1214 (e.g., a hardware accelerator cluster) may include a Programmable Vision Accelerator (PVA), which may alternatively be referred to herein as a computer vision accelerator. The PVA may be designed and configured to accelerate computer vision algorithms for Advanced Driver Assistance Systems (ADAS), autonomous driving, and / or augmented reality (AR) and / or virtual reality (VR) applications. The PVA may provide a balance between performance and flexibility. For example, each PVA may include, for example and without limitation, any number of Reduced Instruction Set Computer (RISC) cores, Direct Memory Access (DMA), and / or any number of vector processors.

[0135] The RISC cores can interact with image sensors (such as the image sensors of any of the cameras described herein), image signal processors, and / or the like. Each of these RISC cores can include any number of memories. Depending on the embodiment, the RISC cores can use any of several protocols. In some examples, the RISC cores can execute a real-time operating system (RTOS). The RISC cores can be implemented using one or more integrated circuit devices, application-specific integrated circuits (ASICs), and / or storage devices. For example, the RISC cores can include an instruction cache and / or tightly coupled RAM.

[0136] The DMA can enable components of the PVA to access system memory independently of the CPU 1206. The DMA can support any number of features used to optimize the PVA, including but not limited to supporting multi-dimensional addressing and / or circular addressing. In some examples, the DMA can support addressing up to six or more dimensions, which can include block width, block height, block depth, horizontal block step, vertical block step, and / or depth step.

[0137] The vector processor can be a programmable processor that can be designed to efficiently and flexibly execute programming for computer vision algorithms and provide signal processing capabilities. In some examples, the PVA can include a PVA core and two vector processing subsystem partitions. The PVA core can include a processor subsystem, one or more DMA engines (such as two DMA engines), and / or other peripherals. The vector processing subsystem can operate as the main processing engine of the PVA and can include a vector processing unit (VPU), an instruction cache, and / or vector memory (such as VMEM). The VPU core can include a digital signal processor, such as a single instruction multiple data (SIMD), very long instruction word (VLIW) digital signal processor. The combination of SIMD and VLIW can enhance throughput and rate.

[0138] Each of the vector processors may include an instruction cache and may be coupled to dedicated memory. As a result, in some examples, each of the vector processors may be configured to execute independently of the other vector processors. In other examples, the vector processors included in a particular PVA may be configured to employ data parallelization. For example, in some embodiments, multiple vector processors included in a single PVA may execute the same computer vision algorithm, but on different regions of an image. In other examples, the vector processors included in a particular PVA may simultaneously execute different computer vision algorithms on the same image, or even execute different algorithms on sequential images or portions of an image. Among other things, any number of PVAs may be included in a hardware accelerator cluster, and any number of vector processors may be included in each of these PVAs. Additionally, the PVA may include additional error correction code (ECC) memory to enhance overall system security.

[0139] The accelerator 1214 (e.g., a hardware accelerator cluster) may include an on-chip computer vision network and SRAM to provide high-bandwidth, low-latency SRAM for the accelerator 1214. In some examples, the on-chip memory may include at least 4MB SRAM consisting of, for example and without limitation, eight field-configurable memory blocks, which may be accessed by both the PVA and the DLA. Each pair of memory blocks may include an advanced peripheral bus (APB) interface, configuration circuitry, a controller, and a multiplexer. Any type of memory may be used. The PVA and the DLA may access the memory via a backbone that provides high-speed memory access to the PVA and the DLA. The backbone may include (e.g., using APB) an on-chip computer vision network that interconnects the PVA and the DLA to the memory.

[0140] The on-chip computer vision network may include an interface that determines that both the PVA and the DLA provide ready and valid signals before transmitting any control signals / address / data. Such an interface may provide separate phases and separate channels for transmitting control signals / address / data, as well as burst communication for continuous data transfer. This type of interface may conform to the ISO 26262 or IEC 61508 standards, but other standards and protocols may also be used.

[0141] In some examples, SoC 1204 can include, for example, a real-time ray tracing hardware accelerator as described in U.S. Patent Application No. 16 / 101,232, filed on August 10, 2018. The real-time ray tracing hardware accelerator can be used to quickly and efficiently determine the position and extent of objects (e.g., within a world model) in order to generate a real-time visualization simulation for RADAR signal interpretation, for sound propagation synthesis and / or analysis, for SONAR system simulation, for general wave propagation simulation, for comparison with LIDAR data for positioning and / or other functional purposes, and / or for other uses. In some embodiments, one or more tree traversal units (TTUs) can be used to perform one or more ray tracing related operations.

[0142] Accelerator 1214 (e.g., a hardware accelerator cluster) has a wide range of autonomous driving applications. The PVA can be a programmable vision accelerator that can be used in key processing stages in ADAS and autonomous vehicles. The capabilities of the PVA are a good match for algorithm domains that require predictable processing, low power, and low latency. In other words, the PVA performs well on semi-dense or dense regular computations, even on small data sets that require predictable runtimes with low latency and low power. Thus, in the context of a platform for autonomous vehicles, the PVA is designed to run classical computer vision algorithms because they are effective in object detection and integer math operations.

[0143] For example, according to one embodiment of the technology, the PVA is used to perform computer stereo vision. In some examples, an algorithm based on semi-global matching can be used, but this is not intended to be limiting. Many applications for level 3 - 5 autonomous driving require instantaneous motion estimation / stereo matching (e.g., structure from motion, pedestrian recognition, lane detection, etc.). The PVA can perform computer stereo vision functions on inputs from two monocular cameras.

[0144] In some examples, the PVA can be used to perform dense optical flow. Process raw RADAR data (e.g., using a 4D fast Fourier transform) to provide processed RADAR. In other examples, the PVA is used for time-of-flight depth processing, which, for example, processes raw time-of-flight data to provide processed time-of-flight data.

[0145] DLA can be used to run any type of network to enhance control and driving safety, including, for example, a neural network that outputs a confidence metric for each object detection. Such confidence values can be interpreted as probabilities or as providing a relative "weight" of each detection compared to other detections. The confidence value enables the system to make further decisions about which detections should be considered true positive detections rather than false positive detections. For example, the system can set a threshold for the confidence and consider only detections that exceed the threshold as true positive detections. In an automatic emergency braking (AEB) system, false positive detections can cause the vehicle to automatically perform emergency braking, which is clearly undesirable. Therefore, only the most confident detections should be considered as triggers for AEB. DLA can run a neural network for regressing confidence values. The neural network can take as its input at least some subset of parameters, such as bounding box dimensions, a ground plane estimate obtained (e.g., from another subsystem), the output of an inertial measurement unit (IMU) sensor 1266 related to the orientation and distance of the vehicle 1200, a 3D position estimate of an object obtained from a neural network and / or other sensors (such as a LIDAR sensor 1264 or a RADAR sensor 1260), etc.

[0146] The SoC 1204 can include one or more data stores 1216 (e.g., memory). The data store 1216 can be on-chip memory of the SoC 1204, which can store neural networks to be executed on the GPU and / or DLA. In some examples, for redundancy and safety, the data store 1216 can be large enough in capacity to store multiple instances of the neural network. The data store 1212 can include an L2 or L3 cache 1212. References to the data store 1216 can include references to memory associated with PVAs, DLAs, and / or other accelerators 1214 as described herein.

[0147] The SoC 1204 may include one or more processors 1210 (e.g., embedded processors). The processor 1210 may include a boot and power management processor, which may be a dedicated processor and subsystem for handling boot power and management functions as well as security implementation related. The boot and power management processor may be part of the SoC 1204 boot sequence and may provide runtime power management services. The boot power and management processor may provide clock and voltage programming, assist system low-power state transitions, SoC 1204 heat and temperature sensor management, and / or SoC 1204 power state management. Each temperature sensor may be implemented as a ring oscillator whose output frequency is proportional to temperature, and the SoC 1204 may use the ring oscillator to detect the temperature of the CPU 1206, GPU 1208, and / or accelerator 1214. If it is determined that the temperature exceeds a threshold, then the boot and power management processor may enter a temperature fault routine and place the SoC 1204 in a lower power state and / or place the vehicle 1200 in a driver-safe parking mode (e.g., safely park the vehicle 1200).

[0148] The processor 1210 may also include a set of embedded processors that can be used as an audio processing engine. The audio processing engine may be an audio subsystem that allows for full hardware support for multi-channel audio over multiple interfaces and a wide and flexible range of audio I / O interfaces. In some examples, the audio processing engine is a dedicated processor core with a digital signal processor with dedicated RAM.

[0149] The processor 1210 may also include an always-on processor engine, which may provide the necessary hardware features to support low-power sensor management and wake-up use cases. The always-on processor engine may include a processor core, tightly coupled RAM, support peripherals (e.g., timers and interrupt controllers), various I / O controller peripherals, and routing logic.

[0150] The processor 1210 may also include a security cluster engine, which includes a dedicated processor subsystem for handling security management of automotive applications. The security cluster engine may include two or more processor cores, tightly coupled RAM, support peripherals (e.g., timers, interrupt controllers, etc.), and / or routing logic. In the security mode, the two or more cores may operate in a lockstep mode and act as a single core with comparison logic for detecting any differences between their operations.

[0151] The processor 1210 may also include a real-time camera engine, which may include a dedicated processor subsystem for handling real-time camera management.

[0152] The processor 1210 may also include a high dynamic range signal processor, which may include an image signal processor, which is a hardware engine that is part of the camera processing pipeline.

[0153] The processor 1210 may include a video image compositor that may be a processing block (e.g., implemented on a microprocessor) that implements the video post-processing functions required for a video playback application to generate the final image for the player window. The video image compositor may perform lens distortion correction on the wide-angle camera 1270, the surround camera 1274, and / or the in-cab monitoring camera sensor. The in-cab monitoring camera sensor is preferably monitored by a neural network running on another instance of the advanced SoC, configured to identify in-cab events and respond accordingly. The in-cab system may perform lip reading to activate mobile phone services and make calls, dictate emails, change the vehicle destination, activate or change the vehicle's infotainment system and settings, or provide voice-activated web surfing. Certain functions are only available to the driver when the vehicle is operating in autonomous mode and are disabled otherwise.

[0154] The video image compositor may include enhanced temporal noise reduction for spatial and temporal noise reduction. For example, in the case of motion in the video, the noise reduction appropriately weights the spatial information, reducing the weight of the information provided by neighboring frames. In the case where the image or a portion of the image does not include motion, the temporal noise reduction performed by the video image compositor may use information from a previous image to reduce the noise in the current image.

[0155] The video image compositor may also be configured to perform stereo correction on input stereo lens frames. When the operating system desktop is in use and the GPU 1208 does not need to continuously render new surfaces, the video image compositor may be further used for user interface composition. Even when the GPU 1208 is powered on and active, performing 3D rendering, the video image compositor may be used to relieve the burden on the GPU 1208 to improve performance and responsiveness.

[0156] The SoC 1204 may also include a Mobile Industry Processor Interface (MIPI) camera serial interface, a high-speed interface, and / or a video input block for receiving video and inputs from cameras and may be used for camera and related pixel input functions. The SoC 1204 may also include an input / output controller that may be software-controlled and may be used to receive I / O signals not committed to a specific role.

[0157] SoC 1204 may also include a wide range of peripheral device interfaces to enable communication with peripheral devices, audio codecs, power management, and / or other devices. SoC 1204 can be used to process data from cameras (connected via Gigabit Multimedia Serial Link and Ethernet), sensors (such as LIDAR sensor 1264, RADAR sensor 1260, etc. that can be connected via Ethernet), data from bus 1202 (such as the speed of vehicle 1200, steering wheel position, etc.), and data from GNSS sensor 1258 (connected via Ethernet or CAN bus). SoC 1204 may also include dedicated high-performance large-capacity storage controllers, which can include their own DMA engines and can be used to free the CPU 1206 from routine data management tasks.

[0158] SoC 1204 can be an end-to-end platform with a flexible architecture that spans automation levels 3 - 5, thus providing an integrated functional safety architecture for a platform that utilizes and efficiently uses computer vision and ADAS technologies to achieve diversity and redundancy, along with deep learning tools to provide a flexible and reliable driving software stack. SoC 1204 can be faster, more reliable, and even more energy-efficient and space-efficient than conventional systems. For example, when combined with CPU 1206, GPU 1208, and data storage 1216, accelerator 1214 can provide a fast and efficient platform for level 3 - 5 autonomous vehicles.

[0159] Thus, this technology provides capabilities and functions that cannot be achieved by conventional systems. For example, computer vision algorithms can be executed on CPUs, which can be configured using high-level programming languages such as the C programming language to perform various processing algorithms across a wide variety of visual data. However, CPUs often cannot meet the performance requirements of many computer vision applications, such as those related to execution time and power consumption. In particular, many CPUs cannot execute complex object detection algorithms in real time, which is a requirement for in-vehicle ADAS applications and for practical level 3 - 5 autonomous vehicles.

[0160] In contrast to conventional systems, the technology described herein allows multiple neural networks to be executed simultaneously and / or sequentially by providing a CPU complex, a GPU complex, and a cluster of hardware accelerators, and combining the results to achieve level 3 - 5 autonomous driving functions. For example, a CNN executed on a DLA or a dGPU (such as GPU 1220) can include text and word recognition, allowing a supercomputer to read and understand traffic signs, including signs for which the neural network has not been specifically trained. The DLA can also include a neural network capable of recognizing, interpreting, and providing semantic understanding of the signs and passing that semantic understanding to a path planning module running on the CPU complex.

[0161] As another example, multiple neural networks can run simultaneously, as required for level 3, 4, or 5 driving. For example, a warning sign consisting of "Caution: Flashing lights indicate icy conditions" together with the electric lights can be interpreted independently or jointly by several neural networks. The sign itself can be recognized as a traffic sign by a first neural network deployed (e.g., a trained neural network), and the text "Flashing lights indicate icy conditions" can be interpreted by a second neural network deployed, which informs the vehicle's path planning software (preferably executed on the CPU complex) that when the flashing lights are detected, there are icy conditions. The flashing lights can be recognized by operating a third neural network deployed over multiple frames, which informs the vehicle's path planning software of the presence (or absence) of the flashing lights. All three neural networks can run simultaneously, for example, within the DLA and / or on the GPU 1208.

[0162] In some examples, the CNNs for face recognition and vehicle owner recognition can use data from the camera sensors to identify the presence of an authorized driver and / or vehicle owner of the vehicle 1200. A processing engine always on the sensor can be used to unlock the vehicle and turn on the lights when the vehicle owner approaches the driver's door, and in a security mode, to disable the vehicle when the vehicle owner leaves the vehicle. In this way, the SoC 1204 provides security against theft and / or carjacking.

[0163] In another example, the CNN for emergency vehicle detection and recognition can use data from the microphone 1296 to detect and recognize an emergency vehicle siren. In contrast to conventional systems that use a general classifier to detect the siren and manually extract features, the SoC 1204 uses the CNN to classify environmental and urban sounds as well as visual data. In a preferred embodiment, the CNN running on the DLA is trained to recognize the relative closing rate of an emergency vehicle (e.g., by using the Doppler effect). The CNN can also be trained to recognize emergency vehicles specific to the local area in which the vehicle is operating, as recognized by the GNSS sensor 1258. Thus, for example, when operating in Europe, the CNN will seek to detect European sirens, and when in the United States, the CNN will seek to recognize only North American sirens. Once an emergency vehicle is detected, with the assistance of the ultrasonic sensor 1262, a control program can be used to execute an emergency vehicle safety routine to slow down the vehicle, drive it to the side of the road, stop the vehicle, and / or idle the vehicle until the emergency vehicle has passed.

[0164] The vehicle may include a CPU 1218 (e.g., a discrete CPU or dCPU) that can be coupled to the SoC 1204 via a high-speed interconnect (e.g., PCIe). The CPU 1218 may include, for example, an X86 processor. The CPU 1218 can be used to perform any of a variety of functions, including, for example, arbitrating potentially inconsistent results between the ADAS sensors and the SoC 1204, and / or monitoring the status and health of the controller 1236 and / or the infotainment SoC 1230.

[0165] The vehicle 1200 may include a GPU 1220 (e.g., a discrete GPU or dGPU) that can be coupled to the SoC 1204 via a high-speed interconnect (e.g., NVIDIA's NVLINK). The GPU 1220 can provide additional artificial intelligence capabilities, for example, by executing redundant and / or different neural networks, and can be used to train and / or update neural networks at least in part based on inputs (e.g., sensor data) from the sensors of the vehicle 1200.

[0166] The vehicle 1200 may further include a network interface 1224, which may include one or more wireless antennas 1226 (e.g., one or more wireless antennas for different communication protocols, such as cellular antennas, Bluetooth antennas, etc.). The network interface 1224 can be used to enable wireless connections to the cloud (e.g., to the server 1278 and / or other network devices), to other vehicles, and / or to computing devices (e.g., the passenger's client device) via the Internet. To communicate with other vehicles, a direct link can be established between the two vehicles, and / or an indirect link (e.g., across a network and via the Internet) can be established. The direct link can be provided using a vehicle-to-vehicle communication link. The vehicle-to-vehicle communication link can provide the vehicle 1200 with information about vehicles approaching the vehicle 1200 (e.g., vehicles in front of, beside, and / or behind the vehicle 1200). This function can be part of the cooperative adaptive cruise control function of the vehicle 1200.

[0167] The network interface 1224 may include an SoC that provides modulation and demodulation functions and enables the controller 1236 to communicate over a wireless network. The network interface 1224 may include a radio frequency front end for upconverting from baseband to radio frequency and downconverting from radio frequency to baseband. The frequency conversion can be performed by a known process and / or can be performed using a super-heterodyne process. In some examples, the radio frequency front end functions can be provided by a separate chip. The network interface may include wireless capabilities for communicating via LTE, WCDMA, UMTS, GSM, CDMA2000, Bluetooth, Bluetooth LE, Wi-Fi, Z-Wave, ZigBee, LoRaWAN, and / or other wireless protocols.

[0168] Vehicle 1200 may also include a data store 1228 that may include off-chip (e.g., outside of SoC 1204) storage devices. The data store 1228 may include one or more storage elements, including RAM, SRAM, DRAM, VRAM, flash memory, hard drives, and / or other components and / or devices that can store at least one bit of data.

[0169] Vehicle 1200 may also include a GNSS sensor 1258. The GNSS sensor 1258 (e.g., GPS, assisted GPS sensor, differential GPS (DGPS) sensor, etc.) is used to assist mapping, perception, occupancy grid generation, and / or path planning functions. Any number of GNSS sensors 1258 may be used, including, for example and without limitation, a GPS using a USB connector with an Ethernet to serial (RS-232) bridge.

[0170] Vehicle 1200 may also include a RADAR sensor 1260. The RADAR sensor 1260 may be used by vehicle 1200 for remote vehicle detection even in dark and / or adverse weather conditions. The RADAR functional safety level may be ASIL B. The RADAR sensor 1260 may use CAN and / or bus 1202 (e.g., to transmit data generated by the RADAR sensor 1260) for control as well as access to object tracking data and, in some examples, accesses Ethernet to access raw data. A variety of RADAR sensor types may be used. For example and without limitation, the RADAR sensor 1260 may be suitable for front, rear, and side RADAR use. In some examples, a pulsed Doppler RADAR sensor is used.

[0171] The RADAR sensor 1260 may include different configurations, such as long-range with a narrow field of view, short-range with a wide field of view, short-range side coverage, and so on. In some examples, long-range RADAR may be used for adaptive cruise control functions. The long-range RADAR system may provide a wide field of view (e.g., within 250 m) achieved through two or more independent scans. The RADAR sensor 1260 may help distinguish between static and moving objects and may be used by the ADAS system for emergency braking assistance and forward collision warning. The long-range RADAR sensor may include a single station multimode RADAR with multiple (e.g., six or more) fixed RADAR antennas and high-speed CAN and FlexRay interfaces. In an example with six antennas, the central four antennas may create a focused beam pattern that is designed to record the surroundings of vehicle 1200 at a higher rate with minimal traffic interference from adjacent lanes. The other two antennas may extend the field of view, making it possible to quickly detect vehicles entering or leaving the lane of vehicle 1200.

[0172] As an example, a mid-range RADAR system can include a range of up to 1260 m (front) or 80 m (rear) and a field of view of up to 42 degrees (front) or 1250 degrees (rear). A short-range RADAR system can include, but is not limited to, RADAR sensors designed to be mounted at both ends of the rear bumper. When mounted at both ends of the rear bumper, such a RADAR sensor system can create two beams that continuously monitor the blind spots behind and beside the vehicle.

[0173] The short-range RADAR system can be used in an ADAS system for blind spot detection and / or lane change assistance.

[0174] Vehicle 1200 can also include ultrasonic sensors 1262. Ultrasonic sensors 1262 that can be placed in the front, rear, and / or sides of vehicle 1200 can be used for parking assistance and / or creating and updating occupancy grids. A variety of ultrasonic sensors 1262 can be used, and different ultrasonic sensors 1262 can be used for different detection ranges (e.g., 2.5 m, 4 m). The ultrasonic sensors 1262 can operate at ASIL B of the functional safety level.

[0175] Vehicle 1200 can include a LIDAR sensor 1264. The LIDAR sensor 1264 can be used for object and pedestrian detection, emergency braking, collision avoidance, and / or other functions. The LIDAR sensor 1264 can be at ASIL B of the functional safety level. In some examples, vehicle 1200 can include multiple LIDAR sensors 1264 (e.g., two, four, six, etc.) that can use Ethernet (e.g., to provide data to a gigabit Ethernet switch).

[0176] In some examples, the LIDAR sensor 1264 may be able to provide a list of objects and their distances for a 360-degree field of view. Commercially available LIDAR sensors 1264 can have, for example, an advertised range of approximately 1200 m, an accuracy of 2 cm - 3 cm, and support a 1200 Mbps Ethernet connection. In some examples, one or more non-protruding LIDAR sensors 1264 can be used. In such examples, the LIDAR sensor 1264 can be implemented as a small device that can be embedded in the front, rear, sides, and / or corners of vehicle 1200. In such examples, the LIDAR sensor 1264 can provide a field of view of up to 120 degrees horizontally and 35 degrees vertically even for low-reflectivity objects, with a range of 200 m. The front-mounted LIDAR sensor 1264 can be configured for a horizontal field of view between 45 degrees and 135 degrees.

[0177] In some examples, LIDAR technologies such as 3D flash LIDAR can also be used. 3D flash LIDAR uses the flash of a laser as the emission source to illuminate the vehicle's surroundings up to about 200m. The flash LIDAR unit includes a receiver that records the laser pulse transmission time and the reflected light on each pixel, which in turn corresponds to the range from the vehicle to the object. Flash LIDAR can allow for the generation of highly accurate and distortion-free images of the surroundings using each laser flash. In some examples, four flash LIDAR sensors can be deployed, one on each side of the vehicle 1200. Available 3D flash LIDAR systems include solid-state 3D staring array LIDAR cameras (e.g., non-scanning LIDAR devices) that have no moving parts other than a fan. The flash LIDAR device can use class I (eye-safe) laser pulses of 5 nanoseconds per frame and can capture the reflected laser in the form of 3D range point clouds and co-registered intensity data. By using flash LIDAR and because flash LIDAR is a solid-state device with no moving parts, the LIDAR sensor 1264 can be less susceptible to motion blur, vibration, and / or shock.

[0178] The vehicle can also include an IMU sensor 1266. In some examples, the IMU sensor 1266 can be located at the center of the rear axle of the vehicle 1200. The IMU sensor 1266 can include, for example and without limitation, accelerometers, magnetometers, gyroscopes, magnetic compasses, and / or other sensor types. In some examples, such as in a six-axis application, the IMU sensor 1266 can include an accelerometer and a gyroscope, while in a nine-axis application, the IMU sensor 1266 can include an accelerometer, a gyroscope, and a magnetometer.

[0179] In some embodiments, the IMU sensor 1266 can be implemented as a miniature high-performance GPS-aided inertial navigation system (GPS / INS) that combines microelectromechanical systems (MEMS) inertial sensors, a high-sensitivity GPS receiver, and an advanced Kalman filtering algorithm to provide estimates of position, velocity, and attitude. Thus, in some examples, the IMU sensor 1266 can enable the vehicle 1200 to estimate the heading by directly observing the change in velocity from the GPS to the IMU sensor 1266 and correlating it without the need for input from a magnetic sensor. In some examples, the IMU sensor 1266 and the GNSS sensor 1258 can be integrated into a single unit.

[0180] The vehicle can include a microphone 1296 placed in and / or around the vehicle 1200. Among other things, the microphone 1296 can be used for emergency vehicle detection and identification.

[0181] The vehicle may also include any number of camera types, including a stereo camera 1268, a wide-angle camera 1270, an infrared camera 1272, a surround camera 1274, a long-range and / or mid-range camera 1298, and / or other camera types. These cameras can be used to capture image data around the entire periphery of the vehicle 1200. The camera types used depend on the embodiment and the requirements of the vehicle 1200, and any combination of camera types can be used to provide the necessary coverage around the vehicle 1200. Additionally, the number of cameras can vary according to the embodiment. For example, the vehicle may include six cameras, seven cameras, ten cameras, twelve cameras, and / or another number of cameras. As an example and without limitation, these cameras can support Gigabit Multimedia Serial Link (GMSL) and / or Gigabit Ethernet. Each of the cameras is described in more detail herein with respect to Figure 12A and Figure 12B is described in more detail.

[0182] The vehicle 1200 may also include a vibration sensor 1242. The vibration sensor 1242 can measure the vibration of components of the vehicle such as an axle. For example, a change in vibration can indicate a change in the road surface. In another example, when two or more vibration sensors 1242 are used, the difference between the vibrations can be used to determine the friction or slip of the road surface (e.g., when there is a vibration difference between a powered drive axle and a free-spinning axle).

[0183] The vehicle 1200 may include an ADAS system 1238. In some examples, the ADAS system 1238 may include a SoC. The ADAS system 1238 may include autonomous / adaptive / auto cruise control (ACC), cooperative adaptive cruise control (CACC), forward collision warning (FCW), automatic emergency braking (AEB), lane departure warning (LDW), lane keeping assist (LKA), blind spot warning (BSW), rear cross traffic warning (RCTW), collision warning system (CWS), lane centering (LC), and / or other features and functions.

[0184] The ACC system can use RADAR sensors 1260, LIDAR sensors 1264, and / or cameras. The ACC system can include longitudinal ACC and / or lateral ACC. Longitudinal ACC monitors and controls the distance to the vehicle immediately in front of the vehicle 1200 and automatically adjusts the vehicle speed to maintain a safe distance from the vehicle in front. Lateral ACC performs distance keeping and, when necessary, advises the vehicle 1200 to change lanes. Lateral ACC is related to other ADAS applications such as LCA and CWS.

[0185] The CACC uses information from other vehicles, which can be received indirectly from other vehicles via the network interface 1224 and / or the wireless antenna 1226 via a wireless link or through a network connection (e.g., via the Internet). The direct link can be provided by a vehicle-to-vehicle (V2V) communication link, while the indirect link can be an infrastructure-to-vehicle (I2V) communication link. Generally, the V2V communication concept provides information about the immediately preceding vehicle (e.g., the vehicle immediately in front of vehicle 1200 and in the same lane as it), while the I2V communication concept provides information about traffic further ahead. The CACC system can include either or both of the I2V and V2V information sources. Given the information of the vehicle in front of vehicle 1200, the CACC can be more reliable, and it has the potential to improve the smoothness of traffic flow and reduce road congestion.

[0186] The FCW system is designed to alert the driver to a danger so that the driver can take corrective measures. The FCW system uses a front camera and / or a RADAR sensor 1260 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback such as a display, a speaker, and / or a vibrating component. The FCW system can provide warnings in the form of, for example, sounds, visual warnings, vibrations, and / or rapid braking pulses.

[0187] The AEB system detects an impending front collision with another vehicle or other object and can automatically apply the brakes if the driver does not take corrective measures within a specified time or distance parameter. The AEB system can use a front camera and / or a RADAR sensor 1260 coupled to a dedicated processor, DSP, FPGA, and / or ASIC. When the AEB system detects a danger, it typically first alerts the driver to take corrective measures to avoid the collision, and if the driver does not take corrective measures, then the AEB system can automatically apply the brakes in an effort to prevent or at least mitigate the impact of the predicted collision. The AEB system can include technologies such as dynamic brake support and / or collision imminent braking.

[0188] The LDW system provides visual, auditory, and / or tactile warnings such as steering wheel or seat vibrations to alert the driver when vehicle 1200 crosses a lane marking. The LDW system is not activated when the driver indicates an intentional lane departure by activating the turn signal. The LDW system can use a front-side-facing camera coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback such as a display, a speaker, and / or a vibrating component.

[0189] The LKA system is a variant of the LDW system. If vehicle 1200 starts to leave a lane, then the LKA system provides a steering input or braking to correct the vehicle 1200.

[0190] The BSW system detects and warns the driver of vehicles in the vehicle's blind spot. The BSW system can provide visual, audible, and / or tactile alerts to indicate that merging or changing lanes is unsafe. The system can provide additional warnings when the driver uses a turn signal. The BSW system can use a rear-facing camera and / or RADAR sensor 1260 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback such as a display, speaker, and / or vibration component.

[0191] The RCTW system can provide visual, audible, and / or tactile notifications when an object is detected outside the rear camera range while the vehicle 1200 is in reverse. Some RCTW systems include AEB to ensure that vehicle brakes are applied to avoid a crash. The RCTW system can use one or more rear RADAR sensors 1260 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback such as a display, speaker, and / or vibration component.

[0192] Conventional ADAS systems may be prone to false positive results, which can be annoying and distracting to the driver, but are typically not catastrophic because the ADAS system alerts the driver and allows the driver to decide whether a safe condition truly exists and act accordingly. However, in an autonomous vehicle 1200, in the case of conflicting results, the vehicle 1200 itself must decide whether to heed the results from the main computer or an auxiliary computer (e.g., the first controller 1236 or the second controller 1236). For example, in some embodiments, the ADAS system 1238 can be a backup and / or auxiliary computer for providing perception information to a redundant computer sanity module. The redundant computer sanity monitor can run redundant and diverse software on hardware components to detect faults in perception and dynamic driving tasks. The output from the ADAS system 1238 can be provided to the supervisory MCU. If the outputs from the main computer and the auxiliary computer conflict, then the supervisory MCU must determine how to reconcile the conflict to ensure safe operation.

[0193] In some examples, the host computer may be configured to provide a confidence score to the supervisory MCU indicating the host computer's confidence in the selected result. If the confidence score exceeds a threshold, then the supervisory MCU may follow the host computer's direction regardless of whether the secondary computer provides conflicting or inconsistent results. In cases where the confidence score does not meet the threshold and where the host computer and the secondary computer indicate different results (e.g., conflict), the supervisory MCU may arbitrate between these computers to determine an appropriate result.

[0194] The supervisory MCU may be configured to run a neural network that is trained and configured to determine conditions under which the secondary computer provides a false alarm, at least in part based on outputs from the host computer and the secondary computer. Thus, the neural network in the supervisory MCU can learn when the output of the secondary computer can be trusted and when it cannot. For example, when the secondary computer is a RADAR-based FCW system, the neural network in the supervisory MCU can learn when the FCW system is identifying a metallic object that is not in fact a danger, such as a drainage grate or manhole cover that triggers an alarm. Similarly, when the secondary computer is a camera-based LDW system, the neural network in the supervisory MCU can learn to disregard the LDW when a cyclist or pedestrian is present and lane departure is actually the safest strategy. In embodiments that include a neural network running on the supervisory MCU, the supervisory MCU may include at least one of a DLA or a GPU suitable for running the neural network with associated memory. In a preferred embodiment, the supervisory MCU may include components of the SoC 1204 and / or be included as a component of the SoC 1204.

[0195] In other examples, the ADAS system 1238 may include a secondary computer that performs ADAS functions using traditional computer vision rules. In this way, the secondary computer may use classical computer vision rules (if-then), and the presence of a neural network in the supervisory MCU can improve reliability, safety, and performance. For example, diverse implementations and intentional non-identity make the overall system more fault-tolerant, especially for failures caused by software (or software-hardware interface) functions. For example, if there is a software vulnerability or error in the software running on the host computer and the non-identical software code running on the secondary computer provides the same overall result, then the supervisory MCU can be more confident that the overall result is correct and that the vulnerability in the software or hardware on the host computer does not cause a substantial error.

[0196] In some examples, the output of the ADAS system 1238 can be fed to the perception block of the main computer and / or the dynamic driving task block of the main computer. For example, if the ADAS system 1238 indicates a forward collision warning due to an object being immediately in front, then the perception block can use this information when identifying the object. In other examples, the auxiliary computer can have its own neural network, which is trained and thus reduces the risk of false positives as described herein.

[0197] The vehicle 1200 can also include an infotainment SoC 1230 (e.g., an in-vehicle infotainment system (IVI)). Although illustrated and described as an SoC, the infotainment system can not be an SoC and can include two or more discrete components. The infotainment SoC 1230 can include a combination of hardware and software that can be used to provide audio (e.g., music, personal digital assistant, navigation instructions, news, radio, etc.), video (e.g., TV, movies, streaming, etc.), phone (e.g., hands-free calling), network connectivity (e.g., LTE, WiFi, etc.), and / or information services (e.g., navigation system, rear parking assistance, radio data system, vehicle-related information such as fuel level, total distance covered, brake fuel level, oil level, door open / close, air filter information, etc.) to the vehicle 1200. For example, the infotainment SoC 1230 can include a radio, disc player, navigation system, video player, USB and Bluetooth connectivity, in-vehicle computer, in-vehicle entertainment, WiFi, steering wheel audio controls, hands-free voice controls, a head-up display (HUD), an HMI display 1234, a telematics device, a control panel (e.g., for controlling various components, features, and / or systems, and / or interacting therewith), and / or other components. The infotainment SoC 1230 can further be used to provide information (e.g., visual and / or auditory) to the user of the vehicle, such as information from the ADAS system 1238, autonomous driving information such as planned vehicle maneuvers, trajectories, surrounding environment information (e.g., intersection information, vehicle information, road information, etc.), and / or other information.

[0198] The infotainment SoC 1230 can include GPU functionality. The infotainment SoC 1230 can communicate with other devices, systems, and / or components of the vehicle 1200 via a bus 1202 (e.g., a CAN bus, Ethernet, etc.). In some examples, the infotainment SoC 1230 can be coupled to a supervisory MCU such that in the event of a failure of the main controller 1236 (e.g., the main and / or standby computer of the vehicle 1200), the GPU of the infotainment system can perform some autonomous driving functions. In such examples, the infotainment SoC 1230 can place the vehicle 1200 in a driver safe parking mode as described herein.

[0199] Vehicle 1200 may also include an instrument cluster 1232 (such as a digital instrument panel, an electronic instrument cluster, a digital instrument surface panel, etc.). The instrument cluster 1232 may include a controller and / or a supercomputer (such as a discrete controller or supercomputer). The instrument cluster 1232 may include a set of instruments, such as a speedometer, fuel level, oil pressure, tachometer, odometer, turn indicator, shift position indicator, seat belt warning light, parking brake warning light, engine malfunction light, airbag (SRS) system information, lighting controls, safety system controls, navigation information, and so on. In some examples, information may be displayed and / or shared between the infotainment SoC 1230 and the instrument cluster 1232. In other words, the instrument cluster 1232 may be included as part of the infotainment SoC 1230, or vice versa.

[0200] Figure 12D A system schematic diagram for communication between a cloud-based server and Figure 12A Example autonomous vehicle 1200 according to some embodiments of the present disclosure. The system 1276 may include a server 1278, a network 1290, and vehicles including the vehicle 1200. The server 1278 may include multiple GPUs 1284(A)-1284(H) (collectively referred to herein as GPUs 1284), PCIe switches 1282(A)-1282(H) (collectively referred to herein as PCIe switches 1282), and / or CPUs 1280(A)-1280(B) (collectively referred to herein as CPUs 1280). The GPUs 1284, CPUs 1280, and PCIe switches may be interconnected by high-speed interconnections such as, for example and without limitation, the NVLink interface 1288 developed by NVIDIA and / or PCIe connections 1286. In some examples, the GPUs 1284 are connected via NVLink and / or an NVSwitch SoC, and the GPUs 1284 and the PCIe switches 1282 are connected via a PCIe interconnect. Although eight GPUs 1284, two CPUs 1280, and two PCIe switches are illustrated, this is not intended to be limiting. Depending on the embodiment, each of the servers 1278 may include any number of GPUs 1284, CPUs 1280, and / or PCIe switches. For example, each of the servers 1278 may include eight, sixteen, thirty-two, and / or more GPUs 1284.

[0201] Server 1278 can receive image data via network 1290 from a vehicle, the image data representing an image showing an unexpected or changed road condition such as a recently started road works. Server 1278 can transmit neural network 1292, updated neural network 1292, and / or map information 1294, including information about traffic and road conditions, via network 1290 to the vehicle. Updates to the map information 1294 can include updates to the HD map 1222, such as information about construction sites, potholes, curves, floods, or other obstacles. In some examples, the neural network 1292, updated neural network 1292, and / or map information 1294 can be represented and / or generated based on data received from new training and / or data from any number of vehicles in the environment and / or experience of training performed at a data center (e.g., using server 1278 and / or other servers).

[0202] Server 1278 can be used to train a machine learning model (e.g., a neural network) based on training data. The training data can be generated by vehicles, and / or can be generated in a simulation (e.g., using a game engine). In some examples, the training data is labeled (e.g., in cases where the neural network benefits from supervised learning) and / or undergoes other preprocessing, while in other examples, the training data is not labeled and / or preprocessed (e.g., in cases where the neural network does not require supervised learning). Training can be performed according to any one or more categories of machine learning techniques, including but not limited to categories such as: supervised training, semi-supervised training, unsupervised training, self-learning, reinforcement learning, federated learning, transfer learning, feature learning (including principal component and clustering analysis), multilinear subspace learning, manifold learning, representation learning (including alternative dictionary learning), rule-based machine learning, anomaly detection, and any variations or combinations thereof. Once the machine learning model is trained, the machine learning model can be used by the vehicle (e.g., transmitted to the vehicle via network 1290), and / or the machine learning model can be used by server 1278 to remotely monitor the vehicle.

[0203] In some examples, server 1278 can receive data from a vehicle and apply the data to a latest real-time neural network for real-time intelligent inference. Server 1278 can include a deep learning supercomputer powered by GPU 1284 and / or a dedicated AI computer, such as DGX and DGX Station machines developed by NVIDIA. However, in some examples, server 1278 can include a deep learning infrastructure of a data center powered only by a CPU.

[0204] The deep learning infrastructure of server 1278 may be capable of fast real-time inference and can use this ability to evaluate and verify the health of the processors, software, and / or associated hardware in vehicle 1200. For example, the deep learning infrastructure may receive periodic updates from vehicle 1200, such as an image sequence and / or objects located in the image sequence that vehicle 1200 has identified (e.g., via computer vision and / or other machine learning object classification techniques). The deep learning infrastructure may run its own neural network to identify the objects and compare them with the objects identified by vehicle 1200. If the results do not match and the infrastructure concludes that the AI in vehicle 1200 has failed, then server 1278 may transmit a signal to vehicle 1200 instructing the fail-safe computer in vehicle 1200 to take control, notify the passengers, and complete a safe parking operation.

[0205] For inference, server 1278 may include GPU 1284 and one or more programmable inference accelerators (e.g., NVIDIA's TensorRT). The combination of GPU-powered servers and inference acceleration can enable real-time response. In other examples, such as when performance is less critical, CPU, FPGA, and other processor-powered servers may be used for inference.

[0206] Example computing device

[0207] Figure 13 is a block diagram of an example computing device 1300 suitable for implementing some embodiments of the present disclosure. Computing device 1300 may include an interconnect system 1302 that directly or indirectly couples the following devices: a memory 1304, one or more central processing units (CPUs) 1306, one or more graphics processing units (GPUs) 1308, a communication interface 1310, input / output (I / O) ports 1312, input / output components 1314, a power supply 1316, one or more rendering components 1318 (e.g., (one or more) displays), and one or more logic units 1320. In at least one embodiment, (one or more) computing devices 1300 may include one or more virtual machines (VMs), and / or any of its components may include virtual components (e.g., virtual hardware components). For a non-limiting example, one or more of GPUs 1308 may include one or more vGPUs, one or more of CPUs 1306 may include one or more vCPUs, and / or one or more of logic units 1320 may include one or more virtual logic units. Thus, (one or more) computing devices 1300 may include discrete components (e.g., full GPUs dedicated to computing device 1300), virtual components (e.g., a portion of a GPU dedicated to computing device 1300), or a combination thereof.

[0208] Although Figure 13 each of the boxes of is shown as being connected via circuitry through the interconnect system 1302, this is not intended to be limiting and is for clarity only. For example, in some embodiments, a presenting component 1318 (such as a display device) may be considered an I / O component 1314 (e.g., if the display is a touch screen). As another example, the CPU 1306 and / or GPU 1308 may include memory (e.g., the memory 1304 may represent a storage device in addition to the memory of the GPU 1308, the CPU 1306, and / or other components). In other words, Figure 13 the computing devices of are illustrative only. No distinction is made between such categories as "workstation", "server", "laptop computer", "desktop computer", "tablet computer", "client device", "mobile device", "handheld device", "gaming console", "electronic control unit (ECU)", "virtual reality system", and / or other device or system types, because all are considered within the scope of Figure 13 the computing devices of.

[0209] The interconnect system 1302 may represent one or more links or buses, such as an address bus, a data bus, a control bus, or a combination thereof. The interconnect system 1302 may include one or more bus or link types, such as an Industry Standard Architecture (ISA) bus, an Extended Industry Standard Architecture (EISA) bus, a Video Electronics Standards Association (VESA) bus, a Peripheral Component Interconnect (PCI) bus, a Peripheral Component Interconnect Express (PCIe) bus, and / or another type of bus or link. In some embodiments, there are direct connections between components. As an example, the CPU 1306 may be directly connected to the memory 1304. Further, the CPU 1306 may be directly connected to the GPU 1308. In cases where there are direct or point-to-point connections between components, the interconnect system 1302 may include a PCIe link to effect the connection. In these examples, a PCI bus need not be included in the computing device 1300.

[0210] The memory 1304 may include any of a variety of computer-readable media. The computer-readable media may be any available media that can be accessed by the computing device 1300. The computer-readable media may include volatile and non-volatile media, as well as removable and non-removable media. By way of example and not limitation, the computer-readable media may include computer storage media and communication media.

[0211] Computer storage media can include volatile and non-volatile media and / or removable and non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules, and / or other data types. For example, memory 1304 can store computer-readable instructions (e.g., representing one or more programs and / or one or more program elements such as an operating system). Computer storage media can include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic tape cartridges, magnetic tape, magnetic disk storage devices or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by computing device 1300. As used herein, computer storage media does not include signals per se.

[0212] Computer storage media can embody computer-readable instructions, data structures, program modules, and / or other data types in a modulated data signal such as a carrier wave or other transmission mechanism and include any information delivery medium. The term "modulated data signal" can refer to a signal that sets or changes one or more of its characteristics in a manner that encodes information in the signal. By way of example and not limitation, computer storage media can include wired media (such as a wired network or direct wired connection) and wireless media (such as acoustic, RF, infrared, and other wireless media). Combinations of any of the above should also be included within the scope of computer-readable media.

[0213] CPU 1306 can be configured to execute at least some of the computer-readable instructions to control one or more components of computing device 1300 to perform one or more of the methods and / or processes described herein. Each of CPU 1306 can include one or more cores (e.g., one, two, four, eight, twenty-eight, seventy-two, etc.) capable of concurrently handling numerous software threads. CPU 1306 can include any type of processor and can include different types of processors depending on the type of computing device 1300 implemented (e.g., a processor with fewer cores for a mobile device and a processor with more cores for a server). For example, depending on the type of computing device 1300, the processor can be an advanced RISC machine (ARM) processor implemented using reduced instruction set computing (RISC) or an x86 processor implemented using complex instruction set computing (CISC). In addition to one or more microprocessors or supplementary co-processors (such as a math co-processor), computing device 1300 can also include one or more CPU 1306.

[0214] In addition to or instead of one or more CPUs 1306, one or more GPUs 1308 may be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 1300 to perform one or more of the methods and / or processes described herein. One or more of the GPUs 1308 may be an integrated GPU (e.g., with one or more of the CPUs 1306) and / or one or more of the GPUs 1308 may be a discrete GPU. In an embodiment, one or more of the GPUs 1308 may be a coprocessor of one or more of the CPUs 1306. The GPUs 1308 may be used by the computing device 1300 to render graphics (e.g., 3D graphics) or perform general-purpose computing. For example, the GPUs 1308 may be used for general-purpose computing on GPUs (GPGPU). The GPUs 1308 may include hundreds or thousands of cores capable of concurrently handling hundreds or thousands of software threads. The GPUs 1308 may generate pixel data of an output image in response to a rendering command (e.g., a rendering command received from the CPU 1306 via a host interface). The GPUs 1308 may include a graphics memory (e.g., display memory) for storing pixel data or any other suitable data (e.g., GPGPU data). The display memory may be included as part of the memory 1304. The GPUs 1308 may include two or more GPUs operating in parallel (e.g., via a link). The link may directly connect the GPUs (e.g., using NVLINK) or may connect the GPUs through a switch (e.g., using NVSwitch). When combined, each GPU 1308 may generate pixel data or GPGPU data for different parts of the output or for different outputs (e.g., a first GPU for a first image and a second GPU for a second image). Each GPU may include its own memory or may share memory with other GPUs.

[0215] In addition to and / or in place of the CPU 1306 and / or GPU 1308, the logic unit 1320 may be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 1300 to perform one or more of the methods and / or processes described herein. In an embodiment, the (one or more) CPUs 1306, the (one or more) GPUs 1308, and / or the (one or more) logic units 1320 may perform any combination of methods, processes, and / or portions thereof discretely or jointly. One or more of the logic units 1320 may be part of one or more of the CPUs 1306 and / or GPUs 1308 and / or integrated in one or more of the CPUs 1306 and / or GPUs 1308 and / or one or more of the logic units 1320 may be discrete components or otherwise external to the CPUs 1306 and / or GPUs 1308. In an embodiment, one or more of the logic units 1320 may be a coprocessor of one or more of the CPUs 1306 and / or one or more of the GPUs 1308.

[0216] Examples of the logic unit 1320 include one or more processing cores and / or their components, such as a data processing unit (DPU), a tensor core (TC), a tensor processing unit (TPU), a pixel vision core (PVC), a vision processing unit (VPU), a graphics processing cluster (GPC), a texture processing cluster (TPC), a streaming multiprocessor (SM), a tree traversal unit (TTU), an artificial intelligence accelerator (AIA), a deep learning accelerator (DLA), an arithmetic logic unit (ALU), an application specific integrated circuit (ASIC), a floating point unit (FPU), an input / output (I / O) element, a peripheral component interconnect (PCI) or a peripheral component interconnect express (PCIe) element, etc.

[0217] The communication interface 1310 may include one or more receivers, transmitters, and / or transceivers that enable the computing device 1300 to communicate with other computing devices via an electronic communication network, including wired and / or wireless communication. The communication interface 1310 may include components and functions that implement communication over any of a plurality of different networks, such as wireless networks (e.g., Wi-Fi, Z-Wave, Bluetooth, Bluetooth LE, ZigBee, etc.), wired networks (e.g., over Ethernet or InfiniBand communication), low-power wide-area networks (e.g., LoRaWAN, SigFox, etc.), and / or the Internet. In one or more embodiments, the logic unit 1320 and / or the communication interface 1310 may include one or more data processing units (DPUs) for directly transferring data received over the network and / or via the interconnect system 1302 to one or more GPUs 1308 (e.g., to the memory thereof).

[0218] The I / O port 1312 may enable the computing device 1300 to be logically coupled to other devices including I / O components 1314, one or more presentation components 1318, and / or other components, some of which may be built into (e.g., integrated in) the computing device 1300. Illustrative I / O components 1314 include a microphone, a mouse, a keyboard, a joystick, a gamepad, a game controller, a satellite dish, a scanner, a printer, a wireless device, etc. The I / O components 1314 may provide a natural user interface (NUI) that processes air gestures, voice, or other physiological inputs generated by a user. In some cases, the input may be transmitted to an appropriate network element for further processing. The NUI may implement any combination of speech recognition, stylus recognition, face recognition, biometric recognition, on-screen and near-screen gesture recognition, air gestures, head and eye tracking, and touch recognition (as described in more detail below) associated with the display of the computing device 1300. The computing device 1300 may include a depth camera for gesture detection and recognition, such as a stereoscopic camera system, an infrared camera system, an RGB camera system, touchscreen technology, and combinations thereof. Additionally, the computing device 1300 may include an accelerometer or a gyroscope (e.g., as part of an inertial measurement unit (IMU)) that enables the detection of motion. In some examples, the computing device 1300 may use the output of the accelerometer or the gyroscope to render immersive augmented reality or virtual reality.

[0219] The power supply 1316 may include a hard-wired power supply, a battery power supply, or a combination thereof. The power supply 1316 may provide power to the computing device 1300 to enable the components of the computing device 1300 to operate.

[0220] The presentation component 1318 may include a display (e.g., a monitor, a touch screen, a television screen, a head-up display (HUD), other display types, or a combination thereof), a speaker, and / or other presentation components. The presentation component 1318 may receive data from other components (e.g., the GPU 1308, the CPU 1306, the DPU, etc.) and output the data (e.g., as an image, a video, a sound, etc.).

[0221] Example data center

[0222] Figure 14 An example data center 1400 that may be used in at least one embodiment of the present disclosure is shown. The data center 1400 may include a data center infrastructure layer 1410, a framework layer 1420, a software layer 1430, and / or an application layer 1440.

[0223] As Figure 14 shown, the data center infrastructure layer 1410 may include a resource coordinator 1412, grouped computing resources 1414, and node computing resources ("node C.R.s") 1416(1)-1416(N), where "N" represents any whole positive integer. In at least one embodiment, the node C.R.s 1416(1)-1416(N) may include, but are not limited to, any number of central processing units (CPUs) or other processors (including DPUs, accelerators, field programmable gate arrays (FPGAs), graphics processors or graphics processing units (GPUs), etc.), memory devices (e.g., dynamic read-only memories), storage devices (e.g., solid-state or disk drives), network input / output (NW I / O) devices, network switches, virtual machines (VMs), power modules, and / or cooling modules, and so on. In some embodiments, one or more of the node C.R.s 1416(1)-1416(N) may correspond to a server having one or more of the above computing resources. Additionally, in some embodiments, the node C.R.s 1416(1)-14161(N) may include one or more virtual components, such as vGPUs, vCPUs, etc., and / or one or more of the node C.R.s 1416(1)-1416(N) may correspond to a virtual machine (VM).

[0224] In at least one embodiment, the grouped computing resources 1414 may include separate groupings of node C.R.s 1416 housed within one or more racks (not shown), or many racks within data centers located at different geographical locations (also not shown). Separate groupings of node C.R.s 1416 within the grouped computing resources 1414 may include grouped computing, network, memory, or storage resources that may be configured or allocated to support one or more workloads. In at least one embodiment, several node C.R.s 1416, including CPUs, GPUs, DPUs, and / or other processors, may be grouped within one or more racks to provide computing resources to support one or more workloads. One or more racks may also include any combination of any number of power modules, cooling modules, and / or network switches.

[0225] The resource coordinator 1412 may configure or otherwise control one or more node C.R.s 1416(1)-1416(N) and / or the grouped computing resources 1414. In at least one embodiment, the resource coordinator 1412 may include a software design infrastructure (SDI) management entity for the data center 1400. The resource coordinator 1412 may include hardware, software, or some combination thereof.

[0226] In at least one embodiment, as Figure 14 shown, the framework layer 1420 may include a job scheduler 1433, a configuration manager 1434, a resource manager 1436, and / or a distributed file system 1438. The framework layer 1420 may include a framework for the software 1432 of the support software layer 1430 and / or one or more applications 1442 of the application layer 1440. The software 1432 or the application 1442 may respectively contain network-based service software or applications, such as those provided by Amazon Web Services, Google Cloud, and Microsoft Azure. The framework layer 1420 may be, but is not limited to, a free and open-source software web application framework (such as Apache Spark) that may utilize the distributed file system 1438 for large-scale data processing (e.g., "big data") TM(hereinafter referred to as "Spark"). In at least one embodiment, the job scheduler 1433 may include a Spark driver to facilitate scheduling of workloads supported by different layers of the data center 1400. The configuration manager 1434 may be able to configure different layers, such as the software layer 1430 and the framework layer 1420 (which includes Spark and the distributed file system 1438 for supporting large-scale data processing). The resource manager 1436 may be able to manage the clustered or grouped computing resources mapped to or allocated for supporting the distributed file system 1438 and the job scheduler 1433. In at least one embodiment, the clustered or grouped computing resources may include the grouped computing resources 1414 in the data center infrastructure layer 1410. The resource manager 1436 may coordinate with the resource coordinator 1412 to manage these mapped or allocated computing resources.

[0227] In at least one embodiment, the software 1432 included in the software layer 1430 may include software used by at least a portion of the node C.R.s 1416(1)-1416(N), the grouped computing resources 1414, and / or the distributed file system 1438 of the framework layer 1420. One or more types of software may include, but are not limited to, Internet web search software, email virus scanning software, database software, and streaming video content software.

[0228] In at least one embodiment, the application 1442 included in the application layer 1440 may include one or more types of applications used by at least a portion of the node C.R.s 1416(1)-1416(N), the grouped computing resources 1414, and / or the distributed file system 1438 of the framework layer 1420. One or more types of applications may include, but are not limited to, any number of genomic applications, cognitive computing, and machine learning applications, including training or inference software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), and / or other machine learning applications used in conjunction with one or more embodiments.

[0229] In at least one embodiment, any one of the configuration manager 1434, the resource manager 1436, and the resource coordinator 1412 may implement any number and type of self-modifying actions based on any amount and type of data obtained in any technically feasible manner. The self-modifying actions may save the data center operator of the data center 1400 from making potentially poor configuration decisions and may avoid underutilization and / or poorly performing portions of the data center.

[0230] According to one or more embodiments described herein, data center 1400 may include tools, services, software, or other resources to train one or more machine learning models or to use one or more machine learning models to predict or infer information. For example, (one or more) machine learning models may be trained by calculating weight parameters according to a neural network architecture by using the software and / or computing resources described above with respect to data center 1400. In at least one embodiment, a trained or deployed machine learning model corresponding to one or more neural networks may be used to infer or predict information by using the resources described above with respect to data center 1400 by using weight parameters calculated by one or more training techniques (such as but not limited to those training techniques described herein).

[0231] In at least one embodiment, data center 1400 may use a CPU, an application specific integrated circuit (ASIC), a GPU, an FPGA, and / or other hardware (or their corresponding virtual computing resources) to perform training and / or inference by using the above resources. In addition, one or more of the software and / or hardware resources described above may be configured to allow a user to train or perform a service for inferring information, such as image recognition, speech recognition, or other artificial intelligence services.

[0232] Example Network Environment

[0233] A network environment suitable for implementing embodiments of the present disclosure may include one or more client devices, servers, network attached storage (NAS), other backend devices, and / or other device types. Client devices, servers, and / or other device types (e.g., each device) may be implemented on one or more instances of Figure 13 the (one or more) computing devices 1300 - for example, each device may include similar components, features, and / or functions of the (one or more) computing devices 1300. In addition, in the case of implementing a backend device (e.g., a server, NAS, etc.), the backend device may be included as part of data center 1400, an example of data center 1400 is described herein with respect to Figure 14 more detail.

[0234] The components of the network environment may communicate with each other via a network, which may be wired, wireless, or both. The network may include multiple networks or one network among multiple networks. For example, the network may include one or more wide area networks (WANs), one or more local area networks (LANs), one or more public networks (such as the Internet and / or the public switched telephone network (PSTN)), and / or one or more private networks. In the case where the network includes a wireless telecommunications network, components such as base stations, communication towers, or even access points (and other components) may provide a wireless connection.

[0235] Compatible network environments can include one or more peer-to-peer network environments (in which case, a server may not be included in the network environment) and one or more client-server network environments (in which case, one or more servers may be included in the network environment). In a peer-to-peer network environment, the functions described herein for a server can be implemented on any number of client devices.

[0236] In at least one embodiment, the network environment can include one or more cloud-based network environments, distributed computing environments, combinations thereof, etc. A cloud-based network environment can include a framework layer, a job scheduler, a resource manager, and a distributed file system implemented on one or more servers, which may include one or more core network servers and / or edge servers. The framework layer can include a framework that supports software layers and / or one or more applications of an application layer. The software or application can respectively include network-based service software or applications. In an embodiment, one or more client devices can use network-based service software or applications (e.g., by accessing service software and / or applications via one or more application programming interfaces (APIs)). The framework layer can be, but is not limited to, a free and open-source software web application framework that can use a distributed file system for large-scale data processing (e.g., "big data").

[0237] A cloud-based network environment can provide any combination of cloud computing and / or cloud storage that performs the computing and / or data storage functions (or one or more portions thereof) described herein. Any of these different functions can be distributed across multiple locations from a central or core server (e.g., one or more data centers in a state, region, country, globally, etc.). If the connection to a user (e.g., a client device) is relatively close to an edge server, the core server can assign at least a portion of the function to the edge server. A cloud-based network environment can be private (e.g., limited to a single organization), public (e.g., available to many organizations), and / or a combination thereof (e.g., a hybrid cloud environment).

[0238] (One or more) client devices can include those described herein with respect to Figure 13At least some of the components, features, and functions of the (one or more) example computing devices 1300 described. By way of example, and not limitation, a client device may be implemented as a personal computer (PC), laptop computer, mobile device, smartphone, tablet computer, smartwatch, wearable computer, personal digital assistant (PDA), MP3 player, virtual reality headset, global positioning system (GPS) or device, video player, camera, surveillance device or system, vehicle, boat, spacecraft, virtual machine, drone, robot, handheld communication device, hospital device, gaming device or system, entertainment system, vehicle computer system, embedded system controller, remote control, appliance, consumer electronic device, workstation, edge device, any combination of these depicted devices, or any other suitable device.

[0239] The present disclosure may be described in the general context of machine-usable instructions or computer code, including computer-executable instructions such as program modules, being executed by a computer or other machine, such as a personal digital assistant or other handheld device. Generally, program modules, including routines, programs, objects, components, data structures, etc., refer to code that performs particular tasks or implements particular abstract data types. The present disclosure may be practiced in a variety of system configurations, including handheld devices, consumer electronics, general-purpose computers, more specialized computing devices, etc. The present disclosure may also be practiced in a distributed computing environment where tasks are performed by remote processing devices linked through a communications network.

[0240] As used herein, the recitation of "and / or" with respect to two or more elements should be construed to refer to only one element or a combination of elements. For example, "element A, element B, and / or element C" may include only element A, only element B, only element C, element A and element B, element A and element C, element B and element C, or elements A, B, and C. Further, "at least one of element A or element B" may include at least one of element A, at least one of element B, or at least one of element A and at least one of element B. Still further, "at least one of element A and element B" may include at least one of element A, at least one of element B, or at least one of element A and at least one of element B.

[0241] The subject matter of the present disclosure is described in detail herein to meet statutory requirements. However, the description itself is not intended to limit the scope of the present disclosure. On the contrary, the inventors have contemplated that the claimed subject matter may also be embodied in other ways, including steps different from those described herein or combinations of steps similar to those described in connection with other current or future technologies. Moreover, although the terms "step" and / or "block" may be used herein to imply different elements of a method employed, these terms should not be construed as implying any particular order among or between the various steps disclosed herein unless the order of the steps is explicitly recited.

[0242] Example paragraph

[0243] A: A method, comprising: using one or more machine learning models and determining a classification corresponding to a portion of the image based at least on image data representative of the image; determining, based at least on map data associated with the environment, that a point within the environment corresponding to the portion of the image is associated with a driving surface; determining, based at least on the classification, whether the driving surface is occluded at the portion of the image; and generating first data indicative of whether the driving surface is occluded at the portion of the image.

[0244] B: The method of paragraph A, wherein determining whether the driving surface is occluded at the portion of the image comprises: determining that the classification does not include one or more surface classifications; and determining, based at least on the classification not including the one or more object classifications, that the driving surface is not occluded at the portion of the image.

[0245] C: The method of paragraph A or paragraph B, wherein determining whether the driving surface is occluded at the portion of the image comprises: determining that the classification includes one or more object classifications; and determining, based at least on the classification including the one or more object classifications, that the driving surface is occluded at the portion of the image by one or more objects corresponding to the one or more object classifications.

[0246] D: The method of any one of paragraphs A - C, wherein determining that the point within the environment is associated with the driving surface comprises: obtaining map data associated with the environment, the map data at least representing a label and a three - dimensional position of the point within the environment; projecting the three - dimensional position onto a two - dimensional position associated with the portion of the image; and determining, based at least on the label, that the point within the environment is associated with the driving surface.

[0247] E: The method of any one of paragraphs A - D, wherein generating the first data includes: generating the first data representing a label associated with the portion of the image, the label indicating one of the following: the driving surface is not occluded in the portion of the image; the driving surface is occluded by a dynamic object in the portion of the image; or the driving surface is occluded by a static object in the portion of the image.

[0248] F: The method of any one of paragraphs A - E, further comprising: using the one or more machine learning models and at least based on the image data, determining a second classification corresponding to a second portion of the image; determining at least based on the map data a second point in the environment corresponding to the second portion of the image and associated with the driving surface; determining at least based on the second classification whether the driving surface is occluded in the second portion of the image; and generating second data indicating whether the driving surface is occluded in the second portion of the image.

[0249] G: The method of any one of paragraphs A - F, further comprising: determining at least based on the map data a first distance associated with the point in the environment; and determining at least based on the point cloud data a second distance associated with the point in the environment, wherein determining whether the driving surface is occluded in the portion of the image is further based at least on the first distance and the second distance.

[0250] H: The method of paragraph G, further comprising: determining whether the second distance is within a threshold distance from the first distance, wherein determining whether the driving surface is occluded in the portion of the image is further based at least on whether the first distance is within the threshold distance from the second distance.

[0251] I: The method of paragraph G, further comprising: generating a first determination of whether the driving surface is occluded in the portion of the image at least based on the classification; and generating a second determination of whether the driving surface is occluded in the portion of the image at least based on the first distance and the second distance, wherein the determination of whether the driving surface is occluded in the portion of the image is based at least on the first determination and the second determination.

[0252] J: A system, comprising: one or more processing units for: determining a classification corresponding to a portion of the image based at least on image data representative of the image; determining, based at least on map data associated with the environment, that a point within the environment corresponding to the portion of the image is associated with a traffic object; determining, based at least on the classification, whether the traffic object is occluded at the portion of the image; and generating first data indicative of whether the traffic object is occluded at the portion of the image.

[0253] K: The system of paragraph J, wherein determining whether the traffic object is occluded at the portion of the image comprises: determining that the classification corresponds to the traffic object; and determining, based at least on the classification corresponding to the traffic object, that the traffic object is not occluded at the portion of the image.

[0254] L: The system of paragraph J or paragraph K, wherein determining whether the traffic object is occluded at the portion of the image comprises: determining that the classification does not correspond to the traffic object; and determining, based at least on the classification not corresponding to the traffic object, that the traffic object is occluded at the portion of the image.

[0255] M: The system of any one of paragraphs J-L, wherein determining that the point within the environment is associated with the traffic object comprises: obtaining map data associated with the environment, the map data representing at least a label and a three-dimensional position of the point within the environment; projecting the three-dimensional position onto a two-dimensional position associated with the portion of the image; and determining, based at least on the label, that the point within the environment is associated with the traffic object.

[0256] N: The system of any one of paragraphs J-M, wherein the first data represents a label associated with the portion of the image, the label indicating one of the following: the traffic object is not occluded at the portion of the image; the traffic object is occluded by a dynamic object at the portion of the image; or the traffic object is occluded by a static object at the portion of the image.

[0257] O: The system of any one of paragraphs J-N, wherein the one or more processing units are further for: determining a second classification corresponding to a second portion of the image based at least on the image data; determining, based at least on the map data, that a second point within the environment corresponding to the second portion of the image is associated with the traffic object; determining, based at least on the second classification, whether the traffic object is occluded at the second portion of the image; and generating second data indicative of whether the traffic object is occluded at the second portion of the image.

[0258] P: A system according to any one of paragraphs J - O, wherein the one or more processing units are further configured to: determine a first distance associated with the point within the environment based at least on the map data; and determine a second distance associated with the point within the environment based at least on the point cloud data, wherein determining whether the traffic object is occluded at the portion of the image is further based at least on the first distance and the second distance.

[0259] Q: A method according to paragraph P, wherein the one or more processing units are further configured to: generate a first determination as to whether the traffic object is occluded at the portion of the image based at least on the classification; and generate a second determination as to whether the traffic object is occluded at the portion of the image based at least on the first distance and the second distance, wherein the determination as to whether the traffic object is occluded at the portion of the image is based at least on the first determination and the second determination.

[0260] R: A system according to any one of paragraphs J - Q, wherein the system is included in at least one of the following: a control system for an autonomous or semi - autonomous machine; a perception system for an autonomous or semi - autonomous machine; a system for performing one or more simulation operations; a system for performing one or more digital twin operations; a system for performing optical transmission simulation; a system for performing collaborative content creation of 3D assets; a system for performing one or more deep learning operations; a system implemented using edge devices; a system implemented using robots; a system for performing one or more generative AI operations; a system for performing one or more operations using large language models; a system for performing one or more conversational AI operations; a system for generating synthetic data; a system for presenting at least one of virtual reality content, augmented reality content, or mixed reality content; a system comprising one or more virtual machines (VMs); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources.

[0261] S: A processor comprising: one or more processing units for generating first data indicative of whether an object or feature is occluded at a portion of an image, wherein a determination as to whether a traffic object is occluded at the portion of the image is generated based at least on a comparison between a first classification and a second classification, the first classification associated with the portion of the image is determined using one or more machine learning models, and the second classification is determined by projecting one or more labels from a map onto the portion of the image.

[0262] T: A processor of paragraph S, where the processor is included in at least one of the following: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing one or more simulation operations; a system for performing one or more digital twin operations; a system for performing optical transmission simulation; a system for performing collaborative content creation of 3D assets; a system for performing one or more deep learning operations; a system implemented using an edge device; a system implemented using a robot; a system for performing one or more generative AI operations; a system for performing one or more operations using a large language model; a system for performing one or more conversational AI operations; a system for generating synthetic data; a system for presenting at least one of virtual reality content, augmented reality content, or mixed reality content; a system including one or more virtual machines (VMs); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources.

[0263] U: A method, including: determining a first distance associated with a point corresponding to a portion of an image within an environment based at least on map data associated with the environment; determining a second distance associated with a point within the environment based at least on point cloud data; determining whether a driving surface is occluded at a portion of the image based at least on the first distance and the second distance; and generating first data indicating whether the driving surface is occluded at a portion of the image.

[0264] V: The method of paragraph U, further including: determining whether the second distance is within a threshold distance from the first distance, where determining whether the driving surface is occluded at a portion of the image is based at least on whether the first distance is within a threshold distance from the second distance.

[0265] W: The method of paragraph V, where determining whether the driving surface is occluded at a portion of the image includes one of the following: determining that the driving surface is not occluded at a portion of the image based at least on the second distance being within a threshold distance from the first distance; or determining that the driving surface is occluded at a portion of the image based at least on the second distance being outside a threshold distance from the first distance.

[0266] X: The method of paragraph V, further including: determining the threshold distance based at least on at least one of an incident angle associated with a point within the environment or a thickness associated with a point on the driving surface.

[0267] Y: The method of any one of paragraphs U-X, further including: obtaining initial point cloud data associated with the environment; obtaining LiDAR data generated using one or more LiDAR sensors; and generating point cloud data associated with the environment based at least on the initial point cloud data and the LiDAR data.

[0268] Z: The method of any one of paragraphs U - Y further includes: obtaining machine - generated image data representing an image; obtaining machine - generated LiDAR data, and obtaining the LiDAR data at least partially while generating the image data; and generating point cloud data associated with the environment based at least on the LiDAR data.

[0269] AA: The method of any one of paragraphs U - Z further includes: determining a third distance associated with a second point corresponding to a second portion of the image within the environment based at least on map data; determining a fourth distance associated with the second point within the environment based at least on the point cloud data; determining whether the driving surface is occluded at the second portion of the image based at least on the third distance and the fourth distance; and generating second data indicating whether the driving surface is occluded at the second portion of the image.

[0270] AB: The method of any one of paragraphs U - AA further includes: determining a classification corresponding to an image portion based at least on the image data representing the image; and determining a point within the environment associated with the driving surface based at least on map data, wherein determining whether the driving surface is occluded at the image portion is further based at least on the classification.

[0271] AC: The method of paragraph AB further includes: a first determination of whether the driving surface is occluded at the image portion based at least on the first distance and the second distance; and a second determination of whether the driving surface is occluded at the image portion based at least on the classification, wherein determining whether the driving surface is occluded at the image portion is based at least on the first determination and the second determination.

[0272] AD: A system includes: one or more processing units for: determining a first distance associated with a point corresponding to a portion of an image within the environment based at least on map data associated with the environment; determining a second distance associated with the point within the environment based at least on the point cloud data; determining whether a traffic object is occluded at the portion of the image based at least on the first distance and the second distance; and generating first data indicating whether the traffic object is occluded at the portion of the image.

[0273] AE: The system of paragraph AD, wherein the one or more processing units are further for: determining whether the second distance is within a threshold distance from the first distance, wherein the determination of whether the traffic object is occluded at the image portion is based at least on whether the first distance is within a threshold distance from the second distance.

[0274] AF: A system as in paragraph AE, wherein determining whether a traffic object is occluded at a portion of an image includes one of the following: determining that the traffic object is not occluded at the portion of the image based at least on a second distance within a threshold distance from a first distance; or determining that the traffic object is occluded at the portion of the image based at least on the second distance outside the threshold distance from the first distance.

[0275] AG: A system as in paragraph AE, wherein one or more processing units are further configured to determine the threshold distance based at least on an incident angle associated with a point in the environment or a thickness associated with a point on a driving surface.

[0276] AH: A system as in any one of paragraphs AD - AG, wherein one or more processing units are further configured to: obtain initial point cloud data associated with the environment; obtain LiDAR data generated using one or more LiDAR sensors; and generate point cloud data associated with the environment based at least on the initial point cloud data and the LiDAR data.

[0277] AI: A system as in any one of paragraphs AD - AH, wherein one or more processing units are further configured to: obtain image data generated by a machine, the image data representing an image; obtain LiDAR data generated by a machine, and obtain the LiDAR data at least while generating the image data; and generate point cloud data associated with the environment based at least on the LiDAR data.

[0278] AJ: A system as in any one of paragraphs AD - AI, wherein one or more processing units are further configured to: determine a classification corresponding to a portion of the image based at least on the image data representing the image; and determine that a point in the environment is associated with a traffic object based at least on map data, wherein determining whether the traffic object is occluded at the portion of the image is further based at least on the classification.

[0279] AK: A system as in paragraph AJ, wherein one or more processing units are further configured to: make a first determination of whether the traffic object is occluded at the portion of the image based at least on a first distance and a second distance; and make a second determination of whether the traffic object is occluded at the portion of the image based at least on the classification, wherein determining whether the traffic object is occluded at the portion of the image is based at least on the first determination and the second determination.

[0280] AL: A system according to any one of paragraphs AD - AK, wherein the system is included in at least one of the following: a control system for an autonomous or semi - autonomous machine; a perception system for an autonomous or semi - autonomous machine; a system for performing one or more simulation operations; a system for performing one or more digital twin operations; a system for performing optical transmission simulation; a system for performing collaborative content creation of 3D assets; a system for performing one or more deep learning operations; a system implemented using edge devices; a system implemented using robots; a system for performing one or more generative AI operations; a system for performing one or more operations using a large language model; a system for performing one or more conversational AI operations; a system for generating synthetic data; a system for presenting at least one of virtual reality content, augmented reality content, or mixed reality content; a system including one or more virtual machines (VMs); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources.

[0281] AM: A processor comprising: one or more processing units for generating first data indicating whether a traffic object is occluded at a portion of an image, wherein the determination of whether the traffic object is occluded at the portion of the image is generated based at least on a first distance associated with a point located in the environment determined using map data and a second distance associated with a point located in the environment determined using point cloud data.

[0282] AN: The processor of paragraph AM, wherein the processor is included in at least one of the following: a control system for an autonomous or semi - autonomous machine; a perception system for an autonomous or semi - autonomous machine; a system for performing one or more simulation operations; a system for performing one or more digital twin operations; a system for performing optical transmission simulation; a system for performing collaborative content creation of 3D assets; a system for performing one or more deep learning operations; a system implemented using edge devices; a system implemented using robots; a system for performing one or more generative AI operations; a system for performing one or more operations using a large language model; a system for performing one or more conversational AI operations; a system for generating synthetic data; a system for presenting at least one of virtual reality content, augmented reality content, or mixed reality content; a system including one or more virtual machines (VMs); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources.

Claims

1. A method comprising: determining, using one or more machine learning models and based on at least image data representing the image, a classification corresponding to a portion of the image; determining, based at least on map data associated with an environment, that a point within the environment corresponding to the portion of the image is associated with a driving surface; determining, based at least on the classification, whether the driving surface is occluded at the portion of the image; as well as First data is generated indicating whether the driving surface is occluded at the portion of the image.

2. The method of claim 1 , wherein determining whether the driving surface is obscured at the portion of the image comprises: determining that the classification does not include one or more surface classifications; as well as Based at least on the classification excluding the one or more object classifications, it is determined that the driving surface is not obstructed at the portion of the image.

3. The method of claim 1 , wherein determining whether the driving surface is obscured at the portion of the image comprises: determining that the classification includes one or more object classifications; as well as Based at least on the classifications including the one or more object classifications, it is determined that the driving surface is occluded at the portion of the image by one or more objects corresponding to the one or more object classifications.

4. The method of claim 1 , wherein determining that the point within the environment is associated with the driving surface comprises: acquiring map data associated with the environment, the map data representing at least a label and a three-dimensional position of the point within the environment; projecting the three-dimensional position to a two-dimensional position associated with the portion of the image; as well as The point within the environment is determined to be associated with the driving surface based at least on the tag.

5. The method according to claim 1, wherein generating the first data comprises: Generating the first data representing a tag associated with the portion of the image, the tag indicating one of: the driving surface is not obscured in the portion of the image; the driving surface is obscured by a dynamic object at the portion of the image; or The driving surface is obscured by a static object at the portion of the image.

6. The method according to claim 1, further comprising: determining, using the one or more machine learning models and based at least on the image data, a second classification corresponding to a second portion of the image; determining, based at least on the map data, that a second point within the environment corresponding to the second portion of the image is associated with the driving surface; determining whether the driving surface is occluded at the second portion of the image based at least on the second classification; as well as Second data is generated indicating whether the driving surface is occluded at the second portion of the image.

7. The method according to claim 1, further comprising: determining a first distance associated with the point within the environment based at least on the map data; as well as determining a second distance associated with the point within the environment based at least on the point cloud data, Wherein, determining whether the driving surface is obscured in the portion of the image is further based on at least the first distance and the second distance.

8. The method according to claim 7, further comprising: determining whether the second distance is within a threshold distance from the first distance, Wherein, determining whether the driving surface is obscured in the portion of the image is further based at least on whether the first distance is within the threshold distance from the second distance.

9. The method according to claim 7, further comprising: generating a first determination whether the driving surface is occluded at the portion of the image based at least on the classification; as well as generating a second determination of whether the traveling surface is obscured at the portion of the image based at least on the first distance and the second distance, Wherein, the determination of whether the driving surface is obscured in the portion of the image is based on at least the first determination and the second determination.

10. A system comprising: One or more processing units for: determining a classification corresponding to a portion of the image based at least on image data representing the image; determining, based at least on map data associated with an environment, that a point within the environment corresponding to the portion of the image is associated with a traffic object; determining, based at least on the classification, whether the traffic object is occluded at the portion of the image; as well as First data is generated indicating whether the traffic object is occluded at the portion of the image.

11. The system of claim 10, wherein determining whether the traffic object is occluded at the portion of the image comprises: determining that the classification corresponds to the traffic object; as well as Based at least on the classification corresponding to the traffic object, it is determined that the traffic object is not occluded at the portion of the image.

12. The system of claim 10, wherein determining whether the traffic object is occluded at the portion of the image comprises: determining that the classification does not correspond to the traffic object; as well as Based at least on the classification not corresponding to the traffic object, it is determined that the traffic object is occluded at the portion of the image.

13. The system of claim 10, wherein determining that the point within the environment is associated with the traffic object comprises: acquiring map data associated with the environment, the map data representing at least a label and a three-dimensional position of the point within the environment; projecting the three-dimensional position to a two-dimensional position associated with the portion of the image; as well as The point within the environment is determined to be associated with the traffic object based at least on the tag.

14. The system of claim 10, wherein the first data represents a label associated with the portion of the image, the label indicating one of: The traffic object is not occluded at the portion of the image; The traffic object is obscured by a dynamic object at the portion of the image; or The traffic object is occluded by a static object at the portion of the image.

15. The system of claim 10, wherein the one or more processing units are further configured to: determining a second classification corresponding to a second portion of the image based at least on the image data; determining, based at least on the map data, that a second point within the environment corresponding to the second portion of the image is associated with the traffic object; determining whether the traffic object is occluded at the second portion of the image based at least on the second classification; as well as Second data is generated indicating whether the traffic object is occluded at a second portion of the image.

16. The system of claim 10, wherein the one or more processing units are further configured to: determining a first distance associated with the point within the environment based at least on the map data; and determining a second distance associated with the point within the environment based at least on the point cloud data, in, Determining whether the traffic object is occluded at the portion of the image is also based on at least the first distance and the second distance.

17. The method of claim 16, wherein the one or more processing units are further configured to: generating a first determination whether the traffic object is occluded at the portion of the image based at least on the classification; and generating a second determination of whether the traffic object is occluded at the portion of the image based at least on the first distance and the second distance, in, The determination of whether the traffic object is occluded at the portion of the image is based on at least the first determination and the second determination.

18. The system of claim 10, wherein the system is included in at least one of the following: control systems for autonomous or semi-autonomous machines; Perception systems for autonomous or semi-autonomous machines; A system for performing one or more simulation operations; A system for performing one or more digital twin operations; A system for performing light transport simulations; A system for performing collaborative content creation of 3D assets; A system for performing one or more deep learning operations; Systems implemented using edge devices; Systems implemented using robots; A system for performing one or more generative AI operations; A system for performing one or more operations using a large language model; A system for performing one or more conversational AI operations; Systems for generating synthetic data; a system for presenting at least one of virtual reality content, augmented reality content, or mixed reality content; A system comprising one or more virtual machines VM; A system implemented at least in part in a data center; or A system implemented at least in part using cloud computing resources.

19. A processor comprising: One or more processing units for generating first data indicating whether an object or feature is occluded at a portion of an image, wherein the determination as to whether a traffic object is occluded at the portion of the image is generated based at least on a comparison between a first classification and a second classification, the first classification associated with the portion of the image being determined using one or more machine learning models, and the second classification being determined by projecting one or more labels from a map onto the portion of the image.

20. The processor of claim 19, wherein the processor is included in at least one of the following: control systems for autonomous or semi-autonomous machines; Perception systems for autonomous or semi-autonomous machines; A system for performing one or more simulation operations; A system for performing one or more digital twin operations; A system for performing light transport simulations; A system for performing collaborative content creation of 3D assets; A system for performing one or more deep learning operations; Systems implemented using edge devices; Systems implemented using robots; A system for performing one or more generative AI operations; A system for performing one or more operations using a large language model; A system for performing one or more conversational AI operations; Systems for generating synthetic data; a system for presenting at least one of virtual reality content, augmented reality content, or mixed reality content; A system comprising one or more virtual machines VM; A system implemented at least in part in a data center; or A system implemented at least in part using cloud computing resources.

Citation Information

Patent Citations

  • Method for programmable timeouts of tree traversal mechanisms in hardware

    US10885698B2