Confidence and visibility modeling in autonomous systems and applications
By generating a visibility confidence model, identifying and processing occlusions and errors in sensor data, sensor blindness, blurring and blocking are solved, improving machine perception capabilities and reliability of control decisions.
Patent Information
- Application Number
- CN202411882648.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-12-20
- Filing Date
- 2024-12-19
- Publication Date
- 2025-06-20
AI Technical Summary
The prior art is difficult to effectively solve sensor blindness, blur and blocking problems, especially in multi-sensor systems, resulting in a decline in machine perception capabilities.
Using a visibility confidence model associated with multiple sensors, the confidence level of sensor data is represented by data structures and visual representations, and occlusions and errors in sensor data are identified and processed.
Improves the system's ability to make informed control decisions in the face of sensor blocking, blurring and occlusion, and enhances reliable perception of the vehicle's surroundings.
Smart Images

Figure CN120180849A_ABST
Abstract
Description
BACKGROUND OF THE INVENTION
[0001] Autonomous driving systems and advanced driver assistance systems (ADAS) can utilize various sensors (e.g., cameras, LiDAR, RADAR, etc.) to perform various tasks such as lane keeping, lane changing, lane assignment, camera calibration, and / or localization. For example, to enable autonomous and ADAS systems to operate independently and efficiently, a real-time or near-real-time understanding of the vehicle's surrounding environment can be generated. To accurately and efficiently understand the vehicle's surrounding environment, it is important for associated sensors to generate available, unobscured sensor data (e.g., images, depth maps, etc.). However, the ability of the corresponding sensors to sense the surrounding environment can be impaired by various sources such as sensor blockage (e.g., from debris, precipitation, etc.), blurring, occlusion, etc., which can lead to sensor blindness. In addition, blockage, blurring, and / or other obscured sensor data associated with a single sensor can reduce the ability of one or more machines to sense the environment, even if the machine includes multiple sensors or can obtain sensor data corresponding to multiple sensors. In some cases, potential causes of sensor blindness can include snow, rain, glare, solar flares, mud, water, signal failures, etc.
[0002] One or more conventional methods for determining sensor blindness, blurring, obstruction, etc. include performing confidence modeling or degradation modeling on a specific sensor, e.g., predicting and / or determining how a sensor will degrade based on time, environmental conditions, etc. However, a limitation of this approach is that other systems (e.g., planning systems, obstacle fusion systems, tracking systems, etc.) may not be able to easily query the respective models corresponding to each sensor before making one or more determinations and / or generating one or more control commands. This limitation increases in systems that include multiple sensors corresponding to multiple respective sensor modalities. SUMMARY OF THE INVENTION
[0003] According to one or more embodiments of the present disclosure, one or more visibility systems, subsystems, machine models, neural networks, etc. can be used to generate one or more visibility confidence models associated with sensor data corresponding to multiple sensors. In some embodiments, the visibility model can include one or more data structures and / or a visual representation of one or more aggregated fields of view. For example, the visibility confidence model can be represented as a top-down view of the aggregated field of view as indicated by the sensor data or using a top-down view of the aggregated field of view as indicated by the sensor data.
[0004] In some embodiments, a visibility confidence model may indicate a confidence level of sensor data, which may correspond to respective sub-sections of an aggregated field of view. In some embodiments, respective fields of view corresponding to the aggregated field of view may each define a potential spatial coverage of sensor data, which may correspond to one or more sensors associated with a machine.
[0005] In some embodiments, a respective confidence level associated with the visibility confidence model may be determined based on one or more faults or errors associated with an individual sensor among one or more sensors corresponding to the machine. In some embodiments, one or more faults or errors may additionally indicate whether a sensor is “healthy” or operating properly. In some embodiments, the presence of one or more errors and / or faults may broadly indicate that a sensor may not be operating properly (e.g., the sensor is not electrically coupled to the machine, the sensor no longer generates sensor data, etc.).
[0006] In some embodiments, a respective confidence level may additionally be determined based on one or more coarse-level degradations and / or obstructions associated with the sensor and / or sensor data corresponding to the sensor. In some embodiments, coarse-level degradations and / or obstructions may include one or more obstructions applicable to all sensor data associated with the sensor. Additionally or alternatively, a respective confidence level may be determined based on one or more fine-level degradations corresponding to the sensor data. Fine-level sensor degradations may include one or more obstructions applicable to a portion of the sensor data generated using the sensor.
[0007] In some embodiments, the method may further include: identifying one or more occlusions that may be present in the sensor data, which may correspond to one or more individual sub-sections of the aggregated field of view. In some embodiments, the method may further include: performing one or more operations based on the visibility confidence model and / or a confidence level corresponding to an individual sub-region of the aggregated field of view.
[0008] Embodiments of the present disclosure may enhance the ability of a system to make one or more control decisions based on reliable sensor data by generating a data structure from which one or more systems, subsystems, control systems, etc. may obtain confidence information corresponding to a particular location. Additionally, in some embodiments, by generating a visibility confidence model of an aggregated field of view centered on an origin or reference frame corresponding to a machine and corresponding to a particular sensor modality, one or more systems may enhance their effectiveness in making informed control determinations in view of potential sensor obstructions, visibility distance degradations, occlusions, and / or the like. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] The present system and method correspond to generating one or more visibility confidence models corresponding to sensor data, wherein:
[0010] Figure 1 is a schematic diagram showing an example environment related to multiple sensors according to one or more embodiments of the present disclosure, the multiple sensors generating sensor data corresponding to respective fields of view or sensing fields;
[0011] Figure 2 shows an example environment for generating a visibility confidence model according to one or more embodiments of the present disclosure;
[0012] Figure 3A shows an example environment according to one or more embodiments of the present disclosure, the example environment showing a visibility system for generating one or more visibility confidence models;
[0013] Figure 3B shows an example environment depicted by an image divided into discrete sub - parts according to one or more embodiments of the present disclosure;
[0014] Figure 3C shows a depiction of an environment having one or more projection rays corresponding to sensor data according to one or more embodiments of the present disclosure;
[0015] Figure 4 is a flowchart showing a method for generating one or more visibility confidence models and generating one or more control determinations based on the one or more visibility confidence models according to one or more embodiments of the present disclosure;
[0016] Figure 5A is an illustration of an example autonomous vehicle according to one or more embodiments of the present disclosure;
[0017] Figure 5B is according to one or more embodiments of the present disclosure Figure 5A an example of the camera positions and fields of view of an example autonomous vehicle;
[0018] Figure 5C is according to one or more embodiments of the present disclosure Figure 5A a block diagram of an example system architecture of an example autonomous vehicle;
[0019] Figure 5D is a system diagram of the communication between a cloud - based server and Figure 5A an example autonomous vehicle according to one or more embodiments of the present disclosure;
[0020] Figure 6is a block diagram of an example computing device suitable for implementing one or more embodiments of the present disclosure; and
[0021] Figure 7 is a block diagram of an example data center suitable for implementing one or more embodiments of the present disclosure. Detailed Description
[0022] One or more embodiments of the present disclosure may relate to generating a data structure, model, and / or another representation that may include information corresponding to a confidence level of sensor data defining one or more fields of view or sensor fields. In non-limiting embodiments, the representation may include or be referred to as a “visibility confidence model.” In some embodiments, one or more visibility systems, subsystems, machine models, neural networks, etc. may be configured to generate one or more visibility confidence models associated with sensor data corresponding to one or more aggregated fields of view.
[0023] In some embodiments, a reference to a particular field of view or sensor field may indicate a potential spatial coverage of sensor data corresponding to an individual sensor. The potential spatial coverage is identified by one or more inherent characteristics of the individual sensor; such as, for example, focal length, lens size, lens type, mirror diameter, aperture size, scan density, resolution, transmit power, antenna size, transmit-to-receive power ratio, etc. For example, a particular image sensor may generate image data representing a particular portion of the environment corresponding to a particular field of view of the particular image sensor (e.g., within the particular field of view of the particular image sensor). Similarly, a particular LiDAR sensor may generate LiDAR data representing or corresponding to a particular portion of the environment within the sensor field or field of view of the LiDAR sensor, which may include a 360-degree scan.
[0024] In some embodiments, an aggregated field of view may be generated based on the potential spatial coverage of sensor data corresponding to multiple sensors. The aggregated field of view may include a combination of the respective fields of view of the sensors such that, in at least some cases, the aggregated field of view may provide a larger field of view compared to any one of the respective fields of view used to generate the aggregated field of view. For example, in the context of a self-machine, the self-machine may include multiple cameras, where each camera may be configured to generate sensor data corresponding to a respective field of view. Continuing with this example, the aggregated field of view may refer to the aggregation of the respective fields of view or sensor fields associated with the respective cameras and / or other sensor modalities corresponding to the self-machine.
[0025] According to one or more embodiments of the present disclosure, one or more visibility systems, subsystems, machine models, neural networks, etc. may be configured to generate one or more visibility confidence models associated with sensor data corresponding to one or more aggregated fields of view. In some embodiments, the visibility confidence model may include one or more data structures and / or a visual representation of one or more aggregated fields of view. For example, the visibility confidence model may include a top-down view or a bird's-eye view (BEV) representation of the aggregated field of view indicated by the sensor data. Additionally or alternatively, embodiments of the present disclosure may include generating queries and / or otherwise requesting and / or retrieving data and / or information from the data structure. Further, the data and / or information that may be received may additionally be used to generate one or more control determinations.
[0026] In some embodiments, one or more visual representations and / or data structures may be subdivided into one or more sub-parts of the aggregated field of view. In some embodiments, one or more sub-parts may represent corresponding parts of the aggregated field of view. In some embodiments, the corresponding parts of the aggregated field of view may include parts of a volume defined by the aggregated field of view. For example, the aggregated field of view may be represented by sensor data corresponding to a region or volume associated with the machine. Continuing with this example, the aggregated field of view may be subdivided into one or more sub-parts, and in some embodiments, the aggregated field of view may be represented by a visibility confidence model that indicates the corresponding confidence of the sensor data corresponding to the corresponding sub-parts of the aggregated field of view.
[0027] In some embodiments, the visibility confidence model may indicate the confidence levels of the sensor data corresponding to the respective sub-parts of the aggregated field of view. In some embodiments, a low confidence level may indicate that the machine may not rely on the data corresponding to a particular sub-part of the aggregated field of view. Additionally or alternatively, a low confidence level may indicate that in the case where the machine is receiving conflicting data regarding a particular sub-part of the aggregated field of view, the machine should rely on the data corresponding to one or more other sensor modalities. In some embodiments, the visibility confidence model may be a data structure that may be queried by one or more other systems, subsystems, processing units, etc.
[0028] One or more embodiments disclosed herein may relate to generating and / or querying a data structure indicative of a visibility level corresponding to an environment in which a self-machine may be located. In some embodiments, the self-machine may be configured to generate the data structure and / or generate one or more queries to the data structure, wherein the data and / or information associated with and / or included in the data structure may assist the self-machine in navigating a particular environment. In some embodiments, the self-machine may include any suitable machine or system capable of performing one or more autonomous and / or semi-autonomous operations. Example self-machines may include, but are not limited to, vehicles (land, sea, space, and / or air), robots, robotic platforms, etc. By way of example, a self-machine computing application may include one or more applications executable by an autonomous vehicle or a semi-autonomous vehicle, such as with respect to Figures 5A to 5D the example autonomous or semi-autonomous vehicle or machine 500 described (alternatively referred to herein as "vehicle 500" or "self-machine 500"). In the present disclosure, references to an "autonomous vehicle" or a "semi-autonomous vehicle" may include any vehicle configured to perform one or more autonomous or semi-autonomous navigation or driving operations. Thus, such vehicles may also include vehicles in which an operator is required or in which an operator may also perform such operations.
[0029] The systems and methods described herein may be used by, but are not limited to, non-autonomous vehicles or machines, semi-autonomous vehicles or machines (e.g., in one or more adaptive driver assistance systems (ADAS)), autonomous vehicles or machines, manned and unmanned robots or robotic platforms, warehouse vehicles, off-road vehicles, vehicles coupled to one or more trailers, airships, vessels, shuttles, emergency response vehicles, motorcycles, electric or motorized bicycles, airplanes, engineering vehicles, submarines, drones, and / or other vehicle types. Additionally, the systems and methods described herein may be used for various purposes, such as, but not limited to, for machine control, machine positioning, machine driving, synthetic data generation, model training, perception, augmented reality, virtual reality, mixed reality, robotics, security and surveillance, simulation and digital twins, autonomous or semi-autonomous machine applications, deep learning, environment simulation, object or actor simulation, and / or digital twins, data center processing, conversational AI, optical transmission simulation (e.g., ray tracing, path tracing, etc.), collaborative content creation of 3D assets, cloud computing, and / or any other suitable application.
[0030] The disclosed embodiments may be included in various different systems, such as automotive systems (e.g., control systems for autonomous or semi-autonomous machines, perception systems for autonomous or semi-autonomous machines), systems implemented using robots, aviation systems, medical systems, boating systems, intelligent area monitoring systems, systems for performing deep learning operations, systems for performing simulation operations, systems for performing digital twin operations, systems implemented using edge devices, systems including one or more virtual machines (VMs), systems for performing synthetic data generation operations, systems implemented at least partially in a data center, systems for performing conversational AI operations, implementing one or more language models (such as one or more large language models (LLMs) that process text, audio, images, sensors, and / or other data types to generate one or more outputs), systems for hosting real-time streaming applications, systems for presenting one or more of virtual reality content, augmented reality content, or mixed reality content, systems for performing optical transmission simulations, systems for performing collaborative content creation of 3D assets, systems implemented at least partially using cloud computing resources, and / or other types of systems.
[0031] These and other embodiments of the present disclosure will be explained with reference to the accompanying drawings. It should be understood that these figures are illustrative and schematic representations of such example embodiments and are not restrictive and not necessarily drawn to scale. In the figures, features with the same numerals indicate the same structure and function unless otherwise noted.
[0032] Now referring to Figure 1 , Figure 1 FIG. is a schematic diagram showing an example environment 100 related to a plurality of sensors 104 according to one or more embodiments of the present disclosure, the plurality of sensors 104 generating sensor data corresponding to respective fields of view (or sensing fields) 106. In some embodiments, the example environment 100 may include a physical area where the ego machine 102 may be located, but in embodiments, a virtual or simulated environment may be used for testing / verification purposes. In some embodiments, the sensors 104 may correspond to the ego machine 102 (e.g., the sensors 104 may be located on or in the ego machine 102) and may generate sensor data corresponding to one or more fields of view 106.
[0033] The ego machine 102 may include one or more machines that may be configured to receive or otherwise obtain sensor data corresponding to one or more sensors 104. In some embodiments, the ego machine 102 may include one or more machines that may navigate a particular environment based at least in part on the sensor data corresponding to one or more sensors 104. In some embodiments, the ego machine 102 may include any suitable machine or system capable of performing one or more autonomous or semi-autonomous operations. For example, the ego machine 102 may include an autonomous vehicle or a semi-autonomous vehicle, such as the example autonomous or semi-autonomous machine or vehicle 500 (also referred to herein as "vehicle 500" or "ego machine 500") described with respect to Figures 5A to 5D the example autonomous or semi-autonomous machine or vehicle 500 (also referred to herein as "vehicle 500" or "ego machine 500"). In the present disclosure, references to an "autonomous vehicle" or a "semi-autonomous vehicle" may include any vehicle that may be configured to perform one or more autonomous or semi-autonomous navigation or driving operations. Thus, such vehicles may also include vehicles in which an operator is required or in which an operator may also perform such operations.
[0034] In some embodiments, the ego machine 102 may include one or more systems, subsystems, machine learning models, neural networks, deep neural networks, etc., that may be configured to determine one or more perception, planning, control, safety, and / or other operations based on data that may be received and / or otherwise obtained. For example, sensor data corresponding to one or more sensors 104 may be obtained or otherwise communicated to the ego machine 102. The ego machine 102 may be configured to determine one or more operations to perform based at least in part on the sensor data generated using one or more corresponding sensors 104. For example, one or more operations may include changing one or more paths through the environment, decelerating, accelerating, turning, changing lanes, performing one or more evasive maneuvers, stopping, relinquishing control to a human operator, etc. In addition to physical movement, the ego machine 102 may perform one or more internal operations based on the sensor data, such as, for example, increasing the power of one or more internal heating elements in response to low temperature, or conversely, increasing the power of one or more cooling elements (e.g., fans, etc.) in response to high temperature, where the temperature may be determined using sensor data collected and / or generated using one or more sensors corresponding to the ego machine 102.
[0035] In some embodiments, the ego machine 102 may include sensors 104 associated with or corresponding to it. For example, the sensors 104 may be disposed on the ego machine 102. In some embodiments, as Figure 1As shown, the ego machine 102 may include a first sensor 104a, a second sensor 104b, and a third sensor 104c, collectively referred to as "sensor 104".
[0036] The sensors 104 may each correspond to a specific sensor modality. A sensor modality may include a specific class of sensors that can be used to detect and measure one or more characteristics of the environment 100. In some embodiments, different sensor modalities may refer to differences in the sensor data corresponding to the sensors 104. For example, the first sensor modality may include sensors that can generate image data, the second sensor modality may include sensors that can generate RADAR data, the third sensor modality may include sensors that can generate LiDAR data, and so on.
[0037] In some embodiments, two or more of the sensors 104 may correspond to the same sensor modality. For example, the first sensor 104a, the second sensor 104b, and the third sensor 104c may be image sensors (e.g., cameras). Additionally or alternatively, two or more of the sensors 104 may correspond to different sensor modalities. For example, the first sensor 104a may be an image sensor, the second sensor 104b may be a RADAR sensor, and the third sensor 104c may be a LiDAR sensor.
[0038] In some embodiments, the sensors 104 may be configured to generate and / or collect sensor data corresponding to the environment 100. For example, in the context of the sensors 104 being image sensors, the image sensors may be configured to generate image data corresponding to the environment. In this example, the image data may be used by the ego machine 102 to sense one or more characteristics, objects, obstacles, and / or other parts of the environment 100 in which the ego machine 102 may be located.
[0039] In some embodiments, the sensors 104 may be configured to generate sensor data corresponding to the respective fields of view (or sensing fields) of the sensors 104 (e.g., represented by the fields of view 106). For example, in some embodiments, the first sensor 104a may be configured to generate sensor data corresponding to a first field of view 106a, the second sensor 104b may be configured to generate sensor data corresponding to a second field of view 106b, and the third sensor 104c may be configured to generate sensor data corresponding to a third field of view 106c. As Figure 1 shown, the set of regions associated with the first field of view 106a, the second field of view 106b, and the third field of view 106c may be collectively referred to as the field of view 106.
[0040] In some embodiments, the respective field of view 106 may define a potential spatial coverage of sensor data corresponding to an individual sensor. In some embodiments, the potential spatial coverage may correspond to a specific plane. For example, the horizon plane corresponds to the ground in the environment 100 in which the ego machine 102 may be located. Additionally or alternatively, the spatial coverage defined by the respective field of view 106 may include a respective volume corresponding to the environment 100, in which sensor data corresponding to the respective sensors 104 may be generated and / or collected.
[0041] In some embodiments, the respective field of view 106 may be defined based on an individual sensor and / or individual sensor characteristics. In some embodiments, one or more intrinsic characteristics of the individual sensor may be used to identify the spatial coverage; such as, for example, focal length, lens size, lens type, mirror diameter, aperture size, scan density, resolution, transmit power, antenna size, transmit-to-receive power ratio, etc. For example, a particular image sensor may generate image data representing a particular portion of the environment corresponding to a particular field of view of the particular image sensor (e.g., within the particular field of view of the particular image sensor).
[0042] Additionally or alternatively, the respective fields of view 106 may be based on one or more mechanical limitations associated with the respective sensors 104 and / or the ego machine 102. In some embodiments, for example, the placement of the sensors relative to the ego machine 102 may affect the corresponding field of view 106 of the sensors 104. For example, in the context of the ego machine 102 being a vehicle, an image sensor may be placed at the center of the vehicle's rear bumper. Accordingly, the field of view of the image sensor may be limited to the environmental area behind the vehicle. Other mechanical limitations that may affect the respective field of view 106 associated with the corresponding sensors 104 may include range of motion, sensor housing, power and energy constraints, environmental conditions, etc. For example, in cases where power savings are desired, the ability of the sensor to tilt, move, translate, etc. may be artificially limited, and thus, the corresponding field of view 106 may be limited accordingly.
[0043] In some embodiments, the respective field of view 106 may be based on the type of sensor 104. For example, in the context of a first sensor 104a being an image sensor, the image sensor may include a series of degrees in the horizontal and vertical directions and a reference distance in front of the image sensor, which may be included in the field of view that may be captured by the image data corresponding to the image sensor (e.g., 50 degrees in the horizontal direction, 50 degrees in the vertical direction, and a distance of 10 meters). As a further example, in the context of a second sensor 104b being a RADAR sensor, the corresponding field of view may include a horizontal field of view of 30 degrees and a vertical field of view of 10 degrees, at a distance of 200 meters, which may be captured by the RADAR sensor data corresponding to the RADAR sensor.
[0044] In some embodiments, the corresponding field of view 106 may be determined, defined, and / or otherwise calculated based on characteristics associated with the sensor 104 itself rather than based on obstacles, obstructions, etc. that may be present in the environment. For example, an image sensor may be configured to generate sensor data corresponding to a field of view that includes a 75-degree horizontal field of view at a distance of 10 meters and a 50-degree vertical field of view at a distance of 10 meters. Continuing with this example, the visibility of the sensor may be limited by obstructions, obstacles, objects, etc., however, the field of view 106 may remain unchanged.
[0045] In some embodiments, the first field of view 106a may be the same as the second field of view 106b and / or the third field of view 106c. Additionally or alternatively, the first field of view 106a may be different from the second field of view 106b and / or the third field of view 106c. In some embodiments, as Figure 1 shown, the fields of view 106 may overlap with each other. For example, the first field of view 106a may overlap completely or partially with the second field of view 106b and / or the second field of view 106b may overlap completely or partially with the third field of view 106c.
[0046] In some embodiments, the first field of view 106a, the second field of view 106b, and the third field of view 106c may be collectively included in an aggregate field of view that may be associated with the ego machine 102. In some embodiments, the aggregate field of view may include a total area and / or volume that may be defined by sensor data generated and / or collected using the sensor 104.
[0047] In some embodiments, the aggregate field of view (or aggregate sensing field) may be aggregated based on sensors in the same sensor modality. For example, in the context where the first sensor 104a and the second sensor 104b are image sensors and the third sensor 104c is a LiDAR sensor, the aggregate field of view may correspond to the image data and may thus include the area and / or volume that may be defined by the first field of view 106a and the second field of view 106b. Additionally or alternatively, the aggregate field of view may include different types of sensors. For example, continuing in the context where the first sensor 104a and the second sensor 104b are image sensors and the third sensor 104c is a LiDAR sensor, the aggregate field of view may include the total area and / or volume defined by the first field of view 106a, the second field of view 106b, and the third field of view 106c.
[0048] In some embodiments, sensors within the field of view 106 may generate and / or collect sensor data that may indicate the presence or absence of one or more objects (e.g., object 108). In some embodiments, object 108 may include one or more objects within the environment 100. For example, object 108 may include one or more static and / or dynamic objects, road hazards, obstacles, items, signs, waiting conditions, traffic signals, construction barriers, boundaries, etc. within the environment 100. For example, in the context where the ego machine 102 is a vehicle, object 108 may include one or more other vehicles. Additionally or alternatively, object 108 may include and / or represent one or more obstacles, dips, speed bumps, potholes or other voids, pedestrians, lanes, etc. In some embodiments, sensor data corresponding to one or more sensors 104 may be used to detect object 108. In some embodiments, sensor data corresponding to one or more sensors 104 within the field of view 106 of one or more individual sensors 104 may be used to detect object 108.
[0049] In some embodiments, sensor data corresponding to the second sensor 104b may be used to capture and / or identify object 108. Additionally or alternatively, sensor data corresponding to the first sensor 104a, the third sensor 104c, and / or the second sensor 104b may be used to capture and / or identify object 108. In some embodiments, it may be determined that object 108 occupies a particular portion of the aggregated field of view defined by the first field of view 106a, the second field of view 106b, and the third field of view 106c.
[0050] In some embodiments, the aggregated field of view may be divided into one or more portions, where one or more portions include one or more sub-regions and / or volumes associated with the aggregated field of view. In some embodiments, one or more portions may be covered by sensor data corresponding to multiple sensors or otherwise detectable using sensor data corresponding to multiple sensors. Additionally or alternatively, one or more portions may be covered by sensor data corresponding to an individual sensor or otherwise detectable using sensor data corresponding to an individual sensor. In some embodiments, it may be determined that a portion of one or more portions of the aggregated field of view may be blurred or occluded. In some embodiments, the ego machine 102 may be configured to assign a lower weight to sensor data corresponding to portions of the aggregated field of view that may be blocked, blurred, or otherwise occluded. Determining whether one or more portions of the aggregated field of view may be blocked, blurred, etc. may be described in the present disclosure such as, for example, with respect to Figure 2 , Figure 3A , Figure 3B , Figure 3C and / or Figure 4Further description and / or illustration.
[0051] Modifications, additions, or omissions can be made to Figure 1 without departing from the scope of the present disclosure. For example, the number of ego machines 102, the number of sensors 104, and the corresponding fields of view 106 can vary. The number, size, type, etc. of objects 108 can vary. The details given and discussed are for the purpose of helping to explain and understand the concepts of the present disclosure and are not meant to be limiting.
[0052] Figure 2 An example environment 200 for generating a visibility confidence model 206 in accordance with one or more embodiments of the present disclosure is shown. In some embodiments, environment 200 can be the same as and / or included in environment 100. For example, the visibility confidence model 206 can include a model indicating the confidence level of sensor data 104 corresponding to the field of view 106.
[0053] In some embodiments, environment 200 can include a visibility system 204, which can be configured to generate a visibility confidence model 206 using sensor data 202. In some embodiments, the visibility system 204 can be included in and / or associated with an ego machine 208 that can be located in environment 200, where the ego machine 208 can be the same as and / or similar to the ego machine 102.
[0054] Sensor data 202 can include data that can be generated and / or collected using one or more sensors. In some embodiments, one or more sensors can be associated with the ego machine 208 that can correspond to the visibility system 204. In some embodiments, sensors that can be used for one or more visibility systems can be used to generate and / or collect sensor data, such as, for example, one or more image sensors, RADAR sensors, LiDAR sensors, sound navigation and ranging (SONAR) sensors, infrared sensors, ultrasonic sensors, proximity sensors, and / or any other type of sensor that can generate and / or collect sensor data 202 that can be used for one or more visibility systems (e.g., visibility system 204). In these or other embodiments, sensors such as, for example, the sensors 104 Figure 1 further described and / or illustrated in the present disclosure can be used to generate and / or collect sensor data 202.
[0055] In some embodiments, the sensor data 202 may correspond to one or more sensors that may generate and / or collect data corresponding to the environment in which the sensors are located. For example, one or more temperature sensors, humidity sensors, barometric pressure sensors (e.g., barometers), rain gauges, ultraviolet sensors, radiation detectors, and / or other sensors that may generate and / or collect data and / or information corresponding to the environment in which the sensors are located may be used to generate and / or collect the sensor data 202.
[0056] In some embodiments, the sensor data 202 may include data and / or information that may be generated using one or more sensors that may be configured to measure one or more movement characteristics. For example, the sensor data 202 may include data and / or information generated using one or more speed sensors, accelerometers, gyroscopes, inertial measurement units (IMUs), global positioning system (GPS) sensors, strain gauges, tilt sensors (e.g., inclinometers), and / or other sensors that may be used to generate and / or collect data corresponding to one or more movement characteristics of one or more machines, systems, devices, etc.
[0057] In some embodiments, the sensor data 202 may correspond to the ego machine 208. For example, the sensor data 202 may include data indicating the speed and / or rate of the ego machine 208, the position and / or orientation of the ego machine 208, the acceleration of the ego machine 208, etc. Additionally, the sensor data 202 may include information corresponding to the environment in which the ego machine 208 may be located. For example, the sensor data 202 may indicate that it may be raining in the environment in which the ego machine 208 is located, there is a certain amount of precipitation, and the environment is at a certain temperature. Additionally or alternatively, the sensor data 202 may be used in one or more visibility systems 204 associated with the ego machine 208. For example, RADAR data may be generated using a RADAR sensor located in or on the ego machine 208 such that the sensor data 202 may correspond to a portion of the environment around the ego machine 208.
[0058] In some embodiments, the sensor data 202 may be associated with a particular field of view or sensing field. For example, the sensor data 202 may correspond to a first sensor that may have a first field of view, the sensor data 202 may correspond to a second sensor that may have a second field of view, and so on. In some embodiments, the sensor data 202 may correspond to a particular data frame, where the data frame includes the sensor data 202 corresponding to a particular field of view at a particular timestamp. In some embodiments, the field of view to which the sensor data 202 may correspond may be the same and / or similar to the field of view 106 further described and / or illustrated in the present disclosure, such as, for example, Figure 1 the same as and / or similar to the field of view 106 further described and / or illustrated in the present disclosure.
[0059] In some embodiments, the sensor data 202 may include metadata. In some embodiments, the metadata may indicate the structure of the sensor data 202, the sensors and / or channels to which the sensor data 202 may correspond, one or more timestamps when the sensor data 202 was captured, data and / or information corresponding to one or more inherent characteristics of the sensors that may generate the sensor data 202 (e.g., focal length, sensitivity, accuracy, resolution, calibration requirements, operating range, etc.). In some embodiments, the sensor data 202 may include the position and orientation of the corresponding sensor relative to the ego machine 208. For example, a camera may generate the sensor data 202. The camera may be located in front of the ego machine 208 and point forward relative to the ego machine 208. Continuing with this example, the camera may generate image data with a field of view that extends one hundred (100) yards and fifteen (15) degrees forward on either side of the centerline of the ego machine 208. Further continuing with this example, the sensor data 202 may include relative information corresponding to the field of view such that the individual fields of view of one or more sensors can be determined. For example, the sensor data 202 may include information that can be used to determine the position at which the individual field of view corresponding to a respective sensor may be located relative to one or more other individual fields of view.
[0060] In some embodiments, the sensor data 202 may include data and / or information corresponding to one or more errors that the sensors that may have generated and / or collected the sensor data 202 may have encountered. For example, one or more electrical errors, disconnections, short circuits, failures to generate data correctly, errors in transmitting or otherwise communicating the sensor data 202, and / or one or more other errors associated with the individual sensors.
[0061] In some embodiments, the sensor data 202 may be transmitted and / or otherwise communicated to the ego machine 208 and / or the visibility system 204. In some embodiments, the ego machine 208 may include one or more systems, robots, vehicles, drones, devices, etc., that may be configured to perform one or more operations based on the sensor data 202. For example, the ego machine 208 may include one or more drones, robots, industrial robots, boats, cars, trucks, other vehicles, ego machines, etc. In some embodiments, the ego machine 208 may be configured to travel from one location to another based on the sensor data 202. In some embodiments, traveling from one location to another may include traveling from one location on the ground plane to a second location on the ground plane. Additionally or alternatively, traveling from one location to another may include the ego machine 208 traveling from a first location and elevation through a particular volume to a second location and elevation.
[0062] In some embodiments, the ego machine 208 may be configured to use the sensor data 202 to traverse an environment and / or a portion of the environment (e.g., from a first location to a second location). For example, the ego machine 208 may include one or more cameras that may collect and / or generate image data corresponding to the environment. Continuing with this example, the ego machine 208 may be configured to use the image data generated using one or more corresponding sensors to navigate from a first portion of the environment to a second portion of the environment. Although this example uses image data, the use of image data is not meant to be limiting; for example, the ego machine 208 may include one or more other sensors, such as, for example, one or more RADAR sensors, LiDAR sensors, SONAR sensors, infrared sensors, and / or other sensors that may generate the sensor data 202.
[0063] In some embodiments, the ego machine 208 may be configured to communicate with one or more other machines, systems, subsystems, etc. to receive information corresponding to the environment. For example, one or more systems external to the ego machine 208 may be configured to transmit and / or otherwise convey the sensor data 202 to the ego machine 208. For example, one or more other systems may be configured to generate weather data, temperature data, humidity data, etc. corresponding to the environment or a portion of the environment. In some embodiments, the ego machine 208 may be configured to receive this information from one or more other sources, systems, devices, machines, servers, edge servers, cloud servers, data centers, etc. Continuing with this example, the ego machine 208 may be configured to adjust and / or change one or more control commands (e.g., change one or more paths through the environment, slow down, turn, change lanes, execute one or more avoidance strategies, etc.) based on the data and / or information that may have been received from one or more other systems.
[0064] In some embodiments, the ego machine 208 may include one or more visibility or perception systems, such as the visibility system 204. In some embodiments, the visibility system 204 may include a system, subsystem, neural network, collection of machine learning models, and / or one or more combinations of the foregoing that may be configured to perform a series of processing operations. In some embodiments, the visibility system 204 may be configured to populate the visibility confidence model 206 with data corresponding to the visibility and / or confidence level (referred to herein as the "confidence level") of sensor data associated with a particular sensor and / or particular sensor modality. In some embodiments, the confidence level may be determined based on sensor data 202 corresponding to individual sensors of a sensor modality. In some embodiments, the visibility system 204 may be included in and / or communicate with one or more perception systems corresponding to the ego machine 208.
[0065] In some embodiments, the visibility system 204 may be configured to populate and / or generate a visibility confidence model 206 corresponding to the environment around the ego machine 208. In some embodiments, the visibility confidence model 206 may include one or more data structures and / or a visual representation of one or more aggregated fields of view. For example, the visibility confidence model 206 may include a top-down view or a representation of a bird's-eye view (BEV) of the aggregated field of view indicated by the sensor data 202. However, in some embodiments, the perspective may be different.
[0066] In some embodiments, the visibility confidence model 206 may correspond to a single sensor modality. For example, a first visibility confidence model 206 may correspond to the total field of view associated with an image sensor corresponding to the ego machine 208. Continuing with this example, a second visibility confidence model 206 may correspond to the total field of view associated with a LiDAR sensor corresponding to the ego machine 208. In some embodiments, the visibility confidence model 206 may include multiple layers, such as, for example, a first layer indicating the respective confidence levels corresponding to image data, a second layer indicating the respective confidence levels corresponding to RADAR data, and so on.
[0067] In some embodiments, the visibility confidence model 206 may include respective confidence levels corresponding to sensor data aggregated across multiple sensor modalities. For example, the visibility confidence model may include a confidence level corresponding to the total field of view associated with sensor data that may be generated using an image sensor, RADAR sensor, SONAR sensor, LiDAR sensor, etc.
[0068] In some embodiments, a visibility confidence model 206 can be generated for each sensor of a sensor modality. Additionally, in some embodiments, the visibility confidence models 206 corresponding to the respective sensors of a sensor modality can be aggregated to include respective confidence levels and / or visibility levels aggregated across all sensors and / or all sensor data 202 corresponding to a particular sensor modality.
[0069] In some embodiments, the visibility confidence model 206 can be organized and / or generated as a queryable data structure. In some embodiments, the ego machine 208 and / or one or more systems, subsystems, etc. that can correspond to the ego machine 208 can be configured to generate one or more queries to determine whether and to what extent sensor data corresponding to a particular location can be weighted for control determination. In some embodiments, the visibility confidence model 206 can provide a data structure that can quickly compare the confidence and / or visibility data corresponding to a first sensor modality at a first location in the environment with the confidence and / or visibility data corresponding to a second sensor modality at the first location in the environment.
[0070] In some embodiments, the data structure can be, for example, one or more radial distance maps (RDMs) corresponding to respective sensors, one RDM including data and / or information corresponding to all sensors in the same sensor modality, a grid-like data structure corresponding to, for example, the ground plane of the environment (where each cell in the grid includes a confidence), and / or visibility data that can be associated with that location on the ground plane. In some embodiments, the visibility confidence model 206 can include a data structure that can include a combination of the foregoing. In these or other embodiments, the visibility confidence model 206 and the corresponding data structure can be further described and / or illustrated in the present disclosure, such as, for example, with respect to Figure 3A 、 Figure 3B and / or Figure 3C as described and / or illustrated.
[0071] In some embodiments, the visual representation and / or data structure can be subdivided into one or more sub-parts of an aggregated field of view. In some embodiments, the one or more sub-parts can represent respective portions of the aggregated field of view, i.e., portions of the volume defined by the aggregated field of view. For example, the aggregated field of view can be represented by sensor data 202 corresponding to a body associated with the ego machine 208. Continuing with this example, the aggregated field of view can be subdivided into one or more sub-parts, and in some embodiments, the aggregated field of view can be represented by the visibility confidence model 206, which indicates the respective confidence of the sensor data 202 corresponding to the respective sub-parts of the aggregated field of view.
[0072] In some embodiments, the visibility confidence model 206 may indicate the confidence levels of the sensor data 202 corresponding to respective sub - parts of the aggregated field of view. In some embodiments, a low confidence level for a particular sub - part of the aggregated field of view may indicate that the ego - machine 208 may not rely on the data corresponding to the particular sub - part. Additionally or alternatively, a low confidence level associated with a particular sensor modality may indicate that in the case where the ego - machine 208 is receiving conflicting data regarding a particular sub - part of the aggregated field of view, the ego - machine 208 should rely on the data corresponding to one or more other sensor modalities. In some embodiments, the visibility confidence model 206 may include a data structure that can be queried by one or more other systems, subsystems, processing units, etc.
[0073] For example, the visibility system 204 may encounter conflicting data corresponding to a particular portion of the aggregated field of view associated with the ego - machine; such as, for example, RADAR data that may indicate the presence of an object and image data that does not indicate the presence of the same object. In response, the visibility subsystem 204 may query the visibility confidence model 206 to determine, for example, the weights assigned to the image data versus the RADAR data corresponding to a particular environment. In response to the query, the visibility system 204 may determine that the image data may not be as reliable as the RADAR data corresponding to that portion of the aggregated field of view and may thus rely on the RADAR data to perform one or more operations in response to the presence of the object.
[0074] In some embodiments, the visibility system 204 may be configured to determine one or more confidence levels based on: (1) one or more faults and / or errors associated with the individual sensors corresponding to the sensor data 202; (2) one or more coarse - level obstructions that may affect the sensor data 202 corresponding to the individual sensors; (3) portions of the sensor data 202 that include degraded sensor data; (4) the presence of one or more occlusions in the sensor data 202; (5) whether the visibility distance of the sensor (e.g., the range or distance from the sensor at which the sensor can be relied upon under normal operating conditions) is impaired or reduced; and / or (6) one or more other evaluations.
[0075] In some embodiments, it may be determined whether one or more faults or errors exist and / or are associated with an individual sensor. The one or more faults or errors may refer to one or more problems corresponding to sensor data 202 collection, transmission, and / or other errors, which may indicate that the sensor data 202 corresponding to the individual sensor is untrustworthy or otherwise unavailable. Additionally or alternatively, the one or more errors or faults may indicate that the individual sensor may be unhealthy or not functioning properly. For example, the sensor may no longer be electrically connected to the ego machine 208, or the connection may be unstable. In some embodiments, the visibility system 204 may be configured to obtain data indicating that the sensor may not be electrically connected to the ego machine 208.
[0076] In some embodiments, in response to determining that one or more errors are associated with an individual sensor, the visibility system 204 may be configured to downgrade or decrease the visibility level associated with the sensor data corresponding to the field of view of the individual sensor. For example, the aggregated field of view corresponding to the ego machine associated with an image sensor may correspond to a particular space, region, or volume around the ego machine. Continuing with this example, the individual image sensor may be configured to generate image data corresponding to a particular sub - portion of the space or volume extending from the rear of the ego machine. Additionally, it may be determined that the individual sensor is not operating or at least not operating properly. Thus, the visibility system 204 may reduce the confidence level of the visibility of the image data corresponding to the particular sub - portion of the space or volume extending from the rear of the ego machine.
[0077] In some embodiments, determining that a sensor is no longer operating properly may end the operations that the visibility system 204 may perform using the individual sensor and / or the sensor data 202 corresponding to the individual sensor. Conversely, in response to not finding, not obtaining, or otherwise not determining one or more faults and / or errors corresponding to an individual sensor, the visibility system 204 may continue to perform one or more operations to determine one or more confidence levels corresponding to the sensor data 202, such as, for example, determining whether one or more coarse - level occlusions may affect the visibility corresponding to the individual sensor.
[0078] In some embodiments, the visibility system 204 may be configured to use data other than the sensor data 202 corresponding to individual sensors to determine whether there is one or more coarse-level obstructions. A coarse-level obstruction may refer to a condition that can affect the sensor data 202 corresponding to an individual sensor. For example, environmental conditions (such as heavy rain, snow, hail, freezing temperatures, extreme heat, dust storms, etc.) may be coarse-level obstructions corresponding to the sensor data 202 associated with an individual sensor. Other examples of coarse-level obstructions may include time of day, time of year, etc. For example, the visibility system 204 may determine that a coarse-level obstruction can be assumed and / or applied to the image data corresponding to an individual image sensor at night. In some embodiments, one or more other sensors (e.g., temperature sensors, humidity sensors, etc.), data corresponding to one or more other systems (e.g., cameras, other machines, other systems, etc.), and other data corresponding to the ego machine 208 and / or the environment in which the ego machine 208 may be located may be used, for example, to determine the determination that there may be one or more coarse-level obstructions in the sensor data 202.
[0079] In some embodiments, in response to determining that a coarse-level obstruction may correspond to the sensor data 202 associated with an individual sensor, the visibility system 204 may be configured to degrade or reduce the visibility level and / or confidence level associated with the sensor data 202 corresponding to the field of view of the individual sensor. In some embodiments, the amount by which the confidence level may be reduced may depend on the type of coarse-level obstruction and / or the degree of sensor data degradation based on the coarse-level obstruction. For example, light rain may result in a smaller reduction in the visibility level corresponding to the sensor data 202 compared to heavy rain.
[0080] In some embodiments, the visibility system 204 can be configured to perform one or more operations using sensor data corresponding to an individual sensor to determine whether there is fine - level occlusion and / or degradation in the sensor data. Fine - level occlusion and / or degradation can include portions of the sensor data 202 that can be occluded, blurred, or otherwise impaired. Coarse - level occlusion can apply to all or substantially all (e.g., 90%) of the sensor data 202 corresponding to an individual sensor. In contrast, more fine - level occlusion and / or degradation can include one or more portions of the sensor data 202 that can be occluded or unreliable. In some embodiments, one or more techniques can be used to determine whether certain portions of the sensor data are degraded, occluded, blurred, or otherwise obscured. For example, in the context where the sensor data 202 is image data, one or more frequency analysis, gradient analysis, object detection and segmentation analysis, and / or other analysis methods can be used to determine whether certain portions of the sensor data are degraded or occluded. As an additional example, in the context of LiDAR data, one or more point cloud analysis, range and intensity analysis, comparative analysis, etc., can be performed to determine whether there is one or more degradation in the LiDAR data and where it exists.
[0081] In some embodiments, in response to determining that there is fine - level occlusion and / or degradation in the sensor data 202 corresponding to an individual sensor, it can be determined whether and / or where the fine - level degradation affects a sub - portion of the aggregated field of view corresponding to the visibility confidence model 206. In some embodiments, one or more other data structures and / or techniques can be used to determine whether the fine - level occlusion and / or degradation intersects with a sub - portion of the aggregated field of view represented by the visibility confidence model 206. Techniques for determining whether one or more fine - level degradations can be included in the sensor data 202 and to what extent can be found in this disclosure such as, for example, with respect to Figure 3AFurther description and / or illustration. Additionally, by way of non-limiting example, one or more techniques for determining blindness, occlusion, blur, visible distance, etc. corresponding to sensor data corresponding to respective sensors may be included in the following documents: U.S. Patent No. 11,508,049, titled "DEEPNEURAL NETWORK PROCESSING FOR SENSOR BLINDNESSDETECTION IN AUTONOMOUS MACHINEAPPLICATIONS", filed on September 13, 2019, and / or U.S. Patent Publication No. US2023 / 0110027, titled "VISIBILITY DISTANCE ESTIMATIONUSING DEEP LEARNING IN AUTONOMOUS MACHINE APPLICATIONS", filed on September 29, 2021, the contents of which are incorporated herein by reference in their entirety.
[0082] In addition to determining whether there is a fine - level occlusion and / or degradation in the sensor data corresponding to a particular sensor, the visibility system 204 may also determine whether and / or where one or more occlusions may be indicated in the sensor data 202. For example, one or more static or dynamic obstacles may occlude one or more other obstacles, objects, or regions such that an individual sensor cannot view and / or detect one or more other obstacles or objects in the field of view. In some embodiments, to determine whether one or more occlusions may be indicated in the sensor data 202, map data corresponding to an environmental map may be used. The map data may include data corresponding to static obstacles and / or objects that may be present within or outside the field of view of an individual sensor. Additionally or alternatively, historical sensor data may be used to determine whether one or more objects in the sensor data corresponding to one or more previous timestamps may cast shadows or otherwise occlude the sensor data 202 corresponding to one or more current timestamps. Techniques for determining whether and to what extent one or more occlusions may be included in the sensor data 202 may be further described and / or illustrated in this disclosure, such as, for example, with respect to Figure 3A Further description and / or illustration.
[0083] In some embodiments, a confidence level corresponding to sensor data 202 associated with each sensor of a corresponding sensor modality may be determined for one or more sensors corresponding to the sensor modality. Additionally, in some embodiments, the confidence levels associated with corresponding sensor data 202 corresponding to corresponding sensors may be aggregated on a visibility confidence model 206. In some embodiments, one or more sub-portions of an aggregated field of view corresponding to the total confidence model 206 may be populated with associated aggregated confidence and / or visibility data and / or levels.
[0084] In some embodiments, the ego machine 208 may perform one or more operations based on the visibility confidence model 206. For example, the ego machine 208 may be configured to generate one or more control commands that may direct the ego machine 208 to perform one or more operations. In some embodiments, the ego machine 208 may plan one or more operations based on sensor data 202 and / or data, values, and / or other information that may be included in one or more visibility confidence models 206. The one or more operations may include decelerating, accelerating, turning, changing lanes, performing one or more avoidance strategies, and the like.
[0085] In some embodiments, the ego machine 208 may be configured to generate one or more queries to obtain information from the visibility confidence model 206. The one or more queries may include location information, sensor information, sensor modality information, and / or other information that may identify one or more sub-portions of the visibility confidence model 206. For example, the ego machine 208 may generate a query that seeks confidence and / or visibility information corresponding to image data associated with a particular location. The generated query may allow the ego machine 208 to locate and / or receive confidence data and / or information corresponding to a particular location, particular sensor, particular sensor data, and the like. In some embodiments, using the confidence information, the ego machine 208 may be configured to make one or more control determinations and / or perform one or more operations.
[0086] In some embodiments, the ego machine 208 may perform one or more operations based on confidence values associated with and / or assigned to sensor data 202 corresponding to one or more sub - portions of the aggregated field of view of the visibility confidence model 206. In some embodiments, for example, the ego machine 208 may be configured to determine weights that can be given to and / or assigned to sensor data 202 (or portions or subsets thereof) corresponding to a particular sub - portion of the visibility confidence model 206 based on these confidence values and / or data associated therewith. A given weight may be used to calculate and / or determine the degree of reliance that the ego machine 208 places on sensor data 202 when generating control commands and / or determining one or more operations to perform.
[0087] In some embodiments, one or more operations may be performed based on a visibility confidence model 206 that may be associated with one sensor modality. For example, the ego machine 208 may include an image sensor that may generate image data. Continuing with this example, one or more sub - portions of the visibility confidence model 206 may indicate that the image data corresponding to a particular sub - portion may have a lower confidence level associated therewith compared to one or more other sub - portions of the visibility confidence model 206. Continuing with this example, the ego machine 208 may be configured to discount reliance on the image data - for example, assign a lower weight to the image data corresponding to a particular sub - portion of the visibility confidence model 206. Further continuing with this example, the ego machine 208 may thus rely more on map data corresponding to an environmental map, plan data that may have been generated, historical sensor data that may have been generated and relied upon at one or more previous timestamps. In the case where the ego machine 208 has no other source of information to rely on, the ego machine may perform one or more avoidance strategies and / or perform operations to safely stop the ego machine 208.
[0088] In some embodiments, one or more operations may be performed based on multiple visibility confidence models 206 associated with multiple respective sensor modalities. In some embodiments, the ego machine 208 may be configured to assign weights to respective subsets of sensor data 202 corresponding to the respective visibility confidence models 206 and rely on sensor data 202 corresponding to elevated confidence levels and / or data indicating elevated confidence levels.
[0089] For example, in the context of a self - vehicle that includes an image sensor and a RADAR sensor that respectively generate image data and RADAR data, two visibility confidence models 206 can be generated. A first visibility confidence model 206 corresponds to the image data, and a second visibility confidence model 206 corresponds to the RADAR data. Continuing with this example, the first visibility confidence model 206 can indicate that the confidence level corresponding to the image data associated with a particular sub - part of the visibility confidence model 206 can be lower than a predetermined threshold (e.g., due to rain, fog, image sensor error, occlusion, etc.). Further continuing with this example, the second visibility confidence model 206 can indicate that the RADAR data corresponding to the same sub - part can have an elevated or increased confidence level compared to the image data. In the case where there is a difference between the image data and the RADAR data, the self - vehicle 208 can rely on the RADAR data rather than the image data to perform one or more operations. For example, if an object is detected in the RADAR data but not in the image data, the self - vehicle 208 can be configured to rely on the RADAR data to decelerate, turn, or perform one or more avoidance strategies to avoid the object.
[0090] Modifications, additions, or omissions can be made Figure 2 without departing from the scope of the present disclosure. For example, the number of self - vehicles 208, the number of sensors that can generate and / or collect sensor data 202 can vary, and the number of visibility confidence models 206 can vary. Additionally, the visibility system 204 can include multiple machine - learning models, neural networks, perception systems, subsystems, etc. Additionally or alternatively, the visibility system 204 can be included in one or more other systems and / or machines in addition to the self - vehicle 208. The details given and discussed are for the purpose of helping to provide an explanation and understanding of the concepts of the present disclosure and are not meant to be limiting.
[0091] Figure 3A An example environment 300 is shown, which shows a visibility system 304 that generates one or more visibility confidence models 306 according to one or more embodiments of the present disclosure. In some embodiments, the environment 300 can be an example of the environment 200 such as further described and / or illustrated in the present disclosure and / or similar to the environment 200. In some embodiments, the visibility system 304 can be the same as and / or similar to the visibility system 204 such as further described and / or illustrated in the present disclosure. In some embodiments, the visibility system 304 can be configured to use sensor data 302 to generate one or more visibility confidence models 306. Figure 2 For example, the environment 200. In some embodiments, the visibility system 304 can be the same as and / or similar to the visibility system 204 such as further described and / or illustrated in the present disclosure. In some embodiments, the visibility system 304 can be configured to use sensor data 302 to generate one or more visibility confidence models 306. Figure 2 For example, the visibility system 204. In some embodiments, the visibility system 304 can be configured to use sensor data 302 to generate one or more visibility confidence models 306.
[0092] In these or other embodiments, the sensor data 302 may be the same as and / or similar to the sensor data 202 described and / or illustrated herein, such as, for example, with respect to Figure 2 Additionally or alternatively, the sensor data 302 may be an example of sensor data generated using one or more sensors that may be associated with one or more self machines (e.g., sensor 104) described and / or illustrated herein, such as, for example, with respect to Figure 1 further.
[0093] In some embodiments, the visibility system 304 may include one or more systems, subsystems, machine learning models, neural networks, large language models (LLMs), DNNs, convolutional neural networks (CNNs), and / or other algorithms that may be configured to determine the visibility of the environment using the sensor data 302. For example, the visibility system 304 may include a neural network 318 that may represent one or more machine learning models and / or neural networks that may be configured to process the sensor data 302 and / or generate one or more visibility confidence models 306.
[0094] In some embodiments, the neural network 318 may include any type of machine learning model, such as a machine learning model using linear regression, logistic regression, decision trees, support vector machines (SVMs), naive Bayes, k-nearest neighbors (Knn), K-means clustering, random forests, dimensionality reduction algorithms, gradient boosting algorithms, neural networks (e.g., autoencoders, convolutional, recurrent, perceptron, long / short-term memory / LSTM, Hopfield, Boltzmann, deep belief, deconvolutional, generative adversarial, liquid machines, transformers, conformers, LLMs, etc.), computer vision algorithms, and / or other types of machine learning models.
[0095] As an example, for instance, when the neural network 318 includes a CNN, the neural network 318 may include any number of layers. For example, one or more layers may include an input layer. The input layer may hold values associated with the sensor data 302. For example, when the sensor data 302 represents an image, the input layer may hold values representing the raw pixel values of the image as a volume (e.g., width, height, and color channels (e.g., RGB), e.g., 32x32x3).
[0096] Additionally or alternatively, one or more layers included in neural network 318 may include a convolutional layer. A convolutional layer may compute the output of neurons connected to local regions in the input layer, where each neuron computes the dot product of its weights with the small region in the input volume to which it is connected. In some embodiments, the result of the convolutional layer may be another volume, where one dimension is based on the number of filters applied (e.g., width, height, and number of filters, e.g., 32x32x12 if the number of filters is 12).
[0097] In some embodiments, one or more layers may include a deconvolutional layer (or transposed convolutional layer). For example, the result of a deconvolutional layer may be another volume with dimensions higher than the input dimensions of the data received at the deconvolutional layer.
[0098] In some embodiments, one or more layers may include a rectified linear unit (ReLU) layer. A ReLU layer may apply an element-wise activation function, such as max(0,x), e.g., thresholding at zero. The resulting volume of the ReLU layer may be the same as the volume of the input to the ReLU layer.
[0099] Additionally or alternatively, one or more layers may include a pooling layer. A pooling layer may perform a downsampling operation along spatial dimensions (e.g., height and width), which may result in a smaller volume than the input to the pooling layer (e.g., from a 32x32x12 input volume to a 16x16x12).
[0100] Additionally or alternatively, one or more layers may include one or more fully connected layers. Each neuron in a fully connected layer may be connected to every neuron in the previous volume. A fully connected layer may compute class scores, and the resulting volume may be 1x1x n, where n equals the number of classes. In some examples, the CNN may include a fully connected layer such that the output of one or more layers of the CNN may be provided as input to the fully connected layer of the CNN. In some examples, one or more convolutional streams may be implemented by neural network 318, and some or all of the convolutional streams may include corresponding fully connected layers.
[0101] In some embodiments, neural network 318 and / or the corresponding neural network may include a series of convolutional and max pooling layers for facilitating image feature extraction, followed by multi-scale dilated convolutional and upsampling layers for facilitating global context feature extraction.
[0102] Although an input layer, convolutional layer, pooling layer, ReLU layer, and fully connected layer are discussed herein with respect to neural network 318, this is not intended to be limiting. For example, additional or alternative layers may be used in neural network 318 and / or the corresponding neural network, such as normalization layers, SoftMax layers, and / or other layer types.
[0103] In addition, some layers may include parameters (e.g., weights and / or biases), such as convolutional layers and fully connected layers, while other layers may not include parameters, such as ReLU layers and pooling layers. In some examples, the parameters may be learned by the neural network 318 during training. Additionally, some layers may include additional hyperparameters (e.g., learning rate, stride, epoch, etc.), such as convolutional layers, fully connected layers, and pooling layers, while other layers may not include additional hyperparameters, such as ReLU layers. In an embodiment where the neural network 318 regresses the visibility distance, the activation function of the last layer of the CNN may include a ReLU activation function.
[0104] In an embodiment where the neural network 318 includes a CNN, different orders and numbers of CNN layers may be used according to the embodiment. In other words, the order and number of the layers of the CNN are not limited to any one architecture.
[0105] For example, in one or more embodiments, the CNN may include an encoder-decoder architecture and / or may include one or more output heads. For example, the CNN may include one or more layers corresponding to the feature detection trunk of the CNN, and one or more output heads may be used to process the output (e.g., feature map) of the feature detection trunk. For example, a first output head (including one or more first layers) may be used to calculate the health status of each sensor, a second output head (including one or more second layers) may be used to calculate and / or determine one or more coarse-level degradations corresponding to the sensor data 302, a third output head (including one or more third layers) may be used to calculate and / or determine one or more fine-level degradations, and / or a fourth output head (including one or more fourth layers) may be used to calculate and / or determine the presence or absence of one or more occlusions corresponding to the sensor data 302. Thus, in the case of using two or more heads, the two or more heads may process the data from the trunk in parallel, and each head may be trained to accurately predict the corresponding output of the output head. However, in other embodiments, a single trunk may be used without separate heads. Although an architecture has been described above for the neural network 318, this architecture is merely exemplary, and there may be other architectures available for performing one or more of the operations corresponding to the neural network 318.
[0106] In some embodiments, the visibility system 304 and / or the corresponding neural network 318 may include one or more heads that may perform one or more operations corresponding to the descriptions of the sensor health module 308, the coarse-level degradation module 310, the fine-level degradation module 312, and / or the occlusion module 314. Additionally or alternatively, the above modules may be included in one or more separate machine learning models, neural networks, systems, subsystems, etc., where one or more outputs corresponding to one or more of the sensor health module 308, the coarse-level degradation module 310, the fine-level degradation module 312, and / or the occlusion module 314 may be used, for example, by the visibility system 304 and / or the neural network 318 to generate one or more visibility confidence models 306.
[0107] In some embodiments, the sensor health module 308, the coarse-level degradation module 310, the fine-level degradation module 312, and / or the occlusion module 314 (collectively referred to as "modules") may represent one or more heads and / or parts of the neural network 318. Additionally or alternatively, the modules may operate independently and / or in conjunction with the neural network 318. For example, the modules may include code and routines configured to allow a computing system to perform one or more operations. Additionally or alternatively, one or more of the modules may be implemented using hardware, including one or more processors, CPUs, graphics processing units (GPUs), data processing units (DPUs), parallel processing units (PPUs), microprocessors (e.g., for performing one or more operations or controlling the execution of one or more operations), field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), accelerators (e.g., deep learning accelerators (DLAs)), and / or other processor types. In these and other embodiments, one or more of the modules may be implemented using a combination of hardware and software. In the present disclosure, the operations described as being performed by the corresponding modules may include operations that one or more of the modules may direct the corresponding computing system to perform. In these or other embodiments, one or more of the modules may be implemented by one or more computing devices, such as the computing devices described in further detail with respect to Figures 5A - 5D , Figure 6 and / or Figure 7 the computing devices described further below.
[0108] In some embodiments, the visibility system 304 and / or the neural network 318 may start with a data structure or a visibility confidence model 306, where the data structure and / or the visibility confidence model 306 includes data indicating that each part of the visibility confidence model 306 is visible. In some embodiments, the fact that each part is visible may indicate that the sensor data 302 corresponding to the aggregated field of view associated with all sensors of the ego machine is not blocked, blurred, obstructed, or otherwise impaired. In some embodiments, each module corresponding to the visibility system 304 and / or the neural network 318 may be configured to determine whether the unblocked sensor data 302 can be used to define the respective parts of the visibility confidence model 306. In some embodiments, it may be determined (e.g., using one or more modules corresponding to the visibility system 304) that one or more parts of the visibility confidence model 306 may include sensor data 302 that may be blurred, blocked, or otherwise obstructed. In embodiments where the sensor data 302 may be blurred, blocked, or otherwise obstructed, the visibility system 304, the neural network 318, and / or the corresponding modules thereof may determine one or more visibility levels, calculate one or more visibility levels, and / or assign one or more visibility levels to the respective parts of the visibility confidence model 306.
[0109] For example, the visibility confidence model 306 may include multiple sub-parts that may be associated with the aggregated field of view corresponding to multiple image sensors of the ego machine. In some embodiments, the visibility confidence model 306 may start, for example, with values assigned to each sub-part indicating that the sub-part is visible. Continuing with this example, the visibility system 304, the neural network 318, and / or the corresponding module may determine whether the sensor data 302 corresponding to the sub-part is blocked, obstructed, blurred, etc. In some embodiments, after determining that the sensor data 302 corresponding to one or more sub-parts of the visibility confidence model 306 is blurred or blocked, the perception system 304, the neural network 318, and / or one or more corresponding modules may be configured to add data and / or information indicating the confidence and / or visibility of the respective sub-parts to the visibility confidence model 306. In some embodiments, the visibility confidence model 306 may be a queryable data structure such that the ego machine may be configured to make one or more control determinations based on the visibility confidence model 306.
[0110] In some embodiments, the sensor health module 308 may be configured to determine whether each sensor from which sensor data 302 can be received is healthy and / or operating properly. In some embodiments, a reference to a sensor being in poor health may indicate, for example, that one or more errors may be associated with a particular sensor, that an individual sensor may not be electrically connected to send sensor data, that each sensor may be collecting data including one or more errors, and / or otherwise indicate that each sensor is not operating properly.
[0111] In some embodiments, the sensor health module 308 may determine that a sensor is not operating properly based on a lack of data received from an expected source. For example, the sensor health module 308 may have received data from a first sensor at one or more previous timestamps. Continuing with this example, for a current timestamp, the sensor data 302 may not include data from the first sensor. The sensor health module 308 may be configured to determine that the first sensor is thus not operating properly, may be unhealthy, or may have encountered one or more errors.
[0112] In some embodiments, the sensor health module 308 may be configured to determine sensor health based on one or more anomalies in the sensor data 302. For example, the sensor data 302 may include one or more error codes that may indicate that one or more errors may have been encountered in generating data using a particular sensor. As an additional example, the sensor health module 308 may be configured to determine that the sensor data 302 corresponding to a first sensor is different from the sensor data 302 previously received from the first sensor. In response to the sensor data 302 from the first sensor being different from the previous data received from the first sensor, the sensor health module 308 may be configured to determine that the first sensor may include one or more errors and / or may not be operating properly. In some embodiments, based on the errors encountered and transmitted using a particular sensor, the sensor health module 308 may be configured to determine whether the particular sensor is operating properly.
[0113] In some embodiments, the sensor health module 308 may be configured to determine whether each sensor is healthy and / or operating properly. In some embodiments, the sensor health module 308 may be configured to iterate through sensor data 302 corresponding to each sensor associated with a particular sensor modality to determine whether each of the individual sensors is healthy and / or operating properly. For example, the ego machine may include ten (10) image sensors that may be used to generate image data corresponding to the environment around the ego machine. Continuing with this example, the sensor health module 308 may be configured to use sensor data 302 corresponding to each of the ten (10) individual image sensors to determine whether each of the ten (10) image sensors is operating properly.
[0114] In some embodiments, in response to the sensor health model 308 determining that an individual sensor is not operating properly, the sensor data 302 corresponding to the field of view associated with the particular sensor may be degraded and / or determined to be invisible to the ego machine. In some embodiments, one and / or more portions of the visibility confidence model 306 may be downgraded to visibility zero (0) or its equivalent value in the visibility confidence model 306.
[0115] In some embodiments, the data structure corresponding to the visibility confidence model 306 may include data and / or information indicating that the portion of the aggregated field of view associated with the individual sensor that is not operating properly is invisible. Additionally or alternatively, the data and / or information may indicate that the portion of the aggregated field of view may include a reduced confidence value, where the confidence value may indicate the amount of weight or confidence that may be assigned to the sensor data 302 corresponding to the portion of the aggregated field of view. In some embodiments, the reduced confidence value and / or data may correspondingly change one or more visual representations of the visibility confidence model 306. For example, the visibility confidence model 306 may depict the environment of the aggregated field of view from top to bottom. In some embodiments, the portion of the aggregated field of view that may correspond to corrupted sensor data may be indicated by a color difference or some other visual indicator indicating a poor or reduced confidence level that a portion of the aggregated field of view is invisible.
[0116] For example, the image data corresponding to a particular image sensor can correspond to a part (e.g., one or more sub-parts) of the visibility confidence model 306. Continuing with this example, the image data can include data and / or information that can indicate that the image sensor may have encountered one or more errors, such that the image data corresponding to the image sensor may be corrupted, damaged, or otherwise untrustworthy. Continuing with this example, as a result, the sensor health module 308 can be configured to degrade or reduce the confidence value or confidence data corresponding to this part of the visibility confidence model 306. Continuing with this example, the sensor health module 308 can be configured to change the data and / or information that may be included in the visibility confidence model 306 to reflect the degradation or reduction in the confidence of the sensor data 302. In some embodiments, in response to the sensor health module 308 determining that a particular sensor may not be operating properly, the sensor data 302 corresponding to that particular sensor may not be transmitted and / or further conveyed to, for example, the coarse-level degradation module 310.
[0117] In some embodiments, in response to the sensor health module 308 determining that an individual sensor may be operating properly, the sensor health module 308 may not change or determine one or more changes to the visibility and / or confidence values corresponding to the visibility confidence model 306 associated with that individual sensor. For example, the visibility confidence model 306 can start with values, data, and / or information corresponding to one or more sub-parts of the aggregated field of view, where the values, data, and / or information can indicate that the sensor data 302 corresponding to each sub-part is trustworthy and / or the environment corresponding to the aggregated field of view is visible. Continuing with this example, in response to the sensor health module 308 determining that the individual sensor and / or the sensor data corresponding to the individual sensor does not include one or more errors, the sensor health module 308 may not change any data and / or information corresponding to the visibility confidence model 306. In this example, the visibility confidence model 306 may have started with the assumption that the field of view is visible, and thus, the sensor health module 308 may not change the data corresponding to the visibility confidence model 306. In some embodiments, the sensor health module 308 can be configured to transmit and / or otherwise convey the sensor data 302 corresponding to the respective field of view to one or more other modules, such as, for example, the coarse-level degradation module 310.
[0118] The coarse - level degradation module 310 can be configured to determine whether one or more coarse - level degradations are included in the sensor data 302 corresponding to a particular sensor and a corresponding field of view. In some embodiments, one or more coarse - level degradations can include blocking, blurring, and / or other degradation factors that can affect all of the sensor data 302 corresponding to a particular sensor and / or multiple sensors. For example, weather can be a factor that affects all or substantially all of the sensor data corresponding to a particular sensor. For example, freezing temperatures can affect and / or degrade sensor quality and, correspondingly, affect and / or degrade the sensor data 302 corresponding to each sensor. As an additional example, heavy rain can affect all of the sensor data 302 corresponding to, for example, an exposed image sensor.
[0119] In some embodiments, the coarse - level degradation module 310 can be configured to determine one or more coarse - level degradations based on data that can be generated and / or collected using one or more other sensors and / or from one or more other systems other than the sensors corresponding to, for example, a self - machine. For example, one or more other sensors and / or systems can be configured to transmit weather data, temperature data, and other data that may be associated with the environment in which the self - machine may be located. For example, one or more edge servers and / or data centers can be configured to send or transmit weather data to the self - machine. Continuing with this example, the coarse - level degradation module 310 can be configured to use the transmitted weather data to determine whether the overall visibility model 306 can be updated to reflect one or more coarse - level degradations that can affect the sensor data 302 corresponding to a particular sensor.
[0120] Additionally or alternatively, the coarse - level degradation module 310 can be configured to determine whether one or more coarse - level degradations will affect the sensor data 302 corresponding to each sensor based on the sensor data 302 itself. In some embodiments, the coarse - level degradation module 310 can be configured to determine the presence of one or more coarse - level degradations in the sensor data 302 based on historical sensor data 316 corresponding to a particular sensor. For example, the sensor data 302 corresponding to a particular sensor can be considered "clear" or at least acceptable for visibility. Continuing with this example, in response to determining that the sensor data 302 has deviated from an acceptable standard established using past data corresponding to the same sensor, the coarse - level degradation module 310 can detect that the sensor data 302 may be blurred, blocked, occluded, or otherwise impaired.
[0121] In some embodiments, in response to detecting and / or determining that one or more gross degradations exist in sensor data 302 corresponding to a particular sensor, the gross degradation module 310 may be configured to degrade or reduce the visibility confidence score corresponding to a sub - portion associated with the entire field of view of the particular sensor. For example, in the context where a visibility confidence score of "1" represents visible and "0" represents completely blocked, in response to determining heavy rain corresponding to the environment, the visibility confidence score associated with the image data corresponding to a particular image sensor may be reduced from 1 to 0.5.
[0122] In some embodiments, the amount of reduction and / or degradation of the visibility confidence score corresponding to sensor data associated with a particular sensor may depend on the sensor type and / or the severity of the gross degradation. For example, in the context of heavy rain corresponding to the environment, in response to the fact that RADAR data is not affected as negatively, blurred, distorted, etc. (e.g., in adverse weather conditions) as image data, the first visibility confidence score and / or value corresponding to the image data generated using an image sensor may be less than the second visibility confidence score and / or value corresponding to the RADAR data generated using a RADAR sensor. As an additional example, the first visibility score corresponding to image data corresponding to an environment with heavy rain may be less than the second visibility score corresponding to image data in a light rain or relatively light rain environment and / or indicate a lower visibility confidence than the second visibility score. Additional examples of environmental conditions that may be included in gross degradation may include excessive sunlight, darkness (e.g., nightfall), other causes of darkness (tunnels, overpasses, etc.), excessive sound or vibration affecting one or more sensors, etc.
[0123] In some embodiments, any blockage, obstruction, blur, etc. in sensor data 302 corresponding to a particular sensor that exceeds a certain threshold may be determined as gross degradation. For example, a blockage or blur corresponding to fifty - one percent (51%) of the sensor data 302 generated by a particular sensor may be a gross - level blockage. In some embodiments, the percentage may be 10%, 20%, 30%, 75%, 80%, 90%, 100%, and any other percentage that can be determined to affect a sufficient amount of sensor data 302 corresponding to a particular sensor and / or field of view to be considered gross degradation. In some embodiments, in response to determining the existence of gross degradation and / or one or more visibility confidence scores on which the gross degradation determination depends, the gross degradation module 310 may be configured to transmit and / or convey information corresponding to the visibility confidence model 306 to one or more other modules, systems, subsystems, etc., such as, for example, the fine - level degradation module 312.
[0124] In some embodiments, the fine - level degradation module 312 may be configured to determine one or more degradations and / or portions of the sensor data 302 that may be partially blocked, blurred, degraded, occluded, etc. In some embodiments, one or more fine - level degradations may not apply to all of the sensor data 302 corresponding to a particular sensor. Instead, the fine - level degradation applies to a portion of the sensor data 302 corresponding to a particular sensor. In some embodiments, the fine - level degradation module 312 may be configured to determine one or more smaller portions of the sensor data 302 corresponding to a particular sensor that may be blocked, blurred, or otherwise obstructed as compared to the coarse - level degradation module. In some embodiments, one or more smaller portions may be predetermined as compared to the coarse - level degradation. For example, the smaller portion may include anything that does not affect all or substantially all (e.g., 95% or more) of the sensor data 302 corresponding to a particular sensor. Additionally or alternatively, the smaller portion may include anything less than 50% of the sensor data 302 corresponding to a particular sensor. In some embodiments, the amount of sensor data 302 that may be included in one or more smaller degraded portions of the sensor data 302 may include any data amount less than the coarse - level degradation (e.g., 1%, 5%, 10%, 25%, 50%, etc. of the total sensor data 302 corresponding to a particular sensor).
[0125] In some embodiments, one or more techniques may be used to determine whether certain portions of the sensor data are degraded, blocked, blurred, or otherwise occluded. For example, in the context where the sensor data is image data, one or more gradient analyses may be used to determine whether one or more portions of the image data are degraded, blurred, blocked, etc. For example, one or more gradients and / or pixel intensity changes in the x - direction and / or y - direction of the image. The x - direction and y - direction refer to the horizontal and vertical distribution of pixels in the image. In some cases, one or more gradient operators (e.g., Sobel, Sharr, and / or Pewitt operators) may be used to determine the gradients and / or pixel intensity changes. Continuing with this example, one or more gradient magnitudes (e.g., for the x - direction and y - direction) corresponding to individual pixels and / or pixel sets may be determined. In certain cases, a threshold may be determined, where a gradient magnitude below the threshold may be considered a blurred region of the image, while a gradient magnitude above the threshold may be a relatively clear portion of the image. In certain cases, the threshold may be determined based on one or more heuristic analyses and / or a loss function corresponding to, for example, a neural network (e.g., neural network 318).
[0126] Additionally or alternatively, one or more other analyses can be used to determine whether one or more portions of the image, for example, include image data that may be blurred, blocked, obstructed, etc. For example, the fine - level degradation module 312 can use one or more frequency analyses, Laplacian edge detection algorithms, contrast analyses, image sharpness metrics, object detection and segmentation analyses, and / or other techniques, algorithms, analyses, etc., which can be used to determine whether portions of the sensor data are degraded, blocked, blurred, etc.
[0127] As an additional example, in the context of LiDAR data, one or more point - cloud analyses can be used to determine whether and / or where one or more degradations may be included in the LiDAR data. For example, a LiDAR sensor can generate LiDAR data in the form of a point cloud, which can correspond to an environment associated with the field of view of the LiDAR sensor. Different densities corresponding to one or more portions and / or regions of the point cloud can be determined and / or calculated. In some cases, large fluctuations, such as one or more low - density portions of the point cloud, can indicate one or more blurred, blocked, or occluded regions of the sensor data (e.g., sensor data 302).
[0128] Additionally or alternatively, one or more other analyses can be used to determine one or more portions of the LiDAR data that may be blurred, blocked, obstructed, etc. For example, range and intensity analyses, comparison analyses, and / or other techniques, algorithms, analyses, etc., can be used to determine whether portions of the LiDAR data are degraded, blocked, blurred, etc. Although examples have been given with respect to image data and LiDAR data, these examples are not meant to be limiting. One or more other analyses or techniques can be used to determine whether one or more degradations exist in the sensor data 302.
[0129] In some embodiments, in response to determining that there may be fine-grained occlusions and / or degradations in the sensor data 302 corresponding to an individual sensor, it can be determined whether and / or where the fine-grained degradation affects a sub-part of the aggregated field of view corresponding to the visibility confidence model 306. In some embodiments, a projection of the sensor data 302 corresponding to the individual sensor can be determined in order to determine the location of the sensor data 302 relative to the overall system degradation. For example, in the context of an autonomous vehicle including multiple image sensors, it can be determined that one or more portions of the image data corresponding to an individual image sensor may be degraded. Continuing with this example, the degradation in the sensor data 302 corresponding to the individual image sensor can be projected onto a "rig-frame" or one or more other virtual environments in order to populate the visibility confidence model 306 with data and / or information associated with the visibility confidence of the sensor data 302 corresponding to the individual sensor.
[0130] In some embodiments, in order to project one or more degradations from the sensor data 302 corresponding to an individual sensor onto a system reference frame that may include the sensor data 302 corresponding to the aggregated field of view, one or more radial distance maps (RDMs) can be generated to determine where one or more degradations in the sensor data 302 may affect the visibility confidence model 306. In some embodiments, the fine-grained degradation module 312 can be configured to project one or more rays from the degraded sensor data 302 onto a ground plane corresponding to the ground. For example, in the context of image data corresponding to an image, the fine-grained degradation module 312 can determine that a portion of the image (e.g., a set of pixels) may be degraded, blurred, occluded, etc. Continuing with this example, the fine-grained degradation module 312 can be configured to project a line or ray from one or more pixels onto the ground plane of the environment that can be depicted using the image data.
[0131] In some embodiments, the fine-grained degradation module 312 can additionally be configured to project one or more virtual objects into the same environment as the rays or lines from the sensor data 302 (e.g., the image data from the above example) that can be projected. In some embodiments, by projecting one or more virtual objects into the environment, it can be determined whether the rays or lines corresponding to the degraded sensor data 302 will affect the perception of the virtual objects and to what extent the degraded sensor data 302 can affect the perception of the virtual objects.
[0132] In some embodiments, the virtual object can be projected as if the virtual object were stationary on the ground plane to determine where and / or if the projected rays corresponding to the degraded sensor data 302 intersect one or more virtual objects; for example, using the projected rays and corresponding azimuth angles. For example, in the context of an autonomous vehicle generating image data corresponding to an image, one or more virtual vehicle-sized objects can be projected onto the ground plane of the environment associated with the sensor data 302. Continuing with this example, one or more rays from the degraded pixels in the image can be projected to determine if the projected rays can intersect one or more virtual vehicle-sized objects. In response to one or more rays intersecting one or more virtual vehicle-sized objects, the confidence level associated with a portion of the visibility confidence model 306 corresponding to the virtual vehicle-sized object can be degraded or otherwise reduced.
[0133] In some embodiments, by generating one or more RDMs corresponding to the sensor data 302 corresponding to one or more individual sensors, the sensor data 302 corresponding to the individual sensors can be projected into the overall system frame such that the respective sub-parts of the visibility confidence model 306 can be populated with visibility confidence data associated with the corresponding sensors of the system.
[0134] Examples of projecting one or more rays from image data into a corresponding 3D environment to determine one or more confidence levels can be described with respect to Figure 3B and Figure 3C FIG. Figure 3B FIG. 320 depicts an example environment depicted by an image 320 subdivided into discrete sub-parts 322, according to one or more embodiments of the present disclosure. In some embodiments, the image 320 can be depicted using image data generated and / or collected using one or more image sensors associated with the ego machine. In some embodiments, the image 320 can include portions corresponding to parts of the ego machine, e.g., in the portion of the image closest to the bottom of the image 320, or, for reference, in the portion of the image closest to the first sub-part 322a. In some embodiments, one or more sub-parts 322 can represent respective portions of the image 320 that represent respective portions of the environment, where the sensor data can be evaluated for confidence in visibility corresponding to the one or more sub-parts 322.
[0135] In some embodiments, the image 320 may include a first sub - part 322a depicting a first part of the environment, a second sub - part 322b depicting a second part of the environment, a third sub - part 322c depicting a third part of the environment, a fourth sub - part 322d depicting a fourth part of the environment, up to and including an nth sub - part 322n. In some embodiments, the number of sub - parts 322 may depend on the size of the environment, the amount of sensor data included in the image 320, the amount of computing and / or processing power available to the system that can evaluate the visibility associated with each sub - part 322, and so on.
[0136] As Figure 3B shown, the nth sub - part 322n may represent one or more sub - parts of the image 320, where the image data corresponding to the nth sub - part may be visible. In Figure 3B , for example, the top two rows and the bottom two rows of sub - parts of the image 320 may be represented by the nth sub - part 322n. In some embodiments, the sub - parts represented by the sub - part 322n may also indicate one or more parts of the image 320 that may be irrelevant to the visibility confidence model (e.g., visibility confidence model 306 and / or visibility confidence model 206). For example, in the context of a self - vehicle, as Figure 3B shown, the bottom two rows may include image data corresponding to the vehicle's hood, while the top two rows of sub - parts of the image 320 may point to the sky, which may be irrelevant to generating one or more visibility confidence models corresponding to the environment around the self - vehicle. In response to one or more sub - parts being irrelevant, the self - machine may save processing and computing power by assuming that the sensor data corresponding to these sub - parts is visible.
[0137] In some embodiments, the relevant individual sub - parts 322 (e.g., 322a - 322d) may be evaluated to determine whether there is one or more parts blocked and / or degraded, as further described in this disclosure such as, for example, with respect to Figure 3A In some embodiments, in response to determining that there may be one or more parts degraded in the sensor data corresponding to the sub - part 322, one or more RDMs may be generated by projecting one or more rays from the image data corresponding to the image 320 onto the ground plane corresponding to the environment that can be depicted by the image 320.
[0138] An example depiction of generating an RDM by projecting the rays corresponding to the degraded sensor data 302 onto the ground plane may be depicted in Figure 3C In Figure 3CDepicts a depiction of an environment 350 according to one or more embodiments of the present disclosure, in which one or more light rays 326 from sensor data corresponding to sensor 324 may be projected. In some embodiments, environment 350 may include a system 330 to which sensor 324 may correspond. For example, in the context of autonomous driving, system 330 may include a self-driving vehicle having a plurality of sensors corresponding thereto.
[0139] In some embodiments, sensor 324 corresponding to system 330 may be configured to generate sensor data. For example, in the context where sensor 324 is an image sensor, the image sensor may be configured to generate image data corresponding to a particular environment. In some embodiments, environment 350 may be, for example, a depiction of the environment captured in image 320 as described Figure 3B above.
[0140] In some embodiments, environment 350 may include one or more light rays 326 that may be projected onto one or more portions of a ground plane corresponding to environment 350 using corresponding azimuth angles. In some embodiments, individual light rays from a particular portion of sensor data may be projected. For example, in the context where sensor 324 generates image data corresponding to Figure 3B image 320 as shown, light rays 326 from a particular portion of the image data corresponding to image 320 may be projected. As shown in environment 350, light rays 326 may correspond to projections of image data corresponding to respective edges of sub-portions 322 of image 320. For example, first light ray 326a may correspond to the edge of first sub-portion 322a of image 320 closest to the vehicle. Additionally, second light ray 326b may correspond to the second edge of first sub-portion 322a furthest from the vehicle and / or the first edge of second sub-portion 322b, and so on.
[0141] In some embodiments, the distance between the light rays 326 can correspond to the distance between the edges of the sub - portion 322 of the image 320 when the image 320 is projected onto the ground plane of the environment 350. In some embodiments, the distance between the light rays 326 can reflect the distance between the edges of the sub - portion 322 when projected from the middle of the sensor 324 onto the ground plane of the environment 350. In some embodiments, the distance between the light rays 326 can be included in one or more regions of interest corresponding to the environment 350. For example, the region between the first light ray 326a and the second light ray 326b can be the first region of interest 332a, the region between the second light ray 326b and the third light ray 326c can be the second region of interest 332b, the distance between the third light ray 326c and the fourth light ray 326d can be the third region of interest 332c, and the region between the fourth light ray 326d and the fifth light ray 326e can be the fourth region of interest 332d, and so on.
[0142] In some embodiments, the environment 350 can include one or more projected objects 328, which can be placed on the ground plane at different distances in front of the system 330. In some embodiments, the one or more projected objects 328 can include one or more objects used as references in the environment 350. For example, in the context where the system 324 is an autonomous vehicle, the one or more projected objects 328 can include one or more vehicles. In some embodiments, the one or more vehicles can simulate vehicles traveling relatively close to the autonomous vehicle. In some embodiments, the size and shape of the projected object 328 can vary. In some embodiments, the distance between the system 324 and the one or more projected objects 328 can vary. As Figure 3C shown, the first object 328a, the second object 328b, and the third object 328c can be included in the environment 350.
[0143] In some embodiments, it can be determined whether one or more light rays 326 and / or the corresponding regions of interest can intersect, overlap, or otherwise be included in the same region as the one or more projected objects 328. In some embodiments, in response to one or more regions of interest intersecting or overlapping with the one or more projected objects 328, a confidence value and / or data associated with the specific region can be determined and / or calculated for the sensor data corresponding to the specific region of interest.
[0144] In some embodiments, confidence values that can be determined can be aggregated over two or more regions of interest 332 that can overlap with one or more projected objects. For example, a first object of interest 328a can intersect a first light ray 326a, a second light ray 326b, and / or a third light ray 326c. In some embodiments, this can indicate that one or more portions of a first region of interest 332a, a second region of interest 332b, and a third region of interest 332c can affect the visibility of the first virtual object 328a. In some embodiments, confidence and / or visibility data, values, information, etc. can be aggregated over the confidence and / or visibility data corresponding to the first region of interest 332a, the second region of interest 332b, and the third region of interest 332c.
[0145] Modifications, additions, or omissions can be made to Figure 3B and / or Figure 3C without departing from the scope of the present disclosure. For example, the number of images 320 corresponding to a particular system, the number of sensors 324 that can generate the images 320, the number of sub-parts 322 corresponding to the images, the number and / or type of systems 330 can vary, the number of projected objects 328, the number of light rays 326 can vary. The details given and discussed aid in the explanation and understanding of the concepts of the present disclosure and are not meant to be limiting.
[0146] Returning to Figure 3A , in some embodiments, the fine-grained degradation module 312 can be configured to generate one or more RDMs. For example, an RDM array can be generated, the RDM array including respective RDMs corresponding to respective sensors. For example, as Figure 3B and Figure 3C shown, an RDM can be generated based on one or more light rays and corresponding azimuth angles that can be associated with degraded sensor data 302 corresponding to a respective sensor. In some embodiments, visibility and / or confidence data can correspond to one or more regions of interest to determine whether respective sensors can be configured to generate reliable and / or accurate sensor data 302 and how good the generated sensor data 302 is.
[0147] Additionally or alternatively, the fine-grained degradation module 312 can be configured to generate respective RDMs using sensor data 302 corresponding to respective individual sensors of a particular sensor modality. For example, respective RDMs can be generated using sensor data 302 corresponding to respective sensors corresponding to a sensor modality. Compared to the above example, the respective RDMs may not be added to a single RDM corresponding to a machine; rather, each RDM can separately represent whether and / or where fine-grained degradation associated with a particular sensor can affect one or more sub-parts of the aggregated field of view corresponding to the visibility confidence model 306.
[0148] In some embodiments, the fine-grained degradation module 312 can be configured to generate a combined RDM using all sensors corresponding to a particular sensor modality. For example, an RDM generated using a single sensor can be added to the combined RDM associated with a particular sensor modality. In some cases, each RDM in the respective RDMs corresponding to the respective sensors can be added to the combined RDM associated with a particular sensor modality corresponding to a system, machine, ego-machine, etc.
[0149] In some embodiments, as Figure 3C shown in the environment 350, one or more objects (e.g., one or more projected objects 328 described with respect to Figure 3C can be projected from the sensor data 302 into the environment to determine one or more regions of interest that may overlap with the projected objects. In some embodiments, the number of rays that can be projected may include only the rays corresponding to sensor data 302 that is ambiguous, blocked, or occluded. In some embodiments, the RDMs can be stored separately, and thus, when determining whether a location may correspond to one or more blocked regions of interest, multiple data structures corresponding to the respective RDMs can be queried.
[0150] Additionally or alternatively, the fine-grained degradation module 312 can be configured to generate a "flat-Earth grid" or Cartesian coordinate grid corresponding to the ground plane in the environment in which the system, machine, ego-machine, etc. may be located. In some embodiments, the grid can be populated by generating a grid representing discrete positions corresponding to the ground plane. In some embodiments, the grid cell size can be determined by projecting an object size (e.g., in the context of an autonomous vehicle, an object of vehicle size) onto the grid. Additionally, the fine-grained degradation module 312 can be configured to determine one or more confidence and / or visibility levels corresponding to the respective cells of the grid. Storing confidence and / or visibility data in this manner can be computationally more expensive than generating, for example, RDMs corresponding to individual sensors. However, in contrast, storing the data in a centralized grid can reduce the complexity and computation of determining whether sensor data is blocked or ambiguous at a particular location.
[0151] In some embodiments, the fine-grained degradation module 312 may be configured to transfer, transmit, and / or send data structures, sensor data 302, and fine-grained degradation data to the occlusion module 314. In some embodiments, the occlusion module 314 may be configured to determine whether one or more occlusions may be included in the sensor data 302. For example, one or more static or dynamic obstacles that may obscure one or more other obstacles, objects, or regions such that an individual sensor may not view and / or detect one or more other obstacles or objects in the field of view. In some embodiments, to determine whether one or more occlusions may be indicated in the sensor data 302, map data corresponding to an environmental map may be used. The map data may include data corresponding to static obstacles and / or objects that may be present within or outside the field of view of an individual sensor.
[0152] For example, in the context of an autonomous vehicle, the autonomous vehicle may include one or more image sensors that may generate image data corresponding to the environment. In some embodiments, the image data may not be able to see one or more static obstacles. For example, buildings, median strips, gates, utility poles, storefronts, etc. In some embodiments, the autonomous vehicle may not be able to see static obstacles because a truck or other obstacle may be located between the vehicle and the static obstacle. Additionally or alternatively, one or more shadows may be located above the static obstacle such that the autonomous vehicle may not be configured to see and / or locate the static obstacle. In some cases, the static obstacle may be around a corner or behind another building, etc. Continuing the above example, the autonomous vehicle may be configured to use HD map data corresponding to the environment to sense static obstacles, and the HD map data may inform the autonomous vehicle of static obstacles in the environment that the autonomous vehicle may not be able to sense using the image data.
[0153] In some embodiments, historical sensor data 316 may be used to determine whether one or more objects in the sensor data corresponding to one or more previous timestamps may cast shadows or otherwise occlude the sensor data 302 corresponding to one or more current timestamps. In some embodiments, the historical sensor data 316 may include sensor data, such as, for example, sensor data 302 that may have been collected and / or generated at one or more previous timestamps. For example, again in the context of an autonomous vehicle generating image data corresponding to the environment. A static obstacle may be visible at time t = 0. However, at t = 2, due to changes in traffic or changes in the position of the autonomous vehicle, etc., the obstacle may no longer be visible. Continuing the example, the autonomous vehicle may be configured to retain some historical memory of the image data that may be included in the historical sensor data 316, and the historical memory of the image data may be used to determine the position corresponding to the static obstacle relative to the autonomous vehicle.
[0154] In some embodiments, determining the presence or absence of occlusion in sensor data 302 may be performed using one or more post - processing and / or rendering techniques corresponding to the sensor data. For example, one or more ray - tracing techniques on a 2D raster, projecting a 3D bounding box into the sensor data, and so on. In some embodiments, the presence or absence of one or more occlusions may change the confidence and / or visibility determination regarding a particular location, region, and / or volume indicated by the sensor data 302.
[0155] In some embodiments, one or more models corresponding to respective sensors may be aggregated to form a visibility confidence model 306. In some embodiments, the visibility confidence model 306 may include multiple models and / or data structures corresponding to respective sensors. Additionally or alternatively, the visibility confidence model 306 may include an aggregation of confidence and / or visibility determinations corresponding to the sensor data 302 of each sensor corresponding to a sensor modality. Additionally or alternatively, the visibility confidence model 306 may include an aggregation of confidence and / or visibility determinations corresponding to the sensor data 302 associated with and across each sensor of multiple sensor modalities.
[0156] In some embodiments, the visibility confidence model 306 may include one or more different data structures, which may include, for example, data and / or information generated using the sensor health module 308, the coarse - level degradation module 310, the fine - level degradation module 312, and / or the occlusion module 314. In some embodiments, the visibility confidence model 306 may include, for example, one or more generated RDMs and / or grid structures, which may include confidence and / or visibility data corresponding to a particular portion of the aggregated field of view. In these or other embodiments, the visibility confidence model 306 may be the same as and / or similar to the visibility confidence model 206 further described and / or illustrated in this disclosure, such as Figure 2 the visibility confidence model 206.
[0157] Modifications, additions, or omissions may be made Figures 3A - 3C without departing from the scope of this disclosure. For example, the quantity and / or type of sensor data 302 may vary, the type of visibility system 304, the quantity and / or type of neural networks 318, the quantity of modules configured to perform various operations, and the quantity of visibility confidence models 306 may vary. The details given and discussed help to explain and understand the concepts of this disclosure and are not meant to be limiting.
[0158] Figure 4is a flowchart showing a method 400 for generating one or more visibility confidence models and generating one or more control determinations based on the one or more visibility confidence models according to one or more embodiments of the present disclosure. Method 400 may include one or more blocks 402, 404, and 406. Although shown as discrete blocks, the operations associated with one or more blocks of method 400 may be divided into additional blocks, combined into fewer blocks, or eliminated according to a particular implementation.
[0159] In some embodiments, method 400 may include block 402. At block 402, a visibility confidence model may be generated, where the visibility confidence model may correspond to an aggregated field of view. In some embodiments, each field of view corresponding to the aggregated field of view may respectively define a potential spatial coverage range of sensor data that may correspond to a plurality of sensors associated with a machine. In some embodiments, the visibility confidence model may indicate a confidence level of sensor data corresponding to each sub - part of the aggregated field of view.
[0160] In some embodiments, a corresponding confidence level may be determined based on one or more faults or errors that may be associated with an individual sensor among the plurality of sensors. In these or other embodiments, one or more techniques and / or analyses such as those further described and / or illustrated in the present disclosure, such as for example Figure 3A the sensor health module 308 may be used to determine one or more faults or errors. Additionally, a corresponding confidence level may be determined based on, for example, one or more coarse - level degradations and / or fine - level degradations as described in Figure 2 and / or Figure 3A Additional or alternatively, a corresponding confidence level may be determined based on one or more fine - level degradations corresponding to the sensor data (such as, for example, the fine - level degradations described in Figure 2 and Figure 3A ). Additional or alternatively, a corresponding confidence level may be determined based on the presence of one or more occlusions in the sensor data corresponding to each sub - part of the aggregated field of view.
[0161] At block 404, the method may further include: in response to a query, sending data corresponding to the corresponding confidence level. In some embodiments, the query may correspond to one or more systems that may navigate an environment corresponding to the sensor data and / or the visibility confidence model. In some embodiments, the query may include one or more specific parts of the environment and a corresponding sensor modality.
[0162] At block 406, the method may further include performing one or more operations based on a visibility confidence model. In some embodiments, the one or more operations may be performed based on one or more confidence levels that may correspond to respective sub-regions of an aggregated field of view.
[0163] Method 400 and / or one or more operations included in method 400 may be modified, added, or omitted without departing from the scope of the present disclosure. For example, the operations corresponding to method 400 may be implemented in a different order. Additionally or alternatively, two or more operations may be performed simultaneously. Further, the operations and actions outlined are provided only as examples, and some of the operations and actions may be optional, combined into fewer operations and actions, or expanded into additional operations and actions without departing from the essence of the described embodiments.
[0164] Example Autonomous Vehicle
[0165] Figure 5A is an illustration of an example autonomous vehicle 500 in accordance with some embodiments of the present disclosure. Autonomous vehicle 500 (alternatively, referred to herein as "vehicle 500") may include, but is not limited to, passenger vehicles such as cars, trucks, buses, emergency vehicles, shuttles, electric or motorized bicycles, motorcycles, fire trucks, police cars, ambulances, boats, construction vehicles, submarines, drones, and / or other types of vehicles (e.g., driverless and / or accommodating one or more passengers). Autonomous vehicles are generally described according to the levels of automation defined by the National Highway Traffic Safety Administration (NHTSA), a division of the United States Department of Transportation, and the Society of Automotive Engineers (SAE) in "Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles" (Standard No. J3016 - 201806, issued June 15, 2018, Standard No. J3016 - 201609, issued September 30, 2016, and prior and future versions of this standard). Vehicle 500 may be capable of implementing one or more functions corresponding to levels 3 - 5 of the autonomous driving levels. Vehicle 500 may be capable of implementing one or more functions corresponding to levels 3 - 5 of the autonomous driving levels. For example, depending on the embodiment, vehicle 500 may be capable of implementing driver assistance (level 1), semi-automation (level 2), conditional automation (level 3), high automation (level 4), and / or full automation (level 5). The term "autonomous" as used herein may include any and / or all types of autonomy of vehicle 500 or other machines, such as fully autonomous, highly autonomous, conditionally autonomous, partially autonomous, providing assistive autonomy, semi-autonomous, primarily autonomous, or other designations.
[0166] Vehicle 500 may include components such as a chassis, a body, wheels (e.g., 2, 4, 6, 8, 18, etc.), tires, axles, and other components of the vehicle. Vehicle 500 may include a propulsion system 550, such as an internal combustion engine, a hybrid power plant, a fully electric motor, and / or another type of propulsion system. Propulsion system 550 may be connected to the driveline of vehicle 500, which may include a transmission, to effect the propulsion of vehicle 500. Propulsion system 550 may be controlled in response to receiving a signal from throttle / accelerator 552.
[0167] A steering system 554, which may include a steering wheel, may be used to steer vehicle 500 (e.g., along a desired path or route) while propulsion system 550 is operating (e.g., while the vehicle is in motion). Steering system 554 may receive a signal from steering actuator 556. For fully autonomous (Level 5) functionality, the steering wheel may be optional.
[0168] A brake sensor system 546 may be used to operate vehicle brakes in response to receiving a signal from brake actuator 548 and / or a brake sensor.
[0169] One or more controllers 536, which may include one or more CPUs, a system-on-chip (SoC) 504 ( Figure 5C ) and / or one or more GPUs, may provide signals (e.g., representing commands) to one or more components and / or systems of vehicle 500. For example, one or more controllers may send signals to operate vehicle brakes via one or more brake actuators 548, to operate steering system 554 via one or more steering actuators 556, and / or to operate propulsion system 550 via one or more throttle / accelerators 552. One or more controllers 536 may include one or more on-board (e.g., integrated) computing devices (e.g., supercomputers) that process sensor signals and output operation commands (e.g., signals representing commands) to effect autonomous driving and / or assist a human driver in driving vehicle 500. One or more controllers 536 may include a first controller 536 for autonomous driving functionality, a second controller 536 for functional safety functionality, a third controller 536 for artificial intelligence functionality (e.g., computer vision), a fourth controller 536 for infotainment functionality, a fifth controller 536 for redundancy in emergency situations, and / or other controllers. In some examples, a single controller 536 may handle two or more of the above functions, two or more controllers 536 may handle a single function, and / or any combination thereof.
[0170] One or more controllers 536 can provide signals for controlling one or more components and / or systems of the vehicle 500 in response to sensor data (e.g., sensor inputs) received from one or more sensors. The sensor data can be received from, for example and without limitation, a global navigation satellite system sensor 558 (e.g., a global positioning system sensor), a RADAR sensor 560, an ultrasonic sensor 562, a LiDAR sensor 564, an inertial measurement unit (IMU) sensor 566 (e.g., an accelerometer, a gyroscope, a magnetic compass, a magnetometer, etc.), a microphone 596, a stereo camera 568, a wide-angle camera 570 (e.g., a fish-eye camera), an infrared camera 572, a surround camera 574 (e.g., a 360-degree camera), a remote and / or mid-range camera 598, a speed sensor 544 (e.g., for measuring the speed of the vehicle 500), a vibration sensor 542, a steering sensor 540, a brake sensor 546 (e.g., as part of a brake sensor system 546), and / or other sensor types.
[0171] One or more of the controllers 536 can receive inputs (e.g., represented by input data) from the instrument cluster 532 of the vehicle 500 and provide outputs (e.g., represented by output data, display data, etc.) via the human-machine interface (HMI) display 534, an audible annunciator, a speaker, and / or via other components of the vehicle 500. These outputs can include messages such as vehicle speed, rate, time, map data (e.g., Figure 5C the HD map 522), location data (e.g., the location of the vehicle 500 on a map, for example), direction, the locations of other vehicles (e.g., occupancy grids), information about objects and object states as perceived by the controller 536, and so on. For example, the HMI display 534 can display information about the presence of one or more objects (e.g., street signs, warning signs, traffic light changes, etc.) and / or information about driving maneuvers that the vehicle has made, is making, or will make (e.g., changing lanes now, exiting 34B in two miles, etc.).
[0172] The vehicle 500 further includes a network interface 524, which can communicate via one or more networks using one or more wireless antennas 526 and / or a modem. For example, the network interface 524 may be capable of communicating via LTE, WCDMA, UMTS, GSM, CDMA2000, etc. One or more wireless antennas 526 can also be used to enable communication between objects (e.g., vehicles, mobile devices, etc.) in an environment using one or more local area networks such as Bluetooth, LE, Z-wave, ZigBee, etc. and / or one or more low-power wide area networks (LPWANs) such as LoRaWAN, SigFox, etc.
[0173] Figure 5BAn example of the camera positions and fields of view for an example autonomous vehicle 500 in accordance with some embodiments of the present disclosure. The cameras and respective fields of view are an example embodiment and are not intended to be limiting. For example, additional and / or alternative cameras may be included, and / or these cameras may be located at different positions on the vehicle 500. Figure 5A The camera types for the cameras may include, but are not limited to, digital cameras that may be adapted to work with components and / or systems of the vehicle 500. The cameras may operate under an automotive safety integrity level (ASIL) B and / or under another ASIL. The camera types may have any image capture rate, such as 60 frames per second (fps), 120 fps, 240 fps, etc., depending on the embodiment. The cameras may be capable of using a rolling shutter, a global shutter, another type of shutter, or a combination thereof. In some examples, the color filter array may include a red-clear-clear-clear (RCCC) color filter array, a red-clear-clear-blue (RCCB) color filter array, a red-blue-green-clear (RBGC) color filter array, a Foveon X3 color filter array, a Bayer sensor (RGGB) color filter array, a monochrome sensor color filter array, and / or another type of color filter array. In some embodiments, clear pixel cameras, such as cameras having RCCC, RCCB, and / or RBGC color filter arrays, may be used in an effort to increase light sensitivity.
[0174] In some examples, one or more of the cameras may be used to perform advanced driver assistance system (ADAS) functions (e.g., as part of a redundant or fail-safe design). For example, a multi-functional monocular camera may be installed to provide functions including lane departure warning, traffic sign assistance, and smart headlight control. One or more of the cameras (e.g., all of the cameras) may simultaneously record and provide image data (e.g., video).
[0175] One or more of the cameras may be mounted in mounting components, such as custom-designed (3D printed) components, to cut off stray light and reflections from within the vehicle (e.g., reflections from the dashboard reflected in the windshield mirror) that may interfere with the image data capture ability of the cameras. Regarding the wing mirror mounting components, the wing mirror components may be custom 3D printed such that the camera mounting plate matches the shape of the wing mirror. In some examples, one or more cameras may be integrated into the wing mirror. For side view cameras, one or more cameras may also be integrated into the four pillars at each corner of the cab.
[0176]
[0177] A camera (e.g., a front camera) having a field of view that includes an environmental portion in front of vehicle 500 can be used for surround view to help identify forward paths and obstacles and, with the help of one or more controllers 536 and / or a control SoC, assist in providing information critical for generating an occupancy grid and / or determining a preferred vehicle path. The front camera can be used to perform many of the same ADAS functions as LiDAR, including emergency braking, pedestrian detection, and collision avoidance. The front camera can also be used for ADAS functions and systems, including lane departure warning ("LDW"), adaptive cruise control ("ACC"), and / or other functions such as traffic sign recognition.
[0178] A variety of cameras can be used in a front-facing configuration, including, for example, a monocular camera platform that includes a CMOS (complementary metal oxide semiconductor) color imager. Another example can be a wide-angle camera 570, which can be used to sense objects (e.g., pedestrians, intersection traffic, or bicycles) entering the field of view from the periphery. Although Figure 5B only one wide-angle camera is illustrated in the figure, any number of wide-angle cameras 570 can be present on vehicle 500. Additionally, a long-range camera 598 (e.g., a long-range stereo camera pair) can be used for depth-based object detection, particularly for objects for which a neural network has not been trained. The long-range camera 598 can also be used for object detection and classification and basic object tracking.
[0179] One or more stereo cameras 568 can also be included in a front-facing configuration. The stereo camera 568 can include an integrated control unit that includes a scalable processing unit that can provide a multi-core microprocessor and programmable logic (FPGA) with an integrated CAN or Ethernet interface on a single chip. Such a unit can be used to generate a 3D map of the vehicle environment, including distance estimates for all points in the image. An alternative stereo camera 568 can include a compact stereo vision sensor that can include two camera lenses (one on the left and one on the right) and an image processing chip that can measure the distance from the vehicle to a target object and use the generated information (e.g., metadata) to activate autonomous emergency braking and lane departure warning functions. Other types of stereo cameras 568 can be used in addition to or in place of those described herein.
[0180] A camera (e.g., a side-view camera) having a field of view that includes an environmental portion on the side of vehicle 500 can be used for surround view, providing information used to create and update an occupancy grid and generate side-impact collision warnings. For example, a surround camera 574 (e.g., as Figure 5BThe four surround cameras 574 shown in [Fig.] can be placed on the vehicle 500. The surround cameras 574 can include wide-angle cameras 570, fisheye cameras, 360-degree cameras, and / or the like. For example, four fisheye cameras can be placed in the front, rear, and sides of the vehicle. In an alternative arrangement, the vehicle can use three surround cameras 574 (e.g., left, right, and rear), and can utilize one or more other cameras (e.g., a forward camera) as the fourth surround camera.
[0181] A camera having a field of view that includes an environmental portion behind the vehicle 500 (e.g., a rearview camera) can be used to assist with parking, surround view, rear collision warning, and creating and updating an occupancy grid. A variety of cameras can be used, including but not limited to cameras that are also suitable as front cameras as described herein (e.g., long-range and / or mid-range cameras 598, stereo cameras 568, infrared cameras 572, etc.).
[0182] Figure 5C For an example autonomous vehicle 500 in accordance with some embodiments of the present disclosure Figure 5A is a block diagram of an example system architecture. It should be understood that this and other arrangements described herein are presented only as examples. Other arrangements and elements (e.g., machines, interfaces, functions, orders, function groupings, etc.) can be used in addition to or instead of those shown, and some elements can be omitted entirely. Further, many of the elements described herein are functional entities that can be implemented as discrete or distributed components or in combination with other components, and in any suitable combination and location. The various functions described herein as being performed by entities can be implemented by hardware, firmware, and / or software. For example, the various functions can be implemented by a processor executing instructions stored in a memory.
[0183] Figure 5C Each of the components, features, and systems in [Fig.] of the vehicle 500 is illustrated as being connected via a bus 502. The bus 502 can include a Controller Area Network (CAN) data interface (alternatively referred to herein as the "CAN bus"). The CAN can be a network within the vehicle 500 used to assist in controlling various features and functions of the vehicle 500, such as driving brakes, acceleration, braking, steering, windshield wipers, etc. The CAN bus can be configured to have dozens or even hundreds of nodes, each with its own unique identifier (e.g., CAN ID). The CAN bus can be read to find the steering wheel angle, ground speed, engine revolutions per minute (RPM), button positions, and / or other vehicle status indicators. The CAN bus can be ASIL B compliant.
[0184] Although the bus 502 is described herein as a CAN bus, this is not intended to be limiting. For example, in addition to or alternatively to the CAN bus, FlexRay and / or Ethernet can be used. Further, although the bus 502 is shown as a single line, this is not intended to be limiting. For example, any number of buses 502 can be present, which can include one or more CAN buses, one or more FlexRay buses, one or more Ethernet buses, and / or one or more other types of buses using different protocols. In some examples, two or more buses 502 can be used to perform different functions, and / or can be used for redundancy. For example, a first bus 502 can be used for a collision avoidance function, and a second bus 502 can be used for drive control. In any example, each bus 502 can communicate with any component of the vehicle 500, and two or more buses 502 can communicate with the same component. In some examples, each SoC 504, each controller 536, and / or each computer within the vehicle can have access to the same input data (e.g., input from sensors of the vehicle 500), and can be connected to a common bus such as a CAN bus.
[0185] The vehicle 500 can include one or more controllers 536, such as those described herein with respect to Figure 5A The controllers 536 can be used for a variety of functions. The controllers 536 can be coupled to any other different components and systems of the vehicle 500, and can be used for control of the vehicle 500, artificial intelligence for the vehicle 500, infotainment for the vehicle 500, and / or the like.
[0186] The vehicle 500 can include one or more system-on-chips (SoC) 504. The SoC 504 can include a CPU 506, a GPU 508, a processor 510, a cache 512, an accelerator 514, a data store 516, and / or other components and features not shown. In a variety of platforms and systems, the SoC 504 can be used to control the vehicle 500. For example, one or more SoC 504 can be combined with an HD map 522 in a system (e.g., a system of the vehicle 500), and the HD map can obtain map refreshes and / or updates from one or more servers (e.g., Figure 5D one or more servers 578) via a network interface 524.
[0187] The CPU 506 may include a CPU cluster or a CPU complex (alternatively referred to herein as a "CCPLEX"). The CPU 506 may include multiple cores and / or L2 caches. For example, in some embodiments, the CPU 506 may include eight cores in a coherent multi-processor configuration. In some embodiments, the CPU 506 may include four dual-core clusters, each with a dedicated L2 cache (e.g., a 2MB L2 cache). The CPU 506 (e.g., CCPLEX) may be configured to support simultaneous cluster operation such that any combination of the clusters of the CPU 506 can be active at any given time.
[0188] The CPU 506 may implement power management capabilities including one or more of the following: each hardware block may automatically perform clock gating when idle to save dynamic power; each core clock may be gated when the core is not actively executing instructions due to the execution of WFI / WFE instructions; each core may be independently power gated; when all cores are clock gated or power gated, each core cluster may be independently clock gated; and / or when all cores are power gated, each core cluster may be independently power gated. The CPU 506 may further implement enhanced algorithms for managing power states, where allowed power states and desired wake-up times are specified, and the hardware / microcode determines the optimal power state for the cores, clusters, and CCPLEX to enter. The processing cores may support a simplified power state entry sequence in software, and this task may be offloaded to the microcode.
[0189] The GPU 508 may include an integrated GPU (alternatively referred to herein as an "iGPU"). The GPU 508 may be programmable and efficient for parallel workloads. In some examples, the GPU 508 may use an enhanced tensor instruction set. The GPU 508 may include one or more streaming microprocessors, where each streaming microprocessor may include an L1 cache (e.g., an L1 cache with at least 96KB of storage capacity), and two or more of these streaming microprocessors may share an L2 cache (e.g., an L2 cache with 512KB of storage capacity). In some embodiments, the GPU 508 may include at least eight streaming microprocessors. The GPU 508 may use a compute application programming interface (API). Additionally, the GPU 508 may use one or more parallel computing platforms and / or programming models (e.g., NVIDIA's CUDA).
[0190] In the case of automotive and embedded use cases, the GPU 508 can be power optimized for best performance. For example, the GPU 508 can be fabricated on fin field-effect transistors (FinFETs). However, this is not intended to be limiting, and the GPU 508 can be fabricated using other semiconductor manufacturing processes. Each streaming microprocessor can incorporate a number of mixed-precision processing cores divided into multiple blocks. By way of example and not limitation, 64 PF32 cores and 32 PF64 cores can be divided into four processing blocks. In such an example, each processing block can be allocated 16 FP32 cores, 8 FP64 cores, 16 INT32 cores, two mixed-precision NVIDIA tensor cores for deep learning matrix arithmetic, an L0 instruction cache, a warp scheduler, a dispatch unit, and / or a 64KB register file. Additionally, the streaming microprocessor can include separate parallel integer and floating-point data paths to enable efficient execution of workloads leveraging a mix of compute and addressing computations. The streaming microprocessor can include separate thread scheduling capabilities to allow for finer-grained synchronization and cooperation between parallel threads. The streaming microprocessor can include a combined L1 data cache and shared memory unit to improve performance while simplifying programming.
[0191] The GPU 508 can include, in some examples, a high-bandwidth memory (HBM) that provides a peak memory bandwidth of approximately 900 GB / s and / or a 16GB HBM2 memory subsystem. In some examples, in addition to or alternatively to HBM memory, synchronous graphics random access memory (SGRAM), such as fifth-generation graphics double data rate synchronous random access memory (GDDR5), can be used.
[0192] The GPU 508 can include unified memory technology that includes access counters to allow memory pages to be more precisely migrated to the processors that most frequently access them, thereby improving the efficiency of the memory range shared between processors. In some examples, address translation service (ATS) support can be used to allow the GPU 508 to directly access the CPU 506 page table. In such an example, when the GPU 508 memory management unit (MMU) experiences a miss, an address translation request can be transmitted to the CPU 506. In response, the CPU 506 can look up the virtual-physical mapping for the address in its page table and transmit the translation back to the GPU 508. In this way, the unified memory technology can allow for a single unified virtual address space for the memory of both the CPU 506 and the GPU 508, thereby simplifying GPU 508 programming and porting applications to the GPU 508.
[0193] In addition, the GPU 508 may include an access counter that can track how frequently the GPU 508 accesses the memory of other processors. The access counter can help ensure that memory pages are moved to the physical memory of the processor that most frequently accesses those pages.
[0194] The SoC 504 may include any number of caches 512, including those described herein. For example, the cache 512 may include an L3 cache that is available to both the CPU 506 and the GPU 508 (e.g., that is connected to both the CPU 506 and the GPU 508). The cache 512 may include a write-back cache that can track the state of lines, for example, by using a cache coherence protocol (such as MEI, MESI, MSI, etc.). Depending on the embodiment, the L3 cache may include 4MB or more, but smaller cache sizes may also be used.
[0195] The SoC 504 may include one or more arithmetic logic units (ALUs) that can be used to perform processing for any of a variety of tasks or operations regarding the vehicle 500, such as processing a DNN. In addition, the SoC 504 may include a floating point unit (FPU) or other math co-processor or digital co-processor type for performing mathematical operations within the system. For example, the SoC 504 may include one or more FPUs integrated within the execution units of the CPU 506 and / or the GPU 508.
[0196] The SoC 504 may include one or more accelerators 514 (e.g., hardware accelerators, software accelerators, or a combination thereof). For example, the SoC 504 may include a hardware acceleration cluster that may include optimized hardware accelerators and / or large on-chip memories. The large on-chip memory (e.g., 4MB SRAM) may enable the hardware acceleration cluster to accelerate neural networks and other computations. The hardware acceleration cluster may be used to supplement the GPU 508 and offload some of the tasks of the GPU 508 (e.g., freeing up more cycles of the GPU 508 for performing other tasks). As an example, the accelerator 514 may be used for targeted workloads that are stable enough to be easily accelerated (such as perception, convolutional neural networks (CNNs), etc.). As used herein, the term "CNN" may include all types of CNNs, including region-based or region convolutional neural networks (RCNNs) and fast RCNNs (e.g., for object detection).
[0197] The accelerator 514 (e.g., a hardware acceleration cluster) may include a Deep Learning Accelerator (DLA). The DLA may include one or more Tensor Processing Units (TPUs) that may be configured to provide an additional one trillion operations per second for deep learning applications and inference. The TPU may be an accelerator configured to perform image processing functions (e.g., for CNN, RCNN, etc.) and optimized for performing image processing functions. The DLA may be further optimized for a specific set of neural network types and floating-point operations, as well as inference. The design of the DLA may provide higher performance per millimeter than a general-purpose GPU and far exceed the performance of a CPU. The TPU may perform several functions, including single-instance convolution functions, supporting INT8, INT16, and FP16 data types for both features and weights, and post-processor functions.
[0198] The DLA may execute neural networks, especially CNNs, quickly and efficiently for any of a variety of functions on processed or unprocessed data, such as, and not limited to: CNNs for object recognition and detection using data from a camera sensor; CNNs for distance estimation using data from a camera sensor; CNNs for emergency vehicle detection and identification and detection using data from a microphone; CNNs for face recognition and vehicle owner recognition using data from a camera sensor; and / or CNNs for security and / or safety-related events.
[0199] The DLA may perform any function of the GPU 508, and by using an inference accelerator, for example, a designer may configure the DLA or the GPU 508 for any function. For example, a designer may focus the processing and floating-point operations of a CNN on the DLA and leave other functions to the GPU 508 and / or other accelerators 514.
[0200] The accelerator 514 (e.g., a hardware acceleration cluster) may include a Programmable Vision Accelerator (PVA), which may alternatively be referred to herein as a computer vision accelerator. The PVA may be designed and configured to accelerate computer vision algorithms for Advanced Driver Assistance Systems (ADAS), autonomous driving, and / or augmented reality (AR) and / or virtual reality (VR) applications. The PVA may provide a balance between performance and flexibility. For example, each PVA may include, for example and not limited to, any number of Reduced Instruction Set Computer (RISC) cores, Direct Memory Access (DMA), and / or any number of vector processors.
[0201] The RISC cores can interact with an image sensor (e.g., the image sensor of any of the cameras described herein), an image signal processor, and / or the like. Each of these RISC cores can include any number of memories. Depending on the embodiment, the RISC cores can use any of several protocols. In some examples, the RISC cores can execute a real-time operating system (RTOS). The RISC cores can be implemented using one or more integrated circuit devices, application-specific integrated circuits (ASICs), and / or storage devices. For example, the RISC cores can include an instruction cache and / or tightly coupled RAM.
[0202] The DMA can enable components of the PVA to access system memory independently of the CPU 506. The DMA can support any number of features used to optimize the PVA, including but not limited to supporting multi-dimensional addressing and / or circular addressing. In some examples, the DMA can support up to six or more dimensions of addressing, which can include block width, block height, block depth, horizontal block stride, vertical block stride, and / or depth stride.
[0203] The vector processor can be a programmable processor that can be designed to efficiently and flexibly execute programming for computer vision algorithms and provide signal processing capabilities. In some examples, the PVA can include a PVA core and two vector processing subsystem partitions. The PVA core can include a processor subsystem, one or more DMA engines (e.g., two DMA engines), and / or other peripherals. The vector processing subsystem can operate as the main processing engine of the PVA and can include a vector processing unit (VPU), an instruction cache, and / or vector memory (e.g., VMEM). The VPU core can include a digital signal processor, such as, for example, a single instruction multiple data (SIMD), very long instruction word (VLIW) digital signal processor. The combination of SIMD and VLIW can enhance throughput and rate.
[0204] Each of the vector processors may include an instruction cache and may be coupled to dedicated memory. As a result, in some examples, each of the vector processors may be configured to execute independently of the other vector processors. In other examples, the vector processors included in a particular PVA may be configured to employ data parallelization. For example, in some embodiments, multiple vector processors included in a single PVA may execute the same computer vision algorithm, but on different regions of an image. In other examples, the vector processors included in a particular PVA may simultaneously execute different computer vision algorithms on the same image, or even execute different algorithms on a sequence of images or portions of an image. Among other things, any number of PVAs may be included in a hardware acceleration cluster, and any number of vector processors may be included in each of these PVAs. Additionally, the PVA may include additional error correction code (ECC) memory to enhance overall system security.
[0205] The accelerator 514 (e.g., a hardware acceleration cluster) may include an on-chip computer vision network and SRAM to provide high-bandwidth, low-latency SRAM for the accelerator 514. In some examples, the on-chip memory may include at least 4MB of SRAM consisting of, for example and without limitation, eight field-configurable memory blocks that may be accessed by both the PVA and the DLA. Each pair of memory blocks may include an advanced peripheral bus (APB) interface, configuration circuitry, a controller, and a multiplexer. Any type of memory may be used. The PVA and the DLA may access the memory via a backbone that provides high-speed memory access to the PVA and the DLA. The backbone may include, for example using the APB, an on-chip computer vision network that interconnects the PVA and the DLA to the memory.
[0206] The on-chip computer vision network may include an interface that determines that both the PVA and the DLA provide ready and valid signals before transmitting any control signals / address / data. Such an interface may provide separate phases and separate channels for transmitting control signals / address / data, as well as burst communication for continuous data transfer. This type of interface may conform to the ISO 26262 or IEC 61508 standards, but other standards and protocols may also be used.
[0207] In some examples, the SoC 504 can include, for example, a real-time ray tracing hardware accelerator as described in U.S. Patent Application No. 16 / 101,232, filed on August 10, 2018. The real-time ray tracing hardware accelerator can be used to quickly and efficiently determine the position and extent of objects (e.g., within a world model) in order to generate a real-time visualization simulation for RADAR signal interpretation, for sound propagation synthesis and / or analysis, for SONAR system simulation, for general wave propagation simulation, for comparison with LiDAR data for positioning and / or other functional purposes, and / or for other uses. In some embodiments, one or more tree traversal units (TTUs) can be used to perform one or more ray tracing-related operations.
[0208] The accelerator 514 (e.g., a hardware accelerator cluster) has a wide range of autonomous driving uses. The PVA can be a programmable vision accelerator, which can be used in key processing stages in ADAS and autonomous vehicles. The capabilities of the PVA are a good match for algorithm domains that require predictable processing, low power, and low latency. In other words, the PVA performs well in semi-dense or dense regular computations, even on small data sets that require predictable runtimes with low latency and low power. Thus, in the context of a platform for autonomous vehicles, the PVA is designed to run classical computer vision algorithms because they are effective in object detection and integer math operations.
[0209] For example, according to one embodiment of the technology, the PVA is used to perform computer stereo vision. In some examples, an algorithm based on semi-global matching can be used, but this is not intended to be limiting. Many applications for level 3-5 autonomous driving require instantaneous motion estimation / stereo matching (e.g., structure from motion, pedestrian recognition, lane detection, etc.). The PVA can perform computer stereo vision functions on inputs from two monocular cameras.
[0210] In some examples, the PVA can be used to perform dense optical flow. Process raw RADAR data (e.g., using a 4D fast Fourier transform) to provide processed RADAR. In other examples, the PVA is used for time-of-flight depth processing, which, for example, processes raw time-of-flight data to provide processed time-of-flight data.
[0211] DLA can be used to run any type of network to enhance control and driving safety, including, for example, a neural network that outputs a confidence metric for each object detection. Such confidence values can be interpreted as probabilities or as providing a relative "weight" of each detection compared to other detections. The confidence value enables the system to make further decisions about which detections should be considered true positive detections rather than false positive detections. For example, the system can set a threshold for the confidence and consider only detections that exceed the threshold as true positive detections. In an automatic emergency braking (AEB) system, false positive detections can cause the vehicle to automatically perform emergency braking, which is clearly undesirable. Therefore, only the most confident detections should be considered as triggers for AEB. DLA can run a neural network for regressing confidence values. The neural network can take as its input at least some subset of parameters, such as bounding box dimensions, a ground plane estimate obtained (e.g., from another subsystem), the output of an inertial measurement unit (IMU) sensor 566 related to the vehicle 500 orientation and distance, a 3D position estimate of an object obtained from the neural network and / or other sensors (such as a LiDAR sensor 564 or a RADAR sensor 560), etc.
[0212] The SoC 504 can include one or more data stores 516 (e.g., memories). The data store 516 can be on-chip memory of the SoC 504, which can store neural networks to be executed on the GPU and / or DLA. In some examples, for redundancy and safety, the data store 516 can be large enough in capacity to store multiple instances of the neural network. The data store 516 can include an L2 or L3 cache 512. References to the data store 516 can include references to memories associated with PVAs, DLAs, and / or other accelerators 514 as described herein.
[0213] The SoC 504 may include one or more processors 510 (e.g., embedded processors). The processor 510 may include a boot and power management processor, which may be a dedicated processor and subsystem for handling boot power and management functions as well as security implementation related. The boot and power management processor may be part of the SoC 504 boot sequence and may provide runtime power management services. The boot power and management processor may provide clock and voltage programming, auxiliary system low power state transitions, SoC 504 thermal and temperature sensor management, and / or SoC 504 power state management. Each temperature sensor may be implemented as a ring oscillator whose output frequency is proportional to temperature, and the SoC 504 may use the ring oscillator to detect the temperature of the CPU 506, GPU 508, and / or accelerator 514. If it is determined that the temperature exceeds a threshold, then the boot and power management processor may enter a temperature fault routine and place the SoC 504 in a lower power state and / or place the vehicle 500 in a driver safety stop mode (e.g., safely stop the vehicle 500).
[0214] The processor 510 may further include a set of embedded processors that can be used as an audio processing engine. The audio processing engine may be an audio subsystem that allows for full hardware support for multi-channel audio over multiple interfaces and a wide range of flexible audio I / O interfaces. In some examples, the audio processing engine is a dedicated processor core with a digital signal processor with dedicated RAM.
[0215] The processor 510 may further include an always-on processor engine, which may provide the necessary hardware features to support low power sensor management and wake-up use cases. The always-on processor engine may include a processor core, tightly coupled RAM, supporting peripherals (e.g., timers and interrupt controllers), various I / O controller peripherals, and routing logic.
[0216] The processor 510 may further include a security cluster engine, which includes a dedicated processor subsystem for handling security management of automotive applications. The security cluster engine may include two or more processor cores, tightly coupled RAM, supporting peripherals (e.g., timers, interrupt controllers, etc.), and / or routing logic. In a security mode, the two or more cores may operate in a lockstep mode and act as a single core with comparison logic for detecting any differences between their operations.
[0217] The processor 510 may further include a real-time camera engine, which may include a dedicated processor subsystem for handling real-time camera management.
[0218] The processor 510 may further include a high dynamic range signal processor, which may include an image signal processor, which is a hardware engine that is part of the camera processing pipeline.
[0219] The processor 510 may include a video image compositor that may be a processing block (e.g., implemented on a microprocessor) that implements the video post-processing functions required for a video playback application to generate the final image for the player window. The video image compositor may perform lens distortion correction on the wide-angle camera 570, the surround camera 574, and / or the in-cab monitoring camera sensor. The in-cab monitoring camera sensor is preferably monitored by a neural network running on another instance of the advanced SoC, configured to identify in-cab events and respond accordingly. The in-cab system may perform lip reading to activate mobile phone services and make calls, dictate emails, change the vehicle destination, activate or change the vehicle's infotainment system and settings, or provide voice-activated web surfing. Certain functions are only available to the driver when the vehicle is operating in autonomous mode and are disabled otherwise.
[0220] The video image compositor may include enhanced temporal noise reduction for spatial and temporal noise reduction. For example, in the case of motion in the video, the noise reduction appropriately weights the spatial information, reducing the weight of the information provided by neighboring frames. In the case where the image or a portion of the image does not include motion, the temporal noise reduction performed by the video image compositor may use information from a previous image to reduce the noise in the current image.
[0221] The video image compositor may also be configured to perform stereo correction on input stereo lens frames. When the operating system desktop is in use and the GPU 508 does not need to continuously render new surfaces, the video image compositor may further be used for user interface composition. Even when the GPU 508 is powered on and active for 3D rendering, the video image compositor may be used to lighten the burden on the GPU 508 to improve performance and responsiveness.
[0222] The SoC 504 may further include a Mobile Industry Processor Interface (MIPI) camera serial interface, a high-speed interface, and / or a video input block for receiving video and inputs from cameras and may be used for camera and related pixel input functions. The SoC 504 may further include an input / output controller that may be software-controlled and may be used to receive I / O signals not committed to a specific role.
[0223] The SoC 504 may further include a wide range of peripheral device interfaces to enable communication with peripheral devices, audio codecs, power management, and / or other devices. The SoC 504 may be used to process data from cameras (connected via Gigabit Multimedia Serial Link and Ethernet), sensors (such as LiDAR sensor 564, RADAR sensor 560, etc. that may be connected via Ethernet), data from bus 502 (such as the speed of vehicle 500, steering wheel position, etc.), and data from GNSS sensor 558 (connected via Ethernet or CAN bus). The SoC 504 may further include dedicated high-performance large-capacity storage controllers, which may include their own DMA engines and which may be used to free the CPU 506 from routine data management tasks.
[0224] The SoC 504 may be an end-to-end platform with a flexible architecture that spans automation levels 3 - 5, thus providing an integrated functional safety architecture for a platform that leverages and efficiently uses computer vision and ADAS technologies to achieve diversity and redundancy, along with deep learning tools to provide a flexible and reliable driving software stack. The SoC 504 may be faster, more reliable, and even more energy-efficient and space-efficient than conventional systems. For example, when combined with the CPU 506, GPU 508, and data storage 516, the accelerator 514 may provide a fast and efficient platform for level 3 - 5 autonomous vehicles.
[0225] Thus, this technology provides capabilities and functions that cannot be achieved by conventional systems. For example, computer vision algorithms may be executed on CPUs that may be configured using high-level programming languages such as the C programming language to perform various processing algorithms across a wide variety of visual data. However, CPUs often cannot meet the performance requirements of many computer vision applications, such as those related to, for example, execution time and power consumption. In particular, many CPUs cannot execute complex object detection algorithms in real time, which is a requirement for in-vehicle ADAS applications and for practical level 3 - 5 autonomous vehicles.
[0226] In contrast to conventional systems, the technology described herein allows multiple neural networks to be executed simultaneously and / or sequentially by providing a CPU complex, a GPU complex, and a hardware acceleration cluster, and combining the results to achieve level 3 - 5 autonomous driving functions. For example, a CNN executed on the DLA or dGPU (such as GPU 520) may include text and word recognition, allowing a supercomputer to read and understand traffic signs, including signs for which the neural network has not been specifically trained. The DLA may further include a neural network capable of recognizing, interpreting, and providing semantic understanding of the sign and passing that semantic understanding to a path planning module running on the CPU complex.
[0227] As another example, multiple neural networks can operate simultaneously, as required for level 3, 4, or 5 driving. For example, a warning sign consisting of "Caution: Flashing lights indicate icy conditions" together with the electric lights can be interpreted independently or jointly by several neural networks. The sign itself can be recognized as a traffic sign by a first neural network (e.g., a trained neural network) deployed, and the text "Flashing lights indicate icy conditions" can be interpreted by a second neural network deployed, which informs the vehicle's path planning software (preferably executed on a CPU complex) that when the flashing lights are detected, there are icy conditions. The flashing lights can be recognized by operating a third neural network deployed on multiple frames, which informs the vehicle's path planning software of the presence (or absence) of the flashing lights. All three neural networks can operate simultaneously, for example, within the DLA and / or on the GPU 508.
[0228] In some examples, the CNNs for face recognition and owner recognition can use data from the camera sensors to identify the presence of an authorized driver and / or owner of the vehicle 500. A processing engine that is always on the sensor can be used to unlock the vehicle and turn on the lights when the owner approaches the driver's door, and in a security mode, to disable the vehicle when the owner leaves the vehicle. In this way, the SoC 504 provides security against theft and / or carjacking.
[0229] In another example, the CNN for emergency vehicle detection and recognition can use data from the microphone 596 to detect and recognize an emergency vehicle siren. In contrast to conventional systems that use a general classifier to detect the siren and manually extract features, the SoC 504 uses the CNN to classify environmental and urban sounds as well as visual data. In a preferred embodiment, the CNN running on the DLA is trained to recognize the relative closing rate of an emergency vehicle (e.g., by using the Doppler effect). The CNN can also be trained to recognize emergency vehicles specific to the local area in which the vehicle is operating, as recognized by the GNSS sensor 558. Thus, for example, when operating in Europe, the CNN will seek to detect European sirens, and when in the United States, the CNN will seek to identify only North American sirens. Once an emergency vehicle is detected, with the assistance of the ultrasonic sensor 562, a control program can be used to execute an emergency vehicle safety routine to slow down the vehicle, pull over to the side of the road, stop the vehicle, and / or idle the vehicle until the emergency vehicle passes.
[0230] The vehicle may include a CPU 518 (e.g., a discrete CPU or dCPU) that may be coupled to the SoC 504 via a high-speed interconnect (e.g., PCIe). The CPU 518 may include, for example, an X86 processor. The CPU 518 may be used to perform any of a variety of functions, including, for example, arbitrating potentially inconsistent results between ADAS sensors and the SoC 504, and / or monitoring the status and health of the controller 536 and / or the infotainment SoC 530.
[0231] The vehicle 500 may include a GPU 520 (e.g., a discrete GPU or dGPU) that may be coupled to the SoC 504 via a high-speed interconnect (e.g., NVIDIA's NVLINK). The GPU 520 may provide additional artificial intelligence capabilities, for example, by executing redundant and / or different neural networks, and may be used to train and / or update neural networks based on inputs (e.g., sensor data) from the sensors of the vehicle 500.
[0232] The vehicle 500 may further include a network interface 524 that may include one or more wireless antennas 526 (e.g., one or more wireless antennas for different communication protocols, such as cellular antennas, Bluetooth antennas, etc.). The network interface 524 may be used to enable wireless connections to the cloud (e.g., to the server 578 and / or other network devices), to other vehicles, and / or to computing devices (e.g., the client devices of passengers) via the Internet. To communicate with other vehicles, a direct link may be established between the two vehicles, and / or an indirect link may be established (e.g., across a network and via the Internet). The direct link may be provided using a vehicle-to-vehicle communication link. The vehicle-to-vehicle communication link may provide the vehicle 500 with information about vehicles approaching the vehicle 500 (e.g., vehicles in front of, to the side of, and / or behind the vehicle 500). This function may be part of the cooperative adaptive cruise control function of the vehicle 500.
[0233] The network interface 524 may include an SoC that provides modulation and demodulation functions and enables the controller 536 to communicate via a wireless network. The network interface 524 may include a radio frequency front end for upconversion from baseband to radio frequency and downconversion from radio frequency to baseband. The frequency conversion may be performed by a known process and / or may be performed using a super-heterodyne process. In some examples, the radio frequency front end functions may be provided by a separate chip. The network interface may include wireless capabilities for communicating via LTE, WCDMA, UMTS, GSM, CDMA2000, Bluetooth, Bluetooth LE, Wi-Fi, Z-Wave, ZigBee, LoRaWAN, and / or other wireless protocols.
[0234] Vehicle 500 may further include a data store 528 that may include off-chip (e.g., outside of SoC 504) storage devices. The data store 528 may include one or more storage elements, including RAM, SRAM, DRAM, VRAM, flash memory, hard drives, and / or other components and / or devices that can store at least one bit of data.
[0235] Vehicle 500 may further include a GNSS sensor 558. The GNSS sensor 558 (e.g., GPS, assisted GPS sensor, differential GPS (DGPS) sensor, etc.) is used to assist mapping, perception, occupancy grid generation, and / or path planning functions. Any number of GNSS sensors 558 may be used, including, for example and without limitation, a GPS using a USB connector with an Ethernet to serial (RS-232) bridge.
[0236] Vehicle 500 may further include a RADAR sensor 560. The RADAR sensor 560 may be used by vehicle 500 for remote vehicle detection even in dark and / or adverse weather conditions. The RADAR functional safety level may be ASIL B. The RADAR sensor 560 may use CAN and / or bus 502 (e.g., to transmit data generated by the RADAR sensor 560) for control as well as access to object tracking data, and in some examples accesses Ethernet to access raw data. A variety of RADAR sensor types may be used. For example and without limitation, the RADAR sensor 560 may be suitable for front, rear, and side RADAR use. In some examples, a pulsed Doppler RADAR sensor is used.
[0237] The RADAR sensor 560 may include different configurations, such as long-range with a narrow field of view, short-range with a wide field of view, short-range side coverage, and so on. In some examples, long-range RADAR may be used for adaptive cruise control functions. The long-range RADAR system may provide a wide field of view (e.g., within 250 m) achieved through two or more independent scans. The RADAR sensor 560 may help distinguish between static and moving objects and may be used by the ADAS system for emergency braking assistance and forward collision warning. The long-range RADAR sensor may include a single station multimode RADAR with multiple (e.g., six or more) fixed RADAR antennas and high-speed CAN and FlexRay interfaces. In an example with six antennas, the central four antennas may create a focused beam pattern that is designed to record the surroundings of vehicle 500 at a higher rate with minimal traffic interference from adjacent lanes. The other two antennas may extend the field of view, making it possible to quickly detect vehicles entering or leaving the lane of vehicle 500.
[0238] As an example, a mid-range RADAR system can include a range of up to 160 m (front) or 80 m (rear) and a field of view of up to 42 degrees (front) or 150 degrees (rear). A short-range RADAR system can include, but is not limited to, RADAR sensors designed to be mounted at both ends of the rear bumper. When mounted at both ends of the rear bumper, such a RADAR sensor system can create two beams that continuously monitor the rear and the blind spots alongside the vehicle.
[0239] The short-range RADAR system can be used in an ADAS system for blind spot detection and / or lane change assistance.
[0240] Vehicle 500 can further include ultrasonic sensors 562. Ultrasonic sensors 562 that can be placed in the front, rear, and / or sides of vehicle 500 can be used for parking assistance and / or creating and updating an occupancy grid. A variety of ultrasonic sensors 562 can be used, and different ultrasonic sensors 562 can be used for different detection ranges (e.g., 2.5 m, 4 m). Ultrasonic sensors 562 can operate at ASIL B for functional safety levels.
[0241] Vehicle 500 can include a LiDAR sensor 564. The LiDAR sensor 564 can be used for object and pedestrian detection, emergency braking, collision avoidance, and / or other functions. The LiDAR sensor 564 can be at ASIL B for functional safety levels. In some examples, vehicle 500 can include multiple LiDAR sensors 564 (e.g., two, four, six, etc.) that can use Ethernet (e.g., to provide data to a gigabit Ethernet switch).
[0242] In some examples, the LiDAR sensor 564 may be able to provide a list of objects and their distances for a 360-degree field of view. Commercially available LiDAR sensors 564 can have, for example, an advertised range of approximately 100 m, an accuracy of 2 cm - 3 cm, and support a 100 Mbps Ethernet connection. In some examples, one or more non-protruding LiDAR sensors 564 can be used. In such examples, the LiDAR sensor 564 can be implemented as a small device that can be embedded in the front, rear, sides, and / or corners of vehicle 500. In such examples, the LiDAR sensor 564 can provide a field of view of up to 120 degrees horizontally and 35 degrees vertically, even for low-reflectivity objects, with a range of 200 m. The front-mounted LiDAR sensor 564 can be configured for a horizontal field of view between 45 degrees and 135 degrees.
[0243] In some examples, LiDAR technologies such as 3D flash LiDAR can also be used. 3D flash LiDAR uses the flash of a laser as the emission source to illuminate the vehicle's surrounding environment up to about 200m. The flash LiDAR unit includes a receiver that records the laser pulse transit time and the reflected light on each pixel, which in turn corresponds to the range from the vehicle to the object. Flash LiDAR can allow for the generation of highly accurate and distortion-free images of the surrounding environment using each laser flash. In some examples, four flash LiDAR sensors can be deployed, one on each side of the vehicle 500. Available 3D flash LiDAR systems include solid-state 3D staring array LiDAR cameras (e.g., non-scanning LiDAR devices) that have no moving parts other than a fan. The flash LiDAR device can use Class I (eye-safe) laser pulses of 5 nanoseconds per frame and can capture the reflected laser in the form of 3D range point clouds and co-registered intensity data. By using flash LiDAR and because flash LiDAR is a solid-state device without moving parts, the LiDAR sensor 564 can be less susceptible to motion blur, vibration, and / or shock.
[0244] The vehicle can further include an IMU sensor 566. In some examples, the IMU sensor 566 can be located at the center of the rear axle of the vehicle 500. The IMU sensor 566 can include, for example and without limitation, accelerometers, magnetometers, gyroscopes, magnetic compasses, and / or other sensor types. In some examples, such as in a six-axis application, the IMU sensor 566 can include an accelerometer and a gyroscope, while in a nine-axis application, the IMU sensor 566 can include an accelerometer, a gyroscope, and a magnetometer.
[0245] In some embodiments, the IMU sensor 566 can be implemented as a miniature high-performance GPS-aided inertial navigation system (GPS / INS) that combines microelectromechanical system (MEMS) inertial sensors, a high-sensitivity GPS receiver, and an advanced Kalman filtering algorithm to provide estimates of position, velocity, and attitude. Thus, in some examples, the IMU sensor 566 can enable the vehicle 500 to estimate the heading without input from a magnetic sensor by directly observing the change in velocity from the GPS to the IMU sensor 566 and correlating them. In some examples, the IMU sensor 566 and the GNSS sensor 558 can be integrated into a single unit.
[0246] The vehicle can include a microphone 596 placed in and / or around the vehicle 500. Among other things, the microphone 596 can be used for emergency vehicle detection and identification.
[0247] The vehicle may further include any number of camera types, including a stereo camera 568, a wide-angle camera 570, an infrared camera 572, a surround camera 574, a long-range and / or mid-range camera 598, and / or other camera types. These cameras can be used to capture image data around the entire periphery of the vehicle 500. The camera types used depend on the embodiment and the requirements of the vehicle 500, and any combination of camera types can be used to provide the necessary coverage around the vehicle 500. Additionally, the number of cameras can vary according to the embodiment. For example, the vehicle can include six cameras, seven cameras, ten cameras, twelve cameras, and / or another number of cameras. As an example and without limitation, these cameras can support Gigabit Multimedia Serial Link (GMSL) and / or Gigabit Ethernet. Each of the cameras is described in more detail herein with respect to Figure 5A and Figure 5B is described in more detail.
[0248] The vehicle 500 may further include a vibration sensor 542. The vibration sensor 542 can measure the vibration of components of the vehicle such as the axles. For example, a change in vibration can indicate a change in the road surface. In another example, when two or more vibration sensors 542 are used, the difference between the vibrations can be used to determine the friction or slip of the road surface (e.g., when there is a vibration difference between a powered drive axle and a free-spinning axle).
[0249] The vehicle 500 may include an ADAS system 538. In some examples, the ADAS system 538 may include a SoC. The ADAS system 538 may include autonomous / adaptive / auto cruise control (ACC), cooperative adaptive cruise control (CACC), forward collision warning (FCW), automatic emergency braking (AEB), lane departure warning (LDW), lane keeping assist (LKA), blind spot warning (BSW), rear cross traffic warning (RCTW), collision warning system (CWS), lane centering (LC), and / or other features and functions.
[0250] The ACC system can use RADAR sensors 560, LiDAR sensors 564, and / or cameras. The ACC system can include longitudinal ACC and / or lateral ACC. Longitudinal ACC monitors and controls the distance to the vehicle immediately in front of the vehicle 500 and automatically adjusts the vehicle speed to maintain a safe distance from the vehicle ahead. Lateral ACC performs distance keeping and, when necessary, advises the vehicle 500 to change lanes. Lateral ACC is related to other ADAS applications such as LCA and CWS.
[0251] The CACC uses information from other vehicles, which can be received indirectly from other vehicles via the network interface 524 and / or the wireless antenna 526 via a wireless link or through a network connection (e.g., via the Internet). The direct link can be provided by a vehicle-to-vehicle (V2V) communication link, while the indirect link can be an infrastructure-to-vehicle (I2V) communication link. Generally, the V2V communication concept provides information about the vehicle immediately in front (e.g., a vehicle immediately in front of vehicle 500 and in the same lane as it), while the I2V communication concept provides information about traffic further ahead. The CACC system can include either or both of the I2V and V2V information sources. Given the information of the vehicle in front of vehicle 500, the CACC can be more reliable, and it has the potential to improve the smoothness of traffic flow and reduce road congestion.
[0252] The FCW system is designed to alert the driver of a hazard so that the driver can take corrective action. The FCW system uses a front camera and / or RADAR sensor 560 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback such as a display, speaker, and / or vibrating component. The FCW system can provide warnings in the form of, for example, sounds, visual warnings, vibrations, and / or rapid braking pulses.
[0253] The AEB system detects an impending front collision with another vehicle or other object and can automatically apply the brakes if the driver does not take corrective action within a specified time or distance parameter. The AEB system can use a front camera and / or RADAR sensor 560 coupled to a dedicated processor, DSP, FPGA, and / or ASIC. When the AEB system detects a hazard, it typically first alerts the driver to take corrective action to avoid the collision, and if the driver does not take corrective action, then the AEB system can automatically apply the brakes in an effort to prevent or at least mitigate the impact of the predicted collision. The AEB system can include technologies such as dynamic brake support and / or collision imminent braking.
[0254] The LDW system provides visual, auditory, and / or tactile warnings such as steering wheel or seat vibrations to alert the driver when vehicle 500 crosses a lane marking. The LDW system is not activated when the driver indicates an intentional lane departure by activating the turn signal. The LDW system can use a front-side facing camera coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback such as a display, speaker, and / or vibrating component.
[0255] The LKA system is a variant of the LDW system. If vehicle 500 starts to leave the lane, then the LKA system provides a steering input or braking to correct the vehicle 500. The BSW system detects and warns the driver of vehicles in the vehicle's blind spot. The BSW system can provide visual, audible, and / or tactile alerts to indicate that merging or changing lanes is unsafe. The system can provide additional warnings when the driver uses the turn signal. The BSW system can use a rear-facing camera and / or RADAR sensor coupled to a dedicated processor, DSP, FPGA, and / or ASIC.
[0256] The RCTW system can provide visual, audible, and / or tactile notifications when an object is detected outside the rear camera range while the vehicle 500 is in reverse. Some RCTW systems include AEB to ensure that the vehicle brakes are applied to avoid a crash. The RCTW system can use one or more rear RADAR sensors 560 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback such as a display, speaker, and / or vibrating component.
[0257] Conventional ADAS systems may be prone to false positive results, which can be annoying and distracting to the driver, but typically are not catastrophic because the ADAS system alerts the driver and allows the driver to decide whether a safe condition truly exists and act accordingly. However, in an autonomous vehicle 500, in the case of conflicting results, the vehicle 500 itself must decide whether to heed the result from the primary computer or the secondary computer (e.g., the first controller 536 or the second controller 536). For example, in some embodiments, the ADAS system 538 can be a backup and / or secondary computer for providing perception information to a backup computer sanity module. The backup computer sanity monitor can run redundant diverse software on hardware components to detect faults in perception and dynamic driving tasks. The output from the ADAS system 538 can be provided to the supervisory MCU. If the outputs from the primary computer and the secondary computer conflict, then the supervisory MCU must determine how to reconcile the conflict to ensure safe operation.
[0258] In some examples, the primary computer can be configured to provide a confidence score to the supervisory MCU, indicating the primary computer's confidence in the selected result. If the confidence score exceeds a threshold, then the supervisory MCU can follow the direction of the primary computer, regardless of whether the secondary computer provides a conflicting or inconsistent result. In the case where the confidence score does not meet the threshold and where the primary computer and the secondary computer indicate different results (e.g., conflict), the supervisory MCU can arbitrate between these computers to determine an appropriate result.
[0259] The supervisory MCU can be configured to run a neural network that is trained and configured to determine, based on outputs from the primary computer and the secondary computer, the conditions under which the secondary computer provides false alarms. Thus, the neural network in the supervisory MCU can learn when the output of the secondary computer can be trusted and when it cannot. For example, when the secondary computer is a RADAR-based FCW system, the neural network in the supervisory MCU can learn when the FCW system is identifying metal objects that are not actually dangerous, such as drain grates or manhole covers that trigger an alarm. Similarly, when the secondary computer is a camera-based LDW system, the neural network in the supervisory MCU can learn to ignore the LDW when a cyclist or pedestrian is present and lane departure is actually the safest strategy. In embodiments that include a neural network running on the supervisory MCU, the supervisory MCU can include at least one of a DLA or a GPU suitable for running the neural network with an associated memory. In a preferred embodiment, the supervisory MCU can include components of the SoC 504 and / or be included as a component of the SoC 504.
[0260] In other examples, the ADAS system 538 can include a secondary computer that performs ADAS functions using traditional computer vision rules. In this way, the secondary computer can use classical computer vision rules (if - then), and the presence of a neural network in the supervisory MCU can improve reliability, safety, and performance. For example, the diverse implementations and intentional non-identity make the overall system more fault-tolerant, especially for failures caused by software (or software-hardware interface) functions. For example, if there is a software vulnerability or error in the software running on the primary computer and the non-identical software code running on the secondary computer provides the same overall result, then the supervisory MCU can be more confident that the overall result is correct and that the vulnerability in the software or hardware on the primary computer does not cause a substantial error.
[0261] In some examples, the output of the ADAS system 538 can be fed to the perception block of the primary computer and / or the dynamic driving task block of the primary computer. For example, if the ADAS system 538 indicates a forward collision warning due to an object being immediately in front, the perception block can use this information when identifying the object. In other examples, the secondary computer can have its own neural network that is trained and thus reduces the risk of false positives as described herein.
[0262] Vehicle 500 may further include an infotainment SoC 530 (e.g., an in-vehicle infotainment (IVI) system). Although illustrated and described as an SoC, the infotainment system may not be an SoC and may include two or more discrete components. The infotainment SoC 530 may include a combination of hardware and software that can be used to provide audio (e.g., music, personal digital assistant, navigation instructions, news, radio, etc.), video (e.g., TV, movies, streaming, etc.), telephone (e.g., hands-free calling), network connectivity (e.g., LTE, Wi-Fi, etc.), and / or information services (e.g., navigation system, rear parking assistance, radio data system, vehicle-related information such as fuel level, total distance covered, brake fuel level, oil level, door open / close, air filter information, etc.) to the vehicle 500. For example, the infotainment SoC 530 may include a radio, a disc player, a navigation system, a video player, USB and Bluetooth connectivity, an in-vehicle computer, in-vehicle entertainment, Wi-Fi, steering wheel audio controls, hands-free voice controls, a head-up display (HUD), an HMI display 534, a telematics device, a control panel (e.g., for controlling various components, features, and / or systems, and / or interacting therewith), and / or other components. The infotainment SoC 530 may further be used to provide information (e.g., visual and / or auditory) to a user of the vehicle, such as information from the ADAS system 538, autonomous driving information such as planned vehicle maneuvers, trajectories, surrounding environment information (e.g., intersection information, vehicle information, road information, etc.), and / or other information.
[0263] The infotainment SoC 530 may include GPU functionality. The infotainment SoC 530 may communicate with other devices, systems, and / or components of the vehicle 500 via a bus 502 (e.g., a CAN bus, Ethernet, etc.). In some examples, the infotainment SoC 530 may be coupled to a supervisory MCU such that, in the event of a failure of the main controller 536 (e.g., the main and / or backup computer of the vehicle 500), the GPU of the infotainment system may perform some autonomous driving functions. In such examples, the infotainment SoC 530 may place the vehicle 500 in the driver safe parking mode as described herein.
[0264] Vehicle 500 may further include an instrument cluster 532 (e.g., a digital instrument panel, an electronic instrument cluster, a digital instrument panel, etc.). The instrument cluster 532 may include a controller and / or a supercomputer (e.g., a discrete controller or supercomputer). The instrument cluster 532 may include a set of instruments, such as a speedometer, fuel level, oil pressure, tachometer, odometer, turn indicator, shift position indicator, seat belt warning light, parking brake warning light, engine fault light, airbag (SRS) system information, lighting controls, safety system controls, navigation information, and so on. In some examples, information may be displayed and / or shared between the infotainment SoC 530 and the instrument cluster 532. In other words, the instrument cluster 532 may be included as part of the infotainment SoC 530, or vice versa.
[0265] Figure 5D A system schematic diagram of communication between a cloud-based server and Figure 5A an exemplary autonomous vehicle 500 according to some embodiments of the present disclosure. The system 576 may include a server 578, a network 590, and vehicles including the vehicle 500. The server 578 may include a plurality of GPUs 584(A)-584(H) (collectively referred to herein as GPUs 584), PCIe switches 582(A)-582(H) (collectively referred to herein as PCIe switches 582), and / or CPUs 580(A)-580(B) (collectively referred to herein as CPUs 580). The GPUs 584, CPUs 580, and PCIe switches may be interconnected by high-speed interconnects such as, for example, and without limitation, the NVLink interface 588 developed by NVIDIA and / or PCIe connections 586. In some examples, the GPUs 584 are connected via NVLink and / or an NVSwitch SoC, and the GPUs 584 and the PCIe switches 582 are connected via a PCIe interconnect. Although eight GPUs 584, two CPUs 580, and two PCIe switches are illustrated, this is not intended to be limiting. Depending on the embodiment, each of the servers 578 may include any number of GPUs 584, CPUs 580, and / or PCIe switches. For example, each of the servers 578 may include eight, sixteen, thirty-two, and / or more GPUs 584.
[0266] Server 578 can receive image data through network 590 and from a vehicle, the image data representing an image showing an unexpected or changed road condition such as a recently started road work. Server 578 can transmit neural network 592, updated neural network 592, and / or map information 594 through network 590 and to the vehicle, including information about traffic and road conditions. Updates to the map information 594 can include updates to the HD map 522, such as information about construction sites, potholes, curves, floods, or other obstacles. In some examples, the neural network 592, updated neural network 592, and / or map information 594 can be represented and / or generated based on data received from new training and / or from any number of vehicles in the environment and / or experience of training performed at a data center (e.g., using server 578 and / or other servers).
[0267] Server 578 can be used to train a machine learning model (e.g., a neural network) based on training data. The training data can be generated by vehicles, and / or can be generated in a simulation (e.g., using a game engine). In some examples, the training data is labeled (e.g., in cases where the neural network benefits from supervised learning) and / or undergoes other preprocessing, while in other examples, the training data is not labeled and / or preprocessed (e.g., in cases where the neural network does not require supervised learning). The training can be performed according to any one or more classes of machine learning techniques, including but not limited to the following classes: supervised training, semi-supervised training, unsupervised training, self-learning, reinforcement learning, joint learning, transfer learning, feature learning (including principal component and clustering analysis), multilinear subspace learning, manifold learning, representation learning (including alternative dictionary learning), rule-based machine learning, anomaly detection, and any variations or combinations thereof. Once the machine learning model is trained, the machine learning model can be used by the vehicle (e.g., transmitted to the vehicle through network 590), and / or the machine learning model can be used by server 578 to remotely monitor the vehicle.
[0268] In some examples, server 578 can receive data from a vehicle and apply the data to the latest real-time neural network for real-time intelligent inference. Server 578 can include a deep learning supercomputer powered by GPU 584 and / or a dedicated AI computer, such as DGX and DGX Station machines developed by NVIDIA. However, in some examples, server 578 can include a deep learning infrastructure of a data center powered only by a CPU.
[0269] The deep learning infrastructure of server 578 may be capable of fast real-time inference and can use this ability to evaluate and verify the health of the processors, software, and / or associated hardware in vehicle 500. For example, the deep learning infrastructure can receive periodic updates from vehicle 500, such as a sequence of images and / or objects located in the image sequence that vehicle 500 has identified (e.g., via computer vision and / or other machine learning object classification techniques). The deep learning infrastructure can run its own neural network to identify the objects and compare them with the objects identified by vehicle 500. If the results do not match and the infrastructure concludes that the AI in vehicle 500 has malfunctioned, then server 578 can transmit a signal to vehicle 500, instructing the fail-safe computer in vehicle 500 to take control, notify the passengers, and complete a safe parking operation.
[0270] For inference, server 578 can include a GPU 584 and one or more programmable inference accelerators (e.g., NVIDIA's TensorRT 3). The combination of a GPU-powered server and inference acceleration can enable real-time response. In other examples, such as when performance is less critical, a CPU, FPGA, and other processor-powered servers can be used for inference.
[0271] Example computing device
[0272] Figure 6 FIG. is a block diagram of an example computing device 600 suitable for implementing some embodiments of the present disclosure. Computing device 600 can include an interconnect system 602 that directly or indirectly couples the following devices: a memory 604, one or more central processing units (CPUs) 606, one or more graphics processing units (GPUs) 608, a communication interface 610, input / output (I / O) ports 612, input / output components 614, a power supply 616, one or more presentation components 618 (e.g., a display), and one or more logic units 620. In at least one embodiment, computing device 600 can include one or more virtual machines (VMs), and / or any of its components can include virtual components (e.g., virtual hardware components). For non-limiting examples, one or more GPUs 608 can include one or more vGPUs, one or more CPUs 606 can include one or more vCPUs, and / or one or more logic units 620 can include one or more virtual logic units. Thus, computing device 600 can include discrete components (e.g., a complete GPU dedicated to computing device 600), virtual components (e.g., a portion of a GPU dedicated to computing device 600), or a combination thereof.
[0273] Although Figure 6The respective blocks are shown as being connected via an interconnection system 602 having circuitry, but this is not intended to be limiting and is for clarity only. For example, in some embodiments, a rendering component 618 such as a display device may be considered an I / O component 614 (e.g., if the display is a touch screen). As another example, the CPU 606 and / or GPU 608 may include memory (e.g., memory 604 may represent a storage device in addition to the memory of the GPU 608, CPU 606, and / or other components). In other words, Figure 6 the computing devices are merely illustrative. No distinction is made between categories such as "workstations", "servers", "laptop computers", "desktop computers", "tablet computers", "client devices", "mobile devices", "handheld devices", "gaming consoles", "electronic control units (ECUs)", "virtual reality systems", and / or other device or system types because all of these are considered within the Figure 6 scope of the computing devices.
[0274] The interconnection system 602 may represent one or more links or buses, such as an address bus, a data bus, a control bus, or a combination thereof. The interconnection system 602 may include one or more types of links or buses, such as an Industry Standard Architecture (ISA) bus, an Extended Industry Standard Architecture (EISA) bus, a Video Electronics Standards Association (VESA) bus, a Peripheral Component Interconnect (PCI) bus, a Peripheral Component Interconnect Express (PCIe) bus, and / or another type of bus or link. In some embodiments, there are direct connections between components. As an example, the CPU 606 may be directly connected to the memory 604. Additionally, the CPU 606 may be directly connected to the GPU 608. In cases where there are direct or point-to-point connections between components, the interconnection system 602 may include a PCIe link to effect the connection. In these examples, a PCI bus need not be included in the computing device 600.
[0275] The memory 604 may include any of a variety of computer-readable media. Computer-readable media can be any available media that can be accessed by the computing device 600. Computer-readable media can include volatile and non-volatile media as well as removable and non-removable media. By way of example and not limitation, computer-readable media can include computer storage media and communication media.
[0276] Computer storage media can include volatile and non-volatile media and / or removable and non-removable media, implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules, and / or other data types. For example, memory 604 can store computer-readable instructions (e.g., which represent programs and / or program elements, such as an operating system). Computer storage media can include, but are not limited to, RAM, ROM, EEPROM, flash memory or other storage technologies, CD-ROM, digital versatile disks (DVDs) or other optical disk storage devices, magnetic tape cartridges, magnetic tape, magnetic disk storage devices or other magnetic storage devices, or any other medium that can be used to store the desired information and that can be accessed by computing device 600. As used herein, computer storage media does not include signals per se.
[0277] Computer storage media can include computer-readable instructions, data structures, program modules, and / or other data types in a modulated data signal, such as a carrier wave, or other transmission mechanism, and include any information conveyance medium. The term "modulated data signal" can refer to a signal that has one or more of its characteristics set or changed in such a manner as to encode information in the signal. By way of example, and not limitation, computer storage media can include wired media such as a wired network or direct wired connection, and wireless media such as sound, RF, infrared, and other wireless media. Any of the foregoing combinations should also be included within the scope of computer-readable media.
[0278] CPU 606 can be configured to execute at least some of the computer-readable instructions to control one or more components of computing device 600 to perform one or more of the methods and / or processes described herein. Each of the CPUs 606 can include one or more cores (e.g., one, two, four, eight, twenty-eight, seventy-two, etc.) capable of simultaneously processing a large number of software threads. CPU 606 can include any type of processor, and can include different types of processors depending on the type of computing device 600 being implemented (e.g., a processor with fewer cores for a mobile device and a processor with more cores for a server). For example, depending on the type of computing device 600, the processor can be an advanced RISC machine (ARM) processor implemented using reduced instruction set computing (RISC) or an x86 processor implemented using complex instruction set computing (CISC). In addition to one or more microprocessors or supplementary coprocessors such as a math coprocessor, computing device 600 can also include one or more CPUs 606.
[0279] In addition to or instead of the CPU 606, the GPU 608 can also be configured to execute at least some computer-readable instructions to control one or more components of the computing device 600 to perform one or more of the methods and / or processes described herein. One or more GPUs 608 can be an integrated GPU (e.g., having one or more CPUs 606) and / or one or more GPUs 608 can be a discrete GPU. In an embodiment, one or more GPUs 608 can be a coprocessor of one or more CPUs 606. The computing device 600 can use the GPU 608 to render graphics (e.g., 3D graphics) or perform general computing. For example, the GPU 608 can be used for general-purpose computing on the GPU (GPGPU). The GPU 608 can include hundreds or thousands of cores capable of processing hundreds or thousands of software threads simultaneously. The GPU 608 can generate pixel data for an output image in response to a rendering command (e.g., a rendering command received from the CPU 606 via a host interface). The GPU 608 can include graphics memory such as display memory for storing pixel data or any other suitable data (e.g., GPGPU data). The display memory can be included as part of the memory 604. The GPU 608 can include two or more GPUs operating in parallel (e.g., via a link). The link can directly connect the GPUs (e.g., using NVLINK) or can connect the GPUs through a switch (e.g., using NVSwitch). When combined, each GPU 608 can generate pixel data or GPGPU data for different portions of the output or for different outputs (e.g., the first GPU for the first image and the second GPU for the second image). Each GPU can include its own memory or can share memory with other GPUs.
[0280] In addition to or instead of the CPU 606 and / or the GPU 608, the logic unit 620 can be configured to execute at least some computer-readable instructions to control one or more components of the computing device 600 to perform one or more of the methods and / or processes described herein. In an embodiment, the CPU 606, the GPU 608, and / or the logic unit 620 can discretely or jointly execute any combination of methods, processes, and / or portions thereof. One or more logic units 620 can be part of and / or integrated into one or more CPUs 606 and / or one or more GPUs 608, and / or one or more logic units 620 can be discrete components of the CPU 606 and / or the GPU 608 or otherwise external to them. In an embodiment, one or more logic units 620 can be a processor of one or more CPUs 606 and / or one or more GPUs 608.
[0281] Examples of the logic unit 620 include one or more processing cores and / or their components, such as a data processing unit (DPU), a tensor core (TC), a tensor processing unit (TPU), a pixel vision core (PVC), a vision processing unit (VPU), a graphics processing cluster (GPC), a texture processing cluster (TPC), a streaming multiprocessor (SM), a tree traversal unit (TTU), an artificial intelligence accelerator (AIA), a deep learning accelerator (DLA), an arithmetic logic unit (ALU), an application specific integrated circuit (ASIC), a floating point unit (FPU), an input / output (I / O) component, a peripheral component interconnect (PCI) or a peripheral component interconnect express (PCIe) component, etc.
[0282] The communication interface 610 may include one or more receivers, transmitters, and / or transceivers that enable the computing device 600 to communicate with other computing devices via an electronic communication network, including wired and / or wireless communication. The communication interface 610 may include components and functions that enable communication over any of several different networks, such as wireless networks (e.g., Wi-Fi, Z-Wave, Bluetooth, Bluetooth LE, ZigBee, etc.), wired networks (e.g., communication via Ethernet or InfiniBand), low power wide area networks (e.g., LoRaWAN, SigFox, etc.), and / or the Internet. In one or more embodiments, one or more of the logic units 620 and / or the communication interface 610 may include one or more data processing units (DPUs) for directly transferring data received over the network and / or the interconnect system 602 to one or more GPUs 608 (e.g., the memory of the GPU 608).
[0283] The I / O port 612 can enable the computing device 600 to be logically coupled to other devices including I / O components 614, presentation components 618, and / or other components, some of which may be built into (e.g., integrated into) the computing device 600. Illustrative I / O components 614 include microphones, mice, keyboards, joysticks, game pads, game controllers, dish satellite antennas, scanners, printers, wireless devices, and the like. The I / O components 614 can provide a natural user interface (NUI) for processing user-generated air gestures, voice, or other physiological inputs. In some instances, the input can be transmitted to appropriate network elements for further processing. The NUI can implement any combination of speech recognition, stylus recognition, face recognition, biometric recognition, on-screen and near-screen gesture recognition, air gestures, head and eye tracking, and touch recognition associated with the display of the computing device 600 (as described in more detail in this disclosure). The computing device 600 can include depth cameras such as stereo camera systems, infrared camera systems, RGB camera systems, touch screen technologies, and combinations thereof for gesture detection and recognition. Additionally, the computing device 600 can include an accelerometer or gyroscope enabling motion detection (e.g., as part of an inertial measurement unit (IMU)). In some examples, the output of the accelerometer or gyroscope can be used by the computing device 600 to render immersive augmented reality or virtual reality.
[0284] The power supply 616 can include a hard-wired power supply, a battery power supply, or a combination thereof. The power supply 616 can power the computing device 600 to enable the components of the computing device 600 to operate.
[0285] The presentation component 618 can include a display (e.g., a monitor, a touch screen, a television screen, a head-up display (HUD), other display types, or a combination thereof), speakers, and / or other presentation components. The presentation component 618 can receive data from other components (e.g., the GPU 608, the CPU 606, etc.) and output the data (e.g., as images, videos, sounds, etc.).
[0286] Example data center
[0287] Figure 7 An example data center 700 is shown, which can be used in at least one embodiment of this disclosure. The data center 700 can include a data center infrastructure layer 710, a framework layer 720, a software layer 730, and an application layer 740.
[0288] As Figure 7As shown, the data center infrastructure layer 710 may include a resource coordinator 712, grouped computing resources 714, and node computing resources ("node C.R.") 716(1)-716(N), where "N" represents any whole positive integer. In at least one embodiment, the node C.R. 716(1)-716(N) may include, but is not limited to, any number of central processing units ("CPU") or other processors (including DPU, accelerator, field programmable gate array (FPGA), graphics processor or graphics processing unit (GPU), etc.), memory devices (such as dynamic read-only memory), storage devices (such as solid state drives or disk drives), network input / output ("NW I / O") devices, network switches, virtual machines ("VM"), power modules, and cooling modules, etc. In some embodiments, one or more of the node C.R. 716(1)-716(N) may correspond to a server having one or more of the above computing resources. Additionally, in some embodiments, the node C.R. 716(1)-716(N) may include one or more virtual components, such as vGPU, vCPU, etc., and / or one or more of the node C.R. 716(1)-716(N) may correspond to a virtual machine (VM).
[0289] In at least one embodiment, the grouped computing resources 714 may include separate groupings (not shown) of the node C.R. 716 housed within one or more racks, or many racks (also not shown) within data centers at various geographical locations. The separate groupings of the node C.R. 716 within the grouped computing resources 714 may include grouped computing, network, memory, or storage resources that may be configured or allocated to support one or more workloads. In at least one embodiment, several node C.R. 716 including CPU, GPU, DPU, and / or other processors may be grouped within one or more racks to provide computing resources to support one or more workloads. One or more racks may also include any number of power modules, cooling modules, and / or network switches in any combination.
[0290] The resource coordinator 712 may configure or otherwise control one or more of the node C.R. 716(1)-716(N) and / or the grouped computing resources 714. In at least one embodiment, the resource coordinator 712 may include a software design infrastructure ("SDI") management entity for the data center 700. The resource coordinator 712 may include hardware, software, or some combination thereof.
[0291] In at least one embodiment, as Figure 7As shown, the framework layer 720 may include a job scheduler 733, a configuration manager 734, a resource manager 736, and a distributed file system 738. The framework layer 720 may include a framework for software 732 that supports the software layer 730 and / or one or more applications 742 of the application layer 740. The software 732 or the application 742 may respectively include web-based service software or applications, such as service software or applications provided by Amazon Web Services, Google Cloud, and Microsoft Azure. The framework layer 720 may be, but is not limited to, a free and open-source software web application framework, such as Apache Spark that can utilize the distributed file system 738 for large-scale data processing (e.g., "big data"). TM (hereinafter referred to as "Spark"). In at least one embodiment, the job scheduler 733 may include a Spark driver for facilitating the scheduling of workloads supported by the various layers of the data center 700. In at least one embodiment, the configuration manager 734 may be able to configure different layers, such as the software layer 730 and the framework layer 720 including Spark and the distributed file system 738 for supporting large-scale data processing. The resource manager 736 is capable of managing the cluster or grouped computing resources mapped to or allocated for supporting the distributed file system 738 and the job scheduler 733. In at least one embodiment, the cluster or grouped computing resources may include grouped computing resources 714 at the data center infrastructure layer 710. The resource manager 736 may coordinate with the resource coordinator 712 to manage these mapped or allocated computing resources.
[0292] In at least one embodiment, the software 732 included in the software layer 730 may include software used by at least a portion of the nodes C.R. 716(1)-716(N), the grouped computing resources 714, and / or the distributed file system 738 of the framework layer 720. One or more types of software may include, but are not limited to, Internet web search software, email virus browsing software, database software, and streaming video content software.
[0293] In at least one embodiment, one or more applications 742 included in the application layer 740 may include one or more types of applications used by at least a portion of nodes C.R. 716(1)-716(N), the grouped computing resources 714, and / or the distributed file system 738 of the framework layer 720. The one or more types of applications may include, but are not limited to, any number of genomics applications, cognitive computing and machine learning applications, including training or inference software, machine learning framework software (such as PyTorch, TensorFlow, Caffe, etc.), and / or other machine learning applications used in conjunction with one or more embodiments.
[0294] In at least one embodiment, any one of the configuration manager 734, the resource manager 736, and the resource coordinator 712 may implement any number and type of self-modifying actions based on any amount and type of data obtained in any technically feasible manner. The self-modifying actions may relieve the data center operator of the data center 700 from making potentially bad configuration decisions and may avoid underutilized and / or poorly performing portions of the data center.
[0295] The data center 700 may include tools, services, software, or other resources for training one or more machine learning models or using one or more machine learning models to predict or infer information in accordance with one or more embodiments described herein. For example, a machine learning model may be trained by calculating weight parameters according to a neural network architecture by using the software and computing resources described in the present disclosure with respect to the data center 700. In at least one embodiment, by using the weight parameters calculated by one or more training techniques, the resources described in the present disclosure with respect to the data center 700 may be used to infer or predict information using the trained machine learning model corresponding to one or more neural networks, such as, but not limited to, those described herein.
[0296] In at least one embodiment, the data center 700 may use a CPU, an application specific integrated circuit (ASIC), a GPU, an FPGA, and / or other hardware (or corresponding virtual computing resources) to perform training and / or inference using the above resources. In addition, one or more software and / or hardware resources described in the present disclosure may be configured as a service to allow a user to train or perform information inference, such as image recognition, speech recognition, or other artificial intelligence services.
[0297] Example Network Environment
[0298] The network environment suitable for implementing the embodiments of the present disclosure may include one or more client devices, servers, network attached storage (NAS), other backend devices, and / or other device types. The client devices, servers, and / or other device types (e.g., each device) may be implemented on one or more instances of the Figure 6 computing device 600 - for example, each device may include similar components, features, and / or functions of the computing device 600. Additionally, in the case of implementing backend devices (e.g., servers, NAS, etc.), the backend devices may be included as part of a data center 700, an example of which is described in more detail herein with respect to Figure 7 more detail.
[0299] The components of the network environment may communicate with each other via a network, which may be wired, wireless, or both. The network may include multiple networks, or networks within multiple networks. By way of example, the network may include one or more wide area networks (WANs), one or more local area networks (LANs), one or more public networks (e.g., the Internet and / or the public switched telephone network (PSTN)), and / or one or more private networks. In the case where the network includes a wireless telecommunications network, components such as base stations, communication towers, or even access points (and other components) may provide a wireless connection.
[0300] A compatible network environment may include one or more peer - to - peer network environments (in which case servers may not be included in the network environment), and one or more client - server network environments (in which case one or more servers may be included in the network environment). In a peer - to - peer network environment, the functions described herein with respect to servers may be implemented on any number of client devices.
[0301] In at least one embodiment, the network environment may include one or more cloud - based network environments, distributed computing environments, combinations thereof, etc. A cloud - based network environment may include a framework layer, a job scheduler, a resource manager, and a distributed file system implemented on one or more servers, which may include one or more core network servers and / or edge servers. The framework layer may include a framework for supporting software layers and / or one or more applications of an application layer. The software or application may respectively include network - based service software or applications. In an embodiment, one or more client devices may use network - based service software or applications (e.g., by accessing the service software and / or applications via one or more application programming interfaces (APIs)). The framework layer may be, but is not limited to, a type of free and open - source software web application framework, for example, which may use a distributed file system for large - scale data processing (e.g., “big data”).
[0302] A cloud-based network environment can provide cloud computing and / or cloud storage that perform any combination of the computing and / or data storage functions (or one or more parts thereof) described herein. Any of these various functions can be distributed across multiple locations from a central or core server (e.g., one or more data centers that can be distributed across states, regions, countries, globally, etc.). If the connection to a user (e.g., a client device) is relatively close to an edge server, the core server can assign at least a portion of the function to the edge server. A cloud-based network environment can be private (e.g., limited to a single organization), public (e.g., available to many organizations), and / or a combination thereof (e.g., a hybrid cloud environment).
[0303] A client device can include at least some of the components, features, and functions of the example computing device 600 described herein with respect to Figure 6 As an example and not a limitation, a client device can be embodied as a personal computer (PC), laptop computer, mobile device, smartphone, tablet computer, smartwatch, wearable computer, personal digital assistant (PDA), MP3 player, virtual reality headset, global positioning system (GPS) or device, video player, camera, surveillance device or system, vehicle, boat, aircraft, virtual machine, drone, robot, handheld communication device, hospital device, gaming device or system, entertainment system, in-vehicle computer system, embedded system controller, remote control, appliance, consumer electronic device, workstation, edge device, any combination of these described devices, or any other suitable device.
[0304] This disclosure can be described in the general context of machine-usable instructions or computer code, including computer-executable instructions such as program modules, executed by a computer or other machine such as a personal digital assistant or other handheld device. Generally, program modules, including routines, programs, objects, components, data structures, etc., refer to code that performs particular tasks or implements particular abstract data types. This disclosure can be practiced in a variety of system configurations, including handheld devices, consumer electronics, general-purpose computers, more specialized computing devices, etc. This disclosure can also be practiced in a distributed computing environment where tasks are performed by remote processing devices linked through a communication network.
[0305] As used herein, the recitation of "and / or" with respect to two or more elements should be construed to refer to only one element or a combination of elements. For example, "element A, element B, and / or element C" may include only element A, only element B, only element C, element A and element B, element A and element C, element B and element C, or elements A, B, and C. Further, "at least one of element A or element B" may include at least one of element A, at least one of element B, or at least one of element A and at least one of element B. Still further, "at least one of element A and element B" may include at least one of element A, at least one of element B, or at least one of element A and at least one of element B. Additionally, the use of the term "based on" should not be construed as "only based on" or "merely based on". Rather, the first element being "based on" the second element includes cases where the first element is based on the second element but may also be based on one or more additional elements.
[0306] The subject matter of the present disclosure is described in detail herein to meet statutory requirements. However, the description itself is not intended to limit the scope of the present disclosure. On the contrary, the inventors have contemplated that the claimed subject matter may also be embodied in other ways, including steps different from those described herein in connection with other current or future technologies or combinations of similar steps. Moreover, although the terms "step" and / or "block" may be used herein to imply different elements of a method employed, these terms should not be construed to imply any particular order among or between the various steps disclosed herein unless the order of the steps is explicitly described.
[0307] For example, the subject technology of the present invention is illustrated in accordance with various aspects described below. For convenience, each example of the various aspects of the subject technology is described as a numbered example (1, 2, 3, etc.). These are provided only as examples and do not limit the subject technology. Unless the context otherwise indicates, aspects of the various implementations described herein may be omitted, replaced with aspects of other implementations, or combined with aspects of other implementations. For example, one or more aspects of Example 1 below may be omitted, replaced with another example (e.g., Example 2) or one or more aspects of additional examples, or combined with aspects of another example. The following is a non-limiting overview of some example implementations presented herein.
[0308] Example 1. A method, comprising:
[0309] generating a visibility confidence model corresponding to an aggregated field of view of a plurality of sensors of a machine, the visibility confidence model indicating confidence levels of sensor data corresponding to respective sub-parts of the aggregated field of view; and
[0310] Perform one or more operations based at least on the visibility confidence model or the confidence levels corresponding to the respective sub - parts of the aggregated field of view.
[0311] According to the method included in Example 1, wherein the plurality of sensors includes sensors corresponding to one or more sensor modalities.
[0312] According to the method included in Example 1, wherein the respective confidence levels are determined based at least on: one or more faults or errors associated with an individual sensor among the plurality of sensors; one or more coarse - level degradations or obstructions; one or more fine - level degradations corresponding to the sensor data; or the presence of one or more occlusions in the sensor data corresponding to the respective sub - parts of the aggregated field of view.
[0313] According to the method included in Example 1, wherein the one or more coarse - level degradations or obstructions are determined based on: weather data; temperature data; time of day; or time of year.
[0314] According to the method included in Example 1, wherein the presence of the one or more occlusions is determined based at least on historical sensor data or map data corresponding to a map.
[0315] According to the method included in Example 1, further comprising: before performing the one or more operations,
[0316] In response to a query, send data corresponding to one or more of the confidence levels in the respective confidence levels based on the query.
[0317] Example 2. A method, comprising:
[0318] Generate one or more visibility confidence models corresponding to one or more fields of view, the one or more fields of view defining a potential spatial coverage of sensor data corresponding to respective sensors associated with a machine;
[0319] Fill one or more parts of the one or more visibility confidence models with confidence data, the confidence data indicating respective confidence levels of sensor data corresponding to the respective sub - parts of the one or more fields of view;
[0320] Aggregate the respective confidence levels of sensor data corresponding to the respective sub - parts of the one or more fields of view to generate a visibility confidence model corresponding to an aggregated field of view of the respective sensors associated with the machine; and
[0321] Use the machine and perform one or more operations based at least on the visibility confidence model.
[0322] According to the method included in Example 2, wherein the corresponding sensor includes one or more sensors of one or more sensor modalities.
[0323] According to the method included in Example 2, wherein the one or more visibility confidence models are generated based at least on:
[0324] One or more faults or errors associated with individual sensors in the corresponding sensor; one or more coarse-level degradations or obstructions; one or more fine-level degradations corresponding to the sensor data; or the presence of one or more occlusions in the sensor data corresponding to respective sub-parts of the aggregated field of view.
[0325] According to the method included in Example 2, wherein the one or more coarse-level degradations or obstructions are determined based at least on environmental conditions affecting substantially all of the sensor data corresponding to a particular sensor.
[0326] According to the method included in Example 2, wherein the presence of the one or more occlusions in the sensor data is determined based at least on historical sensor data or map data corresponding to a map.
[0327] According to the method included in Example 2, further comprising:
[0328] In response to a query, sending data corresponding to one or more of the corresponding confidence levels based on the query.
[0329] According to the method included in Example 2, wherein the query is generated based at least on determining that the perception results corresponding to at least two sensors are inconsistent.
[0330] Example 3. A system, comprising:
[0331] One or more processors, which include processing circuitry for performing operations, the operations including:
[0332] Generating one or more visibility confidence models corresponding to one or more fields of view of one or more sensors associated with a machine;
[0333] Using one or more perception models to generate one or more perception outputs;
[0334] Determining one or more confidences associated with the one or more perception outputs based at least on querying the one or more visibility confidence models in view of the one or more perception outputs; and
[0335] Perform one or more operations using the machine based at least on the one or more confidences and the one or more perception outputs.
[0336] According to the system included in Example 3, wherein the one or more visibility confidence models are stored using a data structure, and a query corresponding to the query corresponds to the data structure.
[0337] According to the system included in Example 3, wherein the one or more visibility confidence models are generated at least based on: one or more faults or errors associated with an individual sensor in the respective sensor; one or more coarse-level degradations or obstructions; one or more fine-level degradations corresponding to sensor data; or the presence of one or more occlusions in the sensor data corresponding to respective sub-parts of the aggregated field of view.
[0338] According to the system included in Example 3, wherein the one or more coarse-level degradations or obstructions are determined at least based on environmental conditions affecting substantially all of the sensor data corresponding to a particular sensor.
[0339] According to the system included in Example 3, wherein the presence of the one or more occlusions in the sensor data is determined at least based on historical sensor data or map data corresponding to a map.
[0340] According to the system included in Example 3, wherein the one or more sensors include sensors corresponding to multiple sensor modalities.
[0341] According to the system included in Example 3, wherein the system is included in at least one of the following:
[0342] A control system for an autonomous or semi-autonomous machine;
[0343] A perception system for an autonomous or semi-autonomous machine;
[0344] A system for performing simulation operations;
[0345] A system for performing digital twin operations;
[0346] A system for performing optical transmission simulation;
[0347] A system for performing collaborative content creation of 3D assets;
[0348] A system for performing deep learning operations;
[0349] A system for presenting at least one of augmented reality content, virtual reality content, or mixed reality content;
[0350] Systems for hosting one or more real-time streaming applications;
[0351] Systems implemented using edge devices;
[0352] Systems implemented using robots;
[0353] Systems for performing conversational AI operations;
[0354] Systems for performing one or more generative AI operations;
[0355] Systems implementing one or more large language models (LLMs);
[0356] Systems for generating synthetic data;
[0357] Systems containing one or more virtual machines (VMs);
[0358] Systems implemented at least in part in a data center; or
[0359] Systems implemented at least in part using cloud computing resources.
Claims
1. A method comprising: generating a visibility confidence model corresponding to an aggregated field of view of a plurality of sensors of a machine, the visibility confidence model indicating confidence levels of sensor data corresponding to respective sub-portions of the aggregated field of view; as well as One or more operations are performed based at least on the visibility confidence model or the confidence levels corresponding to the respective sub-portions of the aggregated field of view. 2 . The method of claim 1 , wherein the plurality of sensors comprises sensors corresponding to one or more sensor modalities.
3. The method of claim 1 , wherein the corresponding confidence level is determined based on at least: one or more faults or errors associated with individual sensors of the plurality of sensors; One or more coarse levels are degraded or blocked; one or more fine-scale degradations corresponding to the sensor data; or One or more occlusions are present in the sensor data corresponding to the respective sub-portions of the aggregate field of view.
4. The method of claim 3, wherein the one or more coarse level degradations or blockages are determined based on: Weather data; Temperature data; time of day; or Time of year. The method of claim 3 , wherein the presence of the one or more occlusions is determined based on at least historical sensor data or map data corresponding to a map.
6. The method according to claim 1, further comprising: Before performing the one or more operations, In response to a query, data corresponding to one or more of the respective confidence levels is sent based on the query.
7. A method comprising: generating one or more visibility confidence models corresponding to one or more fields of view defining potential spatial coverage of sensor data corresponding to respective sensors associated with the machine; populating one or more portions of the one or more visibility confidence models with confidence data indicating respective confidence levels of sensor data corresponding to respective sub-portions of the one or more fields of view; aggregating respective confidence levels of sensor data corresponding to the respective sub-portions of the one or more fields of view to generate a visibility confidence model corresponding to an aggregated field of view of respective sensors associated with the machine; as well as One or more operations are performed using the machine and based at least on the visibility confidence model. The method of claim 7 , wherein the corresponding sensors comprise one or more sensors of one or more sensor modalities.
9. The method of claim 7, wherein the one or more visibility confidence models are generated based on at least: one or more faults or errors associated with individual ones of the corresponding sensors; One or more coarse levels are degraded or blocked; one or more fine-scale degradations corresponding to the sensor data; or One or more occlusions are present in the sensor data corresponding to the respective sub-portions of the aggregate field of view.
10. The method of claim 9, wherein the one or more coarse-level degradations or blockages are determined based at least on environmental conditions that affect substantially all of the sensor data corresponding to a particular sensor. 11 . The method of claim 9 , wherein the presence of the one or more occlusions in the sensor data is determined based on at least historical sensor data or map data corresponding to a map.
12. The method according to claim 7, further comprising: In response to a query, data corresponding to one or more of the corresponding confidence levels is sent based on the query.
13. The method of claim 12, wherein the query is generated based at least on determining that perception results corresponding to at least two sensors are inconsistent.
14. A system comprising: One or more processors comprising processing circuitry configured to perform operations comprising: generating one or more visibility confidence models corresponding to one or more fields of view of one or more sensors associated with the machine; generating one or more perceptual outputs using the one or more perceptual models; determining one or more confidences associated with the one or more perceptual outputs based at least on querying the one or more visibility confidence models in view of the one or more perceptual outputs; and One or more operations are performed using the machine based at least on the one or more confidence levels and the one or more perception outputs. 15 . The system of claim 14 , wherein the one or more visibility confidence models are stored using a data structure, and the query corresponding to the query corresponds to the data structure.
16. The system of claim 14, wherein the one or more visibility confidence models are generated based on at least: one or more faults or errors associated with individual ones of the corresponding sensors; One or more coarse levels are degraded or blocked; one or more fine-scale degradations corresponding to the sensor data; or One or more occlusions are present in the sensor data corresponding to respective sub-portions of the aggregate field of view.
17. The system of claim 16, wherein the one or more coarse-level degradations or blockages are determined based at least on environmental conditions that affect substantially all of the sensor data corresponding to a particular sensor.
18. The system of claim 16, wherein the presence of the one or more occlusions in the sensor data is determined based on at least historical sensor data or map data corresponding to a map.
19. The system of claim 14, wherein the one or more sensors include sensors corresponding to multiple sensor modalities.
20. The system of claim 14, wherein the system is included in at least one of the following: control systems for autonomous or semi-autonomous machines; Perception systems for autonomous or semi-autonomous machines; A system for performing simulation operations; Systems for performing digital twin operations; A system for performing light transport simulations; A system for performing collaborative content creation of 3D assets; Systems for performing deep learning operations; a system for presenting at least one of augmented reality content, virtual reality content, or mixed reality content; A system for hosting one or more real-time streaming applications; Systems implemented using edge devices; Systems implemented using robots; Systems for performing conversational AI operations; A system for performing one or more generative AI operations; A system implementing one or more large language models (LLMs); Systems for generating synthetic data; A system comprising one or more virtual machines VM; A system implemented at least in part in a data center; or A system implemented at least in part using cloud computing resources.
Citation Information
Patent Citations
Method for programmable timeouts of tree traversal mechanisms in hardware
US10885698B2
Visibility distance estimation using deep learning in autonomous machine applications
US20230110027A1