Deep learning-based operational domain verification for autonomous systems and applications using camera-based input.

CN116229390BActive Publication Date: 2026-09-18NVIDIA CORP
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202211106166.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-12-02
Filing Date
2022-09-09
Publication Date
2026-09-18
Estimated Expiration
2042-09-09

AI Technical Summary

Technical Problem

因此,这些专用传感器可能无法区分照明实例(例如,路灯引起的照明与室内停车场光照引起的照明)和不同的路面状况(例如,当道路下雪时与道路无雪但机器的(一个或更多个)传感器上却积了雪时)

Benefits of technology

[0007] Compared to the conventional systems described above, the systems and methods of this disclosure can acquire operational design domain conditions through sensor data, thereby providing accurate and rich information about these conditions similar to human perception. Furthermore, since the systems and methods of this disclosure apply multi-task deep learning, they may require fewer computational resources compared to conventional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116229390B_ABST
    Figure CN116229390B_ABST
Patent Text Reader

Abstract

This application relates to deep learning-based operational domain verification for autonomous systems and applications using camera-based input. In various examples, methods and systems are provided for using machine learning models to determine one or more of the following operational domain conditions relevant to autonomous and / or semi-autonomous machines: camera blindness quantity, blindness classification, lighting level, road surface condition, visibility distance, scene type classification, and distance to the scene. Once one or more of these conditions are determined, the operational level of the machine can be determined, and the machine can be controlled according to the operational level.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications This application relates to U.S. Nonprovisional Application No. 16 / 570,187, filed September 13, 2019, and U.S. Nonprovisional Application No. 17 / 449,306, filed September 29, 2021, the entire contents of which are incorporated herein by reference. Background Technology

[0002] Autonomous and semi-autonomous driving systems (e.g., Advanced Driver Assistance Systems (ADAS)) can use sensors (e.g., cameras, lidar sensors, radar sensors, etc.) to make decisions regarding the performance of various tasks, such as increasing, decreasing, or maintaining a certain speed; blind spot monitoring; automatic emergency braking; lane keeping; object detection; obstacle avoidance; lane changing; lane assignment; camera calibration; and localization. To determine whether and to what extent these autonomous and semi-autonomous systems should operate autonomously, an accurate understanding of the impact of the surrounding environment on these systems can be beneficial. However, the ability of sensors to perceive their surroundings can be affected by various sources, such as weather (e.g., rain, fog, snow, hail, smoke, etc.), traffic conditions, sensor congestion (e.g., from debris, precipitation, etc.), or blurring, which can lead to sensor blindness and / or impaired visibility. Potential causes of sensor blindness and / or impaired visibility may include snow, rain, glare, solar flares, mud, water, signal failure, etc.

[0003] Conventional systems for addressing sensor blindness and impaired visibility have employed feature-level methods to detect individual visual evidence of visibility impairment and sensor blindness, then stitch these features together to determine the presence of sensor blindness and / or impaired visibility events. These conventional methods rely heavily on computer vision techniques, such as detecting potential visibility and sensor blindness problems by analyzing the absence of sharp-edge features (e.g., abrupt changes in gradient, color, intensity) in image regions, using color-based pixel analysis or other low-level feature analyses, and / or binary support vector machine classification with blind and non-blind outputs. However, such feature-based computer vision techniques require individual analysis of each feature—e.g., whether each feature is relevant to visibility and sensor blindness—and analysis of how to combine different features for specific sensor-reduced visibility or blindness conditions. This limits the scalability of such methods due to the inherent complexity of the variability and diversity of conditions and events that can impair data observed using sensors in real-world scenarios. For example, they may be less efficient for real-time or near-real-time deployments due to the computational cost of performing these conventional methods.

[0004] Furthermore, conventional systems may rely on classifying the causes of reduced sensor visibility, such as rain, snow, fog, glare, etc., but may not provide an accurate indication of sensor data availability. For example, a system used to determine whether a corresponding image or a portion thereof is usable for various autonomous or semi-autonomous tasks may fail to recognize rain in an image. In such an example, if it is raining, even if the image clearly depicts the environment within 100 meters of the vehicle, a conventional system may consider the image unusable. Thus, instead of relying on one or more tasks that depend on the image being within the visible range, the image may be incorrectly discarded and one or more tasks may be disabled. Conventional systems may also fail to distinguish between different types of sensor blindness, such as whether an image is blurred or occluded. By treating each type of impaired sensor visibility and sensor blindness with equal or approximately equal weight, less severe or harmful types of impaired sensor visibility and sensor blindness may cause instances of sensor data to be considered unusable, even if the determination may not be entirely accurate (e.g., an image of an environment with light rain may be usable for one or more operations, while an image of an environment with dense fog may be unusable; a blurred image may be usable for some operations, while an occluded image may be unusable).

[0005] Furthermore, by relying on hard-coded computer vision techniques, conventional systems may fail to learn from historical data or learn over time during deployment, thus limiting their ability to adapt to new or sensor-blind situations. Additionally, conventional systems for detecting ambient lighting levels and road surface conditions may rely on specialized sensors, such as lux sensors for measuring light intensity, and / or road surface moisture sensors placed next to tires. Due to the lack of rich data input, these specialized sensors may output low-accuracy and limited data on the surrounding environment. Therefore, these specialized sensors may fail to distinguish between instances of lighting (e.g., illumination from streetlights versus illumination from indoor parking lot lighting) and different road surface conditions (e.g., when the road is snowing versus when the road is snow-free but snow has accumulated on one or more of the machine's sensors). Summary of the Invention

[0006] Embodiments of this disclosure relate to deep learning-based operational domain verification for autonomous systems and applications using camera-based input. Systems and methods are disclosed that use one or more machine learning models (e.g., deep neural networks (DNNs)) to compute outputs indicating one or more of the following operational design domain conditions associated with autonomous and / or semi-autonomous machines: camera blindness quantity; blindness classification corresponding to camera blindness; illumination level; road surface condition; visibility distance; scene type classification; and distance to the scene corresponding to the scene type classification. These outputs can then be used to determine the level of autonomy, such as operational grade, of the autonomous and / or semi-autonomous machine, which can then be used to control the machine in the environment.

[0007] Compared to the conventional systems described above, the systems and methods of this disclosure can acquire operational design domain conditions through sensor data, thereby providing accurate and rich information about these conditions similar to human perception. Furthermore, since the systems and methods of this disclosure apply multi-task deep learning, they may require fewer computational resources compared to conventional methods. Attached Figure Description

[0008] The following describes in detail, with reference to the accompanying drawings, the system and method for deep learning-based operation design domain verification for autonomous machine applications using camera-based input, wherein: Figure 1A This is an example data flow diagram illustrating, according to some embodiments of the present disclosure, an example process for determining lighting levels, road surface conditions, visibility distance, scene classification, distance to the scene, and sensor blindness and its causes. Figure 1B This is an example machine learning model based on some embodiments of the present disclosure for implementing example processes for determining lighting levels, road surface conditions, visibility distance, scene classification, distance to a scene, and sensor blindness and its causes; Figure 2A-2B This includes example visualizations of sensor data at different visibility distance levels, according to some embodiments of this disclosure; Figure 3 Includes example visualizations of sensor data including various sensor types of visual blindness according to some embodiments of this disclosure; Figures 4A-4B Includes example visualizations of sensor data of various scene types according to some embodiments of this disclosure; Figure 5 This is a flowchart illustrating, according to some embodiments of the present disclosure, a method for using a neural network to determine lighting levels, road surface conditions, visibility distance, scene classification, distance to a scene, and sensor blindness and its causes; Figure 6AThese are illustrations of exemplary autonomous vehicles according to some embodiments of the present disclosure; Figure 6B It is according to some embodiments of this disclosure for use Figure 6A An example of the camera position and field of view of an exemplary autonomous vehicle; Figure 6C It is according to some embodiments of this disclosure for use Figure 6A A block diagram of an exemplary system architecture for an exemplary autonomous vehicle; Figure 6D This is according to some embodiments of the present disclosure for use in one or more cloud-based servers and Figure 6A An exemplary system diagram of communication between autonomous vehicles; Figure 7 This is a block diagram of an example computing device applicable to implementing some embodiments of this disclosure; and Figure 8 This is a block diagram of an example data center applicable to implementing some embodiments of this disclosure. Detailed Implementation

[0009] Systems and methods related to deep learning-based operational domain verification for autonomous systems and applications using camera-based input are disclosed. The systems and methods described herein can be used in augmented reality, virtual reality, mixed reality, robotics, security and surveillance, medical imaging, autonomous or semi-autonomous machine applications and / or any other technological space, where it is possible to determine lighting levels, road surface conditions, visibility distance, scene classification, distance to the scene, and sensor blindness and its causes. Although this disclosure may relate to an exemplary autonomous vehicle 600 (which may alternatively be referred to herein as "vehicle 600" or "this machine 600"), examples are directed to... Figures 6A-6D The description is provided below, but is not intended to be limiting. For example, the systems and methods described herein can be used by, but are not limited to, non-autonomous vehicles, semi-autonomous vehicles (e.g., in one or more advanced driver assistance systems (ADAS)), manned and unmanned robots or robotic platforms, warehouse vehicles, off-road vehicles, vehicles connected to one or more trailers, airships, boats, shuttles, emergency response vehicles, motorcycles, electric or motorized bicycles, aircraft, construction vehicles, underwater vehicles, drones, and / or other vehicle types.

[0010] In the methods of this disclosure, a processor including one or more circuits may use a neural network and compute data representing operational domain conditions based at least in part on image data generated using one or more image sensors. In some embodiments, operational domain conditions may include one or more of the following: camera blindness quantity, blindness classification corresponding to the camera blindness quantity, illumination level, road surface condition, visibility distance, scene type classification, and distance to the scene corresponding to the scene type classification.

[0011] Deep neural network (DNN) processing can be used to determine camera blindness quantity, the blindness classification corresponding to the camera blindness quantity, lighting level, road surface condition, visibility distance, scene type classification, and distance to the scene corresponding to the scene type classification. Sensor data (e.g., representing one or more images) can be received from one or more sensors (e.g., cameras) deployed on or otherwise associated with a vehicle. The sensor data can be applied to a neural network (e.g., a DNN, such as a convolutional neural network (CNN)) that can be trained to determine camera blindness quantity, lighting level, road surface condition, visibility distance, scene type classification, and distance to the scene corresponding to the scene type classification associated with the sensor data. Methods and systems for determining camera blindness quantity, the corresponding blindness classification, and visibility distance may be identical or similar to those described in U.S. Non-Provisional Application No. 16 / 570,187, filed September 13, 2019, and U.S. Non-Provisional Application No. 17 / 449,306, filed September 29, 2021, the entire contents of which are incorporated herein by reference.

[0012] Based at least in part on this sensor data, the appropriate operating level for the machine in the current domain or environment can be determined. In an embodiment, the appropriate operating level may correspond to one or more automation levels defined by the Society of Automotive Engineers (SAE), such as Level 0 (non-automation), Level 1 (driver assistance), Level 2 (partial automation), Level 3 (conditional automation), Level 4 (high automation), and Level 5 (full automation).

[0013] The embodiments of this disclosure can be included in a variety of different systems, such as automotive systems (e.g., control systems for autonomous or semi-autonomous machines, perception systems for autonomous or semi-autonomous machines), systems implemented using robots, aerial systems, medical systems, boating systems, intelligent area monitoring systems, systems for performing deep learning operations, systems for performing simulation operations, systems implemented using edge devices, systems containing one or more virtual machines (VMs), systems implemented at least partially in a data center, systems implemented at least partially using cloud computing resources, and / or other types of systems. While specific examples are provided, these examples are generalizable beyond the implementation details.

[0014] Now for reference Figure 1A , Figure 1A This is an example data flow diagram illustrating an example process 100 for determining lighting levels, road surface conditions, visibility distance, scene classification, distance to the scene, and sensor blindness and its causes, according to some embodiments of this disclosure. It should be understood that this and other arrangements described herein are merely illustrative examples. Other arrangements and elements (e.g., machines, interfaces, functions, sequences, functional groupings, etc.) may be used to supplement or replace those shown, and some elements may be omitted entirely. Furthermore, many elements described herein are functional entities that can be implemented as discrete or distributed components, or in combination with other components, and can be implemented in any suitable combination and location. The various functions performed by the entities described herein can be performed by hardware, firmware, and / or software. For example, various functions can be performed by a processor executing instructions stored in memory. In some embodiments, the systems, methods, and processes described herein can be used with… Figures 6A-6D Example of autonomous vehicles 600 Figure 7 Example computing device 700 and / or Figure 8 The example data center 800 uses components, features, and / or functions similar to those of other components, features, and / or functions to perform the same task.

[0015] Process 100 may include generating and / or receiving sensor data 102 from one or more sensors of the machine 600. Sensor data 102 may be used by the machine 600 and, within process 100, to determine, in real-time or near real-time, one or more of the following: camera blindness quantity, blindness classification corresponding to the camera blindness quantity, illumination level, road surface condition, visibility distance, scene type classification, and distance to the scene corresponding to the scene type classification. Sensor data 102 may include, but is not limited to, sensor data 102 from any sensor of the machine 600 (in some examples, and / or other vehicles or objects, such as robotic devices, VR systems, AR systems, etc.). For example, and referring to… Figures 6A-6CSensor data 102 may include, but is not limited to, data generated by one or more Global Navigation Satellite System (GNSS) sensors 658 (e.g., one or more Global Positioning System sensors), one or more radar sensors 660, one or more ultrasonic sensors 662, one or more lidar sensors 664, one or more inertial measurement unit (IMU) sensors 666 (e.g., one or more accelerometers, one or more gyroscopes, one or more magnetic compasses, one or more magnetometers, etc.), one or more microphones 696, one or more stereo cameras 668, one or more wide-angle cameras 670 (e.g., fisheye cameras), one or more infrared cameras 672, one or more omnidirectional cameras 674 (e.g., 360-degree cameras), one or more long-range and / or medium-range cameras 698, one or more speed sensors 644 (e.g., for measuring the speed of vehicle 600), and / or other sensor types. As another example, sensor data 102 may include virtual sensor data generated from any number of sensors of a virtual vehicle or other virtual object. In such an example, the virtual sensor may correspond to a virtual vehicle or other virtual object in a simulated environment (e.g., for testing, training, and / or validating the performance of a neural network), and the virtual sensor data may represent sensor data captured by the virtual sensor in a simulated or virtual environment. Therefore, by using virtual sensor data, simulated data in a simulated environment can be used to test, train, and / or validate one or more machine learning models 104 described herein, which allows for testing more extreme scenarios outside of real-world environments, where such testing might be less safe.

[0016] Sensor data 102 may include image data representing one or more images, image data representing video (e.g., video snapshots), and / or sensor data representing the sensor's sensory field (e.g., depth maps of lidar sensors, value maps of ultrasonic sensors, etc.). Where sensor data 102 includes image data, any type of image data format may be used, such as, but not limited to, compressed images such as Joint Photographic Experts Group (JPEG) or luminance / chrominance (YUV) formats, compressed images as frames derived from compressed video formats (e.g., H.264 / Advanced video Codeing (AVC) or H.265 / High Efficiency video Codeing (HEVC)), raw images (e.g., derived from red-blue (RCCB), red-blue (RCCC), or other types of imaging sensors), and / or other formats. Furthermore, in some examples, sensor data 102 can be used within process 100 without any preprocessing (e.g., in raw or captured format), while in other examples, sensor data 102 can be preprocessed (e.g., noise balancing, demosaicing, scaling, cropping, enhancement, white balance, tone curve adjustment, etc., for example using a sensor data preprocessor (not shown)). As used herein, sensor data 102 may refer to unprocessed sensor data, preprocessed sensor data, or a combination thereof.

[0017] The sensor data preprocessor can use image data (or other data representations) representing one or more images and load the sensor data into memory as a multidimensional array / matrix (in some examples, or tensors, or more specifically, input tensors). The array size can be calculated and / or represented as W×H×C, where W represents the image width in pixels, H represents the height in pixels, and C represents the number of color channels. Other types and orders of the input image components are also possible without loss of generality. Furthermore, when using batch processing, the batch size B can be used as a dimension (e.g., an additional fourth dimension). Batch processing can be used for training and / or inference. Therefore, the input tensor can represent an array of dimensions W×H×C×B. Any ordering of dimensions is possible, which may depend on the specific hardware and software used to implement the sensor data preprocessor. This ordering can be selected to maximize the training and / or inference performance of one or more machine learning models 104.

[0018] In some embodiments, a sensor data preprocessor may use a preprocessing image pipeline to process one or more raw images acquired by one or more sensors (e.g., one or more cameras) and included in sensor data 102 to produce preprocessed image data that may represent one or more input images of one or more input layers of one or more machine learning models 104. An example of a suitable preprocessing image pipeline may be a raw RCCB Bayer (e.g., 1-channel) image from a sensor, which is then converted to an RCB (e.g., 3-channel) planar image stored in a fixed-precision (e.g., 16 bits per channel) format. The preprocessing image pipeline may include decompression, noise reduction, demosaicing, white balance, histogram calculation, and / or adaptive global tone mapping (e.g., in this order or an alternative order).

[0019] When the sensor data preprocessor employs noise reduction, it may include bilateral denoising in the Bayer domain. If the sensor data preprocessor employs demosaicing, it may include bilinear interpolation. In the case where the sensor data preprocessor employs histogram calculation, it may involve calculating the histogram of the C channels, and in some examples may be combined with decompression or noise reduction. In the case where the sensor data preprocessor employs adaptive global tone mapping, it may include performing an adaptive gamma-log transformation. This may include calculating the histogram, obtaining midtone levels, and / or estimating the maximum brightness with midtone levels.

[0020] One or more machine learning models 104 may use one or more images or other data representations (e.g., lidar point clouds, radar representations, etc.) represented by sensor data 102 as input to generate one or more outputs 106. In a non-limiting example, one or more machine learning models 104 may use one or more images (e.g., after preprocessing) represented by sensor data 102 as input to generate one or more outputs 106, such as camera blindness 108, one or more blindness classifications 110, illumination level 112, road surface condition 114, visibility distance, scene type classification 118, and / or distance to scene 120. Although this document describes examples of using neural networks, particularly DNNs, as one or more machine learning models 104, this is not intended to be limiting. For example, but not limited to, one or more machine learning models 104 described herein may include any type of machine learning model, such as machine learning models using linear regression, logistic regression, decision trees, support vector machines (SVM), Naive Bayes, k-nearest neighbors (Knn), k-means clustering, random forests, dimensionality reduction algorithms, gradient boosting algorithms, neural networks (e.g., autoencoders, convolutions, recursions, perceptrons, long / short-term memory (LSTM), Hopfield, Boltzmann, deep belief, deconvolution, generative adversarial, liquid machines, etc.) and / or other types of machine learning models.

[0021] Regarding the computation of camera blindness 108 and blindness classification 110, one or more machine learning models 104 can identify regions of interest within sensor data 102 that relate to sensor blindness, and identify its causes (e.g., blurring, obstruction, etc.). More specifically, one or more machine learning models 104 can be designed to infer blindness markers and output classifications that identify the location of potential sensor blindness in the sensor data, the cause of sensor blindness, and whether the sensor data or a portion thereof is available to the system. In some examples, the neural network can output: a binary decision (e.g., true / false, yes / no, 0 / 1) indicating that the sensor data is at least partially available (e.g., true, yes, 0); and a second decision indicating that the sensor data is unavailable (e.g., false, no, 1). If data is indicated as unavailable, the sensor data can be skipped, ignored, and / or used to determine whether to reduce the operational level to level 0, such as returning control to the driver in autonomous or semi-autonomous applications.

[0022] Visual blindness classification 110 can represent the cause of visual blindness associated with one or more pixels of sensor data 101 that have detected visual blindness. Sensor data 102 can include any number of different visual blindness classifications. For example, without departing from the scope of this disclosure, one or more individual visual blind regions or pixels can have one, two or more associated visual blindness classifications. In some examples, one or more machine learning models 104 can output multiple channels corresponding to multiple classifications that the machine learning models are trained to predict (e.g., each channel can correspond to one visual blindness classification). For example, visual blindness classification 110 can include one or more of obstruction, blur, reflection, openness, vehicle, sky, frame label, etc., and the number of output channels can correspond to the desired number of classifications. Visual blindness classification 110 can include subclasses, including but not limited to rain, glare, broken lens, light, mud, paper, people, etc.

[0023] In the example, the visual blindness classification of "blockage" can be associated with one or more visual blindness subclasses, such as sun, dense fog, water, mist, snow, frozen windowpanes, daytime, nighttime, broken lens, self-glare, mud, paper, leaves, etc. Thus, the visual blindness subclass can further characterize the cause of visual blindness in sensor data 102. In some examples, the visual blindness subclass can include further subcategories of causes of visual blindness, such as those used to indicate the degree of impaired visibility (e.g., severe, moderate, mild, etc.). One or more machine learning models 104 can similarly have as many output channels as the number of visual blindness classifications and / or corresponding visual blindness subclasses that the machine learning model is trained to predict.

[0024] One or more machine learning models 104 can be trained to predict the visibility distance 116 of sensor data 102, for example, to identify the farthest distance of an object or element from the sensor. The machine's operational level can be adjusted based on the predicted visibility distance. For example, if one or more machine learning models 104 determine that the visibility distance is low (e.g., 20 meters or less), the corresponding sensor data can only be used for Level 0 (no automation) or Level 1 (driver assistance) tasks, or only for predictions within 20 meters of the machine (e.g., predictions beyond 20 meters may be ignored, or have lower or no relevant confidence). Similarly, as another example, if the estimated visibility distance is high (e.g., 1000 meters or more), the corresponding sensor data can be relied upon to fully perform Level 3 (conditional driver automation) and Level 4 (high driver automation) tasks, or the corresponding sensor data can be relied upon to make predictions corresponding to locations within 1000 meters of the machine.

[0025] In embodiments, the calculated visibility distance can be categorized as extremely low visibility (e.g., between 10 and 100 meters), low visibility (e.g., between 10 and 100 meters), moderate visibility (e.g., between 100 and 1000 meters), high visibility (e.g., between 1000 and 4000 meters), or clear visibility (e.g., greater than 4000 meters). In other embodiments, extremely low visibility may be associated with extreme fog or blizzards, or high rainfall or heavy snow; low visibility may be associated with foggy weather; moderate visibility may be associated with moderate fog or moderate rainfall or moderate drizzle; high visibility may be associated with foggy weather conditions or light precipitation or drizzle; and clear visibility may be related to the maximum visibility range of the specific sensor that generated the sensor data.

[0026] The predicted visibility distance corresponds to various tasks that the machine can perform. For example, for extremely low visibility distances, as long as the machine is traveling below a threshold speed (e.g., below 25 km / h), Level 0, 1, and 2 low-speed active safety functions (e.g., lane assist, automatic lane keeping, automatic cruise control, automatic emergency braking, etc.) can be fully executed. However, if the vehicle is traveling above the threshold speed, these functions may be disabled – at least for instances with sensor data showing very low disable distances. Regarding low visibility distances, as long as the machine is traveling below the threshold speed, Level 0, 1, and 2 parking functions (e.g., proximity detection, automatic parallel parking alignment, etc.) and low-speed active safety functions can be fully executed. For medium visibility distances, Level 0, 1, 2, and 2+ driving functions can be fully executed. Regarding high visibility distances, in addition to Level 0, 1, 2, and 2+ functions in the first, second, and third bins, Level 3 and 4 parking functions (e.g., automatic parallel parking (APP), MPP, VVP) can be fully executed. Regarding clear visibility, in addition to Level 0, Level 1, Level 2 and Level 2+ functions, it can also fully implement Level 3 autonomous driving highway functions.

[0027] One or more machine learning models 104 can also be trained to predict scene type classification 118 of sensor data 102 and distance to the scene corresponding to scene type classification 120. For example, one or more machine learning models 104 can be trained to predict one or more of the sensor data representing a tunnel, construction area, blocked lane, or toll plaza, and distance to scene type (e.g., affected road area). One or more machine learning models 104 can be further trained to predict road surface conditions 114, such as dry, wet, snowy, icy, etc., and illumination levels 112, such as the amount of light perceived by road users, from sensor data 102.

[0028] In embodiments, such as where one or more machine learning models 104 include DNNs, one or more machine learning models 104 may include any number of layers. One or more layers may include an input layer. The input layer may hold values ​​associated with the sensor data 402. One or more layers may include convolutional layers. The convolutional layers may compute the outputs of neurons connected to local regions in the input layer, with each neuron computing its weights with the dot product between them and the small regions they are connected to in the input volume. The result of the convolutional layers may be another volume, where one dimension is based on the number of filters applied (e.g., if the number of filters is 12, the width, height, and number of filters are such as 32 x 32 x 12).

[0029] One or more layers may include deconvolutional layers (or transposed convolutional layers). For example, the result of a deconvolutional layer may be another volume with a dimension higher than the input dimension of the data received at the deconvolutional layer.

[0030] One or more layers may include rectified linear unit (ReLU) layers. For example, one or more ReLU layers may apply an element-wise activation function such as max(0, x) and threshold at zero. The output volume of a ReLU layer may be the same as the input volume of the ReLU layer.

[0031] One or more layers may include pooling layers. Pooling layers may perform downsampling operations along spatial dimensions (e.g., height and width), which may result in a volume smaller than the input of the pooling layer (e.g., 16 × 16 × 12 from a 32 × 32 × 12 input volume).

[0032] One or more layers may include one or more fully connected layers. Each neuron in a fully connected layer can be connected to every neuron in the previous volume. The fully connected layers can compute class scores, and the resulting volume can be 1×1×n, where n equals the number of classes. In some examples, a CNN may include fully connected layers such that the outputs of one or more layers of the CNN can be fed as input to the fully connected layers of the CNN. In some examples, one or more convolutional streams may be implemented by one or more machine learning models, and some or all of the convolutional streams may include their respective fully connected layers.

[0033] In some non-limiting embodiments, one or more machine learning models 104 may include a series of convolutional layers and max pooling layers to facilitate image feature extraction, followed by multi-scale expanded convolutional layers and upsampling layers to facilitate global contextual feature extraction.

[0034] Although this paper discusses input layers, convolutional layers, pooling layers, ReLU layers, and fully connected layers for one or more machine learning models 104, this is not intended to be limiting. For example, additional or alternative layers, such as normalization layers, SoftMax layers, and / or other layer types, may be used in one or more machine learning models 104.

[0035] Furthermore, some layers may include parameters (e.g., weights and / or biases), such as convolutional and fully connected layers, while other layers may not include parameters, such as ReLU layers and pooling layers. In some examples, parameters may be learned by one or more machine learning models 104 during training. Additionally, some layers may include additional hyperparameters (e.g., learning rate, stride, epochs, etc.), such as convolutional, fully connected, and pooling layers, while other layers may not include them, such as ReLU layers. In an embodiment where one or more machine learning models 104 regress visibility distance, the activation function of the last layer of the CNN may include a ReLU activation function. In an embodiment where one or more machine learning models 104 classify instances of sensor data into distance bins, the activation function of the last layer of the CNN may include a SoftMax activation function. Parameters and hyperparameters are not limited and may vary depending on the embodiment.

[0036] In one or more embodiments of the machine learning model 104 that include CNNs, different orders and numbers of CNN layers may be used depending on the embodiment. In other words, the order and number of CNN layers are not limited to any particular architecture.

[0037] In one embodiment, determining the operational level of the machine (e.g., operation 124) may include weighting camera blindness, blindness classification, lighting level, road surface condition, visibility distance, scene type classification, and distance to the scene according to predefined weights. In other embodiments, camera blindness may be associated with the highest weight among the predefined weights.

[0038] Now for reference Figure 1B and Figure 1A , Figure 1BThis is an example machine learning model 104A according to some embodiments of the present disclosure for implementing an example process for determining illumination level, road surface condition, visibility distance, scene classification, distance to scene, and / or sensor blindness and / or its causes. For example, in one or more embodiments, one or more machine learning models 104A may include a CNN, which may include one or more backbones, an encoder-decoder architecture, and / or one or more output heads (e.g., blindness mask head 142, illumination head 144, road surface head 146, cause of blindness head 148, and scene classification head 150). In the example, the CNN may include one or more convolutional layers corresponding to a feature detection backbone, wherein the feature detection backbone may output intermediate data (e.g., representative feature maps), which may be processed using one or more output heads. For example, using intermediate data, a first output head (e.g., blind hood head 142) can be used to calculate the camera's visual blindness level, a second output head (e.g., visual blindness cause head 148) can be used to calculate the camera's visual blindness classification, a third output head (e.g., lighting head 144) can be used to calculate the lighting level, a fourth output head (e.g., road surface head 146) can be used to calculate the road surface condition, and / or a fifth output head (e.g., scene head 150) can be used to calculate the scene type classification (e.g., the probability of the scene type 152) and / or the distance to the scene corresponding to the scene type classification (e.g., the distance to the scene 154). In other embodiments, the fifth output head can be used to calculate the scene type classification, while the sixth output head can be used to calculate the distance to the scene. When using two or more heads, the two or more heads can process data from the backbone in parallel, and each head can be trained to accurately predict the corresponding output of that output head. However, in other embodiments, a single backbone can be used without separate heads.

[0039] In the example, one or more of the first, second, third, fourth, fifth, and / or sixth heads of the CNN of machine learning model 104 may include at least one fully connected layer. For example, the lighting head 144, road surface head 146, cause of blindness head 148, and scene head 150 may each include at least one fully connected layer. Each neuron in the fully connected layer may be connected to each neuron in the previous volume. The fully connected layers may compute class scores, and the resulting volume may be 1×1×n, where n equals the number of classes. In some non-limiting embodiments, one or more heads of the CNN of one or more machine learning models 104 may include one or more downsampling layers and one or more upsampling layers. For example, the blind hood head 142 of the CNN of one or more machine learning models 104 may include a series of downsampling layers (e.g., pooling layers) to facilitate image feature extraction, followed by upsampling layers to facilitate global contextual feature extraction. In embodiments, one or more pooling layers may perform downsampling operations along spatial dimensions (e.g., height and width), which may result in a volume smaller than the input to the pooling layer (e.g., 16 × 16 × 12 from a 32 × 32 × 12 input volume). In the example, each of the lighting head 144, path surface head 146, cause of visual impairment head 148, and scene head 150 may include a pooling layer. In a further example, scene head 150 may use a sigmoid function to output the probability 152 of the scene type, such as a tunnel, construction area, blocked lane, or tollbooth. In other examples, scene head 150 may also use a ReLU layer to output the distance 154 to the scene.

[0040] Back Figure 1AIn some embodiments, output 106, after being computed by one or more machine learning models 104, may be post-processed by post-processor 122. Post-processor 122 may determine drivability corresponding to sensor data 102. For example, post-processor 122 may perform temporal smoothing of the output 106 of machine learning model 104. In some embodiments, temporal smoothing may be used to improve system stability by reducing false alarms based on a single frame—by incorporating previous predictions of machine learning model 104 corresponding to temporally adjacent frames—to smooth and reduce noise in the output of machine learning model 104. In some examples, the value computed by machine learning model 104 for the current instance of sensor data 102 may be weighted relative to the values ​​computed by machine learning model 104 for one or more previous instances of sensor data 102. In the case where sensor data 102 is image data representing images, for example, output 106 computed by machine learning model 104 for the current or most recent image may be weighted relative to output 106 computed by machine learning model 104 for one or more temporally adjacent (e.g., previous and / or consecutive) images. Therefore, the final value corresponding to an instance of sensor data 102 (e.g., for camera blindness 108, blindness classification 110, illumination level 112, road surface condition 114, visibility distance 116, scene type classification 118, and distance to scene 120) can be determined by weighting previous values ​​associated with one or more other instances of sensor data 102 relative to the current value associated with the instance of sensor data 102. The results of time smoothing can be analyzed to determine drivability relative to sensor data 102.

[0041] One or more operations 124 may be decisions made by the system using sensor data 102 and based on output 106 in real-time or near real-time. For example, some or all of the sensor data 102 may be skipped or ignored in cases where it is not clear or useful enough. In some examples, such as when sensor data 102 is not available for safe operation of vehicle 600, operation 124 may include returning control to the driver (e.g., disengaging autonomous or semi-autonomous operation) or performing an emergency or safety operation (e.g., stopping, pulling over, or a combination thereof). Thus, operation 124 may include suggesting one or more corrective measures for effective and safe driving, such as ignoring certain instances of sensor data.

[0042] In any example, and regarding autonomous or semi-autonomous driving, operation 124 may include any decision corresponding to the drive stack 126 (e.g., the sensor data manager layer of the autonomous driving software stack), the perception layer of the drive stack 126, the world model management layer of the drive stack 126, the planning layer of the drive stack 126, the control layer of the drive stack 126, the obstacle avoidance layer of the drive stack 126, and / or the drive layer of the drive stack 126. Therefore, as described herein, the drivability of sensor data 102 can be determined individually for any number of different operations corresponding to one or more layers of the drive stack 126.

[0043] For example, a first drivability force for object detection operations can be determined for the perception layer of drive stack 126, and a second drivability force for path planning can be determined for the planning layer of drive stack 126. Information representing sensor data 102 can be transmitted to drive stack 126, indicating that one or more features, functions, and / or components of the drive stack can be disabled, enabled, or remain unchanged. Drive stack 126 may include one or more layers, such as a world state manager layer that uses one or more maps (e.g., 3D maps), one or more localization components, one or more perception components, etc., to manage the world state. Furthermore, the autonomous driving software stack may include one or more planning components (e.g., as part of a planning layer), one or more control components (e.g., as part of a control layer), one or more actuation components (e.g., as part of an actuation layer), one or more obstacle avoidance components (e.g., as part of an obstacle avoidance layer), and / or other components (e.g., as part of one or more additional or alternative layers). Thus, operation 124 can provide instructions to any downstream task of drive stack 126 that depends on sensor data 102 in order to manage the appropriate use of sensor data 102.

[0044] Back Figure 1AIn embodiments, one or more machine learning models 104 may be trained using information provided by sensor data 102 and information provided by training engine 128. In an example, training engine 128 may include ground truth data 130. In some embodiments, ground truth data 130 may include labels or annotations corresponding to sensor data 102. For example, for each instance of sensor data 102A, the label or annotation may correspond to one or more of the following: camera blindness measure, blindness classification corresponding to camera blindness, lighting level, road surface condition, visibility distance, scene type classification, and distance to the scene corresponding to the scene type classification. In some examples, the labels and annotations may be generated in a drawing program (e.g., an annotation program), a computer-aided design (CAD) program, a labeling program, another type of program suitable for generating annotations, and / or may be drawn manually. In any example, ground truth data 130 may be generated synthetically (e.g., generated from a computer model or rendering), realistically (e.g., designed and generated from real-world data), machine-automated (e.g., using feature analysis and learning to extract features from the data and then generate labels), human-annotated (e.g., labelers or annotation experts who define the location of labels), and / or a combination thereof (e.g., humans identify visible objects, and computers determine the distance to the objects and the corresponding visibility distance).

[0045] In some embodiments, the determination of the truth label for the distance to the scene can be performed manually. In the example, a frame in which the vehicle is within the scene, such as a frame where the vehicle is in a tunnel, can be selected. Subsequent iterative frames can be marked with a distance of 0 until the scene is no longer visible. Iterative frames preceding the scene can also be marked with a distance of 0 until an entrance to the scene is found or until the scene is no longer visible. IMU and / or GPS data can then be used to calculate the distance between the previous and current frames. This distance can be added to the distance of the previous frame to set the distance of the current frame, thus determining the length to the scene.

[0046] In an embodiment, training engine 128 may include one or more loss functions 132 to compare output 106 with ground truth data 130 (e.g., one or more corresponding to camera blindness, blindness classification corresponding to camera blindness, lighting level, road surface condition, visibility distance, scene type classification, and distance to the scene corresponding to the scene type classification). In an embodiment, sensor data 102 may be used to train one or more machine learning models 104 using multiple iterations until the values ​​of one or more loss functions 132 are below a predetermined threshold. Any type of loss function may be used, such as cross-entropy loss, mean squared error, mean absolute error, mean bias error, and / or other loss function types. In some examples, the gradient of loss function 132 may be computed iteratively relative to training parameters. When training machine learning model 104, an optimizer, such as the Adam optimizer, stochastic gradient descent, or any other type of optimization algorithm, may be used to optimize loss function 132. Machine learning model 104 may be trained iteratively until the training parameters converge to optimal, desired, or acceptable values.

[0047] Now for reference Figure 2A-2B , Figure 2A-2B This is an example visualization of sensor data representations at different visibility distances according to some embodiments of this disclosure. About Figure 2A Assuming that visualization 200A (including heavy rain) corresponds to an instance of sensor data 102A (e.g., an image from a camera on a data collection vehicle), the depth value of vehicle 202A can be known based on outputs from depth sensors (e.g., LiDAR, radar, etc.), outputs from machine learning models or neural networks, outputs from computer vision algorithms, and / or other output types. In such an example, while the data collection vehicle is generating an instance of sensor data 102A including vehicle 202A, the data collection vehicle may have already run one or more low-level processes to determine depth or distance information to objects in the environment. Therefore, if vehicle 202A is the farthest visible object from the data collection vehicle, and the depth to vehicle 202A is known, the visible distance of the instance of sensor data 102A can correspond to the depth or distance to vehicle 202A (e.g., in an embodiment, adding an additional distance visible in the environment outside vehicle 202A, or subtracting some distance where vehicle 202A is slightly visible or blurred). Similarly, for visualization 200B (including light rain conditions), vehicle 202B is visible, so the visibility distance can be determined using the depth or distance value corresponding to vehicle 202B. Although regarding... Figure 2A-2B Rain has been shown and described, but this is not intended to be limiting, and visibility distances can be determined under other conditions such as fog, snow, sleet, hail, lighting, obstruction, and combinations thereof, according to embodiments of this disclosure.

[0048] Now for reference Figure 3 , Figure 3 Examples of sensor data with different levels of sensor blindness according to some embodiments of this disclosure are included in the visualization. For example, machine learning model 104 can use image data representing an input image 310A as input and can output an image mask 310B including image blindness regions 312 and 314. In one or more embodiments, the upper part of image 310A may be covered by precipitation (e.g., Figure 3 The visual blindness is obscured or blurred by debris, glare, or reflections from sunlight. Different classifications can be represented by different RGB colors or indicators. The visual blind area 312 can be classified as an obscured or blurred area, and the visual blind area 314 can also be classified as an obscured or blurred area. Pixels in the visual blind areas 312 and 314 can be similarly classified as obscured or blurred pixels. Visualization 310C shows image 310A overlaid with image mask 310B. In this way, sensor visual blindness can be accurately determined and visualized in a region-specific manner. For example, regarding image 310A, the determination of drivability may be high, making image 310A useful for autonomous or semi-autonomous driving operations. This may be because only a portion of image 310A is blurred, and that portion is located in the sky of an environment where operation 124 may not be affected.

[0049] Now for reference Figures 4A-4B , Figures 4A-4B Examples of visualizations, including illustrations of sensor data corresponding to various scene types according to some embodiments of this disclosure, are provided. Figure 4A As shown, using image data, such as sensor data 102 and / or data provided by training engine 128, machine learning model 104 can output data representing the scene, such as a tunnel. Figure 4B As shown, using image data, such as sensor data 102 and / or data provided by training engine 128, machine learning model 104 can output data representing the scene, such as a tollbooth.

[0050] Now for reference Figure 5 Each block of the method 500 described herein includes a computational process that can be executed using any combination of hardware, firmware, and / or software. For example, various functions can be executed by a processor executing instructions stored in memory. Method 500 can also be implemented as computer-usable instructions stored on a computer storage medium. Method 500 can be provided by a standalone application, service, or managed service (standalone or in combination with another managed service), or a plug-in to another product, etc. Furthermore, as an example, for... Figures 6A-6DThis machine 600 describes method 500. However, method 500 may be performed additionally or alternatively within any process and / or by any system or any combination of processes and systems (including, but not limited to, those described herein).

[0051] Figure 5 This is a flowchart illustrating a method 500 for determining lighting levels, road surface conditions, visibility distance, scene classification, distance to the scene, and sensor blindness and its causes using a neural network according to some embodiments of the present disclosure.

[0052] In box B502, method 500 includes using a neural network and, based at least on image data of the environment generated using one or more image sensors, computing data representing one or more parameters corresponding to the operational domain of the environment. In an example, the one or more parameters include at least one of camera blindness level, blindness classification corresponding to camera blindness, lighting level, road surface condition, visibility distance, scene type classification, or distance to the scene corresponding to the scene type classification. In a further example, the neural network may use sensor data 102 and / or ground truth data 130 and / or loss function 132 to compute data representing camera blindness, blindness classification corresponding to camera blindness, lighting level, road surface condition, visibility distance, scene type classification, and distance to the scene corresponding to the scene type classification.

[0053] In box B504, method 500 includes determining, at least in part, the operating level of the machine corresponding to the environment based on data. For example, the operating level may be determined based on one or more of the following: camera blindness, a blindness classification corresponding to camera blindness, lighting level, road surface condition, visibility distance, scene type classification, and / or distance to a scene corresponding to the scene type classification.

[0054] In box B506, method 500 includes controlling the operation of the machine based on the operation level. For example, if the scene type classification is determined to be a tollbooth and / or the distance to the scene is determined to be 100 yards, according to embodiments of this disclosure, the machine may decelerate. In other examples, if a low visibility distance is determined, the machine may decelerate or perform other safety operations (stop, navigate to a safe or designated area, etc.).

[0055] Example autonomous vehicles Figure 6AThis is an illustration of an exemplary autonomous vehicle 600 according to some embodiments of the present invention. The autonomous vehicle 600 (or referred to herein as “vehicle 600”) may include, but is not limited to, passenger vehicles such as automobiles, trucks, buses, first-response vehicles, shuttles, electric or motorized bicycles, motorcycles, fire trucks, police cars, ambulances, boats, construction vehicles, underwater vehicles, drones, trailer-mounted vehicles, and / or other types of vehicles (e.g., driverless and / or vehicles accommodating one or more passengers). Autonomous vehicles are generally described according to their level of automation, as defined by the National Highway Traffic Safety Administration (NHTSA) of the U.S. Department of Transportation, and by the Society of Automotive Engineers (SAE) “Classification and Definition of Terms Related to Driving Automation Systems for Road Motor Vehicles” (Standard No.: J3016-201806, published June 15, 2018; Standard No.: J3016-201609, published September 30, 2016, and previous and future versions of this standard). Vehicle 600 may implement functions according to one or more of Level 3 to Level 5 of autonomous driving. Vehicle 600 may implement functions according to one or more of Level 1 to Level 5 of autonomous driving. For example, according to an embodiment, vehicle 600 may have driver assistance (Level 1), partial automation (Level 2), conditional automation (Level 3), high automation (Level 4), and / or full automation (Level 5). The term “autonomy” as used herein may include any and / or all types of autonomy of vehicle 600 or other machines, such as full autonomy, high autonomy, conditional autonomy, partial autonomy, assisted autonomy, semi-autonomy, primary autonomy, or other specified autonomy.

[0056] Vehicle 600 may include components such as a chassis, body, wheels (e.g., 2, 4, 6, 8, 18, etc.), tires, axles, and other vehicle parts. Vehicle 600 may include a propulsion system 650, such as an internal combustion engine, a hybrid power plant, an all-electric motor, and / or other propulsion system types. Propulsion system 650 may be connected to the drivetrain of vehicle 600, which may include a transmission, to enable propulsion of vehicle 600. Propulsion system 650 may be controlled in response to receiving signals from a throttle valve / accelerator 652.

[0057] When the propulsion system 650 is in operation (e.g., when the vehicle is moving), a steering system 654, including a steering wheel, can be used to guide the vehicle 600 (e.g., along a desired path or route). The steering system 654 can receive signals from the steering actuator 656. For fully automatic (level 5) functionality, the steering wheel may be optional.

[0058] The brake sensor system 646 can be used to operate the vehicle brakes in response to signals received from the brake actuator 648 and / or the brake sensor.

[0059] One or more controllers 636 may include one or more system-on-chip (SoC) 604 ( Figure 6C One or more controllers and / or one or more GPUs may provide signals (e.g., representing commands) to one or more components and / or systems of vehicle 600. For example, one or more controllers may send signals via one or more brake actuators 648 to operate vehicle brakes, via one or more steering actuators 656 to operate steering system 654, and via one or more throttles / accelerators 652 to operate propulsion system 650. One or more controllers 636 may include one or more onboard (e.g., integrated) computing devices (e.g., supercomputers) that process sensor signals and output operating commands (e.g., signals representing commands) to enable autonomous driving and / or assist a human driver in driving vehicle 600. One or more controllers 636 may include a first controller 636 for autonomous driving functions, a second controller 636 for functional safety functions, a third controller 636 for artificial intelligence functions (e.g., computer vision), a fourth controller 636 for infotainment functions, a fifth controller 636 for redundancy in emergency situations, and / or other controllers. In some examples, a single controller 636 may handle two or more of the functions described above, and two or more controllers 636 may handle a single function and / or any combination thereof.

[0060] One or more controllers 636 may provide signals for controlling one or more components and / or systems of vehicle 600 in response to sensor data (e.g., sensor input) received from one or more sensors. Sensor data may be received from, for example, but not limited to, one or more Global Navigation Satellite System sensors 658 (e.g., one or more Global Positioning System sensors), one or more radar sensors 660, one or more ultrasonic sensors 662, one or more lidar sensors 664, one or more inertial measurement unit (IMU) sensors 666 (e.g., one or more accelerometers, one or more gyroscopes, one or more magnetometers, magnetometers, etc.), one or more microphones 696, one or more stereo cameras 668, one or more wide-angle cameras 670 (e.g., fisheye cameras), one or more infrared cameras 672, one or more surround cameras 674 (e.g., 360-degree cameras), one or more long-range and / or mid-range cameras 698, one or more speed sensors 644 (e.g., for measuring the speed of vehicle 600), one or more vibration sensors 642, one or more steering sensors 640, one or more brake sensors (e.g., as part of brake sensor system 646) and / or other sensor types.

[0061] One or more of the controllers 636 may receive inputs (e.g., represented by input data) from the instrument cluster 632 of the vehicle 600 and provide outputs (e.g., represented by output data, displayed data, etc.) via a human-machine interface (HMI) display 634, an audio signaler, a speaker, etc., and / or via other components of the vehicle 600. Outputs may include information such as vehicle speed, rate, time, map data (e.g., ...). Figure 6C Information such as an HD map 622, location data (e.g., the location of vehicle 600, such as its location on a map), direction, the location of other vehicles (e.g., grid occupancy), and information about objects and their states perceived by one or more controllers 636. For example, an HMI display 634 may display information about the presence of one or more objects (e.g., street signs, warning signs, traffic light changes, etc.), and / or information about driving actions that the vehicle has performed, is performing, or will perform (e.g., changing lanes now, exiting from exit 34B in two miles, etc.).

[0062] The vehicle 600 also includes a network interface 624, which can communicate over one or more networks using one or more wireless antennas 626 and / or a modem. For example, the network interface 624 may be able to communicate via LTE, WCDMA, UMTS, GSM, CDMA2000, etc. One or more wireless antennas 626 may also enable communication between objects in the environment (e.g., vehicles, mobile devices, etc.) using one or more local area networks (e.g., Bluetooth, Bluetooth LE, Z-Wave, ZigBee, etc. and / or low-power wide area networks (LPWANs, such as LoRaWAN, SigFox, etc.).

[0063] Figure 6B According to some embodiments of the present invention Figure 6A An example of the camera position and field of view of an exemplary autonomous vehicle 600. The camera and its respective field of view are an example embodiment and are not intended to be limiting. For example, additional and / or alternative cameras may be included and / or the cameras may be located at different positions on the vehicle 600.

[0064] The camera type may include, but is not limited to, a digital camera, which may be suitable for components and / or systems of vehicle 600. One or more cameras may operate under Automotive Safety Integrity Level (ASIL) B and / or other ASILs. According to embodiments, the camera type may have any image capture rate, such as 60 frames per second (fps), 120 fps, 240 fps, etc. The camera may use a rolling shutter, a global shutter, another type of shutter, or a combination thereof. In some examples, the color filter array may include a red transparent (RCCC) color filter array, a red transparent blue (RCCB) color filter array, a red blue green transparent (RBGC) color filter array, a Foveon X3 color filter array, a Bayer sensor (RGGB) color filter array, a monochrome sensor color filter array, and / or other types of color filter arrays. In some embodiments, a sharp-pixel camera, such as a camera with RCCC, RCCB, and / or RBGC color filter arrays, may be used to improve light sensitivity.

[0065] In some examples, one or more cameras can be used to perform advanced driver assistance system (ADAS) functions (e.g., as part of a redundancy or fail-safe design). For example, a multi-functional single camera can be installed to provide functions such as lane departure warning, traffic sign assistance, and intelligent headlight control. One or more cameras (e.g., all cameras) can simultaneously record and provide image data (e.g., video).

[0066] One or more cameras can be mounted in mounting components, such as custom-designed (3D-printed) components, to cut off stray light and interior reflections (e.g., dashboard reflections from the windshield rearview mirror) that may interfere with the camera's image data capture capability. Regarding wing mirror mounting components, these components can be custom-3D printed so that the camera mounting plate matches the shape of the wing mirror. In some examples, one or more cameras can be integrated into the wing-shaped rearview mirror. For side-view cameras, one or more cameras can also be integrated into the four pillars at each corner of the cab.

[0067] A camera with a field of view including a portion of the environment in front of the vehicle (e.g., a front-facing camera) can be used for surround view to help identify forward paths and obstacles, and, with the assistance of one or more controllers 636 and / or control SOCs, to provide information crucial for generating an occupancy grid and / or determining the preferred vehicle path. The front-facing camera can be used to perform many of the same ADAS functions as lidar, including emergency braking, pedestrian detection, and collision avoidance. The forward-facing camera can also be used in ADAS functions and systems, including lane departure warning (“LDW”), adaptive cruise control (“ACC”), and / or other functions such as traffic sign recognition.

[0068] Various cameras can be used in front-mounted configurations, including, for example, monocular camera platforms that include a CMOS (Complementary Metal-Oxide-Semiconductor) color imager. Another example could be a wide-angle camera 670 that can be used to perceive objects entering the field of view from the periphery (e.g., pedestrians, cross traffic, or bicycles). Although Figure 6B Only one wide-angle camera is shown, but the vehicle 600 may have any number of wide-angle cameras 670. Furthermore, one or more remote cameras 698 (e.g., a pair of long-angle stereo cameras) can be used for depth-based object detection, particularly for objects for which neural networks have not yet been trained. One or more remote cameras 698 can also be used for object detection and classification, as well as basic object tracking.

[0069] One or more stereo cameras 668 may also be included in a front-mounted configuration. One or more stereo cameras 668 may include an integrated control unit comprising a scalable processing unit that can provide programmable logic (FPGA) and a multi-core microprocessor with an integrated CAN or Ethernet interface on a single chip. This unit can be used to generate a three-dimensional map of the vehicle environment, including distance estimates for all points in the image. One or more alternative stereo cameras 668 may include a compact stereo vision sensor that may include two camera lenses (one on each side) and an image processing chip that measures the distance from the vehicle to a target object and uses the generated information (e.g., metadata) to activate automatic emergency braking and lane departure warning functions. In addition to the stereo cameras described herein, or alternatively, other types of stereo cameras 668 may be used.

[0070] A camera with a field of view including a portion of the vehicle's side environment (e.g., a side-view camera) can be used in the surround view to provide information for creating and updating the occupancy grid and generating side collision warnings. For example, one or more surround cameras 674 (e.g., such as...) Figure 6B The four surround cameras 674 shown may be positioned on the vehicle 600. One or more surround cameras 674 may include one or more wide-angle cameras 670, one or more fisheye cameras, one or more 360-degree cameras, etc. For example, four fisheye cameras may be located at the front, rear, and sides of the vehicle. In an alternative arrangement, the vehicle may use three surround cameras 674 (e.g., left, right, and rear) and may utilize one or more other cameras (e.g., a front-facing camera) as a fourth surround-view camera.

[0071] Cameras with a view including a portion of the environment behind the vehicle 600 (e.g., rear-view cameras) can be used for parking assistance, surround view, rear-end collision warning, and creating and updating occupancy grids. A variety of cameras can be used, including but not limited to cameras also suitable as front-facing cameras (e.g., one or more long-range and / or mid-range cameras 698, one or more stereo cameras 668, one or more infrared cameras 672, etc.), as described herein.

[0072] Figure 6C According to some embodiments of the present invention Figure 6A A block diagram of an exemplary system architecture for an exemplary autonomous vehicle 600 is provided. It should be understood that this and other arrangements described herein are merely illustrative. Other arrangements and elements (e.g., machines, interfaces, functions, sequences, functional groups, etc.) may be used in addition to those shown, or alternative arrangements and elements may be used instead of those shown, and some elements may be omitted entirely. Furthermore, many of the elements described herein are functional entities that can be implemented as discrete or distributed components, or in combination with other components, and can be implemented in any suitable combination and location. The various functions described herein as being performed by entities can be performed by hardware, firmware, and / or software. For example, various functions can be performed by a processor executing instructions stored in memory.

[0073] Figure 6C Every component, feature, and system of vehicle 600 is connected via bus 602. Bus 602 may include a Controller Area Network (CAN) data interface (also referred to herein as the "CAN bus"). CAN can be a network within vehicle 600 used to help control various features and functions of vehicle 600, such as the actuation of brakes, acceleration, braking, steering, windshield wipers, etc. The CAN bus can be configured to have dozens or even hundreds of nodes, each with its own unique identifier (e.g., a CAN ID). The CAN bus can be read to locate steering wheel angle, ground speed, engine speed per minute (RPM), button positions, and / or other vehicle status indicators. The CAN bus may conform to the ASIL B standard.

[0074] Although bus 602 is described herein as a CAN bus, this is not intended to be limiting. For example, in addition to or from a CAN bus, FlexRay and / or Ethernet may also be used. Furthermore, although bus 602 is represented by a single line, this is not intended to be limiting. For example, any number of buses 602 may exist, which may include one or more CAN buses, one or more FlexRay buses, one or more Ethernet buses, and / or one or more other types of buses using different protocols. In some examples, two or more buses 602 may be used to perform different functions and / or for redundancy. For example, a first bus 602 may be used for a collision avoidance function, and a second bus 602 may be used for drive control. In any example, each bus 602 may communicate with any component of vehicle 600, and two or more buses 602 may communicate with the same component. In some examples, each SoC 604, each controller 636, and / or each computer within the vehicle may access the same input data (e.g., input from sensors of vehicle 600) and may be connected to a common bus, such as a CAN bus.

[0075] Vehicle 600 may include one or more controllers 636, as described herein. Figure 6A The controller 636 is described above. Controller 636 can be used for various functions. One or more controllers 636 can be coupled to any of the various other components and systems of vehicle 600, and can be used to control vehicle 600, artificial intelligence of vehicle 600, infotainment of vehicle 600, etc.

[0076] Vehicle 600 may include one or more System-on-Chip (SoC) 604. SoC 604 may include one or more CPUs 606, one or more GPUs 608, one or more processors 610, one or more caches 612, one or more accelerators 614, one or more data storage 616, and / or other components and features not shown. One or more SoCs 604 can be used to control vehicle 600 in various platforms and systems. For example, one or more SoCs 604 may be combined with an HD map 622 in a system (e.g., the system of vehicle 600), the HD map 622 being accessible via a network interface 624 from one or more servers (e.g., [server name missing]). Figure 6D Server 678) receives map refresh and / or updates.

[0077] One or more CPUs 606 may include CPU clusters or CPU complexes (or referred to herein as “CCPLEX”). One or more CPUs 606 may include multiple cores and / or a L2 cache. For example, in some embodiments, one or more CPUs 606 may include eight cores in a coherent multiprocessor configuration. In some embodiments, one or more CPUs 606 may include four dual-core clusters, each with a dedicated L2 cache (e.g., 2MB L2 cache). One or more CPUs 606 (e.g., CCPLEX) may be configured to support simultaneous cluster operation, such that any combination of clusters of CPUs 606 is active at any given time.

[0078] One or more CPU 606s can implement power management capabilities including one or more of the following features: automatic clock gating of a single hardware block when idle to conserve dynamic power; clock gating of each core when the core is not actively executing instructions due to executing WFI / WFE instructions; independent power gating of each core; independent clock gating of each core cluster when all cores are clock-gated or power-gated; and / or independent power gating of each core cluster when all cores are power-gated. One or more CPU 606s can further implement enhanced algorithms for managing power states, specifying allowed power states and expected wake-up times, and the hardware / microcode determines the optimal power state for the core, cluster, and CCPLEX to enter. The processing core can support simplified power state input sequences in software and offload the work to the microcode.

[0079] One or more GPUs 608 may include integrated GPUs (or referred to herein as “iGPUs”). GPUs 608 may be programmable and efficient for parallel workloads. In some examples, one or more GPUs 608 may use an enhanced tensor instruction set. One or more GPUs 608 may include one or more streaming microprocessors, where each streaming microprocessor may include a Level 1 cache (e.g., a Level 1 cache with at least 96KB of storage), and two or more streaming microprocessors may share a Level 2 cache (e.g., a Level 2 cache with 512KB of storage). In some embodiments, one or more GPUs 608 may include at least eight streaming microprocessors. One or more GPUs 608 may use one or more computation application programming interfaces (APIs). Furthermore, one or more GPUs 608 may use one or more parallel computing platforms and / or programming models (e.g., NVIDIA’s CUDA).

[0080] One or more GPU 608s can be power-optimized for optimal performance in automotive and embedded use cases. For example, one or more GPU 608s can be fabricated on FinFETs. However, this is not intended to limit, and other semiconductor manufacturing processes can be used to fabricate one or more GPU 608s. Each streaming microprocessor can combine multiple mixed-precision processing cores partitioned into multiple blocks. For example, but not limited to, 64 PF32 cores and 32 PF64 cores can be divided into four processing blocks. In such an example, each processing block can be allocated 16 FP32 cores, 8 FP64 cores, 16 INT32 cores, two mixed-precision NVIDIA tensor cores for deep learning matrix algorithms, an L0 instruction cache, a thread bundle scheduler, a dispatch unit, and / or a 64 KB register file. Furthermore, the streaming microprocessor can include independent parallel integer and floating-point data paths to provide efficient execution of workloads through mixed computation and addressing computation. The streaming microprocessor can include independent thread scheduling capabilities to enable finer-grained synchronization and cooperation between parallel threads. Streaming microprocessors can include a combination of a level-one data cache and a shared memory unit to improve performance while simplifying programming.

[0081] One or more GPUs 608 may include high-bandwidth memory (HBM) and / or a 16 GB HBM2 memory subsystem to provide approximately 900 GB / s peak memory bandwidth in some examples. In some examples, in addition to HBM memory, or optionally from HBM memory, synchronous graphics random access memory (SGRAM), such as fifth-generation graphics dual data rate synchronous random access memory (GDDR5), may be used.

[0082] The fifth-generation GPU 608 may include unified memory technology, which includes access counters to allow more accurate migration of memory pages to the processors that access them most frequently, thereby improving the efficiency of shared memory ranges between processors. In some examples, address translation service (ATS) support may be used to allow one or more GPUs 608 to directly access the page tables of one or more CPUs 606. In such examples, when one or more GPUs 608 memory management units (MMUs) experience a miss, an address translation request may be sent to one or more CPUs 606. In response, one or more CPUs 606 may look up the virtual-to-physical mapping of the address in their page tables and send the translation back to one or more GPUs 608. Therefore, unified memory technology allows for a single unified virtual address space for the memory of both one or more CPUs 606 and one or more GPUs 608, thereby simplifying the programming of one or more GPUs 608 and porting applications to one or more GPUs 608.

[0083] In addition, one or more GPUs 608 may include access counters that track the frequency with which one or more GPUs 608 access the memory of other processors. Access counters help ensure that memory pages are moved to the physical memory of the processor that accesses those pages most frequently.

[0084] One or more SoCs 604 may include any number of caches 612, including the caches 612 described herein. For example, one or more caches 612 may include an L3 cache available to one or more CPUs 606 and one or more GPUs 608 (e.g., connecting both one or more CPUs 606 and one or more GPUs 608). Caches 612 may include write-back caches with traceable thread state, for example, by using a cache coherence protocol (e.g., MEI, MESI, MSI, etc.). Although a small cache size may be used, according to embodiments, the L3 cache may include 4MB or more.

[0085] One or more SoCs 604 may include one or more arithmetic logic units (ALUs) that can be used to perform processing related to various tasks or operations of the vehicle 600, such as processing a DNN. Furthermore, one or more SoCs 604 may include one or more floating-point units (FPUs) or other mathematical coprocessors or digital coprocessor types for performing mathematical operations within the system. For example, one or more SoCs 104 may include one or more FPUs integrated as execution units within a CPU 606 and / or a GPU 608.

[0086] One or more SoCs 604 may include one or more accelerators 614 (e.g., hardware accelerators, software accelerators, or a combination thereof). For example, one or more SoCs 604 may include a hardware acceleration cluster that may include optimized hardware accelerators and / or large on-chip memory. Large on-chip memory (e.g., 4MB of SRAM) can enable the hardware acceleration cluster to accelerate neural networks and other computations. The hardware acceleration cluster can be used to supplement one or more GPUs 608 and offload some tasks from one or more GPUs 608 (e.g., freeing up more cycles from one or more GPUs 608 to perform other tasks). For example, one or more accelerators 614 may be sufficiently stable to be suitable for accelerating target workloads (e.g., perception, convolutional neural networks (CNNs), etc.). The term "CNN" as used herein may include all types of CNNs, including region-based or region-based convolutional neural networks (RCNNs) and fast RCNNs (e.g., for object detection).

[0087] One or more accelerators 614 (e.g., hardware acceleration clusters) may include one or more deep learning accelerators (DLAs). One or more DLAs may include one or more tensor processing units (TPUs) configured to provide an additional trillion operations per second for deep learning applications and inference. TPUs may be accelerators configured to perform image processing functions (e.g., for CNNs, RCNNs, etc.) and optimized for them. One or more DLAs may also be optimized for specific neural network types and floating-point operations and inference. One or more DLAs are designed to provide higher performance per millimeter than general-purpose GPUs and significantly outperform CPUs. One or more TPUs may perform multiple functions, including single-instance convolution functions, such as supporting INT8, INT16, and FP16 data types for features and weights, and post-processor functions.

[0088] One or more DLAs can execute neural networks, especially CNNs, quickly and efficiently on processed or unprocessed data for any of a variety of functions, including but not limited to: CNNs for object recognition and detection using data from camera sensors; CNNs for distance estimation using data from camera sensors; CNNs for emergency vehicle detection and recognition using microphone data; CNNs for facial recognition and vehicle owner recognition using data from camera sensors; and / or CNNs for safety and / or safety-related events.

[0089] One or more DLAs can perform any function of one or more GPUs 608. For example, by using inference accelerators, designers can perform any function for one or more DLAs or one or more GPUs 608. For example, designers can centralize the processing of CNNs and floating-point operations on one or more DLAs and leave other functions to one or more GPUs 608 and / or one or more other accelerators 614.

[0090] One or more accelerators 614 (e.g., hardware acceleration clusters) may include programmable vision accelerators (PVAs), which may also be referred to herein as computer vision accelerators. One or more PVAs may be designed and configured to accelerate computer vision algorithms for advanced driver assistance systems (ADAS), autonomous driving, and / or augmented reality (AR) and / or virtual reality (VR) applications. One or more PVAs may provide a balance between performance and flexibility. For example, each PVA may include, for example, but not limited to, any number of reduced instruction set computer (RISC) cores, direct memory access (DMA), and / or any number of vector processors.

[0091] The RISC core can interact with image sensors (e.g., the image sensor of any camera described herein), one or more image signal processors, etc. Each RISC core may include any amount of memory. The RISC core can use any of a variety of protocols, depending on the implementation. In some examples, the RISC core can run a real-time operating system (RTOS). The RISC core can be implemented using one or more integrated circuit devices, application-specific integrated circuits (ASICs), and / or memory devices. For example, the RISC core may include an instruction cache and / or tightly coupled RAM.

[0092] DMA enables PVA components to access system memory independently of one or more CPUs. DMA can support any features used to optimize PVA, including but not limited to support for multidimensional addressing and / or circular addressing. In some examples, DMA can support addressing in up to six or more dimensions, which may include block width, block height, block depth, horizontal block step, vertical block step, and / or depth step.

[0093] Vector processors can be programmable processors designed to efficiently and flexibly execute computer vision algorithm programming and provide signal processing capabilities. In some examples, a PVA may include a PVA core and two vector processing subsystem partitions. The PVA core may include a processor subsystem, one or more DMA engines (e.g., two DMA engines), and / or other peripherals. The vector processing subsystem may operate as the main processing engine of the PVA and may include a vector processing unit (VPU), an instruction cache, and / or a vector memory (e.g., VMEM). The VPU core may include a digital signal processor, such as a single-instruction, multiple-data (SIMD) or very long instruction word (VLIW) digital signal processor. The combination of SIMD and VLIW can improve throughput and speed.

[0094] Each vector processor may include an instruction cache and may be coupled to dedicated memory. As a result, in some examples, each vector processor may be configured to execute independently of other vector processors. In other examples, vector processors included in a particular PVA may be configured to employ data parallelism. For example, in some embodiments, multiple vector processors included in a single PVA may execute the same computer vision algorithm on different regions of an image. In other examples, vector processors included in a particular PVA may execute different computer vision algorithms simultaneously on the same image, or even different algorithms on a sequence of images or portions of an image. Among other things, any number of PVAs may be included in a hardware-accelerated cluster, and any number of vector processors may be included in each PVA. Furthermore, PVAs may include additional error-correcting code (ECC) memory to enhance overall system security.

[0095] One or more accelerators 614 (e.g., a hardware acceleration cluster) may include on-chip computer vision network and SRAM for providing high-bandwidth, low-latency SRAM for one or more accelerators 614. In some examples, the on-chip memory may include at least 4 MB of SRAM, including but not limited to eight field-configurable memory blocks accessible by the PVA and DLA. Each pair of memory blocks may include an Advanced Peripheral Bus (APB) interface, configuration circuitry, a controller, and a multiplexer. Any type of memory can be used. The PVA and DLA access memory via a backbone that provides high-speed memory access for the PVA and DLA. The backbone may include an on-chip computer vision network that interconnects the PVA and DLA to memory (e.g., using an APB).

[0096] An on-chip computer vision network may include an interface that determines that both the PVA and DLA have provided ready and valid signals before transmitting any control signals / addresses / data. Such an interface can provide independent phases and independent channels for transmitting control signals / addresses / data, as well as burst communication for continuous data transmission. This type of interface may conform to ISO 26262 or IEC 61508 standards, but other standards and protocols may also be used.

[0097] In some examples, one or more SoCs 604 may include a real-time ray tracing hardware accelerator, as described in U.S. Patent Application No. 16 / 101232, filed August 10, 2018. The real-time ray tracing hardware accelerator can be used to rapidly and efficiently determine the location and extent of an object (e.g., within a world model), generate real-time visualization simulations for radar signal interpretation, sound propagation synthesis and / or analysis, sonar system simulation, general wave propagation simulation, comparison with lidar data for localization and / or other functions and / or uses. In some embodiments, one or more tree traversal units (TTUs) may be used to perform one or more ray tracing-related operations.

[0098] One or more accelerators (e.g., hardware accelerator clusters) have wide applications in autonomous driving. A PVA (Programmable Vision accelerator) could be a programmable vision accelerator used in critical processing stages of ADA (Advanced Driver Assistance Systems) and autonomous vehicles. The capabilities of a PVA are well-suited for algorithmic domains requiring predictable processing with low power consumption and low latency. In other words, PVAs perform well on semi-intensive or conventionally intensive computations, even on small datasets that require predictable runtimes with low latency and low power consumption. Therefore, in the context of autonomous vehicle platforms, PVAs are designed to run classic computer vision algorithms, as they are highly efficient in object detection and integer mathematical operations.

[0099] For example, according to one embodiment of this technology, a PVA is used to perform computer stereo vision. In some examples, a semi-global matching-based algorithm can be used, although this is not intended to limit it. Many applications of Level 3-5 autonomous driving require instantaneous motion estimation / stereo matching (e.g., motion structures, pedestrian recognition, lane detection, etc.). A PVA can perform computer stereo vision functions on input from two monocular cameras.

[0100] In some examples, PVA can be used to perform dense optical flow, providing processed radar data based on the raw radar data (e.g., using 4D Fast Fourier Transform). In other examples, PVA is used for time-of-flight depth processing, for example, to provide processed time-of-flight data by processing the raw time-of-flight data.

[0101] DLA can be used to run any type of network to enhance control and driving safety, including neural networks that output a confidence measure for each object detection. Such a confidence value can be interpreted as a probability or to provide a relative “weight” for each detection relative to other detections. This confidence value allows the system to further determine which detections should be considered true positives rather than false positives. For example, the system can set a confidence threshold and only consider detections exceeding the threshold as true positives. In an Automatic Emergency Braking (AEB) system, false positives would cause the vehicle to automatically perform emergency braking, which is clearly undesirable. Therefore, only the most reliable detections should be considered as triggers for AEB. DLA can run a neural network to regress the confidence value. The neural network can take at least a subset of parameters as its input, such as bounding box dimensions, obtained ground plane estimates (e.g., from another subsystem), inertial measurement unit (IMU) sensor 666 outputs related to the vehicle's 600-degree orientation and distance, and three-dimensional position estimates of objects obtained from the neural network and / or other sensors (e.g., lidar sensor 664 or radar sensor 660).

[0102] One or more SoCs 604 may include one or more data storage units 616 (e.g., memory). The data storage unit 616 may be on-chip memory of the SoC 604, which may store neural networks to be executed on the GPU and / or DLA. In some examples, the capacity of the data storage unit 616 may be large enough to store multiple neural network instances for redundancy and security. The data storage unit 616 may include a level 2 or level 3 cache 612. As described herein, references to one or more data storage units 616 may include references to memory associated with the PVA, DLA, and / or one or more other accelerators 614.

[0103] One or more SoCs 604 may include one or more processors 610 (e.g., embedded processors). Processor 610 may include a boot and power management processor, which may be a dedicated processor and subsystem for handling boot power and management functions, as well as related security implementations. The boot and power management processor may be part of a boot sequence for one or more SoCs 604s and may provide runtime power management services. The boot power and management processor may provide clock and voltage programming, system low-power state transition assistance, management of SoC 604 thermal sensors and temperature sensors, and / or management of SoC 604 power states. Each temperature sensor may be implemented as a ring oscillator with an output frequency proportional to temperature, and one or more SoCs 604s may use the ring oscillator to detect the temperature of one or more CPUs 606s, one or more GPUs 608s, and / or one or more accelerators 614s. If a temperature is determined to exceed a threshold, the boot and power management processor may enter a temperature fault routine and place one or more SoCs 604s into a low-power state and / or place vehicle 600 into a driver-safe parking mode (e.g., safely parking vehicle 600).

[0104] One or more processors 610 may also include a set of embedded processors that can serve as an audio processing engine. The audio processing engine may be an audio subsystem capable of providing full hardware support for multi-channel audio through multiple interfaces and a wide range of flexible audio I / O interfaces. In some examples, the audio processing engine is a dedicated processor core of a digital signal processor with dedicated RAM.

[0105] One or more processors 610 may also include a normally-on processor engine that provides the necessary hardware functionality to support low-power sensor management and wake-up use cases. The normally-on processor engine may include a processor core, tightly coupled RAM, peripheral support (such as timers and interrupt controllers), various I / O controller peripherals, and routing logic.

[0106] One or more processors 610 may also include a secure cluster engine, which includes a dedicated processor subsystem for handling security management for automotive applications. The secure cluster engine may include two or more processor cores, tightly coupled RAM, support for peripheral devices (e.g., timers, interrupt controllers, etc.), and / or routing logic. In secure mode, the two or more cores may operate in lockstep mode and function as a single core with comparison logic to detect any differences in their operations.

[0107] One or more processors 610 may also include a real-time camera engine, which may include a dedicated processor subsystem for handling real-time camera management.

[0108] One or more processors 610 may also include a high dynamic range signal processor, which may include an image signal processor, which is a hardware engine as part of the camera processing pipeline.

[0109] One or more processors 610 may include a video image synthesizer, which may be a processing block (e.g., implemented on a microprocessor) that implements the video post-processing functions required by the video playback application to generate the final image of the player window. The video image synthesizer may perform lens distortion correction on one or more wide-angle cameras 670, one or more surround cameras 674, and / or in-cabin monitoring camera sensors. The in-cabin monitoring camera sensors are preferably monitored by a neural network running on another instance of an advanced SoC, configured to recognize in-cabin events and respond accordingly. The in-cabin system may perform lip reading to activate cellular service and make phone calls, dictate emails, change vehicle destination, activate or change the vehicle's infotainment system and settings, or provide voice-activated web browsing. Some functions are only available to the driver when the vehicle is running in automatic mode; otherwise, they are disabled.

[0110] Video image synthesizers may include enhanced temporal noise reduction for both spatial and temporal noise reduction. For example, in the case of motion in a video, noise reduction appropriately weights spatial information, thereby reducing the weight of information provided by adjacent frames. In cases where an image or part of an image does not contain motion, temporal noise reduction performed by the video image synthesizer can use information from the previous image to reduce noise in the current image.

[0111] The video image compositor can also be configured to perform stereoscopic correction on input stereoscopic shot frames. When the operating system desktop is in use, the video image compositor can also be used for user interface compositing, without requiring the GPU 608 to continuously render new surfaces. Even when one or more GPUs 608 are powered on and active during 3D rendering, the video image compositor can be used to offload one or more GPUs 608 to improve performance and responsiveness.

[0112] One or more SoCs 604 may also include a Mobile Industrial Processor Interface (MIPI) camera serial interface, a high-speed interface, and / or a video input block that can be used for receiving video and input from a camera and associated pixel input functions. One or more SoCs 604 may also include one or more input / output controllers that may be software-controlled and can be used to receive I / O signals that are not assigned a specific role.

[0113] One or more SoCs 604 may also include a wide range of peripheral interfaces for communication with peripheral devices, audio codecs, power management and / or other devices. One or more SoCs 604 may be used to process data from cameras (e.g., via gigabit multimedia serial links and Ethernet connections), sensors (e.g., one or more LiDAR sensors 664, one or more radar sensors 660, etc., connected via Ethernet), from bus 602 (e.g., vehicle 600 speed, steering wheel position, etc.), and from one or more GNSS sensors 658 (e.g., via Ethernet or CAN bus connections). One or more SoCs 604 may also include a dedicated high-performance, high-capacity memory controller, which may include its own DMA engine and may be used to free one or more CPUs 606 from routine data management tasks.

[0114] One or more SoCs 604 can form an end-to-end platform with a flexible architecture spanning automation levels 3-5, providing a comprehensive functional safety architecture that leverages and effectively utilizes computer vision and ADAS technologies to achieve diversity and redundancy, providing a platform for a flexible and reliable driver software stack and deep learning tools. Compared to traditional systems, one or more SoCs 604 can be faster, more reliable, and even more energy-efficient and space-saving. For example, when one or more accelerators 614 are combined with one or more CPUs 606, one or more GPUs 608, and one or more data storage units 616, a fast and efficient platform can be provided for Level 3-5 autonomous vehicles.

[0115] Therefore, this technology offers capabilities and functionalities that are unavailable in traditional systems. For example, computer vision algorithms can be executed on a CPU, which can be configured using a high-level programming language (such as C) to execute a wide variety of processing algorithms on diverse visual data. However, CPUs often cannot meet the performance requirements of many computer vision applications, such as those related to execution time and power consumption. In particular, many CPUs cannot execute complex object detection algorithms in real time, which is a requirement for automotive ADAS applications and practical Level 3-5 autonomous vehicles.

[0116] Compared to traditional systems, the techniques described in this paper, by providing CPU complexes, GPU complexes, and hardware acceleration clusters, allow for the simultaneous and / or sequential execution of multiple neural networks and the combination of results to achieve Level 3–5 autonomous driving capabilities. For example, a CNN executing on a DLA or dGPU (e.g., one or more GPU 620s) can include text and character recognition, allowing a supercomputer to read and understand traffic signs, including signs for which the neural network has not yet been specifically trained. The DLA can also include a neural network capable of recognizing, interpreting, and providing semantic understanding of symbols, and passing this semantic understanding to a path planning module running on the CPU complex.

[0117] Another example is the ability to run multiple neural networks simultaneously, as required for Level 3, 4, or 5 driving. For instance, a warning sign consisting of "Warning: Flashing lights indicate icing conditions" and a light can be interpreted independently or jointly by multiple neural networks. The sign itself can be recognized as a traffic sign by a first deployed neural network (e.g., a trained neural network), while the text "Flashing lights indicate icing conditions" can be interpreted by a second deployed neural network, which, when a flashing light is detected, notifies the vehicle routing software (preferably executed on a CPU complex) of the presence of icing conditions. A third deployed neural network can be used to identify the flashing light and notify the vehicle routing software of its presence by operating it across multiple frames. All three neural networks can run simultaneously, for example, within a DLA and / or on one or more GPUs 608.

[0118] In some examples, the CNN used for facial recognition and owner identification can use data from camera sensors to identify the presence of an authorized driver and / or the owner of the vehicle 600. The engine can be unlocked using a normally open sensor when the owner approaches the driver's door and turns on the lights, and the vehicle can be disabled in safe mode when the owner leaves. In this way, one or more SoCs 604 provide anti-theft and / or carjacking protection.

[0119] In another example, the CNN used for emergency vehicle detection and identification can use data from microphone 696 to detect and identify emergency vehicle sirens. Unlike conventional systems that use a general classifier to detect sirens and manually extract features, one or more SoCs(s) 604 use the CNN to classify environmental and urban sounds, as well as visual data. In a preferred embodiment, the CNN running on the DLA is trained to identify the relative closing speed of emergency vehicles (e.g., by using the Doppler effect). The CNN can also be trained to identify emergency vehicles specific to the vehicle's operating area, as identified by one or more GNSS sensors 658. Thus, for example, when operating in Europe, the CNN will seek to detect European sirens, and when operating in the United States, the CNN will seek to identify only North American sirens. Once an emergency vehicle is detected, an emergency vehicle safety routine can be executed using a control program with the aid of ultrasonic sensors 662, causing the vehicle to slow down, pull over, stop, and / or idle until one or more emergency vehicles pass.

[0120] The vehicle may include one or more CPUs 618 (e.g., one or more discrete CPUs or one or more dCPUs) coupled to one or more SoCs 604 via high-speed interconnects (e.g., PCIe). For example, one or more CPUs 618 may include x86 processors. The CPUs 618 may be used to perform any of a variety of functions, including arbitrating potentially inconsistent results between ADAS sensors and SoCs 604, and / or monitoring the status and health of one or more controllers 636 and / or infotainment SoCs 630.

[0121] Vehicle 600 may include one or more GPUs 620 (e.g., one or more discrete GPUs or one or more dGPUs) coupled to SoC 604 via high-speed interconnects (e.g., NVIDIA's NVLINK). One or more GPUs 620 may provide additional artificial intelligence capabilities, such as by executing redundant and / or different neural networks, and may be used to train and / or update neural networks based on inputs from sensors of vehicle 600 (e.g., sensor data).

[0122] Vehicle 600 may also include a network interface 624, which may include one or more wireless antennas 626 (e.g., one or more wireless antennas for different communication protocols, such as cellular antennas, Bluetooth antennas, etc.). Network interface 624 can be used to enable wireless connectivity via the Internet to the cloud (e.g., with one or more servers 678 and / or other network devices), other vehicles, and / or computing devices (e.g., a passenger's client device). For communication with other vehicles, direct and / or indirect links can be established between the two vehicles (e.g., via a network and the Internet). A vehicle-to-vehicle communication link can provide a direct link. A vehicle-to-vehicle communication link can provide vehicle 600 with information about vehicles near vehicle 600 (e.g., vehicles in front, to the side, and / or behind vehicle 600). This functionality may be part of vehicle 600's cooperative adaptive cruise control function.

[0123] Network interface 624 may include a SoC that provides modulation and demodulation functions and enables one or more controllers 636 to communicate over a wireless network. Network interface 624 may include an RF front-end for up-conversion from baseband to RF and down-conversion from RF to baseband. Frequency conversion can be performed using well-known processes and / or using superheterodyne processes. In some examples, the RF front-end functionality may be provided by a separate chip. The network interface may include wireless functions for communication via LTE, WCDMA, UMTS, GSM, CDMA2000, Bluetooth, Bluetooth LE, Wi-Fi, Z-Wave, ZigBee, LoRaWAN, and / or other wireless protocols.

[0124] The vehicle 600 may further include one or more data storage units 628, which may be off-chip (e.g., off-SoC). The data storage unit 628 may include one or more storage elements, including RAM, SRAM, DRAM, VRAM, flash memory, hard disk, and / or other components and / or devices capable of storing at least one bit of data.

[0125] The vehicle 600 may also include one or more GNSS sensors 658. One or more GNSS sensors 658 (e.g., GPS, assisted GPS sensors, differential GPS (DGPS) sensors, etc.) are used to assist in mapping, sensing, occupancy grid generation, and / or path planning functions. Any number of GNSS sensors 658 can be used, including, for example, but not limited to, GPS with a USB connector having an Ethernet-to-serial (RS-232) bridge.

[0126] Vehicle 600 may also include one or more radar sensors 660. Even in dark and / or inclement weather conditions, vehicle 600 can use one or more radar sensors 660 for remote vehicle detection. The radar functional safety level can be ASIL B. One or more radar sensors 660 can use CAN and / or bus 602 (e.g., to transmit data generated by one or more radar sensors 660) for control and access to target tracking data; in some examples, raw data is accessed via Ethernet. Various radar sensor types can be used. For example, but not limited to, one or more radar sensors 660 can be used for front, rear, and side radar applications. In some examples, pulse-Doppler radar sensors are used.

[0127] One or more radar sensors 660 may include different configurations, such as long-range with a narrow field of view, short-range with a wide field of view, short-range side coverage, etc. In some examples, long-range radar can be used for adaptive cruise control functions. Long-range radar systems can provide a wide field of view, for example, within a 250-meter range, achieved through two or more independent scans. One or more radar sensors 660 can help distinguish between static and moving objects and can be used by ADAS systems for emergency braking assistance and forward collision warning. Long-range radar sensors may include monostatic multimode radars with multiple (e.g., six or more) fixed radar antennas and high-speed CAN and FlexRay interfaces. In an example with six antennas, the four central antennas can create a focused beam pattern designed to record 600 elements of the environment around the vehicle at high speed with minimal traffic interference from adjacent lanes. The other two antennas can expand the field of view, enabling rapid detection of vehicles entering or leaving the vehicle within 600 lanes.

[0128] For example, a mid-range radar system may include a range of up to 660 meters (front) or 80 meters (rear), and a field of view of up to 42 degrees (front) or 650 degrees (rear). Short-range radar systems may include, but are not limited to, radar sensors designed to be mounted at both ends of the rear bumper. When mounted at both ends of the rear bumper, such radar sensor systems may generate two beams, continuously monitoring the blind spots behind and beside the vehicle.

[0129] ADAS systems can use short-range radar systems for blind spot detection and / or lane change assistance.

[0130] Vehicle 600 may also include one or more ultrasonic sensors 662. One or more ultrasonic sensors 662 may be located at the front, rear, and / or sides of vehicle 600 and may be used for parking assistance and / or creating and updating occupancy grids. Various ultrasonic sensors 662 may be used, and different ultrasonic sensors 662 may be used for different detection ranges (e.g., 2.5m, 4m). One or more ultrasonic sensors 662 may operate at the ASIL B functional safety level.

[0131] Vehicle 600 may include one or more lidar sensors 664. The one or more lidar sensors 664 may be used for object and pedestrian detection, emergency braking, collision avoidance, and / or other functions. The functional safety level of the one or more lidar sensors 664 may be ASIL B. In some examples, vehicle 600 may include multiple lidar sensors 664 (e.g., two, four, six, etc.) that can use Ethernet (e.g., to provide data to a Gigabit Ethernet switch).

[0132] In some examples, one or more LiDAR sensors 664 can provide a 360-degree field of view of objects and their distance lists. One or more commercially available LiDAR sensors 664 have an advertised range of approximately 600m, an accuracy of 2cm-3cm, and support, for example, a 600Mbps Ethernet connection. In some examples, one or more non-protruding LiDAR sensors 664 can be used. In such examples, one or more LiDAR sensors 664 can be implemented as small devices that can be embedded in the front, rear, sides, and / or corners of a vehicle 600. In such examples, one or more LiDAR sensors 664 can provide a horizontal field of view of up to 120 degrees and a vertical field of view of 35 degrees, with a range of 200m even for low-reflectivity objects. A front-mounted one or more LiDAR sensors 664 can be configured with a horizontal field of view between 45 and 135 degrees.

[0133] In some examples, lidar technology, such as 3D flash lidar, can also be used. 3D flash lidar uses a laser flash as a transmission source, illuminating approximately 200 meters around the vehicle. A flash lidar device includes a receiver that records the laser pulse transmission time and reflected light on each pixel, with each pixel corresponding to the range from the vehicle to the object. Flash lidar can utilize each laser flash to generate a highly accurate, distortion-free image of the environment. In some examples, four flash lidar sensors can be deployed, one on each side of the vehicle 600. Available 3D flash lidar systems include solid-state 3D staring array lidar cameras with no moving parts other than a fan (e.g., a non-scanning lidar device). The flash lidar device can use 5 nanosecond Class I (eye-safe) laser pulses per frame and can capture reflected laser light in the form of a 3D distance point cloud and co-registered intensity data. By using flash lidar, and because flash lidar is a solid-state device with no moving parts, one or more lidar sensors 664 can be less susceptible to motion blur, vibration, and / or shock.

[0134] The vehicle may also include one or more IMU sensors 666. In some examples, one or more IMU sensors 666 may be located at the center of the rear axle of the vehicle 600. One or more IMU sensors 666 may include, for example, but not limited to, one or more accelerometers, one or more magnetometers, one or more gyroscopes, one or more magnetic compasses, and / or other sensor types. In some examples, such as in a six-axis application, one or more IMU sensors 666 may include accelerometers and gyroscopes, while in a nine-axis application, one or more IMU sensors 666 may include accelerometers, gyroscopes, and magnetometers.

[0135] In some embodiments, one or more IMU sensors 666 can be implemented as a miniaturized, high-performance GPS-assisted inertial navigation system (GPS / INS) that combines a microelectromechanical system (MEMS) inertial sensor, a high-sensitivity GPS receiver, and an advanced Kalman filtering algorithm to provide estimates of position, velocity, and attitude. Thus, in some examples, one or more IMU sensors 666 can enable vehicle 600 to estimate heading by directly observing and correlating velocity changes from GPS to one or more IMU sensors 666, without requiring input from magnetic sensors. In some examples, one or more IMU sensors 666 and one or more GNSS sensors 658 can be combined in a single integrated unit.

[0136] The vehicle may include one or more microphones 696 placed inside and / or around the vehicle 600. One or more microphones 696 may be used for emergency vehicle detection and identification, etc.

[0137] The vehicle may also include any number of camera types, including one or more stereo cameras 668, one or more wide-angle cameras 670, one or more infrared cameras 672, one or more surround cameras 674, one or more long-range and / or mid-range cameras 698, and / or other camera types. The cameras can be used to capture image data of the entire perimeter of the vehicle 600. The types of cameras used depend on the embodiment and requirements of the vehicle 600, and any combination of camera types can be used to provide the necessary coverage around the vehicle 600. Furthermore, the number of cameras can vary depending on the embodiment. For example, the vehicle may include six cameras, seven cameras, ten cameras, twelve cameras, and / or other numbers of cameras. As an example, the cameras may support, but are not limited to, Gigabit Multimedia Serial Link (GMSL) and / or Gigabit Ethernet. This document will refer to... Figure 6A and Figure 6B Describe each camera in more detail.

[0138] Vehicle 600 may also include one or more vibration sensors 642. One or more vibration sensors 642 can measure vibrations of vehicle components, such as axles. For example, changes in vibration may indicate changes in road surface. In another example, when two or more vibration sensors 642 are used, differences between vibrations can be used to determine friction or slippage on the road surface (e.g., when the vibration difference is between a driven shaft and a freely rotating shaft).

[0139] Vehicle 600 may include ADAS system 638. In some examples, ADAS system 638 may include SoC. ADAS system 638 may include automatic / adaptive / automatic cruise control (ACC), cooperative adaptive cruise control (CACC), forward collision warning (FCW), automatic emergency braking (AEB), lane departure warning (LDW), lane keeping assist (LKA), blind spot warning (BSW), rear cross traffic warning (RCTW), collision warning system (CWS), lane centering (LC) and / or other features and functions.

[0140] The ACC system may use one or more radar sensors 660, one or more lidar sensors 664, and / or one or more cameras. The ACC system may include longitudinal ACC and / or lateral ACC. Longitudinal ACC monitors and controls the distance to the vehicle directly in front of vehicle 600 and automatically adjusts the vehicle speed to maintain a safe distance. Lateral ACC performs distance keeping and suggests lane changes to vehicle 600 if necessary. Lateral ACC is associated with other ADAS applications such as LCA and CWS.

[0141] CACC uses information from other vehicles, which can be received indirectly from other vehicles via a wireless link or a network connection (e.g., via the Internet) through network interface 624 and / or one or more wireless antennas 626. The direct link can be provided by a vehicle-to-vehicle (V2V) communication link, while the indirect link can be an infrastructure-to-vehicle (I2V) communication link. Generally, the V2V communication concept provides information about vehicles ahead (e.g., vehicles directly in front of vehicle 600 and in the same lane), while the I2V communication concept provides information about traffic ahead. A CACC system may include one or both I2V and V2V information sources. By taking into account information about vehicles ahead of vehicle 600, CACC may be more reliable and has the potential to improve traffic flow smoothness and reduce congestion on the road.

[0142] The Forward-Facing Warning (FCW) system is designed to alert the driver to hazards so that the driver can take corrective action. The FCW system uses a forward-facing camera and / or one or more radar sensors 660 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which are electrically coupled to driver feedback, such as a display, speaker, and / or vibration components. The FCW system can provide warnings such as audible, visual, haptic, and / or rapid braking pulses.

[0143] The AEB system detects an impending forward collision with another vehicle or other object. If the driver does not take corrective action within a specified time or distance parameter, the AEB system may automatically apply the brakes. The AEB system may use one or more front-facing cameras and / or one or more radar sensors 660 coupled to a dedicated processor, DSP, FPGA, and / or ASIC. When the AEB system detects a hazard, it typically first alerts the driver to take corrective action to avoid a collision. If the driver does not take corrective action, the AEB system may automatically apply the brakes to prevent or at least mitigate the effects of the anticipated collision. The AEB system may include technologies such as dynamic brake support and / or collision proximity braking.

[0144] The Lane Departure Warning (LDW) system provides visual, auditory, and / or tactile warnings, such as steering wheel or seat vibrations, to alert the driver when the vehicle crosses lane markings. The LDW system does not activate when the driver indicates intentional lane departure by activating a turn signal. The LDW system may use a front-facing camera coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to actuator feedback, such as a display, speaker, and / or vibration assembly.

[0145] The LKA system is a variant of the LDW system. If the vehicle begins to leave the lane at 60°, the LKA system provides steering input or braking to correct the vehicle's 60° deviation.

[0146] The BSW system detects and warns drivers of vehicles within the blind spot. The BSW system can provide visual, auditory, and / or tactile alerts to indicate unsafe merging or lane changing. Additional warnings may be provided when the driver uses a turn signal. The BSW system may use a rear-facing camera and / or radar sensor 660 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback, such as a display, speaker, and / or vibration assembly.

[0147] When the vehicle 600 is reversing and detects an object outside the range of the rear camera, the RCTW system can provide visual, auditory, and / or tactile notifications. Some RCTW systems include AEB (Autonomous Emergency Braking) to ensure the application of the vehicle's brakes to avoid a collision. The RCTW system may use one or more rear-facing radar sensors 660, which are coupled to a dedicated processor, DSP, FPGA, and / or ASIC, electrically coupled to actuator feedback, such as a display, speaker, and / or vibration assembly.

[0148] Traditional ADAS systems can be prone to false positives, which can be frustrating and distracting for drivers, but usually do not lead to catastrophic consequences because the ADAS system alerts the driver and allows them to determine whether a safe situation truly exists and take appropriate action. However, in an autonomous vehicle 600, in the event of conflicting results, the vehicle 600 itself must decide whether to heed the results from the main computer or the auxiliary computer (e.g., the first controller 636 or the second controller 636). For example, in some embodiments, the ADAS system 638 may be a backup and / or auxiliary computer for providing perception information to a backup computer module. A backup computer rationality monitor may run redundant different software on hardware components to detect faults in perception and dynamic driving tasks. The output from the ADAS system 638 may be provided to a monitoring MCU. If the outputs of the main computer and the auxiliary computer conflict, the monitoring MCU must determine how to reconcile the conflict to ensure safe operation.

[0149] In some examples, the master computer can be configured to provide a confidence score to the monitoring MCU, indicating the master computer's confidence level in the selected result. If the confidence score exceeds a threshold, the monitoring MCU can follow the master computer's direction regardless of whether the auxiliary computer provides conflicting or inconsistent results. If the confidence score does not meet the threshold, and the master computer and the auxiliary computer indicate different results (e.g., conflict), the monitoring MCU can arbitrate between the computers to determine the appropriate result.

[0150] The monitoring MCU can be configured to run one or more trained and configured neural networks to determine the conditions under which the auxiliary computer provides a false alarm based on the outputs of the main computer and the auxiliary computer. Thus, one or more neural networks in the monitoring MCU can learn when the output of the auxiliary computer is reliable and when it is not. For example, when the auxiliary computer is a radar-based FCW system, one or more neural networks in the monitoring MCU can learn when the FCW system identifies a metallic object that is not actually dangerous, such as a drain grille or manhole cover that triggers an alarm. Similarly, when the auxiliary computer is a camera-based LDW system, the neural network in the monitoring MCU can learn to override the LDW when a bicycle or pedestrian is present and lane departure is actually the safest maneuver. In embodiments that include one or more neural networks running on the monitoring MCU, the monitoring MCU may include at least one of a DLA or GPU adapted to run one or more neural networks with associated memory. In a preferred embodiment, the monitoring MCU may include and / or include components that are SoC 604.

[0151] In other examples, ADAS system 638 may include an auxiliary computer that performs ADAS functions using conventional computer vision rules. Therefore, the auxiliary computer can use classic computer vision rules (if-then), and the presence of one or more neural networks in the monitoring MCU can improve reliability, safety, and performance. For example, diverse implementations and intentional non-identification make the entire system more fault-tolerant, especially to failures caused by software (or hardware / software interface) functionality. For instance, if a software defect or error exists in the software running on the main computer, and different software code running on the auxiliary computer provides the same overall result, the monitoring MCU can have greater confidence in the correctness of the overall result, and the software or hardware defect on the main computer did not lead to a major error.

[0152] In some examples, the output of ADAS system 638 may be fed into the perception block and / or the dynamic drive task block of the host computer. For example, if ADAS system 638 indicates a forward collision warning due to an object directly in front, the perception block may use this information when identifying the object. In other examples, as described herein, the secondary computer may have its own trained neural network, thereby reducing the risk of false positives.

[0153] Vehicle 600 may also include an infotainment SoC 630 (e.g., an in-vehicle infotainment system (IVI)). Although shown and described as an SoC, the infotainment system may not be an SoC and may include two or more discrete components. The infotainment SoC 630 may include a combination of hardware and software that can be used to provide audio (e.g., music, personal digital assistant, navigation instructions, news, radio, etc.), video (e.g., television, movies, streaming media, etc.), telephone (e.g., hands-free calling), network connectivity (e.g., LTE, Wi-Fi, etc.), and / or provide information services to vehicle 600 (e.g., navigation system, rear parking assist, radio data system, vehicle-related information such as fuel level, total distance traveled, brake fuel level, engine oil level, door opening / closing, air filter information, etc.). For example, the infotainment SoC 630 may be a radio, disk player, navigation system, video player, USB and Bluetooth connectivity, in-vehicle computer, in-vehicle entertainment, Wi-Fi, steering wheel audio controls, hands-free voice control, head-up display (HUD), HMI display 634, telematics device, control panel (e.g., for controlling and / or interacting with various components, features, and / or systems), and / or other components. The infotainment SoC 630 may also be used to provide information to vehicle users (e.g., visual and / or auditory), such as information from ADAS system 638, autonomous driving information such as planned vehicle maneuvers, trajectory, surrounding environment information (e.g., intersection information, vehicle information, road information, etc.), and / or other information.

[0154] The infotainment SoC 630 may include GPU functionality. The infotainment SoC 630 can communicate with other devices, systems, and / or components of the vehicle 600 via bus 602 (e.g., CAN bus, Ethernet, etc.). In some examples, the infotainment SoC 630 may be coupled to a monitoring MCU, allowing the GPU of the infotainment system to perform self-driving functions in the event of a failure of one or more main controllers 636 (e.g., the main computer and / or backup computer of the vehicle 600). In such an example, the infotainment SoC 630 may place the vehicle 600 into a driver-to-safe parking mode, as described herein.

[0155] Vehicle 600 may also include an instrument cluster 632 (e.g., a digital instrument panel, electronic instrument cluster, etc.). The instrument cluster 632 may include a controller and / or a supercomputer (e.g., a discrete controller or supercomputer). The instrument cluster 632 may include a set of instruments such as a speedometer, fuel level, oil pressure, tachometer, odometer, turn signals, shift position indicator, one or more seatbelt warning lights, one or more parking brake warning lights, one or more engine malfunction lights, airbag (SRS) system information, lighting controls, safety system controls, navigation information, etc. In some examples, information may be displayed and / or shared between the infotainment SoC 630 and the instrument cluster 632. In other words, the instrument cluster 632 may be included as part of the infotainment SoC 630, and vice versa.

[0156] Figure 6D It is one or more cloud-based servers according to some embodiments of this disclosure and Figure 6A A system diagram illustrating communication between exemplary autonomous vehicles 600 is provided. System 676 may include one or more servers 678, one or more networks 690, and vehicles, including vehicle 600. One or more servers 678 may include multiple GPUs 684(a)-684(H) (collectively referred to as GPU 684), PCIe switches 682(a)-682(H) (collectively referred to as PCIe switch 682), and / or CPUs 680(A)-680(B) (collectively referred to herein as CPU 680). GPUs 684, CPUs 680, and PCIe switches may be interconnected via high-speed interconnects, such as, but not limited to, NVIDIA-developed NVLink interface 688 and / or PCIe connection 686. In some examples, GPUs 684 are connected via NVLink and / or NV switch SoCs, and GPUs 684 and PCIe switches 682 are connected via PCIe interconnects. Although eight GPUs 684, two CPUs 680, and two PCIe switches are illustrated, this is not intended to be limiting. According to an embodiment, each of one or more servers 678 may include any number of GPUs 684, CPUs 680, and / or PCIe switches. For example, one or more servers 678 may each include eight, sixteen, thirty-two, and / or more GPUs 684.

[0157] One or more servers 678 may receive image data representing images from vehicles via one or more networks 690, showing unexpected or changed road conditions, such as recently started roadwork. One or more servers 678 may send neural network 692, updated neural network 692, and / or map information 694, including information about traffic and road conditions, to vehicles via one or more networks 690. Updates to map information 694 may include updates to HD map 622, such as information about construction sites, potholes, detours, floods, and / or other obstacles. In some examples, neural network 692, updated neural network 692, and / or map information 694 may originate from new training and / or experience represented in data received from any number of vehicles in the environment, and / or based on training performed in a data center (e.g., using one or more servers 678 and / or other servers).

[0158] One or more servers 678 can be used to train machine learning models (e.g., neural networks) based on training data. Training data may be generated by vehicles and / or generated in simulations (e.g., using a game engine). In some examples, the training data is labeled (e.g., the neural network benefits from supervised learning) and / or undergoes other preprocessing, while in other examples, the training data is unlabeled and / or unprocessed (e.g., the neural network does not require supervised learning). Training can be performed according to any one or more classes of machine learning techniques, including but not limited to: supervised training, semi-supervised training, unsupervised training, self-learning, reinforcement learning, joint learning, transfer learning, feature learning (including principal component analysis and cluster analysis), multilinear subspace learning, manifold learning, representation learning (including alternative dictionary learning), rule-based machine learning, anomaly detection, and any variations or combinations thereof. Once the machine learning model is trained, it can be used by vehicles (e.g., transmitted to vehicles via one or more networks 690, and / or used by one or more servers 678 for remote monitoring of vehicles).

[0159] In some examples, one or more servers 678 may receive data from a vehicle and apply the data to state-of-the-art real-time neural networks for real-time intelligent inference. One or more servers 678 may include a deep learning supercomputer and / or a dedicated AI computer powered by a GPU 684, such as the DGX and DGX Station machines developed by NVIDIA. However, in some examples, one or more servers 678 may include a deep learning infrastructure in a data center using only CPU power.

[0160] The deep learning infrastructure of one or more servers 678 can perform rapid real-time inference and use this capability to assess and verify the health status of the processors, software, and / or associated hardware in vehicle 600. For example, the deep learning infrastructure can receive periodic updates from vehicle 600, such as sequences of images and / or objects located by vehicle 600 in the image sequence (e.g., through computer vision and / or other machine learning object classification techniques). The deep learning infrastructure can run its own neural network to identify objects and compare them with objects identified by vehicle 600. If the results do not match and the infrastructure concludes that the AI ​​in vehicle 600 has malfunctioned, one or more servers 678 can send a signal to vehicle 600 instructing the fail-safe computer of vehicle 600 to take control, notify passengers, and complete a safe stopping operation.

[0161] For inference, one or more servers 678 may include one or more GPUs 684 and one or more programmable inference accelerators (such as NVIDIA's TensorRT). The combination of GPU-powered servers and inference acceleration can enable real-time responses. In other examples, such as where performance is less critical, inference can be performed using servers powered by CPUs, FPGAs, and other processors.

[0162] Example computing device Figure 7 This is a block diagram of an example computing device 700 suitable for implementing some embodiments of the present disclosure. The computing device 700 may include an interconnect system 702 directly or indirectly coupled to: a memory 704, one or more central processing units (CPUs) 706, one or more graphics processing units (GPUs) 708, a communication interface 710, input / output (I / O) ports 712, input / output components 714, a power supply 716, one or more presentation components 718 (e.g., displays), and one or more logic units 720. In at least one embodiment, one or more computing devices 700 may include one or more virtual machines (VMs), and / or any component thereof may include virtual components (e.g., virtual hardware components). For a non-limiting example, one or more GPUs 708 may include one or more vGPUs, one or more CPUs 706 may include one or more vCPUs, and / or one or more logic units 720 may include one or more virtual logic units. Therefore, one or more computing devices 700 may include discrete components (e.g., a complete GPU dedicated to computing device 700), virtual components (e.g., a portion of the GPU dedicated to computing device 700), or a combination thereof.

[0163] although Figure 7The various modules are shown as being connected to lines via interconnect system 702, but this is not intended to be limiting, but merely for clarity. For example, in some embodiments, a presentation component 718 such as a display device can be considered as I / O component 714 (e.g., if the display is a touchscreen). As another example, CPU 706 and / or GPU 708 may include memory (e.g., memory 704 may also represent a storage device in addition to the memory of GPU 708, CPU 706, and / or other components). In other words, Figure 7 The computing devices mentioned are merely illustrative. No distinction is made between "workstation," "server," "laptop," "desktop," "tablet," "client device," "mobile device," "handheld device," "game console," "electronic control unit (ECU)," "virtual reality system," and / or other device or system types, as is the case in [the context of the previous sentence]. Figure 7 As envisioned within the scope of computing devices.

[0164] Interconnect system 702 may represent one or more links or buses, such as address buses, data buses, control buses, or combinations thereof. Interconnect system 702 may include one or more bus or link types, such as Industry Standard Architecture (ISA) buses, Extended Industry Standard Architecture (EISA) buses, Video Electronics Standards Association (VESA) buses, Peripheral Component Interconnect (PCI) buses, Peripheral Component Interconnect Through (PCIE) buses, and / or other types of buses or links. In some embodiments, there is a direct connection between components. For example, CPU 706 may be directly connected to memory 704. Furthermore, CPU 706 may be directly connected to GPU 708. In cases where there is a direct connection or point-to-point connection between components, interconnect system 702 may include a PCIe link for performing the connection. In these examples, a PCI bus is not required in computing device 700.

[0165] The memory 704 may include any of a variety of computer-readable media. Computer-readable media can be any available medium accessible by the computing device 700. Computer-readable media may include volatile and non-volatile media, as well as removable and non-removable media. By way of example and not limitation, computer-readable media may include computer storage media and communication media.

[0166] Computer storage media may include volatile and non-volatile media and / or removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, and / or other data types. For example, memory 704 may store computer-readable instructions (e.g., instructions representing one or more programs and / or one or more program elements), such as an operating system. Computer storage media may include, but is not limited to, RAM, ROM, EEPROM, flash memory or other storage technologies, CD-ROM, digital versatile disk (DVD) or other optical disc storage, cassette tape, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and is accessible by computing device 700. As used herein, computer storage media itself does not include signals.

[0167] Computer storage media can contain computer-readable instructions, data structures, program modules, and / or other data types contained in modulated data signals (such as carrier waves or other transmission mechanisms), and includes any information delivery medium. The term "modulated data signal" can refer to a signal whose one or more characteristics are set or altered in a manner that encodes information in the signal. By way of example and not limitation, computer storage media can include wired media, such as wired networks or direct wired connections, and wireless media, such as acoustic, RF, infrared, and other wireless media. Any combination of the foregoing should also be included within the scope of computer-readable media.

[0168] One or more CPUs 706 may be configured to execute at least some computer-readable instructions to control one or more components of computing device 700 to perform one or more methods and / or processes described herein. Each of the one or more CPUs 706 may include one or more cores capable of processing multiple software threads simultaneously (e.g., one, two, four, eight, twenty-eight, seventy-two, etc.). The one or more CPUs 706 may include any type of processor and may include different types of processors depending on the type of computing device 700 implemented (e.g., processors with fewer cores for mobile devices and processors with more cores for servers). For example, depending on the type of computing device 700, the processor may be an advanced RISC machine (ARM) processor implemented using Reduced Instruction Set Computing (RISC) or an x86 processor implemented using Complex Instruction Set Computing (CISC). In addition to one or more microprocessors or supplementary coprocessors (e.g., math coprocessors), computing device 700 may also include one or more CPUs 706.

[0169] In addition to, or selected from, one or more CPUs 706, one or more GPUs 708 may be configured to execute at least some computer-readable instructions to control one or more components of computing device 700 to perform one or more methods and / or processes described herein. One or more GPUs 708 may be integrated GPUs (e.g., having one or more CPUs 706 and / or one or more GPUs 708 may be discrete GPUs). In embodiments, one or more GPUs 708 may be coprocessors of one or more CPUs 706. Computing device 700 may use one or more GPUs 708 to render graphics (e.g., 3D graphics) or perform general-purpose computing. For example, one or more GPUs 708 may be used for general-purpose computing on a GPU (GPGPU). One or more GPUs 708 may include hundreds or thousands of cores capable of processing hundreds or thousands of software threads simultaneously. One or more GPUs 708 may generate pixel data for an output image in response to rendering commands (e.g., rendering commands received via a host interface from one or more CPUs 706). One or more GPUs 708 may include graphics memory, such as display memory, for storing pixel data or any other suitable data, such as GPGPU data. Display memory may be included as part of memory 704. One or more GPUs 708 may include two or more GPUs operating in parallel (e.g., via a link). The link may directly connect the GPUs (e.g., using NVLINK) or connect the GPUs via a switch (e.g., using NVSwitch). When combined, each GPU 708 may generate pixel data or GPGPU data for different portions of the output or different outputs (e.g., a first GPU for a first image, a second GPU for a second image). Each GPU may include its own memory or may share memory with other GPUs.

[0170] In addition to one or more CPUs 706 and / or one or more GPUs 708, or selected from one or more CPUs 706 and / or one or more GPUs 708, one or more logic units 720 may be configured to execute at least some computer-readable instructions to control one or more components of computing device 700 to perform one or more methods and / or processes described herein. In embodiments, one or more CPUs 706, one or more GPUs 708, and / or one or more logic units 720 may execute any combination of methods, processes, and / or portions thereof discretely or jointly. One or more logic units 720 may be part of and / or integrated into one or more CPUs 706 and / or GPUs 708, and / or one or more logic units 720 may be discrete components or otherwise located external to one or more CPUs 706 and / or one or more GPUs 708. One or more logic units 720 may be coprocessors of one or more CPUs 706 and / or one or more GPUs 708.

[0171] Examples of one or more logic units 720 include one or more processing cores and / or components thereof, such as a data processing unit (DPU), a tensor core (TC), a tensor processing unit (TPU), a pixel vision core (PVC), a vision processing unit (VPU), a graphics processing cluster (GPC), a texture processing cluster (TPC), a streaming multiprocessor (SM), a tree traversal unit (TTU), an artificial intelligence accelerator (AIA), a deep learning accelerator (DLA), an arithmetic logic unit (ALU), an application-specific integrated circuit (ASIC), a floating-point unit (FPU), input / output (I / O) elements, peripheral component interconnect (PCI) or peripheral component interconnect pass-through (PCIe) elements, and / or the like.

[0172] The communication interface 710 may include one or more receivers, transmitters, and / or transceivers that enable the computing device 700 to communicate with other computing devices via an electronic communication network including wired and / or wireless communications. The communication interface 710 may include components and functions to support communication over any of a variety of different networks, such as wireless networks (e.g., Wi-Fi, Z-Wave, Bluetooth, Bluetooth LE, ZigBee, etc.), wired networks (e.g., communication over Ethernet or InfiniBand), low-power wide-area networks (e.g., LoRaWAN, SigFox, etc.), and / or the Internet. In one or more embodiments, the logic unit 720 and / or the communication interface 710 may include one or more data processing units (DPUs) for directly transmitting data received via a network and / or via interconnect system 702 to one or more GPUs 708 (e.g., memory).

[0173] I / O port 712 enables the computing device 700 to be logically coupled to other devices, including I / O components 714, one or more presentation components 718, and / or other components, some of which may be built into (e.g., integrated into) the computing device 700. Illustrative I / O components 714 include microphones, mice, keyboards, joysticks, game pads, game controllers, satellite antennas, scanners, printers, wireless devices, etc. I / O components 714 can provide a natural user interface (NUI) that processes user-generated air gestures, voice, or other physiological input. In some cases, the input can be transmitted to appropriate network elements for further processing. The NUI can implement any combination of voice recognition, stylus recognition, facial recognition, biometrics, on-screen and near-screen gesture recognition, air gestures, head and eye tracking, and touch recognition associated with the display of the computing device 700 (described in more detail below). The computing device 700 may include depth cameras, such as stereo camera systems, infrared camera systems, RGB camera systems, touchscreen technology, and combinations thereof, for gesture detection and recognition. In addition, the computing device 700 may include an accelerometer or gyroscope capable of detecting motion (e.g., as part of an inertial measurement unit (IMU)). In some examples, the output of the accelerometer or gyroscope may be used by the computing device 700 to render immersive augmented reality or virtual reality.

[0174] Power supply 716 may include hard-wired power supply, battery power supply, or a combination thereof. Power supply 716 may supply power to computing device 700 so that components of computing device 700 can operate.

[0175] One or more presentation components 718 may include displays (e.g., monitors, touchscreens, television screens, head-up displays (HUDs), other display types, or combinations thereof), speakers, and / or other presentation components. One or more presentation components 718 may receive data from other components (e.g., one or more GPUs 708, one or more CPUs 706, DPUs, etc.) and output data (e.g., as images, videos, sounds, etc.).

[0176] Example Data Center Figure 8 An example data center 800 that can be used in at least one embodiment of this disclosure is shown. The data center 800 may include a data center infrastructure layer 810, a framework layer 820, a software layer 830, and / or an application layer 840.

[0177] like Figure 8 As shown, the data center infrastructure layer 810 may include a resource coordinator 812, packet computing resources 814, and node computing resources (“nodes CRs”) 816(1)-816(N), where “N” represents any integer. In at least one embodiment, nodes CRs 816(1)-816(N) may include, but are not limited to, any number of central processing units (CPUs) or other processors (including DPUs, accelerators, field-programmable gate arrays (FPGAs), graphics processors or graphics processing units (GPUs), etc.), storage devices (e.g., dynamic read-only memory), and in some embodiments, storage devices (e.g., solid-state or disk drives), network input / output (NW I / O) devices, network switches, virtual machines (“VMs”), power modules and / or cooling modules, etc., and one or more nodes CRs 816(1)-816(N) may correspond to a server having one or more of the aforementioned computing resources. In addition, in some embodiments, nodes CRs 816(1)-8161(N) may include one or more virtual components, such as vGPU, vCPU and / or similar components, and / or one or more nodes CRs 816(1)-816(N) may correspond to virtual machines (VMs).

[0178] In at least one embodiment, the packet computing resource 814 may include individual packets of node CRs 816 located within one or more racks (not shown), or multiple racks located within data centers in different geographical locations (also not shown). Individual packets of node CRs 816 within the packet computing resource 814 may include packet computing, networking, memory, or storage resources that can be configured or allocated to support one or more workloads. In at least one embodiment, multiple node CRs 816, including CPUs, GPUs, DPUs, and / or other processors, may be grouped within one or more racks to provide computing resources to support one or more workloads. One or more racks may also include any combination of any number of power modules, cooling modules, and / or network switches.

[0179] Resource coordinator 812 may be configured or otherwise control one or more nodes CRs 816(1)-816(N) and / or grouped computing resources 814. In at least one embodiment, resource coordinator 812 may include a Software Design Infrastructure (SDI) management entity for data center 800. Resource coordinator 812 may include hardware, software, or some combination thereof.

[0180] In at least one embodiment, such as Figure 8 As shown, framework layer 820 may include job scheduler 833, configuration manager 834, resource manager 836, and / or distributed file system 838. Framework layer 820 may include frameworks for software 832 supporting software layer 830 and / or one or more applications 842 supporting application layer 840. Software 832 or one or more applications 842 may respectively include web-based service software or applications, such as services provided by Amazon Web Services, Google Cloud, and Microsoft Azure. Framework layer 820 may be, but is not limited to, a free and open-source software web application framework, such as Apache Spark™ (hereinafter referred to as "Spark"), which can utilize distributed file system 838 for large-scale data processing (e.g., "big data"). In at least one embodiment, job scheduler 833 may include Spark drivers to facilitate the scheduling of workloads supported by various layers of data center 800. Configuration manager 834 may configure different layers, such as software layer 830 and framework layer 820, including Spark and distributed file system 838, to support large-scale data processing. Resource manager 836 may be able to manage cluster or group computing resources mapped to or allocated to support distributed file system 838 and job scheduler 833. In at least one embodiment, the cluster or group computing resources may include group computing resources 814 at data center infrastructure layer 810. Resource manager 836 may coordinate with resource coordinator 812 to manage these mapped or allocated computing resources.

[0181] In at least one embodiment, the software 832 included in the software layer 830 may include software used in at least a portion of the distributed file system 838 of the nodes CRs 816(1)-816(N), the grouped computing resources 814, and / or the framework layer 820. One or more software may include, but are not limited to, internet web search software, email virus scanning software, database software, and streaming video content software.

[0182] In at least one embodiment, the application 842 included in the application layer 840 may include one or more types of applications used by at least a portion of the nodes CRs 816(1)-816(N), the grouped computing resources 814 and / or the distributed file system 838 of the framework layer 820, but is not limited to any number of genomics applications, perceptual computing and machine learning applications, including training or inference software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.) and / or other machine learning applications used in combination with one or more embodiments.

[0183] In at least one embodiment, any of the configuration manager 834, resource manager 836, and resource coordinator 812 can implement any number and type of self-modification actions based on any amount and type of data acquired in any technically feasible manner. Self-modification actions can protect data center operators of data center 800 from making potentially erroneous configuration decisions and may prevent underutilized and / or poorly performing portions of the data center.

[0184] According to one or more embodiments described herein, data center 800 may include tools, services, software, or other resources for training one or more machine learning models or using one or more machine learning models to predict or infer information. For example, one or more machine learning models may be trained by calculating weight parameters based on a neural network architecture using the software and / or computing resources described above regarding data center 800. In at least one embodiment, a trained or deployed machine learning model corresponding to one or more neural networks may be used to infer or predict information using the resources described above regarding data center 800 by using weight parameters calculated through one or more training techniques (e.g., but not limited to the training techniques described herein).

[0185] In at least one embodiment, the data center 800 may use a CPU, application-specific integrated circuit (ASIC), GPU, FPGA, and / or other hardware (or corresponding virtual computing resources) to perform training and / or inference using the aforementioned resources. Furthermore, one or more of the aforementioned software and / or hardware resources may be configured as services to allow a user to train or perform information inference, such as image recognition, speech recognition, or other artificial intelligence services.

[0186] Example network environment A suitable network environment for implementing embodiments of the present invention may include one or more client devices, servers, network attached storage (NAS), other backend devices, and / or other device types. Client devices, servers, and / or other device types (e.g., each device) may be... Figure 7 Implemented on one or more instances of computing devices 700—for example, each device may include similar components, features, and / or functions of computing device 700. Additionally, backend devices (servers, NAS, etc.) may be included as part of data center 800, examples of which are referred to herein. Figure 8 To describe in more detail.

[0187] Components of a network environment can communicate with each other through one or more networks, whether wired, wireless, or a combination of both. This network can include multiple networks, or networks of networks. For example, the network can include one or more wide area networks (WANs), one or more local area networks (LANs), one or more public networks (such as the Internet and / or the Public Switched Telephone Network (PSTN)), and / or one or more private networks. Where the network includes a wireless telecommunications network, components such as base stations, communication towers, or even access points (and other components) can provide wireless connectivity.

[0188] A compatible network environment may include one or more peer-to-peer network environments—in which case the server may not be included in the network environment—and one or more client-server network environments—in which case one or more servers may be included in the network environment. In a peer-to-peer network environment, the server functionality described herein can be implemented on any number of client devices.

[0189] In at least one embodiment, the network environment may include one or more cloud-based network environments, distributed computing environments, and combinations thereof. The cloud-based network environment may include a framework layer, a job scheduler, a resource manager, and a distributed file system implemented on one or more servers, which may include one or more core network servers and / or edge servers. The framework layer may include a framework supporting software at the software layer and / or application at the application layer. The software or one or more applications may respectively include web-based service software or applications. In embodiments, one or more client devices may use web-based service software or applications (e.g., by accessing the service software and / or applications via one or more application programming interfaces (APIs)). The framework layer may be, but is not limited to, a free and open-source software web application framework, for example, one that can use a distributed file system for large-scale data processing (e.g., "big data").

[0190] A cloud-based network environment can provide cloud computing and / or cloud storage for any combination of the computing and / or data storage functions (or one or more of them) described herein. Any of these various functions can be distributed from a central or core server (e.g., servers in one or more data centers) to multiple locations, which may be located in a state, region, country, globally, etc. If the connection to a user (e.g., a client device) is relatively close to one or more edge servers, then one or more core servers may assign at least a portion of the functionality to one or more edge servers. A cloud-based network environment can be private (e.g., limited to a single organization), public (e.g., available to many organizations), and / or a combination thereof (e.g., a hybrid cloud environment).

[0191] One or more client devices may include the information described in this article. Figure 7 The client device may be, by way of example and not limitation, at least some of the components, features and functions of one or more example computing devices 700. By way of example and not limitation, the client device may be a personal computer (PC), laptop computer, mobile device, smartphone, tablet computer, smartwatch, wearable computer, personal digital assistant (PDA), MP3 player, virtual reality headset, global positioning system (GPS) or device, video player, camera, surveillance equipment or system, vehicle, ship, aircraft, virtual machine, drone, robot, handheld communication device, hospital equipment, gaming device or system, entertainment system, vehicle computer system, embedded system controller, remote control, device, consumer electronics device, workstation, edge device, any combination of these depicted devices or any other suitable device.

[0192] This disclosure can be described in the general context of computer code or machine-usable instructions, including computer-executable instructions, such as program modules, that are executed by a computer or other machine (such as a personal data assistant or other handheld device). Typically, program modules include routines, programs, objects, components, data structures, etc., and refer to code that performs a specific task or implements a specific abstract data type. This disclosure can be implemented in a variety of system configurations, including handheld devices, consumer electronics, general-purpose computers, and more specialized computing devices. The invention can also be implemented in distributed computing environments, where tasks are performed by remote processing devices linked via a communication network.

[0193] As used herein, the phrase “and / or” relating to two or more elements should be interpreted as referring to only one element or a combination of elements. For example, “element A, element B, and / or element C” may include only element A, element B, element C, element A and element B, element A and element C, element B and element C, or element A, element B, and element C. Furthermore, “at least one of element A or element B” may include at least one of element A, at least one of element B, or at least one of element A and at least one of element B. Additionally, “at least one of element A and element B” may include at least one of element A, at least one of element B, or at least one of element A and at least one of element B.

[0194] To meet legal requirements, the subject matter of this disclosure has been described in detail herein. However, the description itself is not intended to limit the scope of this disclosure. Rather, the inventors have envisioned that the claimed subject matter may also be embodied in other ways, including different steps or combinations of steps similar to those described herein, as well as other existing or future techniques. Furthermore, although the terms “step” and / or “block” may be used herein to imply different elements of the method employed, they should not be construed as implying any particular order between the steps disclosed herein unless the order of the individual steps is explicitly described.

Claims

1. A processor, comprising: One or more circuits are used for: Using one or more neural networks and based at least on image data generated using one or more image sensors, calculate first data indicating a reduction in visibility distance associated with the one or more image sensors, the visibility distance being associated with the image data; Based at least on the image data, determine the confidence level associated with the detected object; Based at least on the first data indicating a decrease in the visibility distance, and based on the fact that the distance to the detected object exceeds the visibility distance, a reduction in the confidence associated with the detected object is determined; Based at least in part on the reduced confidence level, determine one or more operations for this machine; and Control the machine according to one or more of the operations.

2. The processor of claim 1, wherein the first data further represents at least one of the following: Visual blindness classification associated with the one or more image sensors; The illumination level associated with the one or more image sensors; or The level of visual impairment associated with the one or more image sensors.

3. The processor of claim 1, wherein the one or more circuits are further configured to: Intermediate data is generated using one or more neural networks; and A subset of the first data is computed based at least on the intermediate data and using at least one head of the one or more neural networks, wherein the subset of the first data represents at least one of the following: Sensor blindness level associated with the one or more image sensors; Visual blindness classification associated with the one or more image sensors; The illumination level associated with the one or more image sensors; or The visibility distance associated with the one or more image sensors.

4. The processor of claim 3, wherein the at least one head comprises one or more downsampling layers and subsequently one or more upsampling layers.

5. The processor of claim 3, wherein the at least one header includes at least one fully connected layer.

6. The processor of claim 2, wherein the determination of the one or more operations of the machine is further based on weighting at least one of the visual impairment level, the visual impairment classification, the illumination level, or the visibility distance according to one or more predefined weights.

7. The processor of claim 6, wherein the visual impairment level has the highest weight among the one or more predefined weights.

8. The processor of claim 1, wherein the one or more operations are associated with an operation level corresponding to a vehicle autonomy level, the vehicle autonomy level including level 0 (L0), level 1 (L1), level 2 (L2), level 3 (L3), level 4 (L4), or level 5 (L5).

9. The processor of claim 1, wherein the determination of the one or more operations is further based at least on sensor data generated using one or more other sensor modes.

10. The processor of claim 1, wherein the processor is included in at least one of the following: Control systems for autonomous or semi-autonomous machines; Sensing systems for autonomous or semi-autonomous machines; A system used to perform simulation operations; A system used to perform deep learning operations; A system containing one or more virtual machines (VMs); Systems that are at least partially implemented in data centers; or A system that utilizes cloud computing resources at least in part.

11. The processor of claim 1, wherein the one or more circuits are further configured to: At least one or more parameters are determined based on the image data, said one or more parameters including at least one of road surface condition, scene type classification, or distance to the scene corresponding to said scene type classification. The determination of one or more operations of the machine is further based at least on one or more parameters.

12. A system comprising: One or more processing units, for: Calculated using one or more neural networks and based at least on image data generated using one or more image sensors: First data indicates whether the visibility distance associated with the one or more image sensors has decreased, the visibility distance being associated with the farthest object that can be identified using the image data; as well as The second data indicates a first visual blindness classification and a second visual blindness classification, wherein the first visual blindness classification indicates that at least a portion of the image data is associated with visual blindness, and the second visual blindness classification indicates at least one of the causes or levels of visual blindness associated with the at least a portion of the image data; The operating level of this machine shall be determined based at least on the first data and the second data. as well as At least one or more control operations are determined based on the operation level.

13. The system of claim 12, wherein the determination of the operation level is further based at least on third data generated using one or more other sensor modes.

14. The system of claim 12, wherein determining the operation level of the machine comprises: The illumination level of the one or more image sensors, the visibility distance, the visual blindness level of the one or more image sensors, the first visual blindness classification, or the second visual blindness classification are weighted according to one or more predefined weights.

15. The system of claim 14, wherein the visual blindness quantity has the highest weight among the one or more predefined weights.

16. The system of claim 12, wherein the system comprises at least one of: Control systems for autonomous or semi-autonomous machines; Sensing systems for autonomous or semi-autonomous machines; A system used to perform simulation operations; A system used to perform deep learning operations; A system containing one or more virtual machines (VMs); Systems that are at least partially implemented in data centers; or A system that utilizes cloud computing resources at least in part.

17. The system of claim 12, wherein the one or more processing units are further configured to: Using the one or more neural networks and at least based on the image data generated using the one or more image sensors, a third data point representing a lighting level is calculated, which is associated with the amount of light perceived by the user and related to the machine. The operating level of the machine is further determined at least based on the third data.

18. The system of claim 12, wherein the one or more processing units are further configured to: Using the one or more neural networks and based at least on the image data generated using the one or more image sensors, calculate third data representing the amount of visual blindness associated with the one or more image sensors. The operating level of the machine is further determined at least based on the third data.

19. The system of claim 12, wherein the one or more processing units are further configured to: Using the one or more neural networks and at least based on the image data generated using the one or more image sensors, a third data point representing a lighting level is calculated, which is associated with the amount of light perceived by the user and related to the machine. The operating level of the machine is further determined at least based on the third data.

20. A method comprising: Using one or more neural networks and based at least on image data generated using one or more image sensors on this machine, calculate at least: The first data represents the visibility distance associated with the one or more image sensors, the visibility distance being associated with the distance to one or more objects that can be detected using the image data; as well as The second data represents the surface condition of the road as indicated by the image data; Based at least on the first data and the second data, determine the operation level associated with the local machine; and The operation of this machine is controlled according to the operation level.

21. The method of claim 20, wherein the surface condition indicates that the road is at least one of dry, wet, icy, or snowy.

Citation Information

Patent Citations

  • Method for programmable timeouts of tree traversal mechanisms in hardware

    US10885698B2

  • Deep neural network processing for sensor blindness detection in autonomous machine applications

    US11508049B2

  • Visibility distance estimation using deep learning in autonomous machine applications

    US20230110027A1

  • System, method, and processor-readable medium for autonomous vehicle reliability assessment

    CN110998470A

  • Multi-view deep neural network for lidar awareness

    CN112904370A