Refining machine learning models to mitigate against attacks in autonomous systems and applications

By fine-tuning the machine learning model by adversarial image dataset, the problem of insufficient robustness of machine learning models for adversarial attacks in the prior art is solved, and more efficient computing and safer systems are achieved.

CN120239876APending Publication Date: 2025-07-01NVIDIA CORP

Patent Information

Application Number
CN202280101744.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2022-11-11
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

The prior art is difficult to effectively improve the robustness of machine learning models to adversarial attacks, resulting in visual driving systems generating incorrect predictions when facing adversarial attacks.

Method used

By fine-tuning the pretrained machine learning model using adversarial images, generate adversarial image datasets and refine the machine learning model to improve the robustness of its adversarial attacks.

Benefits of technology

The machine learning model is achieved with higher robustness to adversarial attacks, reducing computational overhead and resource consumption, and improving system security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120239876A_ABST
    Figure CN120239876A_ABST
Patent Text Reader

Abstract

In various examples, a technique for processing sensor data includes generating, using a machine learning model and based on a first instance of sensor data, a first set of confidence for a set of output types, and a first antagonism confidence representing a likelihood that the first instance of sensor data is antagonistic. The technique further includes determining, based on the first antagonism confidence, that the first sensor data instance is antagonistic. The technique further includes transmitting a first indication to one or more downstream components indicating that the first instance of sensor data is antagonistic, such that the one or more downstream components perform one or more operations based at least on the indication.
Need to check novelty before this filing date? Find Prior Art

Description

BACKGROUND OF THE INVENTION

[0001] Vision-based autonomous or semi-autonomous driving systems analyze sensor data to understand the environment around an autonomous or semi-autonomous vehicle (e.g., a driverless car). These systems typically use machine learning models to process sensor data collected using sensors on the vehicle and use their output to perform various operations, such as detecting, classifying, and / or tracking pedestrians, animals, buildings, road conditions, traffic signs, obstacles, other vehicles, and / or other objects or scenes around the vehicle. This output can be used to guide driving or control decisions to assist the vehicle in operating in a safe manner.

[0002] In some cases, the operation of vision-based driving systems can be hindered by adversarial attacks, causing the machine learning model to generate incorrect, inaccurate, or imprecise predictions. For example, an adversarial attack may involve any of a slight to severe perturbation of the color, shape, text, or other visual attributes of environmental features (e.g., traffic signs, vehicles or other dynamic objects, road markings, etc.). Taking traffic signs as an example, these perturbations may manifest in a form that causes the machine learning model to incorrectly classify or categorize the traffic sign (e.g., by predicting a speed limit significantly higher or lower than the actual speed limit shown on the traffic sign). When such an adversarial attack is successful, the incorrect predictions generated by the vision-based driving system may cause one or more downstream components or systems of the vehicle to make inappropriate decisions.

[0003] Existing methods for reducing the sensitivity of vision-based driving systems to adversarial attacks include augmenting or enhancing the training dataset used to train the machine learning model before it is deployed to the vision-based driving system. These methods can also (or alternatively) use an ensemble of multiple machine learning models to generate predictions, which the vision-based driving system then uses to make or guide driving decisions. However, these techniques do not guarantee that the machine learning model can withstand adversarial attacks manifested in sensor data, which are typically perturbed in a specific way, causing the machine learning model to output incorrect predictions.

[0004] Therefore, more effective techniques are needed to improve the robustness of machine learning models (e.g., models in vision-based driving systems) against adversarial attacks. SUMMARY OF THE INVENTION

[0005] Embodiments of the present disclosure relate to techniques for refining machine learning models to mitigate adversarial attacks. The techniques described herein include using a machine learning model and generating, based on a first sensor data instance, a first set of confidences for a set of output types and a first adversarial confidence representing the probability that the first sensor data instance is adversarial. The techniques further include determining that the first sensor data instance is adversarial based on the first adversarial confidence. The techniques also include transmitting a first indication that the first sensor data instance is adversarial to one or more downstream components such that the one or more downstream components can perform one or more operations based at least in part on the indication.

[0006] One technical advantage of the techniques of the present disclosure over conventional solutions is that the machine learning model can detect, interpret, and assist in the notification of adversarial attacks against the system. Thus, a system using the output of such a machine learning model may be more secure and more resistant to adversarial attacks than a conventional system that includes a machine learning model not trained to identify and / or defend against adversarial attacks. Another technical advantage of the techniques of the present disclosure is that it reduces computational overhead and resource consumption compared to existing technical methods that use a collection of multiple machine learning models to defend against adversarial attacks. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] The present system and method for refining a machine learning model to mitigate adversarial attacks are described in detail below with reference to the accompanying drawings, in which:

[0008] Figure 1 illustrates a computing device configured to implement one or more aspects of various embodiments;

[0009] Figure 2 is a more detailed illustration of an Figure 1 evaluation engine, refinement engine, and execution engine according to various embodiments;

[0010] Figure 3 illustrates a flowchart of a method for fine-tuning a machine learning model using adversarial data according to various embodiments;

[0011] Figure 4 illustrates a flowchart of a method for processing sensor data according to various embodiments;

[0012] Figure 5A is an illustration of an example autonomous vehicle according to some embodiments of the present disclosure;

[0013] Figure 5B is an Figure 5A example of the camera positions and fields of view of an example autonomous vehicle according to some embodiments of the present disclosure;

[0014] Figure 5C is anFigure 5A Block diagram of an example system architecture of an example autonomous vehicle;

[0015] Figure 5D is a system diagram of the communication between a cloud-based server and Figure 5A an example autonomous vehicle according to some embodiments of the present disclosure;

[0016] Figure 6 is a block diagram of an example computing device suitable for implementing some embodiments of the present disclosure; and

[0017] Figure 7 is a block diagram of an example data center suitable for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION

[0018] Systems and methods for mitigating adversarial attacks by refining machine learning models are disclosed. Although the present disclosure may be described in connection with example autonomous or semi-autonomous vehicles or machines 500 (also referred to herein as "vehicle 500" or "self 500"), examples of which are described in connection with Figures 5A to 5D this is not intended to limit the invention. For example, the systems and methods described herein can be used (but are not limited to) non-autonomous vehicles or machines, semi-autonomous vehicles or machines (e.g., in one or more adaptive driver assistance systems (ADAS)), autonomous vehicles or machines, manned and unmanned robots or robotic platforms, warehouse vehicles, off-road vehicles, vehicles coupled to one or more trailers, aircraft, ships, shuttles, emergency response vehicles, motorcycles, electric or motorized bicycles, airplanes, engineering vehicles, underwater vehicles, drones, and / or other types of vehicles. Additionally, although the present disclosure may be described in terms of mitigating adversarial attacks against machine learning models used in autonomous or semi-autonomous vehicles, this is not intended to be limiting, and the systems and methods described herein can be used in augmented reality, virtual reality, mixed reality, robotics, security and surveillance, autonomous or semi-autonomous machine applications, and / or any other technical field where adversarial attacks may occur.

[0019] As described herein, the proper operation of a vision-based driving system in an autonomous or semi-autonomous vehicle can be affected by adversarial attacks that cause one or more machine learning models of the system to generate incorrect predictions. When such adversarial attacks are successful, these incorrect predictions can cause the autonomous or semi-autonomous vehicle or machine (or any other system subject to adversarial attacks) to perform one or more operations incorrectly or to fail to achieve a desired or required level of accuracy or precision.

[0020] To improve the ability of a vision driving system to defend against adversarial attacks, the disclosed techniques use adversarial images (or other types of adversarial data, such as Light Detection and Ranging (LiDAR) data, Radio Detection and Ranging (RADAR) data, ultrasonic data, and / or data generated using any other type of sensor, such as but not limited to the data described herein for Figures 5A to 5D vehicle 500 in

[0021] to fine-tune a pre-trained machine learning model. Adversarial images can be generated by perturbing a set of original images based on one or more types of adversarial attack techniques. After generating the adversarial images, they can be processed using the pre-trained machine learning model, and the predictions generated using the pre-trained machine learning model can be compared with the ground truth data (e.g., labels) corresponding to the original images from which the adversarial images were generated. Based on these comparisons, the adversarial images are divided into a "clean" adversarial image dataset (for which the pre-trained machine learning model generated correct predictions) and an "adversarial" adversarial image dataset (for which the pre-trained machine learning model generated incorrect predictions).

[0022] Then, the adversarial image datasets and the corresponding ground truth data are used to refine the machine learning model. During the refinement process, the adversarial images in the clean dataset can be processed using the machine learning model, and the machine learning model can be trained based on the loss between the predictions generated by the machine learning model based on the adversarial images and the ground truth data of the original images corresponding to the adversarial images.

[0023] After training a machine learning model using adversarial images and corresponding ground truth data, the machine learning model can be deployed (e.g., in a vision driving system) to generate predictions for other images. For example, the machine learning model can be used to detect and / or track objects near an autonomous vehicle during operation of the autonomous vehicle. The output of the machine learning model can be provided to one or more downstream components that generate commands and / or signals for operating the autonomous vehicle.

[0024] One technical advantage of the disclosed technology, relative to traditional solutions, is that the machine learning model is able to detect and (at least in some cases) assist in correcting adversarial data that attempts to cause the machine learning model to operate incorrectly. Thus, a system that uses the output of the machine learning model may be safer and more resistant to adversarial attacks compared to traditional systems that include machine learning models that are not trained to identify and / or defend against adversarial attacks. Another technical advantage of the disclosed technology, compared to traditional methods that use an ensemble of multiple machine learning models to defend against adversarial attacks, is reduced computational overhead and resource consumption.

[0025] Figure 1 A computing device 100 configured to implement one or more aspects of various embodiments is shown. In at least one embodiment, the computing device 100 includes a desktop computer, a laptop computer, a smartphone, a personal digital assistant (PDA), a tablet, a server, one or more virtual machines, and / or any other type of computing device configured to receive input, process data, and optionally display images, and suitable for practicing one or more embodiments. The computing device 100 is configured to run an evaluation engine 122, a refinement engine 124, and an execution engine 126, which may reside in the memory 116. It should be noted that the computing devices described herein are exemplary, and any other technically feasible configuration falls within the scope of the present disclosure. For example, multiple instances of the evaluation engine 122, the refinement engine 124, and / or the execution engine 126 may be executed on a set of nodes in a distributed and / or cloud computing system to implement the functionality of the computing device 100.

[0026] In one embodiment, computing device 100 includes, but is not limited to, an interconnect (bus) 112 that couples one or more processors 102, an input / output (I / O) device interface 104 that couples to one or more input / output (I / O) devices 108, a memory 116, a storage device 114, and / or a network interface 106. Processor 102 may include any suitable processor implemented as a central processing unit (CPU), a graphics processing unit (GPU), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), an artificial intelligence (AI) accelerator, a parallel processing unit (PPU), a data processing unit (DPU), any other type of processing unit, or a combination of different processing units, e.g., a CPU configured to operate in conjunction with a GPU. Generally, processor 102 may include any technically feasible hardware unit capable of processing data and / or executing software applications. Additionally, in the context of the present disclosure, the computing elements shown in computing device 100 may correspond to physical computing systems (e.g., systems in a data center) and / or may correspond to virtual computing instances executing in a computing cloud.

[0027] In at least one embodiment, I / O device 108 includes devices capable of receiving input, such as a keyboard, a mouse, a touchpad, a VR / MR / AR headset, a gesture recognition system, and / or a microphone, and devices capable of providing output, such as a display device and / or a speaker. Additionally, I / O device 108 may also include devices capable of receiving input and providing output, such as a touchscreen, a universal serial bus (USB) port, etc. I / O device 108 may be configured to receive various types of input from an end user (e.g., a designer) of computing device 100 and may also provide various types of output to the end user of computing device 100, such as a displayed digital image, digital video, or text. In some embodiments, one or more I / O devices 108 are configured to couple computing device 100 to network 110.

[0028] In one embodiment, network 110 is any technically feasible communication network that allows for the exchange of data between computing device 100 and internal, local, remote, or external entities or devices (such as a web server or other networked computing device). For example, network 110 may include a wide area network (WAN), a local area network (LAN), a wireless (e.g., WiFi) network, and / or the Internet, etc.

[0029] In at least one embodiment, the storage device 114 includes a non-volatile storage device for storing applications and data, and may include a fixed or removable disk drive, a flash device, and CD-ROM, DVD-ROM, Blu-ray, HD-DVD, or other magnetic, optical, or solid-state storage devices. The evaluation engine 122, the refinement engine 124, and / or the execution engine 126 may be stored in the storage device 114 and loaded into the memory 116 when executed.

[0030] In one embodiment, the memory 116 includes random access memory (RAM) modules, flash memory cells, and / or any other type of memory cells or combinations thereof. The processor 102, the I / O device interface 104, and the network interface 106 may be configured to read data from and write data to the memory 116. The memory 116 may include various software programs executable by the processor 102 and application data related to the software programs, including the evaluation engine 122, the refinement engine 124, and / or the execution engine 126.

[0031] The evaluation engine 122 includes functions for evaluating the performance of a machine learning model in generating predictions based on adversarial examples. For example, the evaluation engine 122 may generate a set of adversarial images by applying various types of perturbations to a set of original images (e.g., captured images, synthetic images, etc.). The evaluation engine 122 may input each adversarial image into a neural network for detecting objects in the image. The evaluation engine 122 may also compare the output generated using the neural network with corresponding ground truth data (e.g., labels, annotations, etc.) corresponding to the original images from which the adversarial images were generated. Then, the evaluation engine 122 may divide the adversarial images into two or more data sets. The first data set may include adversarial images for which the neural network correctly predicted the output for the corresponding original image, and the second data set may include adversarial images for which the neural network did not correctly predict the output for the corresponding original image.

[0032] The refinement engine 124 performs additional training and / or fine-tuning of the machine learning model using the data set generated by the evaluation engine 122. Continuing the above example, the refinement engine 124 can train a neural network (e.g., update one or more parameters of the neural network) to minimize a first loss between a prediction generated using the machine learning model based on an adversarial image in the first data set and the ground truth (label) for the corresponding original image. The refinement engine 124 can train the neural network to predict different labels for different subsets of adversarial images in the second data set. More specifically, the refinement engine 124 can divide the second data set into a first subset of adversarial images associated with a perturbation intensity above a threshold amount and a second subset of adversarial images associated with a perturbation intensity below the threshold amount (which threshold can be the same or different from the threshold for adversarial images). The refinement engine 124 can train the neural network in such a way that a second loss between a prediction generated using the machine learning model from an adversarial image in the first subset and an “adversarial label” that indicates that the image (or image region) should be ignored or otherwise treated differently (e.g., by ignoring the portion of the image identified as adversarial) when determining an action to be performed by one or more downstream components is minimized. The refinement engine 124 can also train the neural network in such a way that a third loss between a prediction generated using the machine learning model from an adversarial image in the second subset and the label or other ground truth data for the corresponding original image is minimized.

[0033] The execution engine 126 uses the trained machine learning model to generate predictions for other data. For example, the execution engine 126 can deploy and execute the trained machine learning model in an autonomous or semi-autonomous software driving stack (“driving stack”) of an autonomous or semi-autonomous machine. During machine operation, the trained machine learning model can be used to perform various operations, such as detecting and tracking objects in images captured using the machine's camera. The output of the machine learning model can be provided to one or more downstream components (e.g., a control component, a behavior planning component, an actuation component, a world model management component, etc.), which use the output to generate one or more commands and / or make other decisions related to the operation of the machine. Thus, the evaluation engine 122, the refinement engine 124, and the execution engine 126 can be used to improve the robustness of the machine learning model and / or (more generally) the driving system against adversarial attacks, as discussed in further detail herein.

[0034] Figure 2 is according to various embodiments Figure 1 A more detailed illustration of the evaluation engine 122, the refinement engine 124, and the execution engine 126 in. As described above, the evaluation engine 122, the refinement engine 124, and the execution engine 126 are used to improve the robustness of the machine learning model 208 against adversarial attacks. Each of these components will be described in more detail herein.

[0035] In some embodiments, the machine learning model 208 includes a pre-trained model for generating predictions related to a set of images 232. For example, the machine learning model 208 may include one or more recurrent neural networks (RNNs), convolutional neural networks (CNNs), deep neural networks (DNNs), deep convolutional neural networks (DCNs), residual neural networks (ResNets), graph neural networks, autoencoders, transformer neural networks, deep stereo geometry networks (DSGNs), stereo R-CNNs, and / or other types of artificial neural networks or components of artificial neural networks. The machine learning model 208 may also (or alternatively) include regression models, support vector machines, decision trees, random forests, gradient boosting trees, naive Bayes classifiers, Bayesian networks, hidden Markov models (HMMs), hierarchical models, ensemble models, clustering techniques, and / or other types of machine learning models that do not use artificial neural network components. The machine learning model 208 may be pre-trained to generate outputs that can be used to perform various operations, such as detecting and / or tracking objects, detecting and / or classifying road conditions, and / or performing other functions related to the images 232. The machine learning model 208 may also (or alternatively) be pre-trained to assist in predicting the trajectories, paths, distances, and / or behaviors of vehicles, pedestrians, and / or other moving or dynamic objects. The machine learning model 208 may also (or alternatively) be pre-trained to assist in generating trajectories for navigating autonomous or semi-autonomous vehicles or machines, given a destination, available routes to the destination, detected obstacles, and / or other criteria.

[0036] As described herein, the evaluation engine 122 analyzes the performance of the machine learning model 208 when processing a set of adversarial images 204. As Figure 2As shown, the evaluation engine 122 generates adversarial images 204 by applying a set of perturbations 202 to a set of original images 210. For example, the evaluation engine 122 can use a set of known adversarial attack techniques (e.g., Fast Gradient Sign Method (FGSM), Projected Gradient Descent (PGD), Limited-memory Broyden-Fletcher-Goldfarb-Shanno (L-BFGS), Carlini & Wagner, Jacobian-based Saliency Map Attack (JSMA), DeepFool, etc.) to apply the perturbations 202 to the original images 210 to generate a corresponding set of adversarial images 204. The evaluation engine 122 can also (or alternatively) use perturbations 202 specified, selected, or defined by one or more users to transform the original images 210 into adversarial images 204. The evaluation engine 122 can also or alternatively apply multiple perturbations 202 to a single original image to generate a corresponding adversarial image, interpolate between two or more perturbations 202 to generate a perturbation applied to the original image, and / or otherwise combine multiple perturbations 202 to generate an adversarial image from the original image.

[0037] The perturbations 202 include changes in the pixel values in the original images 210, which may cause the machine learning model 208 to generate incorrect predictions 212. For example, the perturbations 202 can include changes in the pixel values in the original images 210 that are imperceptible to humans but are intended to cause the machine learning model 208 to generate incorrect or inaccurate outputs, such as incorrectly predicting the location and / or category associated with the object in the original image 210. In another example, the perturbations 202 can include changes in the pixel values in certain "blocks" or "regions" of the original images 210, which may cause the machine learning model 208 to detect ghost objects, misclassify the objects in the original images 210, and / or otherwise generate incorrect or inaccurate predictions 212. In a third example, the perturbations 202 can include changes to the original images 210 that are both perceptible to humans and capable of causing the machine learning model 208 to generate incorrect predictions 212. In a fourth example, the perturbations 202 may include a large number of changes to the original images 210 that will cause both humans and the machine learning model 208 to incorrectly predict the object category, object location, and / or other attributes associated with the objects in the original image 210.

[0038] In one or more embodiments, the original image 210 includes an image from the training dataset of the machine learning model 208. For example, in an automotive use case, the original image 210 can include images depicting roads, traffic signs, vehicles, pedestrians, intersections, lanes, obstacles, and / or other objects near the vehicle. The evaluation engine 122, the refinement engine 124, the execution engine 126, and / or other components can train the machine learning model 208 to detect and / or track objects in the original image 210 based on the bounding shape, class, and / or other labels 206 of the objects.

[0039] Alternatively, the original image 210 can include images that have not been used to train the machine learning model 208. For example, in an automotive use case, the original image 210 can include images of roads, traffic signs, vehicles, pedestrians, intersections, lanes, obstacles, and / or other objects that are different from the objects depicted in the images of the training dataset of the machine learning model 208. The original image 210 can also (or alternatively) include enhanced, synthetic images, and / or other types of images that are visually different from the images in the training dataset.

[0040] After generating the adversarial image 204 that includes the perturbation 202 of the original image 210, the evaluation engine 122 can input or apply the adversarial image 204 (e.g., the image data or sensor data representing it, whether preprocessed or not) to the machine learning model 208. The evaluation engine 122 compares the output generated by the machine learning model 208 from each adversarial image with one or more labels 206 (or other ground truth data) of the original image from which the adversarial image was generated. The evaluation engine 122 uses the comparison results to identify a set of incorrect predictions 212 made using the machine learning model 208 from a first subset of the adversarial images 204, and a set of correct predictions 214 made using the machine learning model 208 from a second subset of the adversarial images.

[0041] For example, the evaluation engine 122 can input a single (or batch) adversarial image from the adversarial image set 204 into the machine learning model 208. The machine learning model 208 can process image or sensor data, and the system can obtain an output from the machine learning model 208 - for example, prediction probabilities representing a set of object classes, as the corresponding output of the machine learning model 208. The evaluation engine 122 can compare the object class with the highest prediction probability to the label (or other ground truth data type) for the corresponding original image. The evaluation engine 122 can also verify whether the highest prediction probability reaches or exceeds a threshold. When the object class with the highest prediction probability matches the label and / or the highest prediction probability reaches or exceeds the threshold, the evaluation engine 122 can determine that the machine learning model 208 has made a correct prediction for the adversarial image. When the object class with the highest prediction probability does not match the label and / or the highest prediction probability does not reach or exceed the threshold, the evaluation engine 122 can determine that the machine learning model 208 has made an incorrect prediction for the adversarial image. Although described as outputting probabilities, other output types are also within the scope of the present disclosure. For example, the machine learning model 208 can output confidence, percentage, binary output (e.g., 0 or 1), and / or can perform regression on one or more output values.

[0042] The evaluation engine 122 uses the incorrect predictions 212 and the correct predictions 214 to generate two data sets associated with the adversarial images 204. As Figure 2 shown, the evaluation engine 122 generates an adversarial data set 216 that includes a set of adversarial images 220 for which the machine learning model 208 has generated incorrect predictions 212. The evaluation engine 122 also generates a clean data set 218 that includes a different set of adversarial images 220 for which the machine learning model has generated correct predictions 214. Thus, the adversarial images 220 and 222 correspond to two disjoint subsets of the adversarial image set 204 input into the machine learning model 208.

[0043] The evaluation engine 122 can populate the clean data set 218 and the adversarial data set 216 with various labels or other ground truth data types associated with the corresponding adversarial images 220 and 222. More specifically, the evaluation engine 122 associates the adversarial images 222 in the clean data set 218 with the original labels 226 for the corresponding original images 210. Thus, the clean data set 218 includes the adversarial images 222 corresponding to the perturbed versions of the subset of the original images 210, and the original labels 226 corresponding to the subset of the labels 206 for these original images 210.

[0044] To add labels to the adversarial dataset 216, the evaluation engine 122 determines the perturbation strength 240 associated with the adversarial images 220 in the adversarial dataset 216. Each perturbation strength represents a measure or indication of the extent to which the original image was perturbed to generate the corresponding adversarial image. For example, the evaluation engine 122 can calculate each perturbation strength as the number or proportion of pixels in the original image that were perturbed to generate the adversarial image, the extent to which the pixel values in the original image were perturbed to generate the adversarial image, and / or another measure of the amount of perturbation applied to the original image to generate the adversarial image. In another example, the evaluation engine 122 can calculate the perturbation strength as a measure of the difference between the prediction output by the machine learning model 208 for a given adversarial image and the label of the corresponding original image. In a third example, the evaluation engine 122 can calculate the perturbation strength as a weighted combination of multiple measures and / or indications of the amount of perturbation associated with the corresponding adversarial image.

[0045] In some embodiments, the perturbation strength 240 is determined based on user input associated with a manual review of the adversarial images 220 and / or the corresponding original images 210. For example, the evaluation engine 122 can present one or more adversarial images 220 and a list of possible object categories for the objects in the adversarial images to one or more users. Each user can select the most likely object category for the object depicted in a given adversarial image. Each user can also (or alternatively) indicate that, given the appearance of the object in the adversarial image, the object category cannot be determined. The object categories correctly identified by the users will indicate a low perturbation strength associated with the adversarial image. The object categories misidentified by the users and / or the objects marked by the users as having an undetermined object category will indicate a high perturbation strength. Given a comparison of the adversarial image with the original image from which the adversarial image was generated, each user can also (or alternatively) provide a score, rating, and / or other input indicating the level of perturbation (e.g., low, medium, high, etc.) associated with the adversarial image. When multiple users provide input related to the perturbation strength associated with a given adversarial image, the input can be averaged and / or otherwise aggregated to form an overall indication of the perturbation strength of the user-specified adversarial image.

[0046] After calculating and / or determining the perturbation strength 240 of the adversarial images 220, the evaluation engine 122 assigns labels to the adversarial images 220 in the adversarial dataset 216 using one or more thresholds of the perturbation strength 240. More specifically, the adversarial dataset 216 includes a first subset of the adversarial images 220 that are associated with a set of original labels 224 for the corresponding original images 210. The adversarial dataset 216 also includes a second subset of the adversarial images 220 that are associated with a set of adversarial labels 228.

[0047] To assign adversarial labels 228 and / or original labels 224 to adversarial images 220, the evaluation engine 122 compares the perturbation intensity of each adversarial image in the adversarial image 220 with a perturbation intensity threshold. If the perturbation intensity does not reach or exceed the threshold, the evaluation engine 122 retrieves the label for the original image from which the adversarial image was generated from the label set 206 for the original image set 210. The evaluation engine 122 adds this label to the adversarial data set 216 (e.g., as part of the original label set 224), and associates this label with the adversarial image (e.g., by mapping the adversarial image to the label in the adversarial data set 216). In other words, the evaluation engine 122 identifies a first subset of the adversarial images 220 associated with the incorrect prediction 212 of the machine learning model 208 and a perturbation intensity 240 that does not reach or exceed the threshold (or another threshold). Then, the evaluation engine 122 adds the first subset of the adversarial images 220 and the original labels 224 including the subset of labels 206 corresponding to the original images 210 to the adversarial data set 216.

[0048] If the perturbation intensity of a given adversarial image in the adversarial image 220 reaches or exceeds the threshold, the evaluation engine 122 assigns an adversarial label (e.g., in the adversarial label 228) to the adversarial image in the adversarial data set 216. This adversarial label corresponds to an adversarial class that indicates the extent to which the adversarial image (or an object in the adversarial image) has been perturbed or corrupted such that the adversarial image (or object) should be ignored or processed differently (e.g., a downstream component can receive an indication of the adversarial nature of the data used to generate an output, and the downstream component can perform an operation based on the indication), including the machine learning model 208.

[0049] In some embodiments, the evaluation engine 122 generates adversarial labels 228 and / or original labels 224 in the adversarial data set 216 based on multiple thresholds of the perturbation intensity 240. For example, the evaluation engine 122 can compare the computed and / or user-generated perturbation intensity 240 of the adversarial image 220 with three or more thresholds or values. If the computed perturbation intensity of the adversarial image is below all thresholds and / or contains a user-specified low perturbation intensity, the evaluation engine 122 can generate a corresponding label that contains the value 1 of the corresponding original label and the value 0 of the adversarial label. If the computed perturbation intensity of the adversarial image is above all thresholds and / or the user-specified perturbation intensity is high, the evaluation engine 122 can generate a label that contains the value 0 of the corresponding original label and the value 1 of the adversarial label.

[0050] Continuing with the above example, if the adversarial image has a computed perturbation strength between two thresholds and / or a user-specified perturbation strength between low and high, the evaluation engine 122 can generate a corresponding label that includes a first non-zero value for the corresponding original label and a second non-zero value for the adversarial label, where the sum of the two non-zero values is 1. Thus, an adversarial image with a "low-medium" perturbation strength may have a label with a value of 0.75 for the corresponding original label and a value of 0.25 for the adversarial label, an adversarial image with a "medium" perturbation strength may have a label with values of 0.5 for both the corresponding original label and the adversarial label, and an adversarial image with a "medium-high" perturbation strength may have a label with a value of 0.25 for the corresponding original label and a value of 0.75 for the adversarial label. Thus, the numerical values assigned to the adversarial label and the original label for a given adversarial image represent the probability or confidence that the object belongs to the corresponding class. In some embodiments, in the case where the output of the machine learning model indicates the degree or value of the adversarial nature of the sensor data, downstream components can operate in different ways based on that degree or value (e.g., more adversarial, ignore the output; less adversarial, use the output but with reduced weighting).

[0051] After generating the clean dataset 218 and the adversarial dataset 216, the refinement engine 124 uses the clean dataset 218 and the adversarial dataset 216 to perform additional training and / or fine-tuning on the machine learning model 208. More specifically, the refinement engine 124 inputs the adversarial images 220 from the adversarial dataset 216 into the machine learning model 208 and obtains the corresponding adversarial dataset predictions 260 from the machine learning model 208. The adversarial dataset predictions 260 include numerical confidences representing the predicted probabilities (or other output types) of the various classes (or other prediction types) represented by the labels 206, as well as numerical confidences representing the predicted probabilities of the adversarial labels for the corresponding adversarial images 220. The refinement engine 124 calculates the mean squared error, cross-entropy loss, and / or one or more other losses 264 between the output confidences and the corresponding adversarial labels 228 and / or original labels 224 of the adversarial images 220. Then, the refinement engine 124 uses training techniques (e.g., gradient descent and backpropagation) to update the model parameters 230 (e.g., weights and biases) of the machine learning model 208 in a way that reduces the computed loss 264.

[0052] The refinement engine 124 also inputs each adversarial image 222 from the clean data set 218 into the machine learning model 208 and obtains corresponding clean data set predictions 262 from the machine learning model 208. The clean data set predictions 262 can also include numerical confidences representing the prediction probabilities of the various classes represented by the labels 206, as well as numerical confidences representing the prediction probabilities of the adversarial labels for the corresponding adversarial images 222. The refinement engine 124 calculates the mean squared error, cross-entropy loss, and / or one or more other losses 266 between the output confidence and the corresponding original label 226 for the adversarial image 222. The refinement engine 124 also uses training techniques (e.g., gradient descent and backpropagation) to update the model parameters 230 of the machine learning model 208 in a manner that reduces the calculated loss 264.

[0053] Accordingly, the refinement engine 124 performs one or more training phases that train the machine learning model 208 to predict the original label 226 associated with the adversarial image 222 in the clean data set 218. This training of the machine learning model 208 using the clean data set 218 allows the machine learning model 208 to continue to generate correct predictions 214 for the adversarial image 222 and / or other images similar to the adversarial image 222.

[0054] The refinement engine 124 also performs one or more training phases that train the machine learning model 208 to predict the original label 224 and / or the adversarial label 228 associated with the adversarial image 220 in the adversarial data set 216. This training of the machine learning model 208 using the adversarial data set 216 allows the machine learning model 208 to learn to predict the original label 224 for the adversarial image 220 associated with the lower perturbation intensity 240, thereby correcting the predictive output of the machine learning model 208 for these types of adversarial images 220 and / or images similar to these types of adversarial images 220. Training the model 208 using the adversarial data set 216 also allows the machine learning model 208 to learn to predict the adversarial label 228 for the adversarial image 220 (and / or other images similar to the adversarial image 220) associated with the higher perturbation intensity 240, thereby enabling the machine learning model 208 to identify adversarial data and flag adversarial data that might otherwise be used to disrupt the operation of the machine learning model 208 and / or downstream components that use the output of the machine learning model 208.

[0055] After the refinement engine 124 has fine-tuned the model parameters 230 of the machine learning model 208 based on the losses 264-266, the execution engine 126 uses the fine-tuned machine learning model 208 to generate adversarial labels 234 and / or non-adversarial labels 236 for additional images 232 that are not in the adversarial dataset 216 or the clean dataset 218. For example, the execution engine 126 can deploy and execute the machine learning model 208 in a corresponding system (such as an autonomous or semi-autonomous driving system). During the operation of an autonomous or semi-autonomous vehicle or machine, images 232 of the environment around the autonomous vehicle can be captured using one or more cameras of the machine or vehicle, as described in further detail herein Figures 5A to 5D as described. The captured images 232 can be input or applied to the machine learning model 208, which can process the data, and the output of the machine learning model 208 can be used to perform various operations, such as detecting and tracking objects near the vehicle or machine. The output of the machine learning model 208 can include non-adversarial labels 236 corresponding to object categories, boundary shapes, semantic segmentation, predicted paths, and / or other output types. The output of the machine learning model 208 can also (or alternatively) include adversarial labels 234 that indicate that the corresponding image 232 and / or region of the image 232 contains adversarial data, and thus can generate an indicator, message, or signal indicating the adversarial data.

[0056] The execution engine 126 also uses the adversarial labels 234 and non-adversarial labels 236 to make certain decisions and / or perform certain actions. Continuing the above example, the execution engine 126 can provide the adversarial labels 234 and non-adversarial labels 236 calculated using the machine learning model 208 to one or more downstream components that use the output (e.g., for generating driving commands and / or making other decisions related to the operation of the autonomous or semi-autonomous machine or vehicle). Since the downstream components are able to sense the images 232 and / or regions of the images 232 that contain adversarial data, the downstream components are able to take into account the adversarial nature of the data and operate based on this information. Thus, the machine learning model 208 can be robust against adversarial attacks.

[0057] Although the above has been mainly described with respect to image data and autonomous or semi-autonomous driving Figure 1the operations of the evaluation engine 122, refinement engine 124, and execution engine 126 in, but it should be appreciated that the evaluation engine 122, refinement engine 124, and execution engine 126 can be used to analyze and improve the performance of the machine learning model 208 against adversarial attacks for other types of data (e.g., LiDAR, RADAR, ultrasonic, etc.), scenarios, and / or use cases. For example, the evaluation engine 122, refinement engine 124, and execution engine 126 can be used to defend against adversarial attacks involving perturbations or modifications to other types of sensor data (as described in further detail below with respect to Figures 5A to 5C ), point cloud data, audio data, video data, text data, time series data, network packets, source code, and / or other types of data. In another example, the evaluation engine 122, refinement engine 124, and execution engine 126 can be used to improve the robustness of the machine learning model 208 and downstream components against adversarial attacks that attempt to cause the machine learning model 208 to generate incorrect predictions 212 related to spam filtering, cybersecurity attack detection, biometric identification, medical imaging, identity fraud, and / or other applications or use cases.

[0058] It should be understood that this and other arrangements described herein are presented only as examples. Other arrangements and elements (e.g., machines, interfaces, functions, sequences, function groupings, etc.) can be used in addition to the shown arrangements and elements, and some elements can be omitted entirely. Moreover, many of the elements described herein are functional entities that can be implemented as discrete or distributed components, or in combination with other components, and can be implemented in any suitable combination and location. The various functions performed by the entities described herein can be executed by hardware, firmware, and / or software. For example, the various functions can be executed by a processor executing instructions stored in a memory. In some embodiments, the systems, methods, and processes described herein can be used with Figures 5A to 5D the exemplary autonomous vehicle 500 in Figure 6 the exemplary computing device 600 in Figure 7 and / or similar components, features, and / or functions in the exemplary data center 700 in

[0059] Now referring to Figures 3 to 4 , each block of the methods 300 and 400 described herein includes a computational process that can be executed using any combination of hardware, firmware, and / or software. For example, the various functions can be executed by a processor executing instructions stored in a memory. These methods can also be embodied as computer-usable instructions stored on a computer storage medium. These methods can be provided by stand-alone applications, services, or hosted services (stand-alone or in combination with other hosted services), or plug-ins of other products, etc. Moreover, herein Figures 1 to 2For example, for the system shown, methods 300 and 400 are described. However, these methods can additionally or alternatively be performed by any one system or any combination of systems, including (but not limited to) the systems described herein. Additionally, without departing from the scope of the present disclosure, the operations in method 300 can be omitted, repeated, and / or performed in any order.

[0060] Figure 3 FIG. shows a flowchart of method 300 for fine-tuning a machine learning model using adversarial data according to various embodiments. As Figure 3 shown, method 300 begins at operation 302, where the evaluation engine 122 applies one or more perturbations to a set of sensor data instances to generate a set of perturbed sensor data instances. For example, the evaluation engine 122 can use one or more known adversarial attack techniques, one or more user-specified adversarial attack techniques, and / or a combination of adversarial attack techniques to perturb a set of "raw images" captured using one or more cameras and / or generated synthetically in a simulated environment of a simulation system. The perturbed images correspond to adversarial images that would cause the machine learning model to generate incorrect predictions.

[0061] In operation 304, the evaluation engine 122 inputs the perturbed sensor data instances into the machine learning model to generate a set of predictions associated with the perturbed sensor data instances. Continuing the above example, the evaluation engine 122 can process the adversarial images using a neural network to generate predictions that include, but are not limited to, object type, object location, object path, behavior associated with the object, and / or other attributes associated with the object depicted in the adversarial image.

[0062] In operation 306, the evaluation engine 122 compares the predictions with the labels for the corresponding sensor data instances. Continuing the above example, the evaluation engine 122 can compare each prediction output derived from the adversarial images using the neural network with the label of the corresponding raw image or other ground truth data type to determine whether the prediction is correct.

[0063] In operation 308, the evaluation engine 122 generates a clean data set that includes the perturbed sensor data instances associated with the correct predictions of the machine learning model and the labels for the corresponding sensor data instances. Continuing the above example, the evaluation engine 122 can populate the clean data set with adversarial images for which the neural network correctly predicted the label of the corresponding raw image. The evaluation engine 122 can also associate the adversarial images with the labels in the clean data set.

[0064] In operation 310, the evaluation engine 122 generates an adversarial dataset that includes perturbed sensor data instances associated with incorrect predictions of the machine learning model. Continuing the above example, the evaluation engine 122 can populate the adversarial dataset with adversarial images for which the neural network fails to correctly predict the label of the corresponding image.

[0065] In operation 312, the evaluation engine 122 determines labels for the perturbed sensor data instances in the adversarial dataset based on the perturbation strength associated with the perturbed sensor data instances. Continuing the above example, the evaluation engine 122 can determine the perturbation strength of each adversarial image in the adversarial dataset based on (but not limited to) the following factors: the number or proportion of pixels in the corresponding original image that have been perturbed, the degree to which the pixels in the original image have been perturbed, the difference between the prediction generated by the neural network based on the adversarial image and the corresponding label, user input indicating the degree to which the pixels in the original image have been perturbed, and / or user input indicating the degree to which the perturbation of the original image affects the ability of a human to correctly predict the label of the original image. The evaluation engine 122 can assign one or more labels to the adversarial images using one or more thresholds and / or values associated with the perturbation strength. When the thresholds and / or values indicate that the adversarial image is associated with a low perturbation strength, the evaluation engine 122 can generate a label for the adversarial image that indicates a high confidence in the label of the original image. As the perturbation strength increases, the evaluation engine 122 can decrease the confidence in the label of the original image and increase the confidence in the adversarial label, which indicates that the adversarial image (or the object in the adversarial image) has been perturbed or corrupted to such an extent that the adversarial image (or object) can be ignored and / or processed in a manner different from when there are no adversarial properties or adversarial features in the image below a certain threshold.

[0066] In operation 314, the refinement engine 124 updates the parameters of the machine learning model based on one or more losses between predictions generated using the machine learning model from perturbed sensor data instances in the clean and adversarial data sets and the corresponding labels. Continuing the above example, the refinement engine 124 may input adversarial images from the two data sets into a neural network. For each adversarial image, the refinement engine 124 may obtain an output from the neural network, which includes (as a non-limiting example) the confidence in the prediction that the object in the adversarial image belongs to various object classes and / or adversarial classes associated with the adversarial labels. The refinement engine 124 may calculate the mean squared error, cross-entropy loss, and / or one or more other losses between the confidence in the output and one or more corresponding labels of the adversarial images in the data set. The refinement engine 124 may then use training techniques (e.g., gradient descent and backpropagation) to update the weights of the neural network in a manner that reduces the calculated losses. The refinement engine 124 may continue to train the neural network using the clean and adversarial data sets until: the weights converge; the loss is below a threshold; a certain number of training steps, iterations, batches, epochs, and / or training phases have been performed; and / or other conditions are met.

[0067] Now refer to Figure 4 , Figure 4 which shows a flowchart of a method 400 for processing sensor data according to various embodiments. As Figure 4 shown, method 400 begins at operation 402, in which the execution engine 126 uses a machine learning model and based on a sensor data instance, generates a set of confidences for a set of output types, and an adversarial confidence representing the likelihood that the sensor data instance is adversarial. For example, the execution engine 126 may input an image into a neural network and receive the set of confidences and the adversarial confidence as the corresponding output of the neural network.

[0068] At operation 404, the execution engine 126 determines the type of prediction for the sensor data instance based on the confidence and the adversarial confidence. For example, the execution engine 126 may set the type of prediction to the output type with the highest confidence. If the adversarial confidence is higher than all other confidences, the execution engine may set the prediction type to "adversarial". The execution engine 126 may also verify whether the highest confidence meets a threshold before determining the type of prediction. If the highest confidence does not meet the threshold, the execution engine 126 may set the type of prediction to "unknown".

[0069] At operation 406, the execution engine 126 transmits an indication of the predicted type to one or more downstream components. For example, the execution engine 126 may transmit the predicted type to downstream components that use the output of the machine learning model to generate driving commands and / or make other decisions related to autonomous vehicle operation. The execution engine 126 may also or alternatively transmit a confidence set and / or adversarial confidence set of the machine learning model output to the downstream components. The downstream components will use the predicted type, the confidence set, and / or the adversarial confidence to evaluate and / or respond to the risk of adversarial attacks against an autonomous or semi-autonomous machine or vehicle.

[0070] The systems and methods described herein may be used by non-autonomous vehicles, semi-autonomous vehicles (e.g., in one or more adaptive driver assistance systems (ADAS)), manned and unmanned robots or robotic platforms, warehouse vehicles, off-road vehicles, vehicles coupled to one or more trailers, airships, boats, shuttles, emergency response vehicles, motorcycles, electric or motorized bicycles, airplanes, engineering vehicles, submarines, drones, and / or other vehicle types, but are not limited thereto. Additionally, the systems and methods described herein may be used for a variety of purposes such as, but not limited to, machine control, machine motion, machine driving, synthetic data generation, model training, perception, augmented reality, virtual reality, mixed reality, robotics, security and surveillance, simulation and digital twins, autonomous or semi-autonomous machine applications, deep learning, environmental simulation, object or actor simulation, and / or digital twins, data center processing, conversational AI, light transport simulation (e.g., ray tracing, path tracing, etc.), collaborative content creation of 3D assets, cloud computing, and / or any other suitable application.

[0071] Embodiments of the present disclosure include in a variety of different systems such as automotive systems (e.g., control systems for autonomous or semi-autonomous machines, perception systems for autonomous or semi-autonomous machines), systems implemented using robots, aviation systems, medical systems, boating systems, intelligent area monitoring systems, systems for performing deep learning operations, systems for performing simulation operations, systems for performing digital twin operations, systems implemented using edge devices, systems including one or more virtual machines (VMs), systems for performing synthetic data generation operations, systems implemented at least partially in a data center, systems for performing conversational AI operations, systems for performing light transport simulation, systems for performing collaborative content creation of 3D assets, systems implemented at least partially using cloud computing resources, and / or other types of systems.

[0072] Example autonomous vehicle

[0073] Figure 5AFIG. is an illustration of an example autonomous vehicle 500 in accordance with some embodiments of the present disclosure. The autonomous vehicle 500 (alternatively, referred to herein as "vehicle 500") can include, but is not limited to, passenger vehicles such as cars, trucks, buses, emergency vehicles, shuttles, electric or motorized bicycles, motorcycles, fire trucks, police cars, ambulances, boats, construction vehicles, submarines, robotic vehicles, drones, airplanes, vehicles coupled to trailers (semi-trailer trucks for hauling cargo) and / or other types of vehicles (e.g., driverless and / or accommodating one or more occupants). Autonomous vehicles are typically described according to the levels of automation defined by the National Highway Traffic Safety Administration (NHTSA), a division of the U.S. Department of Transportation, and the Society of Automotive Engineers (SAE) in "Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles" (Standard No. J3016 - 201806, issued June 15, 2018; Standard No. J3016 - 201609, issued September 30, 2016; and prior and future versions of the standard). The vehicle 500 may be capable of implementing one or more functions corresponding to Levels 3 - 5 of the autonomous driving levels. The vehicle 500 may be capable of implementing one or more functions corresponding to Levels 3 - 5 of the autonomous driving levels. For example, depending on the embodiment, the vehicle 500 may be capable of implementing driver assistance (Level 1), semi-automation (Level 2), conditional automation (Level 3), high automation (Level 4), and / or full automation (Level 5). The term "autonomous" as used herein may include any and / or all types of autonomy of the vehicle 500 or other machines, such as fully autonomous, highly autonomous, conditionally autonomous, partially autonomous, providing assistive autonomy, semi-autonomous, primarily autonomous, or other designations.

[0074] The vehicle 500 may include components such as a chassis, a body, wheels (e.g., 2, 4, 6, 8, 18, etc.), tires, axles, and other components of the vehicle. The vehicle 500 may include a propulsion system 550, such as an internal combustion engine, a hybrid power plant, an all-electric motor, and / or another type of propulsion system. The propulsion system 550 may be connected to the driveline of the vehicle 500, which may include a transmission, to effect the propulsion of the vehicle 500. The propulsion system 550 may be controlled in response to signals received from the throttle / accelerator 552.

[0075] A steering system 554 that may include a steering wheel can be used to steer the vehicle 500 (e.g., along a desired path or route) while the propulsion system 550 is operating (e.g., while the vehicle is in motion). The steering system 554 can receive a signal from a steering actuator 556. For fully autonomous (Level 5) functionality, the steering wheel can be optional.

[0076] A brake sensor system 546 can be used to operate vehicle brakes in response to receiving a signal from a brake actuator 548 and / or a brake sensor.

[0077] One or more controllers 536 that may include one or more system-on-chips (SoCs) 504 ( Figure 5C ) and / or one or more GPUs can provide signals (e.g., representing commands) to one or more components and / or systems of the vehicle 500. For example, one or more controllers can send signals to operate vehicle brakes via one or more brake actuators 548, operate the steering system 554 via one or more steering actuators 556, and operate the propulsion system 550 via one or more throttles / accelerators 552. One or more controllers 536 can include one or more on-board (e.g., integrated) computing devices (e.g., supercomputers) that process sensor signals and output operation commands (e.g., signals representing commands) to enable autonomous driving and / or assist a human driver in driving the vehicle 500. One or more controllers 536 can include a first controller 536 for autonomous driving functionality, a second controller 536 for functional safety functionality, a third controller 536 for artificial intelligence functionality (e.g., computer vision), a fourth controller 536 for infotainment functionality, a fifth controller 536 for redundancy in emergency situations, and / or other controllers. In some examples, a single controller 536 can handle two or more of the above functions, two or more controllers 536 can handle a single function, and / or any combination thereof.

[0078] One or more controllers 536 may provide signals for controlling one or more components and / or systems of vehicle 500 in response to sensor data (e.g., sensor inputs) received from one or more sensors. The sensor data may be received from, for example and without limitation, a Global Navigation Satellite System sensor (“GNSS”) 558 (e.g., a Global Positioning System sensor), a RADAR sensor 560, an ultrasonic sensor 562, a LIDAR sensor 564, an Inertial Measurement Unit (IMU) sensor 566 (e.g., an accelerometer, a gyroscope, a magnetic compass, a magnetometer, etc.), a microphone 596, a stereo camera 568, a wide-angle camera 570 (e.g., a fish-eye camera), an infrared camera 572, a surround camera 574 (e.g., a 360-degree camera), a long-range and / or mid-range camera 598, a speed sensor 544 (e.g., for measuring the speed of vehicle 500), a vibration sensor 542, a steering sensor 540, a brake sensor (e.g., as part of a brake sensor system 546), and / or other sensor types.

[0079] In some embodiments, controller 536 receives sensor data from one or more types of sensors and uses one or more machine learning models to generate predictions related to the sensor data. These machine learning models may be trained and / or improved using Figures 1 to 2 components and / or systems therein to resist adversarial attacks. Controller 536 may use these predictions to generate commands or signals for operating vehicle 500.

[0080] One or more of the controllers 536 may receive inputs (e.g., represented by input data) from the instrument cluster 532 of vehicle 500 and provide outputs (e.g., represented by output data, display data, etc.) via the Human Machine Interface (HMI) display 534, an audible annunciator, a speaker, and / or via other components of vehicle 500. These outputs may include messages such as vehicle speed, rate, time, map data (e.g., Figure 5C a high-definition (“HD”) map 522), location data (e.g., the location of vehicle 500 on a map, for example), direction, the locations of other vehicles (e.g., occupancy grids), information about objects and object states as perceived by controller 536, and so on. For example, the HMI display 534 may display information about the presence of one or more objects (e.g., street signs, warning signs, traffic light changes, etc.) and / or information about driving maneuvers that the vehicle has made, is making, or will make (e.g., changing lanes now, exiting 34B in two miles, etc.).

[0081] Vehicle 500 further includes a network interface 524, which may communicate via one or more networks using one or more wireless antennas 526 and / or a modem. For example, network interface 524 may be capable of communicating via Long Term Evolution (“LTE”), Wideband Code Division Multiple Access (“WCDMA”), Universal Mobile Telecommunications System (“UMTS”), Global System for Mobile Communications (“GSM”), IMT-CDMA Multi-Carrier (“CDMA2000”), etc. One or more wireless antennas 526 may also be used to enable communication between objects (such as vehicles, mobile devices, etc.) in an environment using one or more local area networks such as Bluetooth, Bluetooth Low Energy (“LE”), Z-Wave, ZigBee, etc. and / or one or more low power wide area networks (LPWAN) such as LoRaWAN, SigFox, etc.

[0082] Figure 5B For an example autonomous vehicle 500 according to some embodiments of the present disclosure for Figure 5A example camera positions and fields of view. The cameras and respective fields of view are one example embodiment and are not intended to be limiting. For example, additional and / or alternative cameras may be included, and / or these cameras may be located at different positions on vehicle 500.

[0083] The camera type for the cameras may include, but is not limited to, digital cameras that may be adapted to work with components and / or systems of vehicle 500. The cameras may operate under an Automotive Safety Integrity Level (ASIL) B and / or under another ASIL. The camera type may have any image capture rate, such as 60 frames per second (fps), 120 fps, 240 fps, etc., depending on the embodiment. The cameras may be capable of using a rolling shutter, a global shutter, another type of shutter, or a combination thereof. In some examples, the color filter array may include a Red Clear Clear Clear (RCCC) color filter array, a Red Clear Clear Blue (RCCB) color filter array, a Red Blue Green Clear (RBGC) color filter array, a Foveon X3 color filter array, a Bayer sensor (RGGB) color filter array, a monochrome sensor color filter array, and / or another type of color filter array. In some embodiments, clear pixel cameras such as those with RCCC, RCCB, and / or RBGC color filter arrays may be used in an effort to increase light sensitivity.

[0084] In some examples, one or more of the cameras may be used to perform Advanced Driver Assistance System (ADAS) functions (such as as part of a redundant or fail-safe design). For example, a multi-functional monocular camera may be installed to provide functions including lane departure warning, traffic sign assistance, and intelligent headlight control. One or more of the cameras (such as all of the cameras) may record and provide image data (such as video) simultaneously.

[0085] One or more of the cameras can be mounted in a mounting assembly such as a custom-designed (three-dimensional "3D" printed) component to cut off stray light and reflections from within the vehicle (e.g., reflections from the dashboard reflected in the windshield mirror) that may interfere with the image data capture capabilities of the camera. Regarding the wing mirror mounting assembly, the wing mirror assembly can be custom 3D printed such that the camera mounting plate matches the shape of the wing mirror. In some examples, one or more cameras can be integrated into the wing mirror. For side view cameras, one or more cameras can also be integrated into the four pillars at each corner of the cab.

[0086] A camera (e.g., a front camera) having a field of view that includes an environmental portion in front of the vehicle 500 can be used for surround view to help identify forward paths and obstacles and to assist in providing information crucial for generating an occupancy grid and / or determining a preferred vehicle path with the help of one or more controllers 536 and / or a control SoC. The front camera can be used to perform many of the same ADAS functions as LIDAR, including emergency braking, pedestrian detection, and collision avoidance. The front camera can also be used for ADAS functions and systems, including lane departure warning ("LDW"), adaptive cruise control ("ACC"), and / or other functions such as traffic sign recognition.

[0087] A variety of cameras can be used in a front-facing configuration, including, for example, a monocular camera platform that includes a complementary metal oxide semiconductor ("CMOS") color imager. Another example can be a wide-angle camera 570, which can be used to sense objects (e.g., pedestrians, intersection traffic, or bicycles) entering the field of view from the periphery. Although Figure 5B only one wide-angle camera is illustrated, any number (including zero) of wide-angle cameras 570 can be present on the vehicle 500. Additionally, any number of long-range cameras 598 (e.g., a long-range stereo camera pair) can be used for depth-based object detection, especially for objects for which a neural network has not been trained. The long-range cameras 598 can also be used for object detection and classification and basic object tracking.

[0088] One or more stereo cameras 568 may also be included in a front-facing configuration. In at least one embodiment, one or more of the stereo cameras 568 may include an integrated control unit that includes a scalable processing unit that may provide a multi-core microprocessor and programmable logic ("FPGA") with an integrated controller area network ("CAN") or Ethernet interface on a single chip. Such a unit may be used to generate a 3D map of the vehicle environment, including distance estimates for all points in the image. Alternatively, the stereo camera 568 may include a compact stereo vision sensor that may include two camera lenses (one on the left and one on the right) and an image processing chip that may measure the distance from the vehicle to a target object and activate autonomous emergency braking and lane departure warning functions using the generated information (e.g., metadata). Other types of stereo cameras 568 may be used in addition to or in place of those described herein.

[0089] Cameras having a field of view that includes an environmental portion of the side of the vehicle 500 (e.g., side-view cameras) may be used for surround view, providing information used to create and update an occupancy grid and generate side-impact collision warnings. For example, surround cameras 574 (e.g., four surround cameras 574 as shown in Figure 5B may be disposed on the vehicle 500. The surround cameras 574 may include wide-angle cameras 570, fish-eye cameras, 360-degree cameras, and / or the like. By way of example, four fish-eye cameras may be disposed on the front, rear, and sides of the vehicle. In an alternative arrangement, the vehicle may use three surround cameras 574 (e.g., left, right, and rear) and may utilize one or more other cameras (e.g., a forward-facing camera) as a fourth surround camera.

[0090] Cameras having a field of view that includes an environmental portion of the rear of the vehicle 500 (e.g., rear-view cameras) may be used for assisting with parking, surround view, rear collision warnings, and creating and updating an occupancy grid. A variety of cameras may be used, including but not limited to cameras that are also suitable as front-facing cameras as described herein (e.g., long-range and / or mid-range cameras 598, stereo cameras 568, infrared cameras 572, etc.).

[0091] Figure 5C For use in accordance with some embodiments of the present disclosure Figure 5ABlock diagram of an example system architecture of an example autonomous vehicle 500. It should be understood that this and other arrangements described herein are set forth only as examples. Other arrangements and elements (e.g., machines, interfaces, functions, orders, function groupings, etc.) may be used in addition to or instead of those shown, and some elements may be omitted entirely. Further, many of the elements described herein are functional entities that may be implemented as discrete or distributed components or in combination with other components, and in any suitable combination and location. The various functions described herein as being performed by entities may be implemented by hardware, firmware, and / or software. For example, the various functions may be implemented by a processor executing instructions stored in a memory.

[0092] Figure 5C Each of the components, features, and systems in vehicle 500 is illustrated as being connected via a bus 502. The bus 502 may include a Controller Area Network (CAN) data interface (alternatively referred to herein as the “CAN bus”). The CAN may be a network within vehicle 500 used to assist in controlling various features and functions of vehicle 500, such as driving of brakes, acceleration, braking, steering, windshield wipers, and the like. The CAN bus may be configured to have dozens or even hundreds of nodes, each with its own unique identifier (e.g., CAN ID). The CAN bus may be read to find the steering wheel angle, ground speed, engine revolutions per minute (RPM), button positions, and / or other vehicle status indicators. The CAN bus may be ASIL B compliant.

[0093] Although the bus 502 is described herein as a CAN bus, this is not intended to be limiting. For example, in addition to or instead of the CAN bus, FlexRay and / or Ethernet may be used. Further, although a single line is used to represent the bus 502, this is not intended to be limiting. For example, any number of buses 502 may exist, which may include one or more CAN buses, one or more FlexRay buses, one or more Ethernet buses, and / or one or more other types of buses using different protocols. In some examples, two or more buses 502 may be used to perform different functions, and / or may be used for redundancy. For example, a first bus 502 may be used for collision avoidance functions, and a second bus 502 may be used for drive control. In any example, each bus 502 may communicate with any component of vehicle 500, and two or more buses 502 may communicate with the same component. In some examples, each SoC 504, each controller 536, and / or each computer within the vehicle may have access to the same input data (e.g., input from sensors of vehicle 500), and may be connected to a common bus such as the CAN bus.

[0094] Vehicle 500 may include one or more controllers 536, such as those described herein with respect to Figure 5A the controllers described. The controller 536 may be used for a variety of functions. The controller 536 may be coupled to any other different components and systems of the vehicle 500 and may be used for the control of the vehicle 500, the artificial intelligence of the vehicle 500, the infotainment for the vehicle 500, and / or the like.

[0095] Vehicle 500 may include one or more system-on-chips (SoCs) 504. The SoC 504 may include a CPU 506, a GPU 508, a processor 510, a cache 512, an accelerator 514, a data store 516, and / or other components and features not shown. In a variety of platforms and systems, the SoC 504 may be used to control the vehicle 500. For example, one or more SoCs 504 may be combined with an HD map 522 in a system (such as the system of the vehicle 500), and the HD map may obtain map refreshes and / or updates from one or more servers (such as Figure 5D one or more servers 578) via a network interface 524.

[0096] The CPU 506 may include a CPU cluster or a CPU complex (alternatively referred to herein as "CCPLEX"). The CPU 506 may include multiple cores and / or an L2 cache. For example, in some embodiments, the CPU 506 may include eight cores in a coherent multiprocessor configuration. In some embodiments, the CPU 506 may include four dual-core clusters, each of which has a dedicated L2 cache (such as a 2MB L2 cache). The CPU 506 (such as CCPLEX) may be configured to support simultaneous cluster operation such that any combination of the clusters of the CPU 506 can be active at any given time.

[0097] The CPU 506 may implement power management capabilities including one or more of the following features: each hardware block may automatically perform clock gating when idle to save dynamic power; each core clock may be gated when the core is not actively executing instructions due to the execution of WFI / WFE instructions; each core may be independently power gated; when all cores are clock gated or power gated, each core cluster may be independently clock gated; and / or when all cores are power gated, each core cluster may be independently power gated. The CPU 506 may further implement enhanced algorithms for managing power states, where the allowed power states and the desired wake-up times are specified, and the hardware / microcode determines the best power state for the cores, clusters, and CCPLEX to enter. The processing cores may support a simplified power state entry sequence in software, and this work is offloaded to the microcode.

[0098] The GPU 508 may include an integrated GPU (alternatively referred to herein as "iGPU"). The GPU 508 may be programmable and efficient for parallel workloads. In some examples, the GPU 508 may use enhanced tensor instruction sets. The GPU 508 may include one or more streaming microprocessors, where each streaming microprocessor may include an L1 cache (e.g., an L1 cache with at least 96 KB of storage capacity), and two or more of these streaming microprocessors may share an L2 cache (e.g., an L2 cache with 512 KB of storage capacity). In some embodiments, the GPU 508 may include at least eight streaming microprocessors. The GPU 508 may use a compute application programming interface (API). Additionally, the GPU 508 may use one or more parallel computing platforms and / or programming models (e.g., NVIDIA's CUDA).

[0099] In the case of automotive and embedded use cases, the GPU 508 may be power-optimized for optimal performance. For example, the GPU 508 may be fabricated on fin field-effect transistors (FinFETs). However, this is not intended to be limiting, and the GPU 508 may be fabricated using other semiconductor manufacturing processes. Each streaming microprocessor may incorporate several mixed-precision processing cores partitioned into multiple blocks. For example, and without limitation, 64 PF32 cores and 32 PF64 cores may be partitioned into four processing blocks. In such an example, each processing block may be allocated 16 FP32 cores, 8 FP64 cores, 16 INT32 cores, two mixed-precision NVIDIA tensor cores for deep learning matrix arithmetic, an L0 instruction cache, a warp scheduler, a dispatch unit, and / or a 64 KB register file. Additionally, the streaming microprocessor may include independent parallel integer and floating-point data paths to enable efficient execution of workloads leveraging a mix of computation and addressing computations. The streaming microprocessor may include independent thread scheduling capabilities to allow for finer-grained synchronization and cooperation between parallel threads. The streaming microprocessor may include a combined L1 data cache and shared memory unit to improve performance while simplifying programming.

[0100] The GPU 508 may include a high-bandwidth memory (HBM) and / or a 16 GB HBM2 memory subsystem that provides a peak memory bandwidth of approximately 900 GB / s in some examples. In some examples, in addition to or alternatively to HBM memory, synchronous graphics random access memory (SGRAM), such as fifth-generation graphics double data rate synchronous random access memory (GDDR5), may be used.

[0101] The GPU 508 may include unified memory technology that includes access counters to allow memory pages to be more precisely migrated to the processors that most frequently access them, thereby improving the efficiency of the memory range shared among the processors. In some examples, address translation service (ATS) support may be used to allow the GPU 508 to directly access the CPU 506 page tables. In such examples, when the GPU 508 memory management unit (MMU) experiences a miss, an address translation request may be transmitted to the CPU 506. In response, the CPU 506 may look up the virtual-physical mapping for the address in its page table and transmit the translation back to the GPU 508. In this way, the unified memory technology may allow a single unified virtual address space for the memories of both the CPU 506 and the GPU 508, thus simplifying GPU 508 programming and porting applications to the GPU 508.

[0102] In addition, the GPU 508 may include access counters that may track how frequently the GPU 508 accesses the memories of other processors. The access counters may help ensure that memory pages are moved to the physical memory of the processor that most frequently accesses those pages.

[0103] The SoC 504 may include any number of caches 512, including those described herein. For example, the cache 512 may include an L3 cache available to both the CPU 506 and the GPU 508 (e.g., that is connected to both the CPU 506 and the GPU 508). The cache 512 may include a write-back cache that may track the state of lines, for example, by using a cache coherence protocol (such as MEI, MESI, MSI, etc.). Depending on the embodiment, the L3 cache may include 4MB or more, but smaller cache sizes may also be used.

[0104] The SoC 504 may include one or more arithmetic logic units (ALUs) that may be used to perform processing for any of the various tasks or operations regarding the vehicle 500, such as processing DNNs. In addition, the SoC 504 may include a floating-point unit (FPU) or other math co-processor or digital co-processor type for performing mathematical operations within the system. For example, the SoC 504 may include one or more FPUs integrated as execution units within the CPU 506 and / or the GPU 508.

[0105] The SoC 504 may include one or more accelerators 514 (e.g., hardware accelerators, software accelerators, or a combination thereof). For example, the SoC 504 may include a hardware acceleration cluster, which may include optimized hardware accelerators and / or large on-chip memories. The large on-chip memory (e.g., 4MB SRAM) may enable the hardware acceleration cluster to accelerate neural networks and other computations. The hardware acceleration cluster may be used to supplement the GPU 508 and offload some of the tasks of the GPU 508 (e.g., freeing up more cycles of the GPU 508 for performing other tasks). As an example, the accelerator 514 may be used for targeted workloads that are stable enough to be easily controlled for acceleration (e.g., perception, convolutional neural networks (CNNs), etc.). When used herein, the term "CNN" may include all types of CNNs, including region-based or region convolutional neural networks (RCNNs) and fast RCNNs (e.g., for object detection). Components and / or systems of Figures 1 to 2 may be used to train and / or improve these CNNs to resist adversarial attacks.

[0106] The accelerator 514 (e.g., the hardware acceleration cluster) may include a deep learning accelerator (DLA). The DLA may include one or more tensor processing units (TPUs) that may be configured to provide an additional one trillion operations per second for deep learning applications and inference. The TPU may be an accelerator configured to perform image processing functions (e.g., for CNNs, RCNNs, etc.) and optimized for performing image processing functions. The DLA may be further optimized for a specific set of neural network types and floating-point operations and inference. The design of the DLA may provide higher performance per millimeter than a general-purpose GPU and far exceed the performance of a CPU. The TPU may perform several functions, including single-instance convolution functions, supporting INT8, INT16, and FP16 data types for both features and weights, and post-processor functions.

[0107] The DLA may execute neural networks, especially CNNs, quickly and efficiently for any of a variety of functions on processed or unprocessed data, such as, but not limited to: CNNs for object recognition and detection using data from camera sensors; CNNs for distance estimation using data from camera sensors; CNNs for emergency vehicle detection and identification and detection using data from microphones; CNNs for face recognition and vehicle owner recognition using data from camera sensors; and / or CNNs for security and / or safety-related events.

[0108] The DLA can perform any function of the GPU 508, and by using an inference accelerator, for example, a designer can configure the DLA or the GPU 508 for any function. For example, the designer can focus the processing of CNNs and floating-point operations on the DLA and leave other functions to the GPU 508 and / or other accelerators 514.

[0109] The accelerator 514 (e.g., a hardware acceleration cluster) can include a programmable vision accelerator (PVA), which may alternatively be referred to herein as a computer vision accelerator. The PVA can be designed and configured to accelerate computer vision algorithms for advanced driver assistance systems (ADAS), autonomous driving, and / or augmented reality (AR) and / or virtual reality (VR) applications. The PVA can provide a balance between performance and flexibility. For example, each PVA can include, by way of example and not limitation, any number of reduced instruction set computer (RISC) cores, direct memory access (DMA), and / or any number of vector processors.

[0110] The RISC cores can interact with an image sensor (e.g., the image sensor of any of the cameras described herein), an image signal processor, and / or the like. Each of these RISC cores can include any number of memories. Depending on the embodiment, the RISC cores can use any of several protocols. In some examples, the RISC cores can execute a real-time operating system (RTOS). The RISC cores can be implemented using one or more integrated circuit devices, application specific integrated circuits (ASICs), and / or storage devices. For example, the RISC cores can include an instruction cache and / or tightly coupled RAM.

[0111] The DMA can enable the components of the PVA to access system memory independently of the CPU 506. The DMA can support any number of features used to optimize the PVA, including but not limited to supporting multi-dimensional addressing and / or circular addressing. In some examples, the DMA can support up to six or more dimensions of addressing, which can include block width, block height, block depth, horizontal block step, vertical block step, and / or depth step.

[0112] A vector processor can be a programmable processor that can be designed to efficiently and flexibly execute programming for computer vision algorithms and provide signal processing capabilities. In some examples, the PVA can include a PVA core and two vector processing subsystem partitions. The PVA core can include a processor subsystem, one or more DMA engines (e.g., two DMA engines) and / or other peripherals. The vector processing subsystem can operate as the main processing engine of the PVA and can include a vector processing unit (VPU), an instruction cache, and / or vector memory (e.g., VMEM). The VPU core can include a digital signal processor, such as, for example, a single instruction multiple data (SIMD), very long instruction word (VLIW) digital signal processor. The combination of SIMD and VLIW can enhance throughput and rate.

[0113] Each of the vector processors can include an instruction cache and can be coupled to dedicated memory. As a result, in some examples, each of the vector processors can be configured to execute independently of the other vector processors. In other examples, the vector processors included in a particular PVA can be configured to employ data parallelization. For example, in some embodiments, multiple vector processors included in a single PVA can execute the same computer vision algorithm, but on different regions of an image. In other examples, the vector processors included in a particular PVA can simultaneously execute different computer vision algorithms on the same image, or even execute different algorithms on sequence images or portions of an image. Among other things, any number of PVAs can be included in a hardware acceleration cluster, and any number of vector processors can be included in each of these PVAs. Additionally, the PVA can include additional error correction code (ECC) memory to enhance overall system security.

[0114] The accelerator 514 (e.g., a hardware acceleration cluster) can include an on-chip computer vision network and SRAM to provide high-bandwidth, low-latency SRAM for the accelerator 514. In some examples, the on-chip memory can include at least 4MB SRAM composed of, for example, and not limited to, eight field-configurable memory blocks that can be accessed by both the PVA and the DLA. Each pair of memory blocks can include an advanced peripheral bus (APB) interface, configuration circuitry, a controller, and a multiplexer. Any type of memory can be used. The PVA and the DLA can access the memory via a backbone that provides high-speed memory access to the PVA and the DLA. The backbone can include (e.g., using APB) an on-chip computer vision network that interconnects the PVA and the DLA to the memory.

[0115] The on-chip computer vision network may include an interface that determines that both the PVA and DLA provide ready and valid signals before transmitting any control signals / addresses / data. Such an interface may provide separate phases and separate channels for transmitting control signals / addresses / data, as well as burst communication for continuous data transmission. This type of interface may comply with the ISO 26262 or IEC 61508 standards, but other standards and protocols may also be used.

[0116] In some examples, the SoC 504 may include a real-time ray tracing hardware accelerator, such as that described in U.S. Patent Application No. 16 / 101,232, filed on August 10, 2018. The real-time ray tracing hardware accelerator may be used to quickly and efficiently determine the position and extent of objects (e.g., within a world model) in order to generate a real-time visualization simulation for RADAR signal interpretation, for sound propagation synthesis and / or analysis, for SONAR system simulation, for general wave propagation simulation, for comparison with LIDAR data for positioning and / or other functional purposes, and / or for other uses. In some embodiments, one or more tree traversal units (TTUs) may be used to perform one or more ray tracing-related operations.

[0117] The accelerator 514 (e.g., a hardware accelerator cluster) has a wide range of autonomous driving uses. The PVA may be a programmable vision accelerator that may be used in critical processing stages in ADAS and autonomous vehicles. The capabilities of the PVA are a good match for algorithm domains that require predictable processing, low power, and low latency. In other words, the PVA performs well in semi-dense or dense regular computations, even on small data sets that require predictable runtimes with low latency and low power. Therefore, in the context of a platform for autonomous vehicles, the PVA is designed to run classical computer vision algorithms because they are effective in object detection and integer math operations.

[0118] For example, according to one embodiment of the technology, the PVA is used to perform computer stereo vision. In some examples, an algorithm based on semi-global matching may be used, but this is not intended to be limiting. Many applications for level 3-5 autonomous driving require instantaneous motion estimation / stereo matching (e.g., structure from motion, pedestrian recognition, lane detection, etc.). The PVA may perform computer stereo vision functions on inputs from two monocular cameras.

[0119] In some examples, the PVA may be used to perform dense optical flow. Process raw RADAR data (e.g., using a 4D fast Fourier transform) to provide processed RADAR. In other examples, the PVA is used for time-of-flight depth processing, which, for example, processes raw time-of-flight data to provide processed time-of-flight data.

[0120] DLA can be used to run any type of network to enhance control and driving safety, including, for example, a neural network that outputs a confidence metric for each object detection. Such confidence values can be interpreted as probabilities or as providing a relative "weight" of each detection compared to other detections. The confidence value enables the system to make further decisions about which detections should be considered true positive detections rather than false positive detections. For example, the system can set a threshold for the confidence and consider only detections that exceed the threshold as true positive detections. In an automatic emergency braking (AEB) system, false positive detections can cause the vehicle to automatically perform emergency braking, which is clearly undesirable. Therefore, only the most confident detections should be considered as triggers for AEB. DLA can run a neural network for regressing confidence values. The neural network can take as its input at least some subset of parameters, such as bounding box dimensions, a ground plane estimate obtained (e.g., from another subsystem), the output of an inertial measurement unit (IMU) sensor 566 related to the orientation and distance of the vehicle 500, a 3D position estimate of an object obtained from a neural network and / or other sensors (such as a LIDAR sensor 564 or a RADAR sensor 560), etc.

[0121] The SoC 504 can include one or more data stores 516 (e.g., memories). The data store 516 can be on-chip memory of the SoC 504, which can store neural networks to be executed on the GPU and / or DLA. In some examples, for redundancy and safety, the data store 516 can be large enough in capacity to store multiple instances of the neural network. The data store 512 can include an L2 or L3 cache 512. References to the data store 516 can include references to memories associated with PVAs, DLAs, and / or other accelerators 514 as described herein.

[0122] The SoC 504 may include one or more processors 510 (e.g., embedded processors). The processor 510 may include a boot and power management processor, which may be a dedicated processor and subsystem for handling boot power and management functions as well as security implementation related. The boot and power management processor may be part of the SoC 504 boot sequence and may provide runtime power management services. The boot power and management processor may provide clock and voltage programming, auxiliary system low-power state transitions, SoC 504 heat and temperature sensor management, and / or SoC 504 power state management. Each temperature sensor may be implemented as a ring oscillator whose output frequency is proportional to temperature, and the SoC 504 may use the ring oscillator to detect the temperature of the CPU 506, GPU 508, and / or accelerator 514. If it is determined that the temperature exceeds a threshold, then the boot and power management processor may enter a temperature fault routine and place the SoC 504 in a lower power state and / or place the vehicle 500 in a driver safe stop mode (e.g., safely stop the vehicle 500).

[0123] The processor 510 may further include a set of embedded processors that may serve as an audio processing engine. The audio processing engine may be an audio subsystem that allows for full hardware support for multi-channel audio over multiple interfaces and a wide range of flexible audio I / O interfaces. In some examples, the audio processing engine is a dedicated processor core with a digital signal processor with dedicated RAM.

[0124] The processor 510 may further include an always-on processor engine, which may provide the necessary hardware features to support low-power sensor management and wake-up use cases. The always-on processor engine may include a processor core, tightly coupled RAM, supporting peripherals (e.g., timers and interrupt controllers), various I / O controller peripherals, and routing logic.

[0125] The processor 510 may further include a security cluster engine, which includes a dedicated processor subsystem for handling security management of automotive applications. The security cluster engine may include two or more processor cores, tightly coupled RAM, supporting peripherals (e.g., timers, interrupt controllers, etc.), and / or routing logic. In a security mode, the two or more cores may operate in a lockstep mode and act as a single core with comparison logic for detecting any differences between their operations.

[0126] The processor 510 may further include a real-time camera engine, which may include a dedicated processor subsystem for handling real-time camera management.

[0127] The processor 510 may further include a high dynamic range signal processor, which may include an image signal processor, which is a hardware engine that is part of the camera processing pipeline.

[0128] The processor 510 may include a video image compositor that may be a processing block (e.g., implemented on a microprocessor) that implements the video post-processing functions required for a video playback application to generate the final image for the player window. The video image compositor may perform lens distortion correction on the wide-angle camera 570, the surround camera 574, and / or the in-cab monitoring camera sensor. The in-cab monitoring camera sensor is preferably monitored by a neural network running on another instance of the advanced SoC, configured to identify in-cab events and respond accordingly. The in-cab system may perform lip reading to activate mobile phone services and make calls, dictate emails, change the vehicle destination, activate or change the vehicle's infotainment system and settings, or provide voice-activated web surfing. Certain functions are only available to the driver when the vehicle is operating in autonomous mode and are disabled otherwise.

[0129] The video image compositor may include enhanced temporal noise reduction for spatial and temporal noise reduction. For example, in the case of motion in the video, the noise reduction appropriately weights the spatial information, reducing the weight of the information provided by neighboring frames. In the case where the image or a portion of the image does not include motion, the temporal noise reduction performed by the video image compositor may use information from a previous image to reduce the noise in the current image.

[0130] The video image compositor may also be configured to perform stereo correction on input stereo lens frames. When the operating system desktop is in use and the GPU 508 does not need to continuously render new surfaces, the video image compositor may further be used for user interface composition. Even when the GPU 508 is powered on and active for 3D rendering, the video image compositor may be used to lighten the burden on the GPU 508 to improve performance and responsiveness.

[0131] The SoC 504 may further include a Mobile Industry Processor Interface (MIPI) camera serial interface, a high-speed interface, and / or a video input block for receiving video and inputs from cameras and related pixel input functions. The SoC 504 may further include an input / output controller that may be software-controlled and may be used to receive I / O signals not committed to a specific role.

[0132] The SoC 504 may further include a wide range of peripheral device interfaces to enable communication with peripheral devices, audio codecs, power management, and / or other devices. The SoC 504 can be used to process data from cameras (connected via Gigabit Multimedia Serial Link and Ethernet), sensors (such as LIDAR sensor 564, RADAR sensor 560, etc. that can be connected via Ethernet), data from bus 502 (such as the speed of vehicle 500, steering wheel position, etc.), and data from GNSS sensor 558 (connected via Ethernet or CAN bus). The SoC 504 may further include dedicated high-performance large-capacity storage controllers, which may include their own DMA engines and which can be used to free the CPU 506 from routine data management tasks.

[0133] The SoC 504 can be an end-to-end platform with a flexible architecture that spans automation levels 3 - 5, thus providing an integrated functional safety architecture for a platform that leverages and efficiently uses computer vision and ADAS technologies to achieve diversity and redundancy, along with deep learning tools to provide a flexible and reliable driving software stack. The SoC 504 can be faster, more reliable, and even more energy-efficient and space-efficient than conventional systems. For example, when combined with the CPU 506, GPU 508, and data storage 516, the accelerator 514 can provide a fast and efficient platform for level 3 - 5 autonomous vehicles.

[0134] Thus, this technology provides capabilities and functions that cannot be achieved by conventional systems. For example, computer vision algorithms can be executed on CPUs that can be configured using high-level programming languages such as the C programming language to perform various processing algorithms across a wide variety of visual data. However, CPUs often cannot meet the performance requirements of many computer vision applications, such as those related to execution time and power consumption. In particular, many CPUs cannot execute complex object detection algorithms in real time, which is a requirement for in-vehicle ADAS applications and for practical level 3 - 5 autonomous vehicles.

[0135] In contrast to conventional systems, the technology described herein allows multiple neural networks to be executed simultaneously and / or sequentially by providing a CPU complex, a GPU complex, and a hardware acceleration cluster, and combining the results to achieve level 3 - 5 autonomous driving functions. For example, a CNN executed on a DLA or dGPU (such as GPU 520) can include text and word recognition, allowing a supercomputer to read and understand traffic signs, including signs for which the neural network has not been specifically trained. The DLA may further include a neural network capable of recognizing, interpreting, and providing semantic understanding of the signs and passing that semantic understanding to a path planning module running on the CPU complex.

[0136] As another example, multiple neural networks can run simultaneously, such as required for level 3, 4, or 5 driving. For example, a warning sign consisting of "Caution: Flashing lights indicate icy conditions" together with the electric lights can be interpreted independently or jointly by several neural networks. The sign itself can be recognized as a traffic sign by a first neural network deployed (e.g., a trained neural network), and the text "Flashing lights indicate icy conditions" can be interpreted by a second neural network deployed, which informs the vehicle's path planning software (preferably executed on the CPU complex) that when the flashing lights are detected, there are icy conditions. The flashing lights can be recognized by operating a third neural network on multiple frames, which informs the vehicle's path planning software of the presence (or absence) of the flashing lights. All three neural networks can run simultaneously, for example, within the DLA and / or on the GPU 508. All three neural networks can also use Figures 1 to 2 the systems and / or components in

[0137] In some examples, the CNNs for face recognition and owner recognition can use data from the camera sensors to recognize the presence of an authorized driver and / or owner of the vehicle 500. The processing engine always on the sensor can be used to unlock the vehicle and turn on the lights when the owner approaches the driver's door, and in a security mode, to disable the vehicle when the owner leaves the vehicle. In this way, the SoC 504 provides security against theft and / or carjacking.

[0138] In another example, the CNN for emergency vehicle detection and recognition can use data from the microphone 596 to detect and recognize an emergency vehicle siren. In contrast to conventional systems that use a general classifier to detect the siren and manually extract features, the SoC 504 uses the CNN to classify environmental and urban sounds as well as visual data. In a preferred embodiment, the CNN running on the DLA is trained to recognize the relative closing rate of an emergency vehicle (e.g., by using the Doppler effect). The CNN can also be trained to recognize emergency vehicles specific to the local area in which the vehicle operates, as recognized by the GNSS sensor 558. Thus, for example, when operating in Europe, the CNN will seek to detect European sirens, and when in the United States, the CNN will seek to recognize only North American sirens. Once an emergency vehicle is detected, with the assistance of the ultrasonic sensor 562, the control program can be used to execute an emergency vehicle safety routine to slow down the vehicle, pull over to the side of the road, stop the vehicle, and / or idle the vehicle until the emergency vehicle passes.

[0139] The vehicle may include a CPU 518 (e.g., a discrete CPU or dCPU) that may be coupled to the SoC 504 via a high-speed interconnect (e.g., PCIe). The CPU 518 may include, for example, an X86 processor. The CPU 518 may be used to perform any of a variety of functions, including, for example, arbitrating potentially inconsistent results between the ADAS sensors and the SoC 504, and / or monitoring the status and health of the controller 536 and / or the infotainment SoC 530.

[0140] The vehicle 500 may include a GPU 520 (e.g., a discrete GPU or dGPU) that may be coupled to the SoC 504 via a high-speed interconnect (e.g., NVIDIA's NVLINK). The GPU 520 may provide additional artificial intelligence capabilities, for example, by executing redundant and / or different neural networks, and may be used to train and / or update neural networks based on inputs (e.g., sensor data) from the sensors of the vehicle 500.

[0141] The vehicle 500 may further include a network interface 524, which may include one or more wireless antennas 526 (e.g., one or more wireless antennas for different communication protocols, such as cellular antennas, Bluetooth antennas, etc.). The network interface 524 may be used to enable wireless connections to the cloud (e.g., to the server 578 and / or other network devices), to other vehicles, and / or to computing devices (e.g., the client devices of passengers) via the Internet. To communicate with other vehicles, a direct link may be established between the two vehicles, and / or an indirect link may be established (e.g., across a network and via the Internet). The direct link may be provided using a vehicle-to-vehicle communication link. The vehicle-to-vehicle communication link may provide the vehicle 500 with information about vehicles approaching the vehicle 500 (e.g., vehicles in front of, to the side of, and / or behind the vehicle 500). This function may be part of the cooperative adaptive cruise control function of the vehicle 500.

[0142] The network interface 524 may include an SoC that provides modulation and demodulation functions and enables the controller 536 to communicate via a wireless network. The network interface 524 may include a radio frequency front end for upconversion from baseband to radio frequency and downconversion from radio frequency to baseband. The frequency conversion may be performed by a known process, and / or may be performed using a super-heterodyne process. In some examples, the radio frequency front end functions may be provided by a separate chip. The network interface may include wireless capabilities for communicating via LTE, WCDMA, UMTS, GSM, CDMA2000, Bluetooth, Bluetooth LE, Wi-Fi, Z-Wave, ZigBee, LoRaWAN, and / or other wireless protocols.

[0143] Vehicle 500 may further include a data store 528 that may include off-chip (e.g., outside of SoC 504) storage devices. The data store 528 may include one or more storage elements including RAM, SRAM, DRAM, VRAM, flash memory, hard disks, and / or other components and / or devices that may store at least one bit of data.

[0144] Vehicle 500 may further include a GNSS sensor 558. The GNSS sensor 558 (e.g., GPS, assisted GPS sensor, differential GPS (DGPS) sensor, etc.) is used to assist mapping, perception, occupancy grid generation, and / or path planning functions. Any number of GNSS sensors 558 may be used, including, for example and without limitation, a GPS using a USB connector with an Ethernet to serial (RS-232) bridge.

[0145] Vehicle 500 may further include a RADAR sensor 560. The RADAR sensor 560 may be used by the vehicle 500 for remote vehicle detection even in dark and / or adverse weather conditions. The RADAR functional safety level may be ASIL B. The RADAR sensor 560 may use CAN and / or bus 502 (e.g., to transmit data generated by the RADAR sensor 560) for control as well as to access object tracking data and, in some examples, access Ethernet to access raw data. A variety of RADAR sensor types may be used. For example and without limitation, the RADAR sensor 560 may be suitable for front, rear, and side RADAR use. In some examples, a pulsed Doppler RADAR sensor is used.

[0146] The RADAR sensor 560 may include different configurations, such as long-range with a narrow field of view, short-range with a wide field of view, short-range side coverage, and so on. In some examples, long-range RADAR may be used for adaptive cruise control functions. The long-range RADAR system may provide a wide field of view (e.g., within 250 m) achieved through two or more independent scans. The RADAR sensor 560 may help distinguish between static and moving objects and may be used by the ADAS system for emergency braking assistance and forward collision warning. The long-range RADAR sensor may include a single station multimode RADAR with multiple (e.g., six or more) fixed RADAR antennas and high-speed CAN and FlexRay interfaces. In an example with six antennas, the central four antennas may create a focused beam pattern that is designed to record the surroundings of the vehicle 500 with minimal traffic interference from adjacent lanes at a higher rate. The other two antennas may extend the field of view, making it possible to quickly detect vehicles entering or leaving the lane of the vehicle 500.

[0147] As an example, a mid-range RADAR system can include a range of up to 560m (front) or 80m (rear) and a field of view of up to 42 degrees (front) or 550 degrees (rear). A short-range RADAR system can include, but is not limited to, RADAR sensors designed to be mounted at both ends of the rear bumper. When mounted at both ends of the rear bumper, such a RADAR sensor system can create two beams that continuously monitor the rear and the blind spots alongside the vehicle.

[0148] The short-range RADAR system can be used in an ADAS system for blind spot detection and / or lane change assistance.

[0149] Vehicle 500 can further include ultrasonic sensors 562. Ultrasonic sensors 562 that can be placed at the front, rear, and / or sides of vehicle 500 can be used for parking assistance and / or creating and updating an occupancy grid. A variety of ultrasonic sensors 562 can be used, and different ultrasonic sensors 562 can be used for different detection ranges (e.g., 2.5m, 4m). Ultrasonic sensors 562 can operate at ASIL B, a functional safety level.

[0150] Vehicle 500 can include a LIDAR sensor 564. The LIDAR sensor 564 can be used for object and pedestrian detection, emergency braking, collision avoidance, and / or other functions. The LIDAR sensor 564 can be at ASIL B, a functional safety level. In some examples, vehicle 500 can include multiple LiDAR sensors 564 (e.g., two, four, six, etc.) that can use Ethernet (e.g., to provide data to a gigabit Ethernet switch).

[0151] In some examples, the LIDAR sensor 564 may be able to provide a list of objects and their distances for a 360-degree field of view. Commercially available LIDAR sensors 564 can have, for example, an advertised range of approximately 500m, an accuracy of 2cm - 3cm, and support for a 500Mbps Ethernet connection. In some examples, one or more non-protruding LIDAR sensors 564 can be used. In such examples, the LIDAR sensor 564 can be implemented as a small device that can be embedded into the front, rear, sides, and / or corners of vehicle 500. In such examples, the LiDAR sensor 564 can provide a field of view of up to 120 degrees horizontally and 35 degrees vertically, even for low-reflectivity objects, with a range of 200m. The front-mounted LIDAR sensor 564 can be configured for a horizontal field of view between 45 degrees and 135 degrees.

[0152] In some examples, LIDAR technologies such as 3D flash LIDAR can also be used. 3D flash LIDAR uses the flash of a laser as the emission source to illuminate the vehicle's surrounding environment up to about 200 m. The flash LIDAR unit includes a receiver that records the laser pulse transmission time and the reflected light on each pixel, which in turn corresponds to the range from the vehicle to the object. Flash LIDAR can allow for the generation of highly accurate and distortion-free images of the surrounding environment using each laser flash. In some examples, four flash LIDAR sensors can be deployed, one on each side of the vehicle 500. Available 3D flash LIDAR systems include solid-state 3D staring array LIDAR cameras (e.g., non-scanning LIDAR devices) that have no moving parts other than a fan. The flash LIDAR device can use class I (eye-safe) laser pulses of 5 nanoseconds per frame and can capture the reflected laser in the form of 3D range point clouds and co-registered intensity data. By using flash LIDAR and because flash LIDAR is a solid-state device without moving parts, the LIDAR sensor 564 can be less susceptible to motion blur, vibration, and / or shock.

[0153] The vehicle can further include an IMU sensor 566. In some examples, the IMU sensor 566 can be located at the center of the rear axle of the vehicle 500. The IMU sensor 566 can include, for example and without limitation, an accelerometer, a magnetometer, a gyroscope, a magnetic compass, and / or other sensor types. In some examples, such as in a six-axis application, the IMU sensor 566 can include an accelerometer and a gyroscope, while in a nine-axis application, the IMU sensor 566 can include an accelerometer, a gyroscope, and a magnetometer.

[0154] In some embodiments, the IMU sensor 566 can be implemented as a miniature high-performance GPS-aided inertial navigation system (GPS / INS) that combines microelectromechanical system (MEMS) inertial sensors, a high-sensitivity GPS receiver, and an advanced Kalman filtering algorithm to provide estimates of position, velocity, and attitude. Thus, in some examples, the IMU sensor 566 can enable the vehicle 500 to estimate the heading by directly observing the change in velocity from the GPS to the IMU sensor 566 and correlating it without the need for input from a magnetic sensor. In some examples, the IMU sensor 566 and the GNSS sensor 558 can be integrated into a single unit.

[0155] The vehicle can include a microphone 596 disposed in and / or around the vehicle 500. Among other things, the microphone 596 can be used for emergency vehicle detection and identification.

[0156] The vehicle may further include any number of camera types, including a stereo camera 568, a wide-angle camera 570, an infrared camera 572, a surround camera 574, a long-range and / or mid-range camera 598, and / or other camera types. These cameras can be used to capture image data around the entire periphery of the vehicle 500. The camera types used depend on the embodiment and the requirements of the vehicle 500, and any combination of camera types can be used to provide the necessary coverage around the vehicle 500. Additionally, the number of cameras can vary according to the embodiment. For example, the vehicle may include six cameras, seven cameras, ten cameras, twelve cameras, and / or another number of cameras. As an example and without limitation, these cameras can support Gigabit Multimedia Serial Link (GMSL) and / or Gigabit Ethernet. Each of the cameras is described in more detail herein with respect to Figure 5A and Figure 5B in more detail.

[0157] The vehicle 500 may further include a vibration sensor 542. The vibration sensor 542 can measure the vibration of components of the vehicle such as an axle. For example, a change in vibration can indicate a change in the road surface. In another example, when two or more vibration sensors 542 are used, the difference between the vibrations can be used to determine the friction or slip of the road surface (e.g., when there is a vibration difference between a powered drive axle and a free-rotating axle).

[0158] The vehicle 500 may include an ADAS system 538. In some examples, the ADAS system 538 may include a SoC. The ADAS system 538 may include autonomous / adaptive / auto cruise control (ACC), cooperative adaptive cruise control (CACC), forward collision warning (FCW), automatic emergency braking (AEB), lane departure warning (LDW), lane keeping assist (LKA), blind spot warning (BSW), rear cross-traffic warning (RCTW), collision warning system (CWS), lane centering (LC), and / or other features and functions.

[0159] The ACC system can use RADAR sensors 560, LIDAR sensors 564, and / or cameras. The ACC system can include longitudinal ACC and / or lateral ACC. Longitudinal ACC monitors and controls the distance to the vehicle immediately in front of the vehicle 500 and automatically adjusts the vehicle speed to maintain a safe distance from the vehicle ahead. Lateral ACC performs distance keeping and, when necessary, advises the vehicle 500 to change lanes. Lateral ACC is related to other ADAS applications such as LCA and CWS.

[0160] The CACC uses information from other vehicles, which can be received indirectly from other vehicles via the network interface 524 and / or the wireless antenna 526 over a wireless link or through a network connection (e.g., via the Internet). A direct link can be provided by a vehicle-to-vehicle (V2V) communication link, while an indirect link can be an infrastructure-to-vehicle (I2V) communication link. Generally, the V2V communication concept provides information about the vehicle immediately ahead (e.g., a vehicle immediately in front of vehicle 500 and in the same lane as it), while the I2V communication concept provides information about traffic further ahead. The CACC system can include either or both of the I2V and V2V information sources. Given information about the vehicle ahead of vehicle 500, the CACC can be more reliable, and it has the potential to improve the smoothness of traffic flow and reduce road congestion.

[0161] The FCW system is designed to alert the driver of a hazard so that the driver can take corrective action. The FCW system uses a front camera and / or RADAR sensor 560 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback such as a display, speaker, and / or vibration component. The FCW system can provide warnings in the form of, for example, audible, visual warnings, vibrations, and / or rapid braking pulses.

[0162] The AEB system detects an impending forward collision with another vehicle or other object and can automatically apply the brakes if the driver does not take corrective action within specified time or distance parameters. The AEB system can use a front camera and / or RADAR sensor 560 coupled to a dedicated processor, DSP, FPGA, and / or ASIC. When the AEB system detects a hazard, it typically first alerts the driver to take corrective action to avoid the collision, and if the driver does not take corrective action, then the AEB system can automatically apply the brakes in an effort to prevent or at least mitigate the impact of the predicted collision. The AEB system can include technologies such as dynamic brake support and / or collision imminent braking.

[0163] The LDW system provides visual, audible, and / or tactile warnings such as steering wheel or seat vibrations to alert the driver when vehicle 500 crosses a lane marking. The LDW system is not activated when the driver indicates an intentional lane departure by activating the turn signal. The LDW system can use a front-side-facing camera coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback such as a display, speaker, and / or vibration component.

[0164] The LKA system is a variant of the LDW system. If the vehicle 500 starts to leave the lane, then the LKA system provides a steering input or braking to correct the vehicle 500.

[0165] The BSW system detects and warns the driver of vehicles in the blind spot of the vehicle. The BSW system can provide visual, audible, and / or tactile alerts to indicate that merging or changing lanes is unsafe. The system can provide additional warnings when the driver uses the turn signal. The BSW system can use a rear-facing camera and / or RADAR sensor 560 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback such as a display, speaker, and / or vibration component.

[0166] The RCTW system can provide visual, audible, and / or tactile notifications when an object is detected outside the rear camera range while the vehicle 500 is reversing. Some RCTW systems include AEB to ensure that the vehicle brakes are applied to avoid a crash. The RCTW system can use one or more rear RADAR sensors 560 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback such as a display, speaker, and / or vibration component.

[0167] Conventional ADAS systems may be prone to false positive results, which can be annoying and distracting to the driver, but are typically not catastrophic because the ADAS system alerts the driver and allows the driver to decide whether a safe condition actually exists and act accordingly. However, in an autonomous vehicle 500, in the case of conflicting results, the vehicle 500 itself must decide whether to heed the results from the main computer or an auxiliary computer (e.g., the first controller 536 or the second controller 536). For example, in some embodiments, the ADAS system 538 can be a backup and / or auxiliary computer for providing perception information to a backup computer rationality module. The backup computer rationality monitor can run redundant and diverse software on hardware components to detect faults in perception and dynamic driving tasks. The output from the ADAS system 538 can be provided to the supervisory MCU. If the outputs from the main computer and the auxiliary computer conflict, then the supervisory MCU must determine how to reconcile the conflict to ensure safe operation.

[0168] In some examples, the host computer may be configured to provide a confidence score to the supervisory MCU indicating the host computer's confidence in the selected result. If the confidence score exceeds a threshold, then the supervisory MCU may follow the direction of the host computer regardless of whether the secondary computer provides conflicting or inconsistent results. In cases where the confidence score does not meet the threshold and where the host computer and the secondary computer indicate different results (e.g., conflict), the supervisory MCU may arbitrate between these computers to determine an appropriate result.

[0169] The supervisory MCU may be configured to run a neural network that is trained and configured to determine conditions under which the secondary computer provides a false alarm based on outputs from the host computer and the secondary computer. Thus, the neural network in the supervisory MCU can learn when the output of the secondary computer can be trusted and when it cannot. For example, when the secondary computer is a RADAR-based FCW system, the neural network in the supervisory MCU can learn when the FCW system is identifying a metallic object that is not in fact dangerous, such as a drain grate or manhole cover that triggers an alarm. Similarly, when the secondary computer is a camera-based LDW system, the neural network in the supervisory MCU can learn to disregard the LDW when a cyclist or pedestrian is present and lane departure is actually the safest strategy. In embodiments that include a neural network running on the supervisory MCU, the supervisory MCU may include at least one of a DLA or a GPU suitable for running the neural network with an associated memory. In a preferred embodiment, the supervisory MCU may include components of the SoC 504 and / or be included as a component of the SoC 504.

[0170] In other examples, the ADAS system 538 may include a secondary computer that performs ADAS functions using traditional computer vision rules. In this way, the secondary computer may use classical computer vision rules (if - then), and the presence of a neural network in the supervisory MCU can improve reliability, safety, and performance. For example, diverse implementations and intentional non-identity make the overall system more fault-tolerant, especially for failures caused by software (or software-hardware interface) functions. For example, if there is a software vulnerability or error in the software running on the host computer and the non-identical software code running on the secondary computer provides the same overall result, then the supervisory MCU can be more confident that the overall result is correct and that the vulnerability in the software or hardware on the host computer does not cause a substantial error.

[0171] In some examples, the output of the ADAS system 538 can be fed to the perception block of the main computer and / or the dynamic driving task block of the main computer. For example, if the ADAS system 538 indicates a forward collision warning due to an object being immediately in front, the perception block can use this information when identifying the object. In other examples, the auxiliary computer can have its own neural network, which is trained and thus reduces the risk of false positives as described herein.

[0172] The vehicle 500 can further include an infotainment SoC 530 (e.g., an in-vehicle infotainment system (IVI)). Although illustrated and described as an SoC, the infotainment system can not be an SoC and can include two or more discrete components. The infotainment SoC 530 can include a combination of hardware and software that can be used to provide audio (e.g., music, personal digital assistant, navigation instructions, news, radio, etc.), video (e.g., TV, movies, streaming, etc.), telephone (e.g., hands-free calling), network connectivity (e.g., LTE, Wi-Fi, etc.), and / or information services (e.g., navigation system, rear parking assistance, radio data system, vehicle-related information such as fuel level, total distance covered, brake fuel level, oil level, door open / close, air filter information, etc.) to the vehicle 500. For example, the infotainment SoC 530 can include a radio, a disc player, a navigation system, a video player, USB and Bluetooth connectivity, an in-vehicle computer, in-vehicle entertainment, Wi-Fi, steering wheel audio controls, hands-free voice controls, a head-up display (HUD), an HMI display 534, a telematics device, a control panel (e.g., for controlling various components, features, and / or systems, and / or interacting therewith), and / or other components. The infotainment SoC 530 can further be used to provide information (e.g., visual and / or auditory) to the user of the vehicle, such as information from the ADAS system 538, autonomous driving information such as planned vehicle maneuvers, trajectories, surrounding environment information (e.g., intersection information, vehicle information, road information, etc.), and / or other information.

[0173] The infotainment SoC 530 can include GPU functionality. The infotainment SoC 530 can communicate with other devices, systems, and / or components of the vehicle 500 via a bus 502 (e.g., a CAN bus, Ethernet, etc.). In some examples, the infotainment SoC 530 can be coupled to a supervisory MCU such that in the event of a failure of the main controller 536 (e.g., the main and / or backup computer of the vehicle 500), the GPU of the infotainment system can perform some self-driving functions. In such examples, the infotainment SoC 530 can place the vehicle 500 in a driver safe parking mode as described herein.

[0174] Vehicle 500 may further include an instrument cluster 532 (such as a digital instrument panel, an electronic instrument cluster, a digital instrument surface panel, etc.). The instrument cluster 532 may include a controller and / or a supercomputer (such as a discrete controller or supercomputer). The instrument cluster 532 may include a set of instruments, such as a speedometer, a fuel level, an oil pressure, a tachometer, an odometer, a turn indicator, a shift position indicator, a seat belt warning light, a parking brake warning light, an engine malfunction light, a safety airbag (SRS) system information, lighting controls, safety system controls, navigation information, etc. In some examples, information may be displayed and / or shared between the infotainment SoC 530 and the instrument cluster 532. In other words, the instrument cluster 532 may be included as part of the infotainment SoC 530, or vice versa.

[0175] Figure 5D A system schematic diagram for communication between a cloud-based server and Figure 5A an example autonomous vehicle 500 in accordance with some embodiments of the present disclosure. The system 576 may include a server 578, a network 590, and vehicles including the vehicle 500. The server 578 may include a plurality of GPUs 584(A)-584(H) (collectively referred to herein as GPUs 584), PCIe switches 582(A)-582(H) (collectively referred to herein as PCIe switches 582), and / or CPUs 580(A)-580(B) (collectively referred to herein as CPUs 580). The GPUs 584, CPUs 580, and PCIe switches may be interconnected by high-speed interconnects such as, for example, and without limitation, an NVLink interface 588 developed by NVIDIA and / or a PCIe connection 586. In some examples, the GPUs 584 are connected via NVLink and / or an NVSwitch SoC, and the GPUs 584 and the PCIe switches 582 are connected via a PCIe interconnect. Although eight GPUs 584, two CPUs 580, and two PCIe switches are illustrated, this is not intended to be limiting. Depending on the embodiment, each of the servers 578 may include any number of GPUs 584, CPUs 580, and / or PCIe switches. For example, each of the servers 578 may include eight, sixteen, thirty-two, and / or more GPUs 584.

[0176] Server 578 can receive image data via network 590 and from a vehicle, the image data representing an image showing an unexpected or changed road condition such as a recently started roadwork. Server 578 can transmit neural network 592, updated neural network 592, and / or map information 594 via network 590 and to the vehicle, including information about traffic and road conditions. Updates to the map information 594 can include updates to the HD map 522, such as information about construction sites, potholes, curves, floods, or other obstacles. In some examples, the neural network 592, updated neural network 592, and / or map information 594 can be represented and / or generated based on data from new training and / or data received from any number of vehicles in the environment and / or experience from training performed at a data center (e.g., using server 578 and / or other servers). For example, the neural network 592, updated neural network 592, and / or map information 594 can be trained, improved, and / or updated to detect and defend against adversarial attacks using the systems and / or techniques described above with respect to Figures 1 to 4 to detect and defend against adversarial attacks.

[0177] Server 578 can be used to train a machine learning model (e.g., Figure 2 machine learning model 208, neural network, etc.) based on training data. The training data can be generated by vehicles and / or can be generated in simulation (e.g., using a game engine). In some examples, the training data is labeled (e.g., in cases where the neural network benefits from supervised learning) and / or undergoes other preprocessing, while in other examples, the training data is not labeled and / or preprocessed (e.g., in cases where the neural network does not require supervised learning). In some embodiments, the training data includes Figure 2 clean data set 218 and / or adversarial data set 216. The training can be performed according to any one or more classes of machine learning techniques, including but not limited to the following classes: supervised training, semi-supervised training, unsupervised training, self-learning, reinforcement learning, federated learning, transfer learning, feature learning (including principal component and clustering analysis), multilinear subspace learning, manifold learning, representation learning (including alternative dictionary learning), rule-based machine learning, anomaly detection, adversarial attack recognition, and any variations or combinations thereof. Once the machine learning model is trained, the machine learning model can be used by the vehicle (e.g., transmitted to the vehicle via network 590), and / or the machine learning model can be used by server 578 to remotely monitor the vehicle.

[0178] In some examples, the server 578 can receive data from the vehicle and apply the data to a latest real-time neural network for real-time intelligent inference. The server 578 can include a deep learning supercomputer powered by the GPU 584 and / or a dedicated AI computer, such as the DGX and DGX Station machines developed by NVIDIA. However, in some examples, the server 578 can include a deep learning infrastructure of a data center powered only by a CPU.

[0179] The deep learning infrastructure of the server 578 may be capable of fast real-time inference and can use this ability to evaluate and verify the health of the processors, software, and / or associated hardware in the vehicle 500. For example, the deep learning infrastructure can receive periodic updates from the vehicle 500, such as an image sequence and / or objects located in the image sequence that the vehicle 500 has identified (e.g., via computer vision and / or other machine learning object classification techniques). The deep learning infrastructure can run its own neural network to identify the objects and compare them with the objects identified by the vehicle 500. If the results do not match and the infrastructure concludes that the AI in the vehicle 500 has failed, then the server 578 can transmit a signal to the vehicle 500, instructing the fail-safe computer of the vehicle 500 to take control, notify the passengers, and complete a safe parking operation.

[0180] For inference, the server 578 can include the GPU 584 and one or more programmable inference accelerators (such as NVIDIA's TensorRT 3). The combination of a GPU-powered server and inference acceleration can enable real-time response. In other examples, such as when performance is less important, a CPU, FPGA, and other processor-powered servers can be used for inference.

[0181] Example computing device

[0182] Figure 6Block diagram of an example computing device 600 suitable for use in implementing some embodiments of the present disclosure. The computing device 600 may include an interconnect system 602 that directly or indirectly couples the following devices: a memory 604, one or more central processing units (CPUs) 606, one or more graphics processing units (GPUs) 608, a communication interface 610, input / output (I / O) ports 612, input / output components 614, a power supply 616, one or more presentation components 618 (such as a display), and one or more logic units 620. In at least one embodiment, the computing device 600 may include one or more virtual machines (VMs), and / or any of its components may include virtual components (e.g., virtual hardware components). For non-limiting examples, one or more of the GPUs 608 may include one or more vGPUs, one or more of the CPUs 606 may include one or more vCPUs, and / or one or more of the logic units 620 may include one or more virtual logic units. Thus, the computing device 600 may include discrete components (e.g., a complete GPU dedicated to the computing device 600), virtual components (e.g., a portion of a GPU dedicated to the computing device 600), or a combination thereof.

[0183] Although Figure 6 each of the boxes of is shown as being connected via the interconnect system 602 having lines, this is not intended to be restrictive and is merely for clarity. For example, in some embodiments, a presentation component 618 such as a display device may be considered an I / O component 614 (e.g., if the display is a touch screen). As another example, the CPU 606 and / or GPU 608 may include memory (e.g., the memory 604 may represent a storage device in addition to the memory of the GPU 608, CPU 606, and / or other components). In other words, Figure 6 the computing devices of are merely illustrative. No distinction is made between categories such as "workstation", "server", "laptop", "desktop", "tablet", "client device", "mobile device", "handheld device", "gaming console", "electronic control unit (ECU)", "virtual reality system", and / or other device or system types, as all of these are considered within the scope of Figure 6 the computing devices of.

[0184] The interconnect system 602 can represent one or more links or buses, such as an address bus, a data bus, a control bus, or a combination thereof. The interconnect system 602 can include one or more types of links or buses, such as an Industry Standard Architecture (ISA) bus, an Extended Industry Standard Architecture (EISA) bus, a Video Electronics Standards Association (VESA) bus, a Peripheral Component Interconnect (PCI) bus, a Peripheral Component Interconnect Express (PCIe) bus, and / or another type of bus or link. In some embodiments, there are direct connections between components. As an example, the CPU 606 can be directly connected to the memory 604. Additionally, the CPU 606 can be directly connected to the GPU 608. In cases where there are direct or point-to-point connections between components, the interconnect system 602 can include a PCIe link to perform the connection. In these examples, the computing device 600 does not need to include a PCI bus.

[0185] The memory 604 can include any of a variety of computer-readable media. Computer-readable media can be any available media that can be accessed by the computing device 600. Computer-readable media can include volatile and non-volatile media as well as removable and non-removable media. By way of example and not limitation, computer-readable media can include computer storage media and communication media.

[0186] Computer storage media can include volatile and non-volatile media and / or removable and non-removable media, implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules, and / or other data types. For example, the memory 604 can store computer-readable instructions (such as those representing programs and / or program elements, such as an operating system). Computer storage media can include, but are not limited to, RAM, ROM, EEPROM, flash memory, or other storage technologies, CD-ROM, digital versatile disks (DVDs), or other optical disk storage devices, magnetic tape cartridges, tapes, magnetic disk storage devices, or any other medium that can be used to store the desired information and that can be accessed by the computing device 600. As used herein, computer storage media does not include signals per se.

[0187] A computer storage medium can include computer-readable instructions, data structures, program modules, and / or other data types in a modulated data signal, such as a carrier wave, or other transmission mechanism, and includes any information conveyance medium. The term "modulated data signal" can refer to a signal that has one or more of its characteristics set or changed in such a manner as to encode information in the signal. By way of example, and not limitation, computer storage media can include wired media such as a wired network or direct wired connection, and wireless media such as sound, RF, infrared, and other wireless media. Any combination of the foregoing should also be included within the scope of computer-readable media.

[0188] The CPU 606 can be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 600 to perform one or more of the methods and / or processes described herein. Each of the CPUs 606 can include one or more cores (e.g., one, two, four, eight, twenty-eight, seventy-two, etc.) capable of simultaneously processing a large number of software threads. The CPU 606 can include any type of processor and can include different types of processors depending on the type of computing device 600 being implemented (e.g., a processor with fewer cores for a mobile device and a processor with more cores for a server). For example, depending on the type of computing device 600, the processor can be an Advanced RISC Machine (ARM) processor implemented using Reduced Instruction Set Computing (RISC) or an x86 processor implemented using Complex Instruction Set Computing (CISC). In addition to one or more microprocessors or supplementary coprocessors such as a math coprocessor, the computing device 600 can also include one or more CPUs 606.

[0189] In addition to or instead of the CPU 606, the GPU 608 can also be configured to execute at least some computer-readable instructions to control one or more components of the computing device 600 to execute one or more of the methods and / or processes described herein. One or more GPUs 608 can be an integrated GPU (e.g., having one or more CPUs 606) and / or one or more GPUs 608 can be a discrete GPU. In an embodiment, one or more GPUs 608 can be a coprocessor of one or more CPUs 606. The computing device 600 can use the GPU 608 to render graphics (e.g., 3D graphics) or perform general computing. For example, the GPU 608 can be used for general-purpose computing on the GPU (GPGPU). The GPU 608 can include hundreds or thousands of cores capable of processing hundreds or thousands of software threads simultaneously. The GPU 608 can generate pixel data for an output image in response to a rendering command (e.g., a rendering command from the CPU 606 received via the host interface). The GPU 608 can include graphics memory such as display memory for storing pixel data or any other suitable data (e.g., GPGPU data). The display memory can be included as part of the memory 604. The GPU 608 can include two or more GPUs operating in parallel (e.g., via a link). The link can directly connect the GPUs (e.g., using NVLINK) or can connect the GPUs through a switch (e.g., using NVSwitch). When combined, each GPU 608 can generate pixel data or GPGPU data for different parts of the output or for different outputs (e.g., the first GPU for the first image and the second GPU for the second image). Each GPU can include its own memory or can share memory with other GPUs.

[0190] In addition to or instead of the CPU 606 and / or the GPU 608, the logic unit 620 can be configured to execute at least some computer-readable instructions to control one or more components of the computing device 600 to execute one or more of the methods and / or processes described herein. In an embodiment, the CPU 606, the GPU 608, and / or the logic unit 620 can execute any combination of methods, processes, and / or portions thereof discretely or jointly. One or more logic units 620 can be part of and / or integrated in one or more CPUs 606 and / or one or more GPUs 608, and / or one or more logic units 620 can be discrete components of the CPU 606 and / or the GPU 608 or otherwise external to them. In an embodiment, one or more logic units 620 can be a processor of one or more CPUs 606 and / or one or more GPUs 608.

[0191] Examples of the logic unit 620 include one or more processing cores and / or their components, such as data processing units (DPUs), tensor cores (TCs), tensor processing units (TPUs), pixel vision cores (PVCs), vision processing units (VPUs), graphics processing clusters (GPCs), texture processing clusters (TPCs), streaming multiprocessors (SMs), tree traversal units (TTUs), artificial intelligence accelerators (AIAs), deep learning accelerators (DLAs), arithmetic logic units (ALUs), application specific integrated circuits (ASICs), floating point units (FPUs), input / output (I / O) elements, peripheral component interconnect (PCI) or peripheral component interconnect express (PCIe) elements, etc.

[0192] The communication interface 610 may include one or more receivers, transmitters, and / or transceivers that enable the computing device 600 to communicate with other computing devices via an electronic communication network, including wired and / or wireless communication. The communication interface 610 may include components and functionality that enable communication over any of several different networks, such as wireless networks (e.g., Wi-Fi, Z-Wave, Bluetooth, Bluetooth LE, ZigBee, etc.), wired networks (e.g., communication via Ethernet or InfiniBand), low power wide area networks (e.g., LoRaWAN, SigFox, etc.), and / or the Internet. In one or more embodiments, one or more of the logic units 620 and / or the communication interface 610 may include one or more data processing units (DPUs) for directly transferring data received over the network and / or the interconnect system 602 to one or more GPUs 608 (e.g., the memory of the GPU 608).

[0193] The I / O port 612 can enable the computing device 600 to be logically coupled to other devices including I / O components 614, presentation components 618, and / or other components, some of which may be built into (e.g., integrated into) the computing device 600. Illustrative I / O components 614 include microphones, mice, keyboards, joysticks, game pads, game controllers, dish satellite antennas, scanners, printers, wireless devices, and so on. The I / O components 614 can provide a natural user interface (NUI) that processes user-generated air gestures, voice, or other physiological inputs. In some instances, the input can be transmitted to appropriate network elements for further processing. The NUI can implement any combination of speech recognition, stylus recognition, face recognition, biometric recognition, on-screen and near-screen gesture recognition, air gestures, head and eye tracking, and touch recognition associated with the display of the computing device 600 (described in more detail below). The computing device 600 can include depth cameras such as stereo camera systems, infrared camera systems, RGB camera systems, touch screen technologies, and combinations thereof for gesture detection and recognition. Additionally, the computing device 600 can include an accelerometer or gyroscope (e.g., as part of an inertial measurement unit (IMU)) that enables motion detection. In some examples, the output of the accelerometer or gyroscope can be used by the computing device 600 to render immersive augmented reality or virtual reality.

[0194] The power supply 616 can include a hard-wired power supply, a battery power supply, or a combination thereof. The power supply 616 can supply power to the computing device 600 to enable the components of the computing device 600 to operate.

[0195] The presentation component 618 can include a display (e.g., a monitor, a touch screen, a television screen, a head-up display (HUD), other display types, or a combination thereof), speakers, and / or other presentation components. The presentation component 618 can receive data from other components (e.g., the GPU 608, the CPU 606, the DPU, etc.) and output the data (e.g., as images, videos, sounds, etc.).

[0196] Example data center

[0197] Figure 7 An example data center 700 is shown, which can be used in at least one embodiment of the present disclosure. The data center 700 can include a data center infrastructure layer 710, a framework layer 720, a software layer 730, and an application layer 740.

[0198] As Figure 7As shown, the data center infrastructure layer 710 may include a resource coordinator 712, grouped computing resources 714, and node computing resources (“node C.R.”) 716(1)-716(N), where “N” represents any whole positive integer. In at least one embodiment, the node C.R. 716(1)-716(N) may include, but is not limited to, any number of central processing units (“CPU”) or other processors (including DPU, accelerator, field programmable gate array (FPGA), graphics processor or graphics processing unit (GPU), etc.), memory devices (such as dynamic read-only memory), storage devices (such as solid state drives or disk drives), network input / output (“NW I / O”) devices, network switches, virtual machines (“VM”), power modules, and cooling modules, etc. In some embodiments, one or more of the node C.R. 716(1)-716(N) may correspond to a server having one or more of the above computing resources. Additionally, in some embodiments, the node C.R. 716(1)-716(N) may include one or more virtual components, such as vGPU, vCPU, etc., and / or one or more of the node C.R. 716(1)-716(N) may correspond to a virtual machine (VM).

[0199] In at least one embodiment, the grouped computing resources 714 may include separate groupings (not shown) of node C.R. 716 housed within one or more racks, or many racks (also not shown) within data centers located in various geographical locations. Separate groupings of node C.R. 716 within the grouped computing resources 714 may include grouped computing, network, memory, or storage resources that may be configured or allocated to support one or more workloads. In at least one embodiment, several node C.R. 716 including CPUs, GPUs, DPUs, and / or other processors may be grouped within one or more racks to provide computing resources to support one or more workloads. One or more racks may also include any number of power modules, cooling modules, and / or network switches in any combination.

[0200] The resource coordinator 712 may configure or otherwise control one or more of the node C.R. 716(1)-716(N) and / or the grouped computing resources 714. In at least one embodiment, the resource coordinator 712 may include a software design infrastructure (“SDI”) management entity for the data center 700. The resource coordinator 712 may include hardware, software, or some combination thereof.

[0201] In at least one embodiment, as Figure 7As shown, the framework layer 720 may include a job scheduler 733, a configuration manager 734, a resource manager 736, and a distributed file system 738. The framework layer 720 may include a framework for software 732 that supports the software layer 730 and / or one or more applications 742 of the application layer 740. The software 732 or the application 742 may respectively include web-based service software or applications, such as service software or applications provided by Amazon Web Services, Google Cloud, and Microsoft Azure. The framework layer 720 may be, but is not limited to, a free and open-source software web application framework, such as Apache Spark that can utilize the distributed file system 738 for large-scale data processing (e.g., "big data"). TM (hereinafter referred to as "Spark"). In at least one embodiment, the job scheduler 733 may include a Spark driver for facilitating the scheduling of workloads supported by the various layers of the data center 700. In at least one embodiment, the configuration manager 734 may be capable of configuring different layers, such as the software layer 730 and the framework layer 720 including Spark and the distributed file system 738 for supporting large-scale data processing. The resource manager 736 is capable of managing the cluster or grouped computing resources mapped to or allocated for supporting the distributed file system 738 and the job scheduler 733. In at least one embodiment, the cluster or grouped computing resources may include grouped computing resources 714 at the data center infrastructure layer 710. The resource manager 736 may coordinate with the resource coordinator 712 to manage these mapped or allocated computing resources.

[0202] In at least one embodiment, the software 732 included in the software layer 730 may include software used by at least a portion of the nodes C.R. 716(1)-716(N), the grouped computing resources 714, and / or the distributed file system 738 of the framework layer 720. One or more types of software may include, but are not limited to, Internet web search software, email virus browsing software, database software, and streaming video content software.

[0203] In at least one embodiment, one or more applications 742 included in the application layer 740 may include one or more types of applications used by at least a portion of nodes C.R. 716(1)-716(N), grouped computing resources 714, and / or the distributed file system 738 of the framework layer 720. The one or more types of applications may include, but are not limited to, any number of genomics applications, cognitive computing and machine learning applications, including training or inference software, machine learning framework software (such as PyTorch, TensorFlow, Caffe, etc.), and / or other machine learning applications used in conjunction with one or more embodiments.

[0204] In at least one embodiment, any one of the configuration manager 734, the resource manager 736, and the resource coordinator 712 may implement any number and type of self-modifying actions based on any amount and type of data obtained in any technically feasible manner. The self-modifying actions may relieve the data center operator of the data center 700 from making potentially bad configuration decisions and may avoid underutilized and / or poorly performing portions of the data center.

[0205] The data center 700 may include tools, services, software, or other resources for training one or more machine learning models or using one or more machine learning models to predict or infer information in accordance with one or more embodiments described herein. For example, a machine learning model may be trained by calculating weight parameters according to a neural network architecture by using the software and computing resources described above with respect to the data center 700. The machine learning model may also or alternatively be improved to detect and / or resist adversarial attacks. In at least one embodiment, by using weight parameters calculated by one or more training techniques, the resources described above with respect to the data center 700 may be used to infer or predict information using a trained machine learning model corresponding to one or more neural networks, such as, but not limited to, those described herein.

[0206] In at least one embodiment, the data center 700 may use a CPU, an application specific integrated circuit (ASIC), a GPU, an FPGA, and / or other hardware (or corresponding virtual computing resources) to perform training and / or inference using the above resources. In addition, one or more of the above software and / or hardware resources may be configured as a service to allow a user to train or perform information inference, such as image recognition, speech recognition, or other artificial intelligence services.

[0207] Example Network Environment

[0208] The network environment suitable for implementing the embodiments of the present disclosure may include one or more client devices, servers, network-attached storage (NAS), other backend devices, and / or other device types. The client devices, servers, and / or other device types (e.g., each device) may be implemented on one or more instances of the Figure 6 computing device 600 - for example, each device may include similar components, features, and / or functions of the computing device 600. Additionally, in the case of implementing backend devices (e.g., servers, NAS, etc.), the backend devices may be included as part of a data center 700, an example of which is described herein with respect to Figure 7 more detail.

[0209] The components of the network environment may communicate with each other via a network, which may be wired, wireless, or both. The network may include multiple networks, or networks within multiple networks. By way of example, the network may include one or more wide area networks (WANs), one or more local area networks (LANs), one or more public networks (e.g., the Internet and / or the public switched telephone network (PSTN)), and / or one or more private networks. In the case where the network includes a wireless telecommunications network, components such as base stations, communication towers, or even access points (and other components) may provide wireless connectivity.

[0210] A compatible network environment may include one or more peer-to-peer network environments (in which case servers may not be included in the network environment), and one or more client-server network environments (in which case one or more servers may be included in the network environment). In a peer-to-peer network environment, the functions described herein with respect to servers may be implemented on any number of client devices.

[0211] In at least one embodiment, the network environment may include one or more cloud-based network environments, distributed computing environments, combinations thereof, etc. A cloud-based network environment may include a framework layer, a job scheduler, a resource manager, and a distributed file system implemented on one or more servers, which may include one or more core network servers and / or edge servers. The framework layer may include a framework for supporting one or more applications of a software layer and / or an application layer. The software or application may respectively include network-based service software or applications. In an embodiment, one or more client devices may use network-based service software or applications (e.g., by accessing the service software and / or applications via one or more application programming interfaces (APIs)). The framework layer may be, but is not limited to, a type of free and open-source software web application framework, for example, which may use a distributed file system for large-scale data processing (e.g., "big data").

[0212] A cloud-based network environment can provide cloud computing and / or cloud storage that perform any combination of the computing and / or data storage functions (or one or more portions thereof) described herein. Any one of these various functions can be distributed across multiple locations from a central or core server (e.g., one or more data centers that can be distributed across states, regions, countries, the globe, etc.). If the connection to a user (e.g., a client device) is relatively close to an edge server, the core server can assign at least a portion of the function to the edge server. A cloud-based network environment can be private (e.g., limited to a single organization), public (e.g., available to many organizations), and / or a combination thereof (e.g., a hybrid cloud environment).

[0213] The client device can include at least some of the components, features, and functions of the example computing device 600 described herein with respect to Figure 6 As an example and not a limitation, the client device can be embodied as a personal computer (PC), laptop computer, mobile device, smartphone, tablet computer, smartwatch, wearable computer, personal digital assistant (PDA), MP3 player, virtual reality headset, global positioning system (GPS) or device, video player, camera, surveillance device or system, vehicle, boat, aircraft, virtual machine, drone, robot, handheld communication device, hospital device, gaming device or system, entertainment system, in-vehicle computer system, embedded system controller, remote control, appliance, consumer electronic device, workstation, edge device, any combination of these described devices, or any other suitable device.

[0214] In summary, the disclosed techniques use adversarial images (or other types of adversarial data) to fine-tune a pre-trained machine learning model. Adversarial images can be generated by perturbing a set of original images based on one or more adversarial attack techniques. After generating the adversarial images, the adversarial images are input into the pre-trained machine learning model, and the predictions generated by the pre-trained machine learning model based on the input adversarial images are compared with the labels of the original images from which the adversarial images were generated. Based on these comparisons, the adversarial images are divided into a "clean" adversarial image dataset for which the pre-trained machine learning model generates correct predictions and an "adversarial" adversarial image dataset for which the pre-trained machine learning model generates incorrect predictions.

[0215] Then, each adversarial image dataset and the corresponding set of labels are used to refine the machine learning model. During the refinement process, the adversarial images in the clean dataset are input into the machine learning model, and the machine learning model is trained based on the loss between the predictions generated by the machine learning model based on the input adversarial images and the labels of the original images corresponding to the input adversarial images.

[0216] Adversarial images in the adversarial dataset are also input into the machine learning model, and the machine learning model is trained based on the loss between the prediction generated by the machine learning model according to the input adversarial image and the label determined based on the perturbation strength associated with the input adversarial image. More specifically, adversarial images in the adversarial dataset that are associated with a perturbation strength less than a threshold amount are assigned the label corresponding to the original image, while adversarial images in the adversarial dataset that are associated with a perturbation strength greater than the threshold amount are assigned an "adversarial label" that indicates that the corresponding image (or image region) should be ignored when determining the action to be taken based on the output of the machine learning model. Then, the machine learning model is trained based on the loss between the prediction generated by the machine learning model according to the input adversarial image and the corresponding label.

[0217] After training the machine learning model using the adversarial images and the corresponding labels, the machine learning model can be used in vision ADSs and / or other types of systems to generate predictions for additional images. For example, the machine learning model can be used to detect and / or track objects in the vicinity of an autonomous vehicle during its operation. The output of the machine learning model can be provided to one or more downstream components that generate commands and / or signals for operating the autonomous vehicle.

[0218] Compared with the prior art, one technical advantage of the disclosed technology is that the machine learning model can detect and correct adversarial data that attempts to cause the machine learning model to operate incorrectly. Therefore, a system that utilizes the output of the machine learning model is safer and more resistant to adversarial attacks compared to traditional systems that include machine learning models not trained to identify and / or defend against adversarial attacks. Another technical advantage of the disclosed technology is that it reduces computational overhead and resource consumption compared to prior art methods that use an ensemble of multiple machine learning models to defend against adversarial attacks. These technical advantages provide one or more technical improvements over prior art methods.

[0219] 1. In some embodiments, a method includes: generating, using a machine learning model and based at least on a sensor data instance, one or more base outputs and one or more adversarial outputs, the one or more adversarial outputs representing a likelihood that the sensor data instance is adversarial; determining, based on the one or more adversarial outputs, that the sensor data instance is adversarial; and sending an indication that the sensor data instance is adversarial to one or more downstream components to cause the one or more downstream components to perform one or more operations on the one or more base outputs and according to the indication.

[0220] 2. The method according to clause 1 further includes: using the machine learning model and at least based on a second sensor data instance, generating one or more second base outputs and one or more second adversarial outputs, where the one or more second adversarial outputs represent the possibility that the second sensor data instance is adversarial; using the one or more second base outputs, determining an output type corresponding to the second sensor data instance at least based on the one or more second adversarial outputs; and sending a second indication of the output type corresponding to the second sensor data instance to the one or more downstream components.

[0221] 3. The method according to any one of clauses 1-2, wherein determining the output type includes: matching the highest confidence included in the one or more second base outputs and the one or more second adversarial outputs with the output type.

[0222] 4. The method according to any one of clauses 1-3, wherein the machine learning model is trained at least based on: updating one or more parameters of the machine learning model at least based on one or more losses associated with a set of adversarial sensor data instances processed using the machine learning model.

[0223] 5. The method according to any one of clauses 1-4, wherein the one or more losses are calculated using a set of outputs generated using the machine learning model based on the set of adversarial sensor data instances and a set of labels corresponding to the set of adversarial sensor data instances, where the set of labels includes: (1) the original labels of the original sensor data instances that have been perturbed to generate the first adversarial sensor data instances included in the set of adversarial sensor data instances; and (2) adversarial labels, which indicate that the second adversarial sensor data instances included in the set of adversarial sensor data instances correspond to adversarial attacks.

[0224] 6. The method according to any one of clauses 1-5, wherein determining that the sensor data instance is adversarial includes: determining that at least one of the one or more adversarial outputs exceeds a threshold.

[0225] 7. The method according to any one of clauses 1-6, wherein determining that the sensor data instance is adversarial includes: determining that the one or more adversarial outputs include a confidence level higher than the confidence levels associated with the one or more base outputs.

[0226] 8. The method according to any one of clauses 1-7, wherein the sensor data instance includes one or more images.

[0227] 9. A method according to any one of clauses 1 - 8, wherein the one or more base outputs correspond to a set of classes associated with an object.

[0228] 10. A method according to any one of clauses 1 to 9, wherein the one or more downstream components are included in at least one of the following: a control system for an autonomous or semi - autonomous machine; a sensing system for an autonomous or semi - autonomous machine; a system for performing simulation operations; a system for performing digital twin operations; a system for performing optical transmission simulation; a system for performing 3D asset collaborative content creation; a system for performing deep learning operations; a system implemented using an edge device; a system for generating or presenting at least one of virtual reality content, augmented reality content, or mixed reality content; a system implemented using a robot; a system for performing conversational artificial intelligence operations; a system for generating synthetic data; a system comprising one or more virtual machines (VMs); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources.

[0229] 11. In some embodiments, a method includes: generating, using a machine learning model and based at least on a perturbed sensor data instance, an output that includes a base output and an adversarial output, the adversarial output corresponding to a likelihood that the perturbed sensor data instance is adversarial; calculating one or more loss values based at least on the base output, the adversarial output, and ground - truth data indicating that the perturbed sensor data instance is adversarial; and updating one or more parameters of the machine learning model based at least on the one or more loss values.

[0230] 12. The method according to clause 11, further comprising: generating, using the machine learning model and based at least on a second perturbed sensor data instance, a second base output and a second adversarial output, the second adversarial output corresponding to a likelihood that the second perturbed sensor data instance is adversarial; calculating one or more second loss values based at least on the second base output, the second adversarial output, and second ground - truth data corresponding to an original sensor data instance that was perturbed to generate the second perturbed sensor data instance; and updating at least one of the one or more parameters of the machine learning model or one or more other parameters based at least on the one or more second loss values.

[0231] 13. The method according to any one of clauses 11 - 12, further comprising: assigning the second ground - truth data to the second perturbed sensor data instance based at least on a match between the second ground - truth data and a previous output generated using the machine learning model based at least on the second perturbed sensor data instance.

[0232] 14. The method according to any one of clauses 11-13 further includes: determining the second ground truth data based at least on the perturbation intensity associated with the second perturbed sensor data instance being lower than a threshold perturbation intensity.

[0233] 15. The method according to any one of clauses 11-14, wherein calculating the one or more second loss values is further based at least on the second base output and third ground truth data indicating that the second perturbed sensor data instance corresponds to an adversarial attack, wherein the second ground truth data includes a first target probability of an output type of the second base output, and the third ground truth data includes a second target probability of the second adversarial output.

[0234] 16. The method according to any one of clauses 11-15 further includes: determining the ground truth data based at least on the perturbation intensity associated with the perturbed sensor data instance exceeding a threshold perturbation intensity.

[0235] 17. The method according to any one of clauses 11-16 further includes: applying one or more perturbations to an original sensor data instance to generate the perturbed sensor data instance.

[0236] 18. The method according to any one of clauses 11-17, wherein the one or more perturbations include at least one of the following: Fast Gradient Sign Method (FGSM), Projected Gradient Descent (PGD) technique, Limited-memory Broyden-Fletcher-Goldfarb-Shanno (L-BFGS) technique, Carlini & Wagner technique, Jacobian-based Saliency Map Attack (JSMA), or DeepFool technique.

[0237] 19. In some embodiments, a system includes: one or more processing units configured to perform one or more operations on a base output of a neural network at least based on an indication of an adversarial attack, the indication of the adversarial attack being generated at least based on an adversarial output of the neural network different from the base output.

[0238] 20. The system according to claim 19, wherein the system is included in at least one of the following: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing simulation operations; a system for performing digital twin operations; a system for performing optical transmission simulation; a system for performing 3D asset collaborative content creation; a system for performing deep learning operations; a system implemented using an edge device; a system for generating or presenting at least one of virtual reality content, augmented reality content, or mixed reality content; a system implemented using a robot; a system for performing conversational AI operations; a system for generating synthetic data; a system including one or more virtual machines (VMs); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources.

[0239] The present disclosure may be described in the general context of machine-usable instructions or computer code, including computer-executable instructions such as program modules, executed by a computer or other machine, such as a personal digital assistant or other handheld device. Generally, program modules, including routines, programs, objects, components, data structures, etc., refer to code that performs particular tasks or implements particular abstract data types. The present disclosure may be practiced in a variety of system configurations, including handheld devices, consumer electronics, general-purpose computers, more specialized computing devices, etc. The present disclosure may also be practiced in a distributed computing environment where tasks are performed by remote processing devices linked through a communication network.

[0240] As used herein, the recitation of "and / or" with respect to two or more elements should be construed to refer to only one element or a combination of elements. For example, "element A, element B, and / or element C" may include only element A, only element B, only element C, element A and element B, element A and element C, element B and element C, or element A, B, and C. Additionally, "at least one of element A or element B" may include at least one of element A, at least one of element B, or at least one of element A and at least one of element B. Further, "at least one of element A and element B" may include at least one of element A, at least one of element B, or at least one of element A and at least one of element B.

[0241] This disclosure describes the subject matter of the present disclosure in detail to meet statutory requirements. However, the description itself is not intended to limit the scope of the present disclosure. On the contrary, the inventors have contemplated that the claimed subject matter may also be embodied in other ways, including steps different from those described herein or combinations of similar steps in conjunction with other current or future technologies. Moreover, although the terms "step" and / or "block" may be used herein to imply different elements of a method employed, these terms should not be construed as implying any particular order among or between the various steps disclosed herein unless the order of the steps is expressly described.

Claims

1. A method, comprising: generating, using a machine learning model and based at least on sensor data instances, one or more base outputs and one or more adversarial outputs, the one or more adversarial outputs indicating a likelihood that the sensor data instances are adversarial; determining, based on the one or more adversarial outputs, that the sensor data instances are adversarial; and sending an indication that the sensor data instances are adversarial to one or more downstream components to cause the one or more downstream components to perform one or more operations on the one or more base outputs and in accordance with the indication.

2. The method according to claim 1, further comprising: generating, using the machine learning model and based at least on second sensor data instances, one or more second base outputs and one or more second adversarial outputs, the one or more second adversarial outputs indicating a likelihood that the second sensor data instances are adversarial; determining, using the one or more second base outputs, an output type corresponding to the second sensor data instances based at least on the one or more second adversarial outputs; and sending a second indication of the output type corresponding to the second sensor data instances to the one or more downstream components.

3. The method according to claim 2, wherein determining the output type comprises: Matching the highest confidence included in the one or more second base outputs and the one or more second adversarial outputs with the output type.

4. The method according to claim 1, wherein the machine learning model is trained at least based on: updating one or more parameters of the machine learning model based at least on one or more losses associated with a set of adversarial sensor data instances processed using the machine learning model.

5. The method according to claim 4, wherein the one or more losses are calculated using a set of outputs generated using the machine learning model based on the set of adversarial sensor data instances and a set of labels corresponding to the set of adversarial sensor data instances, wherein the set of labels includes: (1) The original labels of the original sensor data instances that have been perturbed to produce the first adversarial sensor data instances included in the set of adversarial sensor data instances; and (2) adversarial labels that indicate that the second adversarial sensor data instances included in the set of adversarial sensor data instances correspond to adversarial attacks.

6. The method according to claim 1, wherein determining that the sensor data instance is adversarial comprises: Determining that at least one of the one or more adversarial outputs exceeds a threshold.

7. The method according to claim 1, wherein determining that the sensor data instance is adversarial comprises: Determining that the one or more adversarial outputs include a confidence that is higher than the confidence associated with the one or more base outputs.

8. The method according to claim 1, wherein the sensor data instances include one or more images.

9. The method according to claim 1, wherein the one or more base outputs correspond to a set of classes associated with an object.

10. The method according to claim 1, wherein the one or more downstream components are included in at least one of the following: A control system for an autonomous or semi-autonomous machine; A perception system for an autonomous or semi-autonomous machine; A system for performing simulation operations; A system for performing digital twin operations; A system for performing optical transmission simulation; A system for performing 3D asset collaborative content creation; A system for performing deep learning operations; A system implemented using an edge device; A system for generating or presenting at least one of virtual reality content, augmented reality content, or mixed reality content; A system implemented using a robot; A system for performing conversational artificial intelligence operations; A system for generating synthetic data; A system including one or more virtual machines (VMs); A system implemented at least partially in a data center; or A system implemented at least partially using cloud computing resources.

11. A method, comprising: Generating an output using a machine learning model and based at least on a perturbed sensor data instance, the output including a base output and an adversarial output, the adversarial output corresponding to a likelihood that the perturbed sensor data instance is adversarial; Calculating one or more loss values based at least on the base output, the adversarial output, and ground truth data indicating that the perturbed sensor data instance is adversarial; And Updating one or more parameters of the machine learning model based at least on the one or more loss values.

12. The method according to claim 11, further comprising: Generating a second base output and a second adversarial output using the machine learning model and based at least on a second perturbed sensor data instance, the second adversarial output corresponding to a likelihood that the second perturbed sensor data instance is adversarial; Calculating one or more second loss values based at least on the second base output, the second adversarial output, and second ground truth data corresponding to an original sensor data instance perturbed to generate the second perturbed sensor data instance; and Updating at least one of the one or more parameters of the machine learning model or one or more other parameters based at least on the one or more second loss values.

13. The method according to claim 12 further comprises: Assigning the second ground truth data to the second perturbed sensor data instance based at least on a match between the second ground truth data and a previous output generated using the machine learning model and based at least on the second perturbed sensor data instance.

14. The method according to claim 12, further comprising: Determining the second ground truth data based at least on a perturbation intensity associated with the second perturbed sensor data instance being lower than a threshold perturbation intensity.

15. The method according to claim 12, wherein calculating the one or more second loss values is further based at least on the second base output and third ground truth data indicating that the second perturbed sensor data instance corresponds to an adversarial attack, wherein the second ground truth data includes a first target probability of an output type of the second base output, and the third ground truth data includes a second target probability of the second adversarial output.

16. The method according to claim 11 further comprises: Determining the ground truth data based at least on a perturbation intensity associated with the perturbed sensor data instance exceeding a threshold perturbation intensity.

17. The method according to claim 11 further comprises: Applying one or more perturbations to an original sensor data instance to generate the perturbed sensor data instance.

18. The method according to claim 17, wherein the one or more perturbations include at least one of the following: Fast Gradient Sign Method (FGSM), Projected Gradient Descent (PGD) technique, Limited-memory Broyden-Fletcher-Goldfarb-Shanno (L-BFGS) technique, Carlini & Wagner technique, Jacobian-based Saliency Map Attack (JSMA), or DeepFool technique.

19. A system comprising: one or more processing units configured to perform one or more operations on a base output of a neural network based at least on an indication of an adversarial attack, the indication of the adversarial attack being generated at least based on an adversarial output of the neural network that is different from the base output.

20. The system according to claim 19, wherein the system is included in at least one of the following: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing simulation operations; a system for performing digital twin operations; a system for performing optical transmission simulation; a system for performing 3D asset collaborative content creation; a system for performing deep learning operations; a system implemented using edge devices; a system for generating or presenting at least one of virtual reality content, augmented reality content, or mixed reality content; a system implemented using a robot; a system for performing conversational AI operations; a system for generating synthetic data; a system comprising one or more virtual machines (VMs); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources.

Citation Information

Patent Citations

  • Method for programmable timeouts of tree traversal mechanisms in hardware

    US10885698B2

Cited By

  • Contragitive texture generation method, device and product

    CN121190642A