OBSCRIBING IMAGE CONTENT DURING CODING FOR AUTOMOTIVE SYSTEMS AND APPLICATIONS
By employing an encoder to obscure sensitive image content using fixed quantization parameters and altered residual information, the system efficiently addresses the challenge of protecting privacy in vehicle-generated image data, reducing computational resources and processing time.
Patent Information
- Application Number
- DE102024134074
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-20
- Filing Date
- 2024-11-20
- Publication Date
- 2025-05-22
AI Technical Summary
Existing systems lack efficient methods to obscure sensitive image content, such as license plates and pedestrian faces, in image data generated by vehicles, which is important for privacy reasons.
The use of an encoder to obscure image portions representing sensitive content during the encoding process, by employing techniques such as fixed quantization parameters and altering residual information, without requiring additional hardware or software components.
This approach effectively obscures sensitive image content, reducing the computational resources needed and processing time, while ensuring that the obscured content cannot be retrieved or relayed.
Smart Images

Figure 00000044_0000 
Figure 00000045_0000 
Figure 00000046_0000
Abstract
Description
BACKGROUND
[0001] In certain circumstances, it may be important or desirable to obscure image portions that depict certain types of content—such as personal or private content—so that the content is unidentifiable and / or cannot be retrieved or shared. For example, a vehicle or other machine may contain one or more cameras that generate image data depicting an environment around the vehicle. The vehicle may then send the image data to one or more systems that perform one or more tasks using the image data, such as processing the image data, creating simulations using the image data, creating maps using the image data, training one or more machine learning models using the image data, and / or so on.However, the images generated by the vehicle may contain sensitive, private, or personal information, such as license plates of other vehicles and / or the faces of nearby pedestrians. Therefore, it may be important and / or necessary (e.g., for privacy reasons) for the vehicle to blur this sensitive information before sending the image data to a data center, another vehicle or machine, a user device, and / or another system. SUMMARY
[0002] The invention is defined by the claims. To illustrate the invention, aspects and embodiments are described herein, which may or may not fall within the scope of the claims.
[0003] In various examples, the obscuring of image content during encoding for automotive systems and applications is described herein. Systems and methods are disclosed that utilize an encoder to obscure image portions representing certain types of content during the process of encoding image data representing the images. For example, the image may be processed to determine an image portion associated with a certain type of content (e.g., a representation), such as sensitive or confidential information. The encoder may then obscure the image portion during the process of encoding the image data. As described herein, the encoder may employ various techniques to obscure the image portion, such as by encoding a portion of the image data using specified quantization parameters (e.g.,a maximum quantization value) and / or by changing the residual information associated with the part of the image data.
[0004] Embodiments of the present disclosure relate to the obscuring of image content during encoding for automotive systems and applications. Systems and methods are disclosed that utilize an encoder to obscure image portions representing certain types of content during the process of encoding image data representing the images. For example, for an image, the image may be processed to determine an image portion associated with a certain type of content (e.g., a representation), such as sensitive, confidential, private, and / or personal information. Data indicative of the image portion may then be generated, such asData representing a boundary shape indicating the portion(s) of the image, data representing a feature map indicating the portion(s) of the image, data representing a segmentation mask indicating the portion(s) of the image, and / or another indicator or representation of the portion(s) of the image that is / are to be obscured, blurred, or otherwise made less recognizable or completely unrecognizable. The encoder may then use the data to obscure the image portion during the process of encoding the image data. As described herein, the encoder may use various techniques to obscure the image portion, such as by encoding a portion of the image data using specified quantization parameters (e.g., a maximum and / or predetermined quantization value) and / or by altering the residual information associated with the portion of the image data.
[0005] By performing the processes described herein, current systems in some embodiments are capable of blurring portions of images with the encoder and at least partially during encoding of the image data. Therefore, current systems may not require additional components, such as additional hardware components and / or software components, to blur portions of the images. This may, in some embodiments, improve current systems by saving the amount of computing resources required to blur portions of the images and / or reducing the overall time required to process the image data.
[0006] The disclosure extends to any new aspects or features described and / or illustrated herein.
[0007] Further features of the disclosure are characterized by the independent and dependent claims.
[0008] Any feature of one aspect of the disclosure may be applied to other aspects of the disclosure, in any suitable combination. In particular, method aspects may be applied to device or system aspects, and vice versa.
[0009] Furthermore, functions implemented in hardware may also be implemented in software, and vice versa. Any reference to software and hardware features in this document should be interpreted accordingly.
[0010] Any system or device feature described herein may also be provided as a method feature, and vice versa. System and / or device aspects described functionally (including means-plus-functional features) may alternatively be expressed in terms of their corresponding structure, e.g., an appropriately programmed processor and associated memory.
[0011] It should also be appreciated that certain combinations of the various features described and defined in all aspects of the disclosure may be implemented and / or provided and / or used independently of one another.
[0012] The disclosure also provides computer programs and computer program products comprising software code that, when executed on a data processing device, is capable of performing any of the methods described herein and / or embodying any of the device and system features described herein, including all or part of the component steps of a method.
[0013] The disclosure also provides a computer or computer system (including networked or distributed systems) having an operating system that supports a computer program for performing any of the methods described herein and / or for embodying any of the device or system features described herein.
[0014] The disclosure also provides a computer-readable medium on which one or more of the above-mentioned computer programs are stored.
[0015] The disclosure also provides a signal carrying one or more of the above-mentioned computer programs.
[0016] The disclosure extends to methods and / or devices and / or systems as described herein with reference to the accompanying drawings.
[0017] Aspects and embodiments of the disclosure will now be described, by way of example only, with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The present systems and methods for obscuring image content during coding for automotive systems and applications are described in detail below with reference to the accompanying drawing figures, wherein: Fig. 1 shows an exemplary data flow diagram for a process of obscuring image content during encoding in accordance with some embodiments of the present disclosure; The Fig. 2A and Fig. 2B show examples of images representing types of content that may be obscured in accordance with some embodiments of the present disclosure; The Fig. 3A and Fig. 3B illustrates an example of limiting shape information associated with an image portion representing a type of content, in accordance with some embodiments of the present disclosure; The Fig. 4A-4B illustrate an example of generating feature maps associated with an image in accordance with some embodiments of the present disclosure; The Fig. 5A and Fig. 5B illustrates an example of encoding image data using a fixed quantization parameter value to obscure an image portion according to some embodiments of the present disclosure; The Fig. 6A-6B illustrate an example of encoding image data by updating residual information to obscure an image portion according to some embodiments of the present disclosure; Fig. 7 is a flowchart illustrating a method for using an encoder to obscure an image portion associated with a type of content, in accordance with some embodiments of the present disclosure; Fig. 8 is a flowchart illustrating a method for using an encoder to blur an image portion based at least on a quantization parameter value, in accordance with some embodiments of the present disclosure; Fig. 9 is a flowchart illustrating a method for using an encoder to blur an image patch based at least on updating residual information, in accordance with some embodiments of the present disclosure; Fig. 10A is an illustration of an example of an autonomous vehicle according to some embodiments of the present disclosure; Fig. 10B is an example of camera positions and fields of view for the autonomous vehicle from Fig. 10A, in accordance with some embodiments of the present disclosure; Fig. 10C is a block diagram of an example system architecture for the example autonomous vehicle of Fig. 10A, in accordance with some embodiments of the present disclosure; Fig. 10D is a system diagram for the communication between the cloud-based server(s) and the autonomous example vehicle of Fig. 10A, in accordance with some embodiments of the present disclosure; Fig. 11 is a block diagram of an example computing device suitable for use in implementing some embodiments of the present disclosure; and Fig. 12 is a block diagram of a data center suitable for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION
[0019] Systems and methods for obscuring image content during encoding for automotive systems and applications are disclosed. Although the present disclosure is described with respect to an example of an autonomous or semi-autonomous vehicle or machine 1000 (also referred to herein as "vehicle 1000," "ego-vehicle 1000," "machine 1000," or "ego-machine 1000," an example of which is described with respect to the Fig. 10A-10D), this is not intended to be limiting. For example, the systems and methods described herein may be used, without limitation, by non-autonomous vehicles or machines, semi-autonomous vehicles or machines (e.g., in one or more adaptive driver assistance systems (ADAS)), autonomous vehicles or machines, guided and unguided robots or robotic platforms, warehouse vehicles, off-highway vehicles, vehicles coupled to one or more trailers, flying vessels, boats, shuttles, emergency medical vehicles, motorcycles, electric or motorized bicycles, aircraft, construction vehicles, underwater vehicles, drones, and / or other types of vehicles.Although the present disclosure is described with respect to object recognition and / or map creation, this is not to be construed as limiting, and the systems and methods described herein may be used in the fields of augmented reality, virtual reality, mixed reality, robotics, security and surveillance, autonomous or semi-autonomous machine applications, and / or any other technology area in which object recognition and / or map creation may be used.
[0020] For example, one or more systems may receive image data generated using one or more image sensors, such as one or more cameras of one or more machines (e.g., a vehicle). In some examples, the system(s) may be internal to the machine; in other examples, the system(s) may be external to the machine and receive the image data over one or more networks. The system(s) may then process image data representing at least one image to identify at least one image portion that corresponds to (e.g., represents) a particular type of content. As described herein, in some examples, the specific type of content may include sensitive, confidential, personal, or private information, such as privacy information (e.g., identification information) associated with vehicles, people, buildings, accounts, and / or so forth.For example, when a machine is navigating an environment, the image may depict sensitive information such as one or more license plates of one or more vehicles, one or more pedestrian faces (and / or portions of pedestrian faces), one or more unique vehicle identifiers, one or more logos, and / or any other type of sensitive information. However, in other examples, the specific type of content may include any other type of predetermined or preselected content that the system(s) are designed to obscure, as described in detail herein.
[0021] The system(s) may then generate data (referred to as "position data" in some examples) indicating the image patch corresponding to the type of content. In some examples, the position data may represent a bounding shape (e.g., a bounding box, a polygon, a circle, etc.) that at least partially encloses the type of content. For example, the position data may represent vertex positions within the image for the bounding shape, where a vertex position may include a pixel position, coordinate positions (e.g., an x-coordinate position, a y-coordinate position, etc.), and / or any other type of information identifying a position of a vertex. Additionally or alternatively, in some examples, the position data may represent a feature map associated with the image, where the feature map indicates the image patch.For example, the feature map may be divided into blocks, with each block representing a region of the image (e.g., a 4x4 pixel region, an 8x8 pixel region, a 16x16 pixel region, etc.). One or more blocks associated with the image patch may contain an indication of the type of content, which is further described herein. In other examples, a segmentation mask may be used to identify pixels to be blurred and pixels not to be blurred (e.g., using a binary 0 (no adjustment of QP or other parameters due to the nature of the content) or 1 (maximum QP or other parameters that result in the content being blurred) to indicate where blurring should be implemented, or using a range of values between 0 and 1, where the larger the value, the more opaque the underlying blurred feature).
[0022] The system(s) may then use an encoder to obscure (e.g., remove, block, blur, alter, etc.) the image patch so that the image patch is less identifiable or unidentifiable and / or cannot be retrieved or passed on, e.g., by a decoder. For example, the encoder may receive at least the image data representing the image and the position data indicating the image patch. The encoder may then use the image data and the position data to obscure the image patch using one or more techniques. In a first example, the encoder may use a quantization parameter that may be programmed, set, and / or determined when encoding a portion of the image data corresponding to the image patch.As described herein, the quantization parameter may be equal to or greater than a threshold to increase the quantization associated with the portion of the image data. For example, the encoder may use a maximum value for the quantization parameter to maximize the quantization associated with the image portion, thereby minimizing the quality associated with the image portion.
[0023] In a second example, during the encoding process, the encoder may generate residual data associated with the image, such as a residual image containing residual information associated with the image. As such, the encoder may further process a portion of the residual data associated with the image portion, e.g., by reducing values (e.g., pixel values and / or coefficient values) associated with the portion of the residual data so that they are less than or equal to a threshold. As described herein, in some examples, the encoder may reduce the values to a minimum value, such as 0. Furthermore, in some examples, the encoder may perform intra-prediction associated with the image, which is described in more detail herein.In some examples, these processes may ensure that additional blocks associated with encoding the image data cannot receive the data associated with the image slice, which may result in the data not being able to be recaptured.
[0024] While the examples above describe how these processes are performed to obscure a portion of an image, other examples may use similar processes to obscure more than one portion of the image. For example, if the image contains two parts related to the type of content, similar processes may be used to obscure both parts of the image. While the examples above describe how these processes are performed for a single image, other examples may use similar processes for any number of images represented by the image data.
[0025] As described herein, after encoding the image data to generate the encoded image data, the system(s) may send the encoded image data to one or more additional systems. The additional system(s) may then store the encoded image data, process the encoded image data, generate a map using the encoded image data, train one or more machine learning models using the encoded image data, and / or perform any other task using the encoded image data. However, because the system(s) performed the processes described herein to initially obscure the portions of the images corresponding to the nature of the content, the additional system(s) may not be able to identify and / or recapture the content.For example, when the additional system(s) decodes the encoded image data, the images represented by the decoded image data may no longer represent the content, such as the license plates and / or the faces of pedestrians.
[0026] While the examples described herein describe techniques for encoding image data to obscure at least a portion of the image data, in other examples, similar methods may be used to obscure at least a portion of other data types. For example, similar techniques may be used to obscure at least a portion of LiDAR data, RADAR data, audio data, etc., during an encoding process.
[0027] The systems and methods described herein may be used, without limitation, by non-autonomous vehicles or machines, semi-autonomous vehicles or machines (e.g., in one or more adaptive driver assistance systems (ADAS)), autonomous vehicles or machines, guided and unguided robots or robotic platforms, warehouse vehicles, off-road vehicles, vehicles coupled to one or more trailers, flying vessels, boats, shuttles, emergency medical vehicles, motorcycles, electric or motorized bicycles, aircraft, construction vehicles, underwater vehicles, drones, and / or other types of vehicles.Furthermore, the systems and methods described herein may be used for a variety of purposes, including, without limitation, machine control, machine locomotion, machine propulsion, synthetic data generation, model training, perception, augmented reality, virtual reality, mixed reality, robotics, security and surveillance, simulation and digital twinning, autonomous or semi-autonomous machine applications, deep learning, environmental simulation, object or actor simulation and / or digital twinning, data center processing, conversational AI, light transport simulation (e.g., ray tracing, path tracing, etc.), collaborative content creation for 3D assets, cloud computing, and / or other suitable applications.
[0028] The presented embodiments may be included in a variety of different systems, such as automotive systems (e.g., a control system for an autonomous or semi-autonomous machine, a perception system for an autonomous or semi-autonomous machine), systems implemented using a robot, aviation systems, medical systems, boat systems, intelligent area monitoring systems, systems for performing deep learning operations, systems for performing simulation operations, systems for performing digital twin operations, systems implemented using an edge device, systems implementing large language models (LLMs), systems containing one or more virtual machines (VMs), systems for performing operations for generating synthetic data, systems implemented at least partially in a data center,Systems for performing conversational AI operations, systems for performing light transport simulations, systems for performing collaborative content creation for 3D assets, systems for performing generative AI operations, systems implemented at least in part using cloud computing resources, and / or other types of systems.
[0029] With reference to Fig. 1 shows Fig. 1 illustrates an exemplary dataflow diagram for a process 100 for obscuring image content during encoding in accordance with some embodiments of the present disclosure. It should be understood that these and other arrangements described herein are examples only. Other arrangements and elements (e.g., engines, interfaces, functions, arrangements, groupings of functions, etc.) may be used in addition to or in place of those shown, and some elements may be omitted entirely. Furthermore, many of the elements described herein are functional units that may be implemented as individual or distributed components, or in conjunction with other components, and in any suitable combination and location. Various functions performed by units described herein may be performed by hardware, firmware, and / or software.For example, various functions may be performed by a processor executing instructions stored in memory. In some embodiments, the systems, methods, and processes described herein may be implemented using similar components, features, and / or functions as those of the example autonomous vehicle 1000 of FIG. Fig. 10A-10D, the example computing device 1100 of Fig. 11 and / or the example data center 1200 of Fig. 12.
[0030] The method 100 may include a blurring component 102 that receives image data 104 generated using one or more image sensors 106. As described herein, in some examples, the blurring component 102 and / or the image sensor(s) 106 may be connected to a machine, such as the vehicle 1000. For example, the machine may use the image sensor(s) 106 when navigating an environment to generate the image data 104 that represents one or more images depicting the environment. Furthermore, the image(s) represented by the image data 104 may represent a particular type of content that the blurring component 102 is configured to blur within the image(s). As a first example, if the image data 104 is generated using a machine navigating an environment, the image(s) mayThe images may depict sensitive information (e.g., a type of content), such as license plates of other vehicles, pedestrian faces, and / or other information that is considered sensitive and / or must be obscured for privacy reasons. A second example: If the image data 104 depicts one or more images of personal documents, the images may depict personal information (e.g., a type of content), such as social security numbers, passwords, and / or other information that is considered sensitive and / or must be obscured.
[0031] In the Fig. 2A and Fig. 2B, examples of images are shown that depict a type of content that should be obscured according to some embodiments of the present disclosure. In the example of Fig. 2A, an image 202 may depict an environment 204 comprising three vehicles 206(1)-(3) (also referred to individually as "vehicle 206" or plurally as "vehicles 206") traveling along a road. As shown, the image 202 depicts at least one license plate 208 associated with the vehicle 206(1), wherein the information of the license plate 208 (e.g., numbers) may be considered sensitive information (e.g., one type of content), while other portions of the image 202 depict other types of content (e.g., the vehicles 206 themselves, roads, etc.) that are not considered sensitive information. Furthermore, an image 210, as in the example of Fig. 2B, an environment 212 may depict two pedestrians 214(1)-(2) (also individually referred to as "pedestrian 214" or plurally as "pedestrians 214"). As shown, the image 210 depicts the faces 216(1)-(2) (also individually referred to as "face 216" or plurally as "faces 216") of the pedestrians 214(1)-(2), where the faces 216 may be considered sensitive information (e.g., some type of content).
[0032] Returning to the example of Fig. 1, the process 100 may include the obscuring component 102 using an inference component 108 to identify one or more portions of the image(s) that correspond to the type of content. The inference component 108 may, for example, include object detection, segmentation, and / or classification functions. For example, for an image, the inference component 108 may process the image data 104 representing the image using one or more techniques, such as using one or more machine learning models, one or more neural networks, one or more perceptual models, one or more object detection models, and / or another technique. At least based on the processing, the inference component 108 may determine that a portion of the image corresponds to (e.g., represents) the type of content. For example, the inference component 108 maydetermine the portion of the image that represents the type of content. The inference component 108 can then generate position data 110 that represents information indicating at least the image section.
[0033] As described herein, in some examples, the position data 110 may represent a bounding shape (e.g., a bounding box, a bounding triangle, a bounding polygon, a bounding circle, etc.) associated with the image, where the bounding shape includes at least the image portion associated with the type of content. In such examples, the position data 110 may include any type of data indicating the position of the bounding shape within the image. For example, the position data 110 may represent vertex positions associated with the bounding shape, where a vertex position may include a pixel position, coordinate positions (e.g., x-coordinate position, y-coordinate position, etc.), and / or any other type of position associated with the image.For example, if the bounding shape comprises a bounding box, the position data 110 may represent a first pixel and / or coordinate location associated with a first vertex within the image, a second pixel and / or coordinate location associated with a second vertex within the image, a third pixel and / or coordinate location associated with a third vertex within the image, and a fourth pixel and / or coordinate location associated with a fourth vertex within the image.
[0034] For example, the Fig. 3A-3B illustrate an example of boundary shape information associated with a type of content corresponding to image 202, in accordance with some embodiments of the present disclosure. As in the example of Fig. 3A, the inference component 108 may process image data representing the image 202 and, at least based on the processing, generate a boundary shape 302 associated with the license plate 208, which may include the type of content. While the example in Fig. 3A shows the bounding shape 302 as a bounding box, in other examples, the bounding shape 302 may comprise any other type of shape. Furthermore, as in the example of Fig. 3B, generate position data 304 representing the position of the bounding shape 302 within the image 202. For example, the position data 304 may include at least a first position 306(1) associated with a first identifier 308(1) of a first vertex of the bounding shape 302, a second position 306(2) associated with a second identifier 308(2) of a second vertex of the bounding shape 302, a third position 306(3) associated with a third identifier 308(3) of a third vertex of the bounding shape 302, and a fourth position 306(4) associated with a fourth identifier 308(4) of a fourth vertex of the bounding shape 302.
[0035] To return to the example of Fig. 1, in addition to or alternatively to the position data 110 representing the bounding shape, the position data 110 may, in some examples, represent a feature map associated with the image. For example, the inference component 108 may generate the feature map, where the feature map is divided into blocks representing different regions of the image. For example, one or more blocks (e.g., each block) may represent a particular number of pixels associated with the image, such as a 4x4 pixel region, an 8x8 pixel region, a 16x16 pixel region, and / or a different sized region of the image. The inference component 108 may then determine one or more blocks of the feature map associated with the image portion corresponding to the type of content.In some examples, the inference component 108 may make the determination by determining which of the blocks are at least partially contained within the bounding shape. In some examples, the inference component 108 may make the determination by determining which of the blocks contain pixels representing the type of content. In some examples, the inference component 108 may also make the determination using a different technique.
[0036] The inference component 108 may then generate position data 110 representing the block(s) associated with the image patch. In some examples, the inference component 108 may generate the position data 110 to also include a feature map, where one or more blocks associated with the image patch include one or more first values (e.g., 1) and one or more blocks not associated with the image patch include one or more second values (e.g., 0). Additionally or alternatively, in some examples, the inference component 108 may generate the position data 110 to represent one or more identifiers for the block(s) associated with the image patch and / or one or more identifiers for the block(s) not associated with the image patch.While these are just some examples of information that may be used to identify the block(s) associated with the image section, in other examples, the position data 110 may represent additional and / or alternative information that identifies the block(s) associated with the image section.
[0037] For example, the Fig. 4A-4B illustrate an example of generating feature maps associated with image 210 in accordance with some embodiments of the present disclosure. As in the example of Fig. 4A, the inference component 108 may process image data representing the image 210 and, at least based on the processing, generate a first feature map 402 associated with the image 210. As shown, the first feature map 402 may be divided into a number of blocks 404(1)-(64) (also referred to individually as a "block 404" or plurally as "blocks 404"), with each block 404 representing a particular region of the image 210. For example, each block may represent a 4x4 pixel region, an 8x8 pixel region, a 16x16 pixel region, and / or a different sized region of the image 210. The first feature map 402 may further display the blocks 404 associated with faces 216 of the pedestrians 214, which may include the type of content and are indicated by the gray shading.For example, blocks 404(18), 404(26) and 404(27) may be associated with face 216(1) of pedestrian 214(1) and blocks 404(15), 404(22) and 404(23) may be associated with face 216(2) of pedestrian 214(2).
[0038] Next, the inference component 108, as in the example of Fig. 4B, generate a second feature map 406 associated with the image 210. For example, to generate the second feature map 406, the inference component 108 may cause the blocks 404 associated with the faces 216 (e.g., the type of content) to include a first value 408 and the blocks 404 not associated with the faces 216 (e.g., not the type of content) to include a second value 410. In some examples, the first value 408 may include the value 1 and the second value 410 may include the value 0, but in other examples, the first value 408 and / or the second value 410 may include any other value. In some examples, by generating the second feature map 406 that includes only the first value 408 and the second value 410 for the blocks 404, an encoding component can quickly determine which of the blocks 404 should and should not be obscured during encoding, as described in more detail herein.
[0039] Returning to the example of Fig. 1, in some examples, by using the feature maps instead of the bounding shapes to indicate the portions of the images that correspond to the type of content, the inference component 108 may reduce the overall area that is later obscured by the obscuring component 102, described in more detail herein. For example, for an image patch, the inference component 108 may be able to more accurately specify the patch using the feature map instead of the bounding box. For example, if an image shows a license plate oriented at a 45-degree angle in the image, then a bounding shape indicating the patch showing the license plate may include the patch associated with the license plate, as well as one or more additional portions of the image surrounding the license plate, based on the orientation of the bounding shape (e.g.,the boundary shape is oriented at a 0-degree angle with respect to the image). However, a feature map can specify the image section more precisely, as only the blocks in which the license plate is depicted can be identified.
[0040] In some examples, the inference component 108 may further optimize the position data 110 using one or more of the feature maps described herein. For example, the inference component 108 may convert the feature map into a new format, such as a (run, level) format. For example, the inference component 108 may determine a first number of blocks to hide and / or a second number of blocks not to hide. As described herein, the inference component 108 may determine the first number of blocks and / or the second number of blocks using the values associated with the feature map, such that the first number of blocks includes the blocks with a first value (e.g., 1) and the second number of blocks includes the blocks with a second value (e.g., 0).The inference component 108 then groups the first number of blocks into one or more first groups and the second number of blocks into one or more second groups, where the groups are then represented by the position data 110. The position data 110 may, for example, represent values indicative of the groups. In some examples, by executing such features to optimize the feature map, the obfuscation component 102 may obfuscate the blocks in groups, rather than obfuscating the blocks individually.
[0041] As in the example of Fig. 1, the process 100 may include an encoding component 112 that receives the image data 104, the position data 110, and / or the configuration data 114. As described herein, in some examples, the encoding component 112 may include one or more encoders configured to encode the image data 104 to generate encoded image data 116. An encoder may include, but is not limited to, an H.264 encoder, an FFmpeg encoder, a shutter encoder, a MediaCoder, and / or any other type of encoder. Furthermore, the configuration data 114 may be used to configure the encoder to perform at least some of the processes described herein. For example, the configuration data 114 may represent one or more parameters for encoding the image data 104 to obscure the image portion.As described herein, the parameters may include, but are not limited to, a type of content to be obscured (e.g., sensitive information, unimportant information, etc.), a maximum number of parts to be obscured per image (e.g., 1 part, 2 parts, 5 parts, etc.), a quantization parameter (QP) value to be used to obscure the content, and / or any other parameter.
[0042] The encoding component 112 may then use the position data 110 and / or the configuration data 114 to encode the image data 104, with the encoding component 112 further obscuring the image patch while performing the encoding. As described herein, in some examples, the encoding component 112 may obscure the image patch using the QP value represented by the configuration data 114. For example, the configuration data 114 may specify a QP value that meets (e.g., is equal to or greater than) at least a QP threshold, where the larger the QP threshold, the greater the obscuration of the image patch. The larger the QP value during encoding, the more the image patch is compressed, which may also result in a reduction in the quality of the image patch.For example, in some examples, configuration data 114 may specify a maximum QP value associated with the type of encoder used by encoding component 112 to perform encoding. For example, if the encoder uses H.264 encoding, which has a QP value range between 0 and 51, then configuration data 114 may specify a QP value of 51.
[0043] The encoding component 112 may then encode the image data 104 by determining at least a portion of the image data 104 corresponding to the image section to be obscured. In some examples, the portion of the image data 104 may correspond to one or more blocks associated with the image, where the block(s) are associated with values (e.g., pixel values, frequency component values, etc.). The encoding component 112 may then encode the image data 104 such that at least the portion of the image data 104 is encoded using the QP value represented by the configuration data 114. In some examples, the encoding component 112 first generates a quantization matrix associated with the encoding, where the QP values of the quantization matrix associated with the portion of the image data 104 include the adjusted QP value.The coding component 112 then uses the quantization matrix to encode the image data 104. By using the set QP value to encode the part of the image data 104, wherein the set QP value may in turn contain a maximum QP value, the quality associated with the image section may be reduced such that the image section becomes blurred.
[0044] For example, the Fig. 5A-5B illustrate an example of encoding the image 202 using a fixed QP value to obscure an image portion 202 according to some embodiments of the present disclosure. As in the example of Fig. 5A, the coding component 112 may receive at least image data 502 (which may represent and / or include the image data 104) representing the image 202, as well as position data 504 representing the image section 202 to be obscured. For example, the position data 504 may represent at least the boundary shape 302 associated with the image 202, wherein the boundary shape 302 indicates the image section 202 to be obscured.
[0045] The encoding component 112 may then use the position data 504 to generate quantization data 506 for encoding the image data 502. For example, the quantization data 506 may represent a quantization matrix in which the QP values associated with the portion of the image data 502 representing the image section 202 include the adjusted QP value. As described herein, in some examples, the adjusted QP value may include a maximum QP value associated with the encoding component 112. Additionally, other portions of the quantization matrix associated with other portions of the image 202 that are not to be obscured may include different QP values than the adjusted QP values. For example, the other portions of the quantization matrix may include QP values that are smaller than the adjusted QP value.The encoding component 112 may then use at least the quantization data 506 when encoding the image data 502. As shown, the encoding component 112 may generate encoded image data 508 (which may represent and / or include the encoded image data 116) based at least on the performance of the encoding.
[0046] As in the example of Fig. 5B, the encoded image data 508 may represent an image 510 similar to the image 202, at least based on the performance of the encoding, but the image portion 510 associated with the license plate now includes a blurred area 512. By performing the processes described herein, the blurring component 102 is able to blur the image portion 202 associated with the type of content (e.g., the sensitive information) without blurring other portions of the image 202. Furthermore, the blurring component 102 may blur the image portion 202 during a normal encoding process associated with the image 202.
[0047] Returning to the example of Fig. 1, in addition to or alternatively to using the adjusted QP value to obscure the image patch, in other examples, the encoding component 112 may update residual values associated with the portion of the image data 104 to meet (e.g., be less than or equal to) a threshold residual value. In some examples, the encoding component 112 may update the residual values such that the residual values are 0. By updating the residual values associated with the portion of the image data 104 to 0, the encoding component 112 may remove the residual information associated with the portion of the image data 104, which may result in the image patch becoming obscured. In some examples, the encoding component 112 may perform additional processes when updating the residual values, e.g., when the image is associated with a type of image.For example, if the image includes an interframe, the coding component 112 may also restrict the prediction associated with the image to an intraframe prediction. In such an example, the coding component 112 may effect the restriction to prevent other regions of one or more reference images (e.g., one or more reference frames) from being used to reconstruct the image section (e.g., to reconstruct the nature of the content).
[0048] For example, the Fig. 6A-6B illustrate an example of encoding an image (e.g., a frame of a video) by updating residual information to obscure an image portion according to some embodiments of the present disclosure. More specifically, Fig. 6A illustrates an exemplary data flow diagram for a process 600 for obscuring image content by updating residual information, according to some embodiments of the present disclosure. In some examples, the process 600 may be performed by the encoding component 112 of the example of Fig. 1. However, in other examples, at least a portion of process 600 may be performed by one or more additional and / or alternative components. Furthermore, an image associated with process 600 may correspond to an image as described herein.
[0049] As shown, in block B602, process 600 may determine whether an image (e.g., image 210) that may be represented by image data 604 (which may represent and / or include image data 104) includes a portion to be obscured. For example, encoding component 112 may determine whether a portion of the image should be obscured. As described herein, in some examples, encoding component 112 may determine whether there is a portion of the image to be obscured based at least on location data (e.g., location data 110) associated with image data 604.In a first example, the coding component 112 may determine that an image section to be obscured is present at least based on the receipt of the position data associated with the image data 604, or may determine that no image section to be obscured is present at least based on the non-receipt of the position data associated with the image data 604. In a second example, the coding component 112 may determine that an image section to be obscured is present at least based on the position data associated with the image data 604 indicating the part, or may determine that no image section to be obscured is present at least based on the position data associated with the image data 604 indicating no part.
[0050] If it is determined at block B602 that no image portion is to be blurred, the process 600 may proceed to block B606, which is described in more detail below. However, if it is determined at block B602 that an image portion is to be blurred, then the process 600 may determine at block B608 whether the image is associated with an intraframe. As described herein, an intraframe (e.g., i-frame) may include a frame for which encoding is performed using only the data associated with the frame itself. Furthermore, an inter-frame (e.g., p-frame or b-frame) may include a frame for which encoding is performed using the frame itself and at least one other frame, such as an intra-frame or another inter-frame.
[0051] If, at block B608, it is determined that the image is associated with an intra-frame, the process 600 may proceed to block B610, which is described in more detail below. However, if, at block B608, it is determined that the image is not associated with an intra-frame (e.g., the image is associated with an inter-frame), then, at block B612, the process 600 may include triggering an intra-frame prediction mode. As described herein, in some examples, the intra-frame prediction mode may restrict the prediction of the image so that only intra-prediction may be used. For example, if the coding component 112 is performing a prediction for an intermediate image (e.g., a picture), the coding component 112 may perform the prediction using only one or more intermediate images.In such examples, this can ensure that prediction from other regions in the reference image does not contribute to the reconstruction of the obscured content.
[0052] The process 600 may then include, at block B612, updating the residual information associated with the image. During encoding, the encoding component 112 may, for example, generate various types of data, such as prediction data and residual data 614. As described herein, the residual data 614 may represent residual information associated with the image. For example, the residual data 614 may represent coefficient values associated with different parts of the image, such as different blocks (e.g., pixel blocks described herein) associated with the image. As such, updating the residual information may include updating the coefficient values associated with the image portion to be obscured to be less than or equal to a threshold coefficient value. In some examples and as described herein, the threshold coefficient may be 0.By updating the coefficient values to 0 within the part of the residual information associated with the image section to be obscured, it is ensured that the data cannot be reconstructed.
[0053] The process 600 may then include, at block B606, encoding the image data 604. For example, the block B606 may include at least performing a transformation and / or quantization in conjunction with the encoding of the image data 604. Additionally, the process 600 may include, at block B616, performing entropy encoding of the image data 604. Based at least on the performance of the process 600, the encoding component 112 may output encoded image data 618 (which may represent and / or include the encoded image data 116).
[0054] As in the example of Fig. 6B, at least based on the performance of the encoding, the encoded image data 618 may represent an image 620 corresponding to the image 210, but with the portions of the image 620 associated with the pedestrian faces 216 now including obscured regions 622(1)-(2). By performing the processes described herein, the obscuring component 102 is able to obscure the portions of the image 210 associated with the type of content (e.g., the sensitive information) without obscuring other portions of the image 210. Furthermore, the obscuring component 102 may obscure the portions of the image during a normal encoding process associated with the image 210.
[0055] Returning to the example of Fig. 1, the process 100 includes the output of the encoded image data 116 by the obfuscation component 102. For example, if the obfuscation component 102 is connected to a machine, such as the vehicle 1000, the output may include the machine sending the encoded image data 116 to one or more other computing devices, such as the data center 1200. In such examples, the data center 1200 may then perform one or more processes associated with the encoded image data 116, such as storing the encoded image data 116, processing the encoded image data 116, generating a map using the encoded image data 116, training one or more machine learning models using the encoded image data 116, and / or performing another task using the encoded image data 116.
[0056] While the examples herein describe the machine processing the image data 104 using the blurring component 102 to blur portions of the image, in other examples, the machine may process the image data 104 using one or more additional and / or alternative methods. For example, the machine may process the image data 104 using one or more of the image data processing techniques described herein with respect to the vehicle 1000 to locate the machine in an environment, determine locations of objects in the environment, determine paths for the machine to navigate in the environment, and / or perform any other process.In other words, the machine may process the image data 104 using first processes to navigate the environment and may also process the image data 104 using second processes to generate the encoded image data 116 for transmission to one or more computing devices.
[0057] While the example in Fig. 1 shows that the obfuscation component 102 includes the inference component 108. In other examples, the inference component 108 may be separate from the obfuscation component 102. In such examples, the inference component 108 may send the position data 110 to the obfuscation component 102.
[0058] Each block of the methods 700, 800, and 900 described herein comprises a computing process that may be performed using any combination of hardware, firmware, and / or software (see Fig. ). For example, various functions may be performed by a processor executing instructions stored in memory. Methods 700, 800, and 900 may also be stored as computer-usable instructions on computer storage media. Methods 700, 800, and 900 may be provided by a standalone application, a service, or a hosted service (standalone or in combination with another hosted service), or a plug-in for another product, to name a few. Furthermore, methods 700, 800, and 900 are described by way of example with respect to Fig. 1. However, these methods 700, 800, and 900 may additionally or alternatively be performed by any system or combination of systems, including, but not limited to, the systems described herein.
[0059] Fig. 7 is a flowchart illustrating a method 700 for using an encoder to obscure an image portion associated with a type of content, in accordance with some embodiments of the present disclosure. The method 700 may include, at block B702, determining that an image portion is associated with a type of content based at least on image data representative of an image. For example, the obscuring component 102 (e.g., the inference component 108) may determine that the image portion is associated with the type of content, such as sensitive information. As described herein, in some examples, the image data 104 representing the image may be generated using the image sensor(s) 106, such as the image sensor(s) 106 of a machine.
[0060] The method 700 may include, at block B704, generating first data indicative of the image portion. For example, based at least on the determination that the image portion is associated with the type of content, the obscuring component 102 (e.g., the inference component 108) may generate the location data 110 indicative of the image portion. As described herein, in some examples, the location data 110 may represent a bounding shape indicative of at least the image portion. Additionally or alternatively, in some examples, the location data 110 may represent a feature map indicative of at least the image portion. These are just a few examples of information that may be represented by the location data 110. In other examples, the location data 110 may represent any type of information indicative of the image portion associated with the type of content.
[0061] The method 700 may include, at block B706, generating encoded image data using an encoder and at least based on the first data by encoding the image data at least to obscure the image portion. For example, the obscuring component 102 (e.g., the encoding component 112) may encode the portion of the image data 104 to obscure the image portion. As described herein, the obscuring component 102 may use one or more techniques to perform the encoding. In a first example, during encoding, the obscuring component 102 may encode the portion of the image data 104 with a QP value (e.g., a maximum QP value) configured to obscure the content. In a second example, during encoding, the obfuscation component 102 may update residual information associated with the portion of the image data 104, e.g.by updating one or more coefficient values to include 0. In both examples, the blurring component 102 may determine the image portion to be blurred using the position data 110.
[0062] The method 700 may include, at block B708, sending the encoded image data to one or more computing devices. For example, the obfuscation component 102 may send the encoded image data 116 to one or more computing devices, such as a system. For example, if the obfuscation component 102 is connected to a vehicle, such as vehicle 1000, the vehicle may use these processes to obscure sensitive information before uploading the image data 104 to a system.
[0063] Fig. 8 is a flowchart illustrating a method 800 for using an encoder to obscure an image portion based at least on a quantization parameter value, in accordance with some embodiments of the present disclosure. For example, at block B802, the method 800 may include determining a first quantization parameter value associated with obscuring content. For example, the obscuring component 102 may determine the first QP value, such as by receiving configuration data 114 representing the first QP value. As described herein, the first QP value may satisfy (e.g., be equal to or greater than) a QP threshold. For example, the first QP value may include a maximum QP value associated with an encoder.
[0064] The method 800 may include, at block B804, determining, based at least on image data representative of an image, that a portion of the image data is associated with a type of content. For example, the obfuscation component 102 (e.g., the inference component 108) may determine that the portion of the image data 104 is associated with the type of content, such as sensitive information. The obfuscation component 102 may then generate the location data 110 indicative of the image portion. In some examples, the location data 110 may represent a bounding shape indicative of at least the image portion. In some examples, the location data 110 may represent a feature map indicative of at least the image portion.
[0065] The method 800 may include, at block B806, generating encoded image data by encoding at least the first portion of the image data using the first quantization parameter value and a second portion of the image data using a second quantization parameter value. For example, the obfuscation component 102 (e.g., the encoding component 112) may generate the encoded image data 116 by encoding the first portion of the image data 104 using the first QP value and a second portion of the image data 104 using a second QP value. As described herein, in some examples, the second QP value may be less than the first QP value such that the content associated with the first portion of the image data is obscured and the content associated with the second portion of the image data 104 is not obscured.
[0066] Fig. 9 is a flowchart illustrating a method 900 for using an encoder to obscure an image patch based at least on updating residual information, in accordance with some embodiments of the present disclosure. The method 900 may, at block B902, determine, based at least on image data representative of an image, that a portion of the image data is associated with a type of content. For example, the obscuring component 102 (e.g., the inference component 108) may determine that the portion of the image data 104 is associated with the type of content, such as sensitive information. The obscuring component 102 may then generate the position data 110 indicative of the image patch. In some examples, the position data 110 may represent a bounding shape indicative of at least the image patch.In some examples, the position data 110 may represent a feature map that indicates at least the image section.
[0067] The method 900 may include, at block B904, generating residual data associated with the information. For example, the obscuring component 102 (e.g., the encoding component 112) may generate the residual data associated with the image data 104. As described herein, the residual data may represent residual information associated with the image represented by the image data 104. For example, the residual data may represent coefficient values associated with different parts of the image, such as different blocks (e.g., pixel blocks described herein) associated with the image.
[0068] The method 900 may include, at block B906, updating one or more first coefficient values associated with a portion of the residual data associated with the portion of the image data to include one or more second coefficient values. For example, the obscuring component 102 (e.g., the encoding component 112) may update the portion of the residual data associated with the portion of the image data 104. As described herein, the updating may include decreasing the coefficient values associated with the portion of the residual data such that the coefficient values are less than or equal to a threshold coefficient value. In some examples, the threshold coefficient value may include 0.
[0069] The method 900 may include, at block B908, generating encoded image data based at least on the image data and the updated residual data. For example, the obscuring component 102 (e.g., the encoding component 112) may generate the encoded image data 116 by further processing the image data 104 and / or the updated residual data using one or more further encoding processes, such as transformation, quantization, entropy, and / or the like. EXAMPLE AUTONOMOUS VEHICLE
[0070] Fig. 10A is an illustration of an example of an autonomous vehicle 1000, in accordance with some embodiments of the present disclosure. The autonomous vehicle 1000 (alternatively referred to herein as "vehicle 1000") may be, without limitation, a passenger vehicle, such as a car, a truck, a bus, a first responder vehicle, a shuttle, an electric or motorized bicycle, a motorcycle, a fire engine, a police vehicle, an ambulance, a boat, a construction vehicle, an underwater vehicle, a robotic vehicle, a drone, an aircraft, a vehicle coupled to a trailer (e.g., a semi-trailer used to transport goods), and / or another type of vehicle (e.g., an unmanned vehicle and / or a vehicle with one or more passengers).Autonomous vehicles are generally described in terms of automation levels defined by the National Highway Traffic Safety Administration (NHTSA), a division of the U.S. Department of Transportation, and the Society of Automotive Engineers (SAE) "Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles" (Standard No. J3016-201806, published June 15, 2018, Standard No. J3016-201609, published September 30, 2016, and prior and future versions of this standard). The vehicle 1000 may be capable of performing functions in accordance with one or more of the levels 3 through 5 of the autonomous driving levels. The vehicle 1000 may be capable of performing one or more of the levels 1 through 5 of the autonomous driving levels.For example, depending on the embodiment, the vehicle 1000 may provide driver assistance (Level 1), partial automation (Level 2), conditional automation (Level 3), high automation (Level 4), and / or full automation (Level 5). The term "autonomous" as used herein may encompass any and / or all types of autonomy for the vehicle 1000 or other machine, such as fully autonomous, highly autonomous, conditionally autonomous, partially autonomous, assisted autonomy, semi-autonomous, primarily autonomous, or another designation.
[0071] The vehicle 1000 may include components such as a chassis, a vehicle body, wheels (e.g., 2, 4, 6, 8, 18, etc.), tires, axles, and other components of a vehicle. The vehicle 1000 may include a propulsion system 1050, such as an internal combustion engine, a hybrid electric power plant, a pure electric motor, and / or another type of propulsion system. The propulsion system 1050 may be connected to a drivetrain of the vehicle 1000, which may include a transmission to enable propulsion of the vehicle 1000. The propulsion system 1050 may be controlled in response to receiving signals from the throttle / accelerator pedal 1052.
[0072] A steering system 1054, which may include a steering wheel, may be used to steer the vehicle 1000 (e.g., along a desired path or route) when the propulsion system 1050 is operating (e.g., when the vehicle is moving). The steering system 1054 may receive signals from a steering actuator 1056. The steering wheel may optionally be used for full automation (Level 5).
[0073] The brake sensor system 1046 may be used to apply the vehicle brakes in response to receiving signals from the brake actuators 1048 and / or brake sensors.
[0074] Controller 1036, which controls one or more System-on-Chips (SoCs) 1004 ( Fig. 10C) and / or GPU(s), may send signals (e.g., representative of commands) to one or more components and / or systems of the vehicle 1000. For example, the controller(s) may send signals to actuate the vehicle brakes via one or more brake actuators 1048, to actuate the steering system 1054 via one or more steering actuators 1056, and to actuate the propulsion system 1050 via one or more throttle / accelerator pedals 1052. The controller(s) 1036 may include one or more built-in (e.g., integrated) computing devices (e.g., supercomputers) that process sensor signals and issue operational commands (e.g., signals representing commands) to enable autonomous driving and / or to assist a human driver in driving the vehicle 1000.The controller(s) 1036 may include a first controller 1036 for autonomous driving functions, a second controller 1036 for functional safety functions, a third controller 1036 for artificial intelligence functions (e.g., computer vision), a fourth controller 1036 for infotainment functions, a fifth controller 1036 for redundancy under emergency conditions, and / or other controllers. In some examples, a single controller 1036 may perform two or more of the above functions, two or more controllers 1036 may perform a single function, and / or any combination thereof.
[0075] The controller(s) 1036 may provide the signals to control one or more components and / or systems of the vehicle 1000 in response to sensor data received from one or more sensors (e.g., sensor inputs). The sensor data may, for example and without limitation, be obtained from Global Navigation Satellite System (“GNSS”) sensor(s) 1058 (e.g., Global Positioning System sensor(s)), RADAR sensor(s) 1060, ultrasonic sensor(s) 1062, LIDAR sensor(s) 1064, Inertial Measurement Unit (IMU) sensor(s) 1066 (e.g., accelerometer(s), gyroscope(s), magnetic compass(es), magnetometer(s), etc.), microphone(s) 1096, stereo camera(s) 1068, wide-angle camera(s) 1070 (e.g., fisheye cameras), infrared camera(s) 1072, environmental camera(s) 1074 (e.g., 360-degree cameras), long-range and / or medium-range camera(s) 1098, speed sensor(s) 1044 (e.g.for measuring the speed of the vehicle 1000), vibration sensor(s) 1042, steering sensor(s) 1040, brake sensor(s) (e.g., as part of the brake sensor system 1046), and / or other sensor types.
[0076] One or more controllers 1036 may receive inputs (e.g., in the form of input data) from an instrument cluster 1032 of the vehicle 1000 and provide outputs (e.g., in the form of output data, display data, etc.) via a human-machine interface (HMI) display 1034, an audible annunciator, a speaker, and / or via other components of the vehicle 1000. The outputs may include information such as vehicle speed, RPM, time, map data (e.g., the high-definition ("HD") map 1022 of Fig. 10C), location data (e.g., the location of the vehicle 1000, e.g., on a map), direction, location of other vehicles (e.g., an occupancy grid), information about objects and the status of objects as perceived by the controller(s) 1036, etc. For example, the HMI display 1034 may display information about the presence of one or more objects (e.g., a road sign, a warning sign, a changing traffic light, etc.) and / or information about maneuvers the vehicle has performed, is currently performing, or will perform (e.g., change lanes now, take exit 34B in two miles, etc.).
[0077] The vehicle 1000 further includes a network interface 1024 that may utilize one or more wireless antenna(s) 1026 and / or modem(s) to communicate over one or more networks. For example, the network interface 1024 may be capable of communicating over Long-Term Evolution ("LTE"), Wideband Code Division Multiple Access ("WCDMA"), Universal Mobile Telecommunications System ("UMTS"), Global System for Mobile Communication ("GSM"), IMT-CDMA Multi-Carrier ("CDMA2000"), etc. The wireless antenna(s) 1026 may also enable communication between objects in the environment (e.g., vehicles, mobile devices, etc.) using local area networks such as Bluetooth, Bluetooth Low Energy ("LE"), Z-Wave, ZigBee, etc., and / or low-power wide area networks ("LPWANs") such as LoRaWAN, SigFox, etc.
[0078] Fig. 10B is an example of camera positions and fields of view for the autonomous vehicle 1000 of Fig. 10A, in accordance with some embodiments of the present disclosure. The cameras and respective fields of view are illustrative and not limiting. For example, additional and / or alternative cameras may be included and / or the cameras may be located at various locations on the vehicle 1000.
[0079] The camera types for the cameras may include, but are not limited to, digital cameras that can be adapted for use with the components and / or systems of the vehicle 1000. The camera(s) may operate at Security Level B (ASIL) and / or another ASIL. The camera types may have any frame rate, e.g., 60 frames per second (fps), 120 fps, 240 fps, etc., depending on the embodiment. The cameras may use rolling shutter, global shutter, another shutter type, or a combination thereof. In some examples, the color filter array may include a red-clear-clear-clear color filter array (RCCC), a red-clear-clear-blue color filter array (RCCB), a red-blue-green-clear color filter array (RBGC), a Foveon X3 color filter array, a Bayer sensor color filter array (RGGB), a monochrome sensor color filter array, and / or another type of color filter array.In some embodiments, cameras with clear pixels, such as cameras with an RCCC, an RCCB, and / or an RBGC color filter array, may be used to increase light sensitivity.
[0080] In some examples, one or more of the cameras can be used to perform advanced driver assistance systems (ADAS) (e.g., as part of a redundant or fail-safe design). For example, a multifunction mono camera can be installed to provide features such as lane departure warning, traffic sign assist, and intelligent headlight control. One or more of the cameras (e.g., all cameras) can simultaneously record and provide image data (e.g., video).
[0081] One or more of the cameras may be mounted in a mounting fixture, such as a custom-designed (three-dimensionally ("3D") printed) fixture, to eliminate stray light and reflections from inside the vehicle (e.g., reflections from the dashboard reflected in the windshield mirrors) that can interfere with the camera's image data acquisition. Regarding the mounting of the exterior mirrors, the exterior mirrors may be custom-3D printed so that the camera mounting plate conforms to the shape of the exterior mirror. In some examples, the cameras may be integrated into the exterior mirror. For side-mounted cameras, the cameras may also be integrated into the four pillars at each corner of the cabin.
[0082] Cameras with a field of view that includes portions of the environment in front of the vehicle 1000 (e.g., forward-facing cameras) can be used for the surrounding view to help identify forward paths and obstacles, as well as, with the assistance of one or more controllers 1036 and / or control SoCs, to provide information critical for establishing an occupancy grid and / or determining preferred vehicle paths. Forward-facing cameras can be used to perform many of the same ADAS functions as LIDAR, including emergency braking, pedestrian detection, and collision avoidance. Forward-facing cameras can also be used for ADAS features and systems such as lane departure warnings ("LDW"), autonomous cruise control ("ACC"), and / or other features such as traffic sign recognition.
[0083] A variety of cameras may be used in a forward-facing configuration, such as a monocular camera platform that includes a CMOS color imager. Another example is one or more wide-angle cameras 1070, which may be used to detect objects that enter the field of view from the periphery (e.g., pedestrians, crossing traffic, or bicycles). Although in Fig. 10B illustrates only one wide-angle camera, the vehicle 1000 may be equipped with any number (including zero) of wide-angle cameras 1070. Furthermore, any number of wide-angle cameras 1098 (e.g., a wide-angle stereo camera pair) may be used for depth-based object detection, particularly for objects for which a neural network has not yet been trained. The long-range camera(s) 1098 may also be used for object detection and classification, as well as basic object tracking.
[0084] Any number of stereo cameras 1068 may also be included in a forward-facing configuration. In at least one embodiment, one or more of the stereo cameras 1068 may include an integrated control unit comprising a scalable processing unit capable of providing a field-programmable gate array (“FPGA”) and a multi-core microprocessor with an integrated controller area network (“CAN”) or Ethernet interface on a single chip. Such a unit may be used to create a 3D map of the vehicle's surroundings, including a distance estimate for all points in the image. An alternative stereo camera(s) 1068 may include a compact stereo vision sensor(s) including two camera lenses (one each on the left and right) and an image processing chip that measures the distance between the vehicle and the target object and processes the generated information (e.g.,Metadata) is used to activate the autonomous emergency braking and lane departure warning functions. In addition to or as an alternative to the stereo cameras described here, other types of stereo cameras 1068 may also be used.
[0085] Cameras with a field of view that includes parts of the environment to the side of the vehicle 1000 (e.g., side cameras) can be used for the environmental view and provide information used to create and update the occupancy grid and to generate side impact warnings. For example, the environmental camera(s) 1074 (e.g., four environmental cameras 1074, as in Fig. 10B) may be positioned on the vehicle 1000. The surround camera(s) 1074 may include wide-angle camera(s) 1070, fisheye camera(s), 360-degree camera(s), and / or the like. For example, four fisheye cameras may be mounted on the front, rear, and sides of the vehicle. In an alternative arrangement, the vehicle may utilize three surround cameras 1074 (e.g., left, right, and rear) and utilize one or more other cameras (e.g., a forward-facing camera) as a fourth surround camera.
[0086] Cameras with a field of view that includes portions of the environment behind the vehicle 1000 (e.g., rearview cameras) may be used for parking assistance, surrounding view, rear collision warnings, and creating and updating the occupancy grid. A variety of cameras may be used, including, but not limited to, cameras that are also suitable as forward-facing cameras (e.g., long-range and / or mid-range camera(s) 1098, stereo camera(s) 1068, infrared camera(s) 1072, etc.), as described herein.
[0087] Fig. 10C is a block diagram of an example system architecture for the example autonomous vehicle 1000 of Fig. 10A, in accordance with some embodiments of the present disclosure. It should be understood that these and other arrangements described herein are presented only as examples. Other arrangements and elements (e.g., engines, interfaces, functions, arrangements, groupings of functions, etc.) may be used in addition to or in place of those illustrated, and some elements may be omitted entirely. Furthermore, many of the elements described herein are functional units that may be implemented as individual or distributed components, or in conjunction with other components, and in any suitable combination and location. Various functions performed by units described herein may be performed by hardware, firmware, and / or software. For example, various functions may be performed by a processor executing instructions stored in memory.
[0088] All components, features and systems of the vehicle 1000 in Fig. 10C are shown connected via bus 1002. Bus 1002 may include a Controller Area Network (CAN) data interface (alternatively referred to herein as a "CAN bus"). A CAN bus may be a network within vehicle 1000 used to support the control of various features and functions of vehicle 1000, such as brake application, acceleration, braking, steering, windshield wipers, etc. A CAN bus may be configured to have dozens or even hundreds of nodes, each with its own unique identifier (e.g., a CAN ID). The CAN bus may be read to determine steering wheel angle, vehicle speed, engine speed (RPM), button positions, and / or other vehicle status indicators. The CAN bus may be ASIL B compliant.
[0089] Although bus 1002 is described herein as a CAN bus, this is not a limitation. For example, FlexRay and / or Ethernet may be used in addition to or as an alternative to the CAN bus. Although a single wire is used to represent bus 1002, this is not a limitation. For example, there may be any number of buses 1002, which may include one or more CAN buses, one or more FlexRay buses, one or more Ethernet buses, and / or one or more other types of buses using a different protocol. In some examples, two or more buses 1002 may be used to perform different functions and / or may be used for redundancy. For example, a first bus 1002 may be used for collision avoidance functionality and a second bus 1002 may be used for actuation control.In each example, each bus 1002 may communicate with any component of the vehicle 1000, and two or more buses 1002 may communicate with the same components. In some examples, each SoC 1004, each controller 1036, and / or each computer within the vehicle may have access to the same input data (e.g., inputs from sensors of the vehicle 1000) and be connected to a common bus, such as the CAN bus.
[0090] The vehicle 1000 may include one or more controllers 1036 as described herein with respect to Fig. 10A. The controller(s) 1036 may be used for a variety of functions. The controller(s) 1036 may be coupled to the various other components and systems of the vehicle 1000 and may be used for controlling the vehicle 1000, the artificial intelligence of the vehicle 1000, the infotainment for the vehicle 1000, and / or the like.
[0091] The vehicle 1000 may include one or more systems on a chip (SoC) 1004. The SoC 1004 may include CPU(s) 1006, GPU(s) 1008, processor(s) 1010, cache(s) 1012, accelerators 1014, data storage 1016, and / or other components and features not shown. The SoC(s) 1004 may be used to control the vehicle 1000 in a variety of platforms and systems. For example, the SoC(s) 1004 may be combined in a system (e.g., the system of the vehicle 1000) with an HD map 1022 that may receive map refreshes and / or updates via a network interface 1024 from one or more servers (e.g., server(s) 1078 of Fig. 10D).
[0092] The CPU(s) 1006 may comprise a CPU cluster or CPU complex (also referred to herein as a "CCPLEX"). The CPU(s) 1006 may include multiple cores and / or L2 caches. For example, in some embodiments, the CPU(s) 1006 may comprise eight cores in a coherent multiprocessor configuration. In some embodiments, the CPU(s) 1006 may comprise four dual-core clusters, each cluster having a dedicated L2 cache (e.g., a 2 MB L2 cache). The CPU(s) 1006 (e.g., the CCPLEX) may be configured to support concurrent cluster operation, so that any combination of the clusters of the CPU(s) 1006 may be active at any time.
[0093] The CPU(s) 1006 may implement power management features that include one or more of the following: individual hardware blocks may be automatically clocked when idle to conserve dynamic power; each core clock may be controlled when the core is not actively executing instructions due to the execution of WFI / WFE instructions; each core may be independently power controlled; each core cluster may be independently clocked when all cores are clocked or power controlled; and / or each core cluster may be independently power controlled when all cores are power controlled. The CPU(s) 1006 may also implement an enhanced power state management algorithm in which allowable power states and expected wake-up times are established, and the hardware / microcode determines the best power state for the core, cluster, and CCPLEX.The processor cores can support simplified sequences for entering the power state in software, offloading the work to the microcode.
[0094] The graphics processor(s) 1008 may include an integrated graphics processor (alternatively referred to herein as an "iGPU"). The GPU(s) 1008 may be programmable and efficient for parallel workloads. The GPU(s) 1008 may, in some examples, utilize an extended Tensor instruction set. The GPU(s) 1008 may include one or more streaming microprocessors, where each streaming microprocessor may include an L1 cache (e.g., an L1 cache with at least 96 KB of memory capacity) and two or more of the streaming microprocessors may share an L2 cache (e.g., an L2 cache with 512 KB of memory capacity). In some embodiments, the GPU(s) 1008 may include at least eight streaming microprocessors. The GPU(s) 1008 may use application programming interface(s) (API(s)) for computations.In addition, the GPU(s) 1008 may utilize one or more parallel computing platforms and / or programming models (e.g., NVIDIA's CUDA).
[0095] The graphics processor(s) 1008 may be optimized for best performance in automotive and embedded use cases. For example, the GPU(s) 1008 may be fabricated on a fin field-effect transistor (FinFET). However, this is not a limitation, and the GPU(s) 1008 may also be fabricated using other semiconductor fabrication techniques. Each streaming microprocessor may include a number of mixed-precision compute cores divided into multiple blocks. For example, 64 PF32 cores and 32 PF64 cores may be divided into four processing blocks, but are not limited to this. In such an example, each processing block can be assigned 16 FP32 cores, 8 FP64 cores, 16 INT32 cores, two NVIDIA TENSOR COREs with mixed precision for deep learning matrix arithmetic, an L0 instruction cache, a warp scheduler, a dispatch unit, and / or a 64 KB register file.Additionally, streaming microprocessors can include independent parallel integer and floating-point datapaths to enable efficient execution of workloads with a mix of computations and addressing calculations. Streaming microprocessors can include independent thread scheduling to enable finer synchronization and collaboration between parallel threads. Streaming microprocessors can include a combined L1 data cache and shared memory unit to improve performance while simplifying programming.
[0096] The graphics processor(s) 1008 may include high-bandwidth memory (HBM) and / or a 16 GB HBM2 memory subsystem to provide, in some examples, a peak memory bandwidth of approximately 900 GB / second. In some examples, synchronous graphics random-access memory (SGRAM), such as synchronous graphics double-data-rate random-access memory type 5 (GDDR5), may be used in addition to or as an alternative to HBM memory.
[0097] The GPU(s) 1008 may include unified memory technology with access counters to enable more accurate migration of memory pages to the processor that accesses them most frequently, thereby improving the efficiency of memory regions shared between the processors. In some examples, address translation services (ATS) support may be used to allow the GPU(s) 1008 to directly access the page tables of the CPU(s) 1006. In such examples, when the memory management unit (MMU) of the GPU(s) 1008 detects a fault, an address translation request may be transmitted to the CPU(s) 1006. In response, the CPU(s) 1006 may look up the virtual-physical mapping for the address in its page tables and transmit the translation back to the GPU(s) 1008.For example, unified memory technology can enable a single, unified virtual address space for the memory of both the CPU(s) 1006 and the GPU(s) 1008, thereby simplifying the programming of the GPU(s) 1008 and the porting of applications to the GPU(s) 1008.
[0098] Additionally, the GPU(s) 1008 may include an access counter that can track the frequency of access by the GPU(s) 1008 to the memory of other processors. The access counter can help ensure that memory pages are moved to the physical memory of the processor that accesses the pages most frequently.
[0099] The SoC(s) 1004 may include any number of cache(s) 1012, including those described herein. For example, the cache(s) 1012 may include an L3 cache available to both the CPU(s) 1006 and the GPU(s) 1008 (e.g., connected to both the CPU(s) 1006 and the GPU(s) 1008). The cache(s) 1012 may include a write-back cache that can track the state of the lines, e.g., by using a cache coherence protocol (e.g., MEI, MESI, MSI, etc.). The L3 cache may be 4 MB or more, depending on the embodiment, although smaller cache sizes may be used.
[0100] The SoC(s) 1004 may include an arithmetic logic unit(s) (ALU(s)) that may be utilized in performing processing related to any variety of tasks or operations of the vehicle 1000, such as processing DNNs. Additionally, the SoC(s) 1004 may include a floating-point unit(s) (FPU(s))—or other mathematical coprocessors or numerical coprocessor types—for performing mathematical operations within the system. For example, the SoC(s) 1004 may include one or more FPUs integrated as execution units within a CPU(s) 1006 and / or GPU(s) 1008.
[0101] The SoC(s) 1004 may include one or more accelerators 1014 (e.g., hardware accelerators, software accelerators, or a combination thereof). For example, the SoC(s) 1004 may include a hardware acceleration cluster, which may include optimized hardware accelerators and / or large on-chip memory. The large on-chip memory (e.g., 4 MB SRAM) may enable the hardware acceleration cluster to accelerate neural networks and other computations. The hardware acceleration cluster may be used to complement the GPU(s) 1008 and offload some tasks from the GPU(s) 1008 (e.g., to free up more cycles of the GPU(s) 1008 to perform other tasks). For example, the accelerator(s) 1014 may be used for targeted workloads (e.g., perception, convolutional neural networks (CNNs), etc.) that are stable enough to be accelerated.The term “CNN” as used here can encompass all types of CNNs, including region-based or regional convolutional neural networks (RCNNs) and fast RCNNs (e.g., for object detection).
[0102] The accelerator(s) 1014 (e.g., the hardware acceleration cluster) may include a deep learning accelerator (DLA). The DLA(s) may include one or more tensor processing units (TPUs), which may be configured to provide an additional tens of trillion operations per second for deep learning applications and inferencing. The TPUs may be accelerators configured and optimized to perform image processing functions (e.g., for CNNs, RCNNs, etc.). The DLA(s) may further be optimized for a specific set of neural network types and floating-point operations, as well as for inferencing. The design of the DLA(s) may provide more performance per millimeter than a general-purpose GPU, far exceeding the performance of a CPU. The TPU(s) may perform multiple functions, including a single-instance convolution function, for example,INT8, INT16 and FP16 data types are supported for both features and weights as well as post-processor functions.
[0103] The DLA(s) can quickly and efficiently execute neural networks, in particular CNNs, on processed or unprocessed data for a variety of functions, including, for example and without limitation: a CNN for object identification and recognition using data from camera sensors; a CNN for distance estimation using data from camera sensors; a CNN for emergency vehicle detection and identification and recognition using data from microphones; a CNN for facial recognition and vehicle owner identification using data from camera sensors; and / or a CNN for safety-related events.
[0104] The DLA(s) can perform any function of the GPU(s) 1008, and by using an inference accelerator, for example, a developer can use either the DLA(s) or the GPU(s) 1008 for any function. For example, the developer can concentrate the processing of CNNs and floating-point operations on the DLA(s) and leave other functions to the GPU(s) 1008 and / or other accelerators 1014.
[0105] The accelerator(s) 1014 (e.g., the hardware acceleration cluster) may include a programmable image processing accelerator (PVA), which may alternatively be referred to herein as a computer vision accelerator. The PVA(s) may be designed and configured to accelerate computer vision algorithms for advanced driver assistance systems (ADAS), autonomous driving, and / or augmented reality (AR) and / or virtual reality (VR) applications. The PVA(s) may provide a balance between performance and flexibility. For example, and without limitation, each PVA may include any number of reduced instruction set computers (RISC) cores, direct memory access (DMA), and / or any number of vector processors.
[0106] The RISC cores may interact with image sensors (e.g., the image sensors of one of the cameras described herein), image signal processors, and / or the like. Each of the RISC cores may include any amount of memory. The RISC cores may use any number of protocols, depending on the embodiment. In some examples, the RISC cores may execute a real-time operating system (RTOS). The RISC cores may be implemented with one or more integrated circuits, application-specific integrated circuits (ASICs), and / or memory devices. For example, the RISC cores may include an instruction cache and / or tightly coupled RAM.
[0107] The DMA may enable components of the PVA(s) to access system memory independently of the CPU(s) 1006. The DMA may support any number of features used to optimize the PVA, including, but not limited to, support for multi-dimensional addressing and / or circular addressing. In some examples, the DMA may support up to six or more dimensions of addressing, which may include block width, block height, block depth, horizontal block gradation, vertical block gradation, and / or depth gradation.
[0108] The vector processors may be programmable processors that can be designed to efficiently and flexibly execute the programming of computer vision algorithms and provide signal processing functions. In some examples, the PVA may include a PVA core and two vector processing subsystem partitions. The PVA core may include a processor subsystem, DMA engine(s) (e.g., two DMA engines), and / or other peripherals. The vector processing subsystem may act as the primary processing unit of the PVA and may include a vector processing unit (VPU), an instruction cache, and / or a vector memory (e.g., VMEM). A VPU core may include a digital signal processor, such as a single instruction multiple data (SIMD) and very long instruction word (VLIW) digital signal processor. The combination of SIMD and VLIW can increase throughput and speed.
[0109] Each of the vector processors may include an instruction cache and be connected to dedicated memory. Therefore, in some examples, each of the vector processors may be configured to operate independently of the other vector processors. In other examples, the vector processors included in a particular PVA may be configured to use data parallelism. Thus, in some embodiments, the majority of vector processors included in a single PVA may execute the same computer vision algorithm, but for different regions of an image. In other examples, the vector processors included in a particular PVA may concurrently execute different computer vision algorithms for the same image, or even different algorithms for consecutive images or portions of an image.Among other things, any number of PVAs can be included in the hardware acceleration cluster, and any number of vector processors can be included in each of the PVAs. Furthermore, the PVA(s) can contain additional ECC (Error Correcting Code) memory to increase overall system security.
[0110] The accelerator(s) 1014 (e.g., the hardware acceleration cluster) may include an on-chip computer vision network and SRAM to provide high-bandwidth, low-latency SRAM to the accelerator(s) 1014. In some examples, the on-chip memory may include at least 4 MB of SRAM, consisting of, for example, and without limitation, eight field-configurable memory blocks accessible by both the PVA and the DLA. Each pair of memory blocks may include an Advanced Peripheral Bus (APB) interface, configuration circuitry, a controller, and a multiplexer. Any type of memory may be used. The PVA and the DLA may access the memory through a backbone that provides the PVA and the DLA with high-speed access to the memory. The backbone may include an on-chip computer vision network that connects the PVA and the DLA to the memory (e.g.,using the APB).
[0111] The on-chip computer vision network can include an interface that verifies that both the PVA and the DLA are delivering ready and valid signals before transmitting control signals / addresses / data. Such an interface can provide separate phases and channels for transmitting control signals / addresses / data, as well as burst communication for continuous data transmission. This type of interface can comply with ISO 26262 or IEC 61508, but other standards and protocols can also be used.
[0112] In some examples, the SoC(s) 1004 may include a real-time ray tracing hardware accelerator, as described in U.S. Patent Application No. 16 / 101,232, filed August 10, 2018. The real-time ray tracing hardware accelerator may be used to quickly and efficiently determine the positions and extents of objects (e.g., within a world model) to generate real-time visualization simulations, for radar signal interpretation, for sound propagation synthesis and / or analysis, for simulation of sonar systems, for general wave propagation simulation, for comparison with lidar data for localization, and / or for other functions, and / or for other purposes. In some embodiments, one or more tree traversal units (TTUs) may be used to perform one or more operations related to ray tracing.
[0113] The accelerator(s) 1014 (e.g., the hardware accelerator cluster) can be used in a variety of ways for autonomous driving. The PVA can be a programmable image processing accelerator that can be used for critical processing steps in ADAS and autonomous vehicles. The capabilities of the PVA are well suited to algorithmic areas that require predictable processing with low power and low latency. In other words, the PVA is well suited for semi-dense or dense regular computations, even on small datasets, that require predictable runtimes with low latency and low power. In the context of autonomous vehicle platforms, the PVAs are therefore designed to execute classical computer vision algorithms because of their efficiency in object detection and processing integer mathematical data.
[0114] For example, in one embodiment of the technology, the PVA is used to perform computer stereo vision. In some examples, a semi-global adaptation-based algorithm may be used, although this is not intended as a limitation. Many Level 3-5 autonomous driving applications require motion estimation / stereo matching while driving (e.g., structure from motion, pedestrian detection, lane detection, etc.). The PVA can perform a computer stereo vision function for inputs from two monocular cameras.
[0115] In some examples, PVA may be used to perform dense optical flow, such as processing raw radar data (e.g., using a 4D Fast Fourier Transform) to obtain processed radar data. In other examples, PVA is used for processing time-of-flight depth data, such as processing raw time-of-flight data to obtain processed time-of-flight data.
[0116] Any type of network can be powered by the DLA to improve control and driving safety, such as a neural network that outputs a confidence score for each object detection. Such a confidence score can be interpreted as a probability or as the relative "weight" of each detection compared to other detections. This confidence score allows the system to make further decisions about which detections should be considered true positives and which should be considered false positives. For example, the system can set a confidence threshold and only consider detections that exceed this threshold as true positives. In an automatic emergency braking (AEB) system, false positives would cause the vehicle to automatically perform emergency braking, which is obviously undesirable.Therefore, only the most confident detections should be considered as triggers for AEB. The DLA can employ a neural network to regress the confidence score. The neural network can take as input at least a subset of parameters, such as the dimensions of the bounding box, the ground plane estimate obtained (e.g., from another subsystem), the output of the IMU sensor 1066 correlated with the orientation of the vehicle 1000, the distance, the 3D pose estimate of the object obtained from the neural network and / or other sensors (e.g., LIDAR sensor(s) 1064 or RADAR sensor(s) 1060), and others.
[0117] The SoC(s) 1004 may include data storage 1016 (e.g., memory). The data storage(s) 1016 may be on-chip memory of the SoC(s) 1004 that may store neural networks to be executed on the GPU and / or the DLA. In some examples, the capacity of the data storage(s) 1016 may be large enough to store multiple instances of neural networks for redundancy and security. The data storage(s) 1012 may include L2 or L3 cache(s) 1012. The reference to the data storage(s) 1016 may also include a reference to the memory associated with the PVA, DLA, and / or other accelerators 1014, as described herein.
[0118] The SoC(s) 1004 may include one or more processor(s) 1010 (e.g., embedded processors). The processor(s) 1010 may include a boot and power management processor, which may be a dedicated processor and subsystem to handle boot power and management functions and associated security enforcement. The boot and power management processor may be part of the boot sequence of the SoC(s) 1004 and may provide runtime power management services. The boot power and management processor may provide clock and voltage programming, support for low-power state transitions, management of SoC(s) 1004 temperatures and temperature sensors, and / or management of the SoC(s) 1004 power states.Each temperature sensor may be implemented as a ring oscillator whose output frequency is proportional to temperature, and the SoC(s) 1004 may use the ring oscillators to sense the temperatures of the CPU(s) 1006, GPU(s) 1008, and / or accelerator 1014. If temperatures are determined to exceed a threshold, the boot and power management processor may enter a temperature fault routine and place the SoC(s) 1004 into a lower power state and / or place the vehicle 1000 into a chauffeur-to-safe-stop mode (e.g., bring the vehicle 1000 to a safe stop).
[0119] The processor(s) 1010 may also include a number of embedded processors that can serve as an audio processing module. The audio processing engine may be an audio subsystem that enables full hardware support for multi-channel audio across multiple interfaces, as well as a wide and flexible range of audio I / O interfaces. In some examples, the audio processing engine is a dedicated processor core with a digital signal processor with dedicated RAM.
[0120] The processor(s) 1010 may also include an always-on processor engine that provides the necessary hardware features to support low-power sensor management and wake-up use cases. The always-on processor engine may include a processor core, tightly coupled memory, supporting peripherals (e.g., timers and interrupt controllers), various I / O controller peripherals, and routing logic.
[0121] The processor(s) 1010 may also include a security cluster engine, which includes a dedicated processor subsystem for security management of automotive applications. The security cluster engine may include two or more processor cores, tightly coupled memory, supporting peripherals (e.g., timers, an interrupt controller, etc.), and / or routing logic. In a security mode, the two or more cores may operate in a lockstep mode, functioning as a single core with comparison logic to detect differences between their operations.
[0122] The processor(s) 1010 may also include a real-time camera engine, which may include a dedicated processor subsystem for managing the real-time camera.
[0123] The processor(s) 1010 may also include a high dynamic range signal processor, which may include an image signal processor that is a hardware engine that is part of the camera processing pipeline.
[0124] The processor(s) 1010 may include a video image compositor, which may be a processing block (e.g., implemented on a microprocessor) that implements video post-processing functions required by a video playback application to generate the final image for the player window. The video image compositor may perform lens distortion correction on the wide-angle camera(s) 1070, the surround camera(s) 1074, and / or the in-cabin surveillance camera sensors. The in-cabin surveillance camera sensor is preferably monitored by a neural network running on another instance of the Advanced SoC and configured to detect and respond to events in the cabin.An in-cabin system can perform lip reading to activate cellular service and place a call, dictate emails, change the vehicle's destination, activate or change the vehicle's infotainment system and settings, or enable voice-activated web browsing. Certain features are available to the driver only when the vehicle is operating in autonomous mode and are disabled otherwise.
[0125] The video image compositor can incorporate enhanced temporal noise reduction for both spatial and temporal noise reduction. For example, if motion occurs in a video, the noise reduction weights the spatial information accordingly, reducing the weight of information coming from neighboring frames. If an image or a portion of an image contains no motion, the temporal noise reduction performed by the video compositor can use information from the previous frame to reduce noise in the current frame.
[0126] The video compositor may also be configured to perform stereo image distortion correction on input stereo image frames. The video compositor may also be used for user interface design when the operating system desktop is in use and the GPU(s) 1008 do not need to constantly render new surfaces. Even when the graphics processor(s) 1008 are turned on and actively performing 3D rendering, the video compositor may be used to offload the graphics processor(s) 1008, thereby improving performance and responsiveness.
[0127] The SoC(s) 1004 may also include a MIPI serial camera interface for receiving video and inputs from cameras, a high-speed interface, and / or a video input block that may be used for camera and related pixel input functions. The SoC(s) 1004 may also include one or more input / output controllers that may be controlled by software and used to receive I / O signals that are not associated with a specific role.
[0128] The SoC(s) 1004 may also include a wide range of peripheral interfaces to enable communication with peripheral devices, audio codecs, power management, and / or other devices. The SoC(s) 1004 may be used to process data from cameras (e.g., connected via Gigabit Multimedia Serial Link and Ethernet), sensors (e.g., LIDAR sensor(s) 1064, RADAR sensor(s) 1060, etc., which may be connected via Ethernet), data from bus 1002 (e.g., speed of vehicle 1000, steering wheel position, etc.), data from GNSS sensor(s) 1058 (e.g., connected via Ethernet or CAN bus). The SoC(s) 1004 may also include dedicated high-performance mass storage controllers, which may include their own DMA engines and which may be used to offload routine data management tasks from the CPU(s) 1006.
[0129] The SoC(s) 1004 may be an end-to-end platform with a flexible architecture covering automation levels 3 through 5, thereby providing a comprehensive functional safety architecture, leveraging computer vision and ADAS techniques for diversity and redundancy, and providing a platform for a flexible, reliable driving software stack along with deep learning tools. The SoC(s) 1004 may be faster, more reliable, and even more power and space efficient than conventional systems. For example, the accelerator(s) 1014, in combination with the CPU(s) 1006, the GPU(s) 1008, and the data memory(s) 1016, may form a fast, efficient platform for Level 3-5 autonomous vehicles.
[0130] The technology thus offers capabilities and functions that cannot be achieved with conventional systems. For example, computer vision algorithms can be executed on CPUs, which can be configured using high-level languages such as the C programming language to execute a variety of processing algorithms on a wide variety of visual data. However, CPUs are often unable to meet the performance requirements of many image processing applications, such as execution time and power consumption. In particular, many CPUs are unable to execute complex object detection algorithms in real time, which is a prerequisite for in-vehicle ADAS applications and a requirement for practical Level 3-5 autonomous vehicles.
[0131] In contrast to conventional systems, the technology described here, by providing a CPU complex, a GPU complex, and a hardware acceleration cluster, enables multiple neural networks to be executed simultaneously and / or sequentially, and the results to be combined to enable Level 3-5 autonomous driving functions. For example, a CNN running on the DLA or dGPU (e.g., GPU(s) 1020) may include text and word recognition, allowing the supercomputer to read and understand traffic signs, even those for which the neural network has not been specifically trained. The DLA may further include a neural network capable of identifying, interpreting, and semantically understanding the traffic sign and passing this semantic understanding to the path planning modules running on the CPU complex.
[0132] Another example is that multiple neural networks can operate simultaneously, as required for Level 3, 4, or 5 driving. For example, a warning sign reading "Caution: Flashing lights indicate black ice" along with an electric light can be interpreted independently or jointly by multiple neural networks. The sign itself can be identified as a traffic sign by a first deployed neural network (e.g., a trained neural network), and the text "Flashing lights indicate black ice" can be interpreted by a second deployed neural network, which informs the vehicle's path-planning software (preferably running on the CPU complex) that black ice is present upon detection of flashing lights.The flashing light can be identified by running a third neural network over multiple frames, which informs the vehicle's path planning software of the presence (or absence) of flashing lights. All three neural networks can run simultaneously, e.g., within the DLA and / or on the GPU(s) 1008.
[0133] In some examples, a facial recognition and vehicle owner identification CNN may use data from camera sensors to detect the presence of an authorized driver and / or owner of the vehicle 1000. The "always on" sensor processing engine may be used to unlock the vehicle when the owner approaches the driver's door and turns on the lights, and to disable the vehicle in security mode when the owner exits the vehicle. In this way, the SoC(s) 1004 provide security against theft and / or carjacking.
[0134] In another example, an emergency vehicle detection and identification CNN may use data from microphones 1096 to detect and identify emergency vehicle sirens. Unlike conventional systems that use general classifiers to detect sirens and manually extract features, the SoC(s) 1004 utilize the CNN to classify environmental and urban sounds, as well as visual data. In a preferred embodiment, the CNN running on the DLA is trained to detect the relative approach speed of the emergency vehicle (e.g., by using the Doppler effect). The CNN may also be trained to identify emergency vehicles specific to the local area the vehicle is traveling in, as identified by GNSS sensor(s) 1058.For example, the CNN will attempt to detect European sirens when deployed in Europe and only North American sirens when deployed in the United States. Once an emergency vehicle is detected, a control program can be used to execute an emergency vehicle safety routine, slow the vehicle, pull over to the side of the road, park the vehicle, and / or idle the vehicle, using ultrasonic sensors 1062, until the emergency vehicle(s) have passed.
[0135] The vehicle may include one or more CPU(s) 1018 (e.g., discrete CPU(s) or dCPU(s)) that may be connected to the SoC(s) 1004 via a high-speed connection (e.g., PCIe). The CPU(s) 1018 may include, for example, an x86 processor. The CPU(s) 1018 may be used to perform a variety of functions, including reconciling potentially inconsistent results between ADAS sensors and the SoC(s) 1004 and / or monitoring the status and health of the controller(s) 1036 and / or the infotainment SoC 1030, for example.
[0136] The vehicle 1000 may include one or more GPU(s) 1020 (e.g., discrete GPU(s) or dGPU(s)) that may be coupled to the SoC(s) 1004 via a high-speed interconnect (e.g., NVIDIA's NVLINK). The GPU(s) 1020 may provide additional artificial intelligence capabilities, e.g., by executing redundant and / or distinct neural networks, and may be used to train and / or update neural networks based on inputs (e.g., sensor data) from sensors of the vehicle 1000.
[0137] The vehicle 1000 may further include the network interface 1024, which may include one or more wireless antennas 1026 (e.g., one or more wireless antennas for various communication protocols, such as a cellular antenna, a Bluetooth antenna, etc.). The network interface 1024 may be used to enable wireless communication via the Internet with the cloud (e.g., with the server(s) 1078 and / or other network devices), with other vehicles, and / or with computing devices (e.g., passenger client devices). To communicate with other vehicles, a direct connection between the two vehicles and / or an indirect connection may be established (e.g., via networks and the Internet). Direct connections may be established via a vehicle-to-vehicle communication link.The vehicle-to-vehicle communication link may provide information to the vehicle 1000 about vehicles in the vicinity of the vehicle 1000 (e.g., vehicles in front of, beside, and / or behind the vehicle 1000). This function may be part of a cooperative adaptive cruise control function of the vehicle 1000.
[0138] The network interface 1024 may include an SoC that provides modulation and demodulation functions and enables the controller(s) 1036 to communicate over wireless networks. The network interface 1024 may include a radio frequency front-end for upconversion from baseband to radio frequency and downconversion from radio frequency to baseband. The frequency conversions may be performed using known methods and / or superheterodyne techniques. In some examples, the radio frequency front-end functionality may be provided by a separate chip. The network interface may include wireless functions for communication over LTE, WCDMA, UMTS, GSM, CDMA2000, Bluetooth, Bluetooth LE, Wi-Fi, Z-Wave, ZigBee, LoRaWAN, and / or other wireless protocols.
[0139] The vehicle 1000 may further include one or more data storage devices 1028 that may be stored off-chip (e.g., outside the SoC(s) 1004). The data storage device(s) 1028 may include one or more memory elements such as RAM, SRAM, DRAM, VRAM, flash, hard disks, and / or other components and / or devices capable of storing at least one bit of data.
[0140] The vehicle 1000 may also include GNSS sensor(s) 1058. The GNSS sensor(s) 1058 (e.g., GPS, assisted GPS sensors, differential GPS (DGPS), etc.) assist in mapping, sensing, occupancy grid creation, and / or path planning functions. Any number of GNSS sensors 1058 may be used, including, but not limited to, a GPS using a USB port with an Ethernet-to-serial (RS-232) bridge.
[0141] The vehicle 1000 may also include RADAR sensor(s) 1060. The RADAR sensor(s) 1060 may be used by the vehicle 1000 for long-range vehicle detection, even in darkness and / or adverse weather conditions. The RADAR sensor(s) 1060 may use the CAN bus and / or bus 1002 (e.g., to transmit data generated by the RADAR sensor(s) 1060) for control and to access object tracking data, with raw data being accessed via Ethernet in some examples. A variety of RADAR sensor types may be used. For example, and without limitation, the RADAR sensor(s) 1060 may be suitable for front, rear, and side RADAR deployment. In some examples, pulse-Doppler RADAR sensors are used.
[0142] The RADAR sensor(s) 1060 may include various configurations, such as long range with a narrow field of view, short range with a wide field of view, short range side coverage, etc. In some examples, long range RADAR may be used for adaptive cruise control. Long range RADAR systems may provide a wide field of view realized by two or more independent scans, e.g., within a range of 250 m. The RADAR sensor(s) 1060 may help distinguish between static and moving objects and may be used by ADAS systems for emergency braking and forward collision warning. Long range RADAR sensors may include monostatic multimodal RADAR with multiple (e.g., six or more) fixed RADAR antennas and a high-speed CAN and FlexRay interface.In a six-antenna example, the middle four antennas can create a focused beam pattern designed to detect the vehicle's surroundings at higher speeds with minimal interference from traffic in adjacent lanes. The other two antennas can expand the field of view, allowing vehicles entering or exiting the vehicle's lane to be quickly detected.
[0143] For example, medium-range radar systems can have a range of up to 1,060 m (front) or 80 m (rear) and a field of view of up to 42 degrees (front) or 1,050 degrees (rear). Short-range radar systems can include radar sensors mounted on both ends of the rear bumper. When mounted on both ends of the rear bumper, such a radar sensor system can create two beams that continuously monitor the blind spot behind and to the side of the vehicle.
[0144] Short-range radar systems can be used in an ADAS system for blind spot detection and / or as a lane change assistant.
[0145] The vehicle 1000 may also include ultrasonic sensor(s) 1062. The ultrasonic sensor(s) 1062, which may be mounted at the front, rear, and / or sides of the vehicle 1000, may be used for parking assistance and / or for creating and updating an occupancy grid. A variety of ultrasonic sensors 1062 may be used, and different ultrasonic sensors 1062 may be used for different detection ranges (e.g., 2.5 m, 4 m). The ultrasonic sensor(s) 1062 may operate at functional safety levels of ASIL B.
[0146] The vehicle 1000 may include LIDAR sensor(s) 1064. The LIDAR sensor(s) 1064 may be used for object and pedestrian detection, emergency braking, collision avoidance, and / or other functions. The LIDAR sensor(s) 1064 may be ASIL B functional safety rated. In some examples, the vehicle 1000 may include multiple LIDAR sensors 1064 (e.g., two, four, six, etc.) that may use Ethernet (e.g., to provide data to a Gigabit Ethernet switch).
[0147] In some examples, the LIDAR sensor(s) 1064 may be capable of providing a list of objects and their distances for a 360-degree field of view. Commercially available LIDAR sensors 1064 may have a range of approximately 1000 m, with an accuracy of 2 cm to 3 cm, and with support for a 1000 Mbps Ethernet connection, for example. In some examples, one or more non-protruding LIDAR sensors 1064 may be used. In such examples, the LIDAR sensor(s) 1064 may be implemented as a small device that may be embedded in the front, rear, sides, and / or corners of the vehicle 1000. In such examples, the LIDAR sensor(s) 1064 may provide a horizontal field of view of up to 120 degrees and a vertical field of view of up to 35 degrees with a range of 200 m even for objects with low reflectivity.The front-mounted LIDAR sensor(s) 1064 can be configured for a horizontal field of view between 45 degrees and 135 degrees.
[0148] In some examples, LIDAR technologies such as 3D flash LIDAR may also be used. 3D flash LIDAR uses a laser flash as a transmission source to illuminate the vehicle's surroundings up to a distance of approximately 200 m. A flash LIDAR unit includes a receiver that records the time of flight of the laser pulse and the reflected light at each pixel, which in turn corresponds to the distance between the vehicle and the objects. Flash LIDAR can produce highly accurate and distortion-free images of the surroundings with each laser flash. In some examples, four flash LIDAR sensors may be deployed, one on each side of the vehicle. Available 3D flash LIDAR systems include a solid-state 3D star array LIDAR camera, which has no moving parts other than a fan (e.g., a non-scanning LIDAR device).The flash lidar device can use a 5-nanosecond Class I (eye-safe) laser pulse per image and capture the reflected laser light as 3D range point clouds and co-registered intensity data. By using flash lidar, and because flash lidar is a solid-state device with no moving parts, the lidar sensor(s) 1064 can be less susceptible to motion blur, vibration, and / or shock.
[0149] The vehicle may also include IMU sensor(s) 1066. The IMU sensor(s) 1066 may, in some examples, be located at the center of the rear axle of the vehicle 1000. The IMU sensor(s) 1066 may, for example and without limitation, include accelerometer(s), magnetometer(s), gyroscope(s), magnetic compass(es), and / or other sensor types. In some examples, such as in six-axis applications, the IMU sensor(s) 1066 may include accelerometers and gyroscopes, while in nine-axis applications, the IMU sensor(s) 1066 may include accelerometers, gyroscopes, and magnetometers.
[0150] In some embodiments, the IMU sensor(s) 1066 may be implemented as a miniaturized, high-performance GPS-based inertial navigation system (GPS / INS) that combines microelectromechanical inertial sensors (MEMS), a high-sensitivity GPS receiver, and advanced Kalman filter algorithms to provide estimates of position, velocity, and attitude. Thus, in some examples, the IMU sensor(s) 1066 may enable the vehicle 1000 to estimate heading without requiring input from a magnetic sensor by directly observing velocity changes from GPS and correlating them with the IMU sensor(s) 1066. In some examples, the IMU sensor(s) 1066 and the GNSS sensor(s) 1058 may be combined into a single integrated unit.
[0151] The vehicle may include one or more microphones 1096 mounted in and / or around the vehicle 1000. The microphone(s) 1096 may be used, among other things, for detecting and identifying emergency vehicles.
[0152] The vehicle may further include any number of camera types, including stereo camera(s) 1068, wide-angle camera(s) 1070, infrared camera(s) 1072, surround camera(s) 1074, long-range and / or medium-range camera(s) 1098, and / or other camera types. The cameras may be used to capture image data around the entire perimeter of the vehicle 1000. The camera types used depend on the embodiments and requirements for the vehicle 1000, and any combination of camera types may be used to provide the required coverage around the vehicle 1000. Furthermore, the number of cameras may vary depending on the embodiment. For example, the vehicle may include six cameras, seven cameras, ten cameras, twelve cameras, and / or another number of cameras. The cameras may support, for example and without limitation, Gigabit Multimedia Serial Link (GMSL) and / or Gigabit Ethernet.Each of the cameras is described herein with respect to . Fig. 10A and Fig. 10B is described in more detail.
[0153] The vehicle 1000 may also include one or more vibration sensors 1042. The vibration sensor(s) 1042 may measure vibrations of components of the vehicle, such as the axle(s). For example, changes in vibration may indicate a change in the road surface. In another example, when using two or more vibration sensors 1042, the differences between the vibrations may be used to determine the friction or slippage of the road surface (e.g., if the difference in vibration is between a driven axle and a free-spinning axle).
[0154] The vehicle 1000 may include an ADAS system 1038. The ADAS system 1038 may include an SoC in some examples. The ADAS system 1038 may include autonomous / adaptive / automatic cruise control (ACC), cooperative adaptive cruise control (CACC), forward collision warning (FCW), automatic emergency braking (AEB), lane departure warning (LDW), lane keep assist (LKA), blind spot warning (BSW), rear cross traffic alert (RCTW), forward collision warning (CWS), lane centering (LC), and / or other features and functions.
[0155] The ACC systems may use RADAR sensor(s) 1060, LIDAR sensor(s) 1064, and / or camera(s). The ACC systems may include longitudinal ACC and / or lateral ACC. Longitudinal ACC monitors and controls the distance to the vehicle immediately in front of vehicle 1000 and automatically adjusts vehicle speed to maintain a safe distance from preceding vehicles. Lateral ACC monitors the distance and instructs vehicle 1000 to change lanes if necessary. Lateral ACC is related to other ADAS applications such as LCA and CWS.
[0156] CACC utilizes information from other vehicles, which may be received via the network interface 1024 and / or the wireless antenna(s) 1026 from other vehicles over a wireless connection or indirectly via a network connection (e.g., over the Internet). Direct connections may be provided by a vehicle-to-vehicle (V2V) communication link, while indirect connections may be an infrastructure-to-vehicle (I2V) communication link. In general, the V2V communication concept provides information about the immediately preceding vehicles (e.g., vehicles immediately in front of and in the same lane as vehicle 1000), while the I2V communication concept provides information about traffic further ahead. CACC systems may include both I2V and V2V information sources.Given the information about the 1000 vehicles ahead of the vehicle, the CACC system can be more reliable and has the potential to improve traffic flow and reduce congestion on the road.
[0157] FCW systems are designed to warn the driver of a hazard so they can take corrective action. FCW systems utilize a forward-facing camera and / or radar sensor(s) 1060 coupled with a dedicated processor, DSP, FPGA, and / or ASIC that is electrically connected to feedback to the driver, such as a display, speaker, and / or vibrating component. FCW systems can provide a warning, such as a sound, a visual warning, a vibration, and / or a rapid braking pulse.
[0158] AEB systems detect an impending collision with another vehicle or object and can automatically apply the brakes if the driver does not take corrective action within a specified time or distance parameter. AEB systems may use forward-facing camera(s) and / or radar sensor(s) 1060 connected to a dedicated processor, DSP, FPGA, and / or ASIC. When the AEB system detects a hazard, it typically first alerts the driver to take corrective action to avoid the collision. If the driver does not take corrective action, the AEB system can automatically apply the brakes to prevent or at least mitigate the effects of the predicted collision. AEB systems may incorporate techniques such as dynamic brake support and / or crash-imminent braking.
[0159] Lane departure warning systems warn the driver through visual, audible, and / or tactile signals, such as vibrations in the steering wheel or seat, when the vehicle crosses lane markings. A lane departure warning system will not activate if the driver indicates an intentional lane departure by activating the turn signal. LDW systems may use forward-facing cameras connected to a dedicated processor, DSP, FPGA, and / or ASIC that is electrically coupled to driver feedback, such as a display, speaker, and / or vibrating component.
[0160] LKA systems are a variant of LDW systems. LKA systems provide steering intervention or braking to correct the vehicle 1000 when the vehicle 1000 begins to leave the lane.
[0161] BSW systems detect and warn the driver of vehicles in the vehicle's blind spot. BSW systems can provide a visual, audible, and / or tactile warning signal to indicate that merging or changing lanes is unsafe. The system can provide an additional warning if the driver activates a turn signal. BSW systems can utilize rear-facing camera(s) and / or radar sensor(s) 1060 coupled with a dedicated processor, DSP, FPGA, and / or ASIC that is electrically coupled to driver feedback, such as a display, speaker, and / or vibrating component.
[0162] RCTW systems may provide a visual, audible, and / or tactile notification when an object is detected outside the range of the rear camera while the vehicle 1000 is reversing. Some RCTW systems include AEB to ensure the vehicle brakes are applied to avoid a crash. RCTW systems may utilize one or more rear-facing RADAR sensors 1060 coupled to a dedicated processor, DSP, FPGA, and / or ASIC that is electrically coupled to driver feedback, such as a display, speaker, and / or vibrating component.
[0163] Conventional ADAS systems can produce false positive results, which can be annoying and distracting for the driver, but are generally not catastrophic because ADAS systems warn the driver and give them the opportunity to decide whether a safety condition truly exists and act accordingly. However, in an autonomous vehicle 1000, in the event of conflicting results, the vehicle 1000 must decide for itself whether to consider the result of a primary computer or a secondary computer (e.g., a first controller 1036 or a second controller 1036). In some embodiments, the ADAS system 1038 may, for example, be a backup and / or secondary computer that provides perception information to a rationality module of the backup computer.The backup computer's rationality monitor can run redundant, diverse software on hardware components to detect errors in perception and dynamic driving tasks. The outputs of the ADAS system 1038 can be forwarded to a higher-level MCU. If the outputs of the primary and secondary computers conflict, the higher-level MCU must determine how to resolve the conflict to ensure safe operation.
[0164] In some examples, the primary computer may be configured to provide the parent MCU with a confidence score indicating the primary computer's confidence in the chosen outcome. If the confidence score exceeds a threshold, the supervising MCU may follow the primary computer's instruction regardless of whether the secondary computer provides a conflicting or inconsistent result. If the confidence score does not meet the threshold and the primary and secondary computers provide different results (e.g., a conflict), the parent MCU may mediate between the computers to determine the appropriate outcome.
[0165] The monitoring MCU can be configured to run a neural network(s) trained and configured to determine, based on the outputs of the primary computer and the secondary computer, the conditions under which the secondary computer triggers false alarms. This allows the neural network(s) in the monitoring MCU to learn when the output of the secondary computer can and cannot be trusted. For example, if the secondary computer is a radar-based FCW system, a neural network in the parent MCU can learn when the FCW system identifies metallic objects that are not actually hazardous, such as a drain grate or manhole cover, which triggers an alarm.If the secondary computer is a camera-based lane departure warning system, a neural network in the monitoring MCU can learn to override the lane departure warning system when cyclists or pedestrians are present and lane departure is actually the safest maneuver. In embodiments where a neural network(s) runs on the monitoring MCU, the monitoring MCU can include at least one DLA or a graphics processor suitable for operating the neural network(s) with associated memory. In preferred embodiments, the monitoring MCU can comprise and / or be included as a component of the SoC(s) 1004.
[0166] In other examples, the ADAS system 1038 may include a secondary computer that performs the ADAS functions using conventional computer vision rules. Thus, the secondary computer may use classical computer vision rules (if-then), and the presence of a neural network(s) in the parent MCU may improve reliability, safety, and performance. For example, the diverse implementation and intentional non-identity make the overall system more fault-tolerant, particularly against errors caused by software functions (or software-hardware interfaces).For example, if a software bug occurs in the software running on the primary computer and the non-identical software code on the secondary computer produces the same overall result, the monitoring MCU can assume with greater confidence that the overall result is correct and that the bug in the primary computer's software or hardware does not cause a significant failure.
[0167] In some examples, the output of the ADAS system 1038 may be fed into the perception block of the primary computer and / or the dynamic driving task block of the primary computer. For example, if the ADAS system 1038 displays a collision warning due to an object immediately ahead, the perception block may use this information in identifying objects. In other examples, the secondary computer may have its own neural network trained, thus reducing the risk of false alarms, as described herein.
[0168] The vehicle 1000 may further include the infotainment SoC 1030 (e.g., an in-vehicle infotainment (IVI) system). Although the infotainment system is illustrated and described as an SoC, it need not be an SoC, but may also include two or more discrete components. The infotainment SoC 1030 may include a combination of hardware and software used to provide audio (e.g., music, a personal digital assistant, navigation instructions, news, radio, etc.), video (e.g., TV, movies, streaming, etc.), phone (e.g., hands-free calling), network connectivity (e.g., LTE, Wi-Fi, etc.), and / or information services (e.g., navigation systems, rear parking assist, a radio data system, vehicle-related information such as fuel level, total distance traveled, brake fuel level, oil level, door open / close, air filter information, etc.) to the vehicle 1000.The infotainment SoC 1030 may include, for example, radios, record players, navigation systems, video players, USB and Bluetooth connectivity, car computers, in-car entertainment, Wi-Fi, steering wheel audio controls, hands-free calling, a heads-up display (HUD), an HMI display 1034, a telematics device, a control panel (e.g., for controlling and / or interacting with various components, functions, and / or systems), and / or other components. The infotainment SoC 1030 may further be used to provide information (e.g., visual and / or audible) to one or more users of the vehicle, such as information from the ADAS system 1038, autonomous driving information such as planned vehicle maneuvers, trajectories, environmental information (e.g., intersection information, vehicle information, road information, etc.), and / or other information.
[0169] The infotainment SoC 1030 may include GPU functionality. The infotainment SoC 1030 may communicate with other devices, systems, and / or components of the vehicle 1000 via the bus 1002 (e.g., CAN bus, Ethernet, etc.). In some examples, the infotainment SoC 1030 may be coupled to a supervisory MCU so that the infotainment system's GPU may perform some self-driving functions in the event that the primary control unit(s) 1036 (e.g., the primary and / or backup computers of the vehicle 1000) fail. In such an example, the infotainment SoC 1030 may place the vehicle 1000 in a chauffeur mode until a safe stop, as described herein.
[0170] The vehicle 1000 may further include an instrument cluster 1032 (e.g., a digital dashboard, an electronic instrument cluster, a digital instrument panel, etc.). The instrument cluster 1032 may include a controller and / or a supercomputer (e.g., a discrete controller or a supercomputer). The instrument cluster 1032 may include a number of instruments, such as a speedometer, fuel level, oil pressure, tachometer, odometer, turn signals, shift position indicator, seat belt warning light(s), parking brake warning light(s), engine malfunction light(s), airbag system (SRS) information, lighting controls, safety system controls, navigation information, etc. In some examples, information may be displayed and / or shared between the infotainment SoC 1030 and the instrument cluster 1032. In other words, the instrument cluster 1032 may be part of the infotainment SoC 1030, or vice versa.
[0171] Fig. 10D is a system diagram for the communication between the cloud-based server(s) and the example autonomous vehicle 1000 of Fig. 10A, in accordance with some embodiments of the present disclosure. The system 1076 may include the server(s) 1078, the network(s) 1090, and the vehicles, including the vehicle 1000. The server(s) 1078 may include a plurality of GPUs 1084(A)-1084(H) (collectively referred to herein as GPUs 1084), PCIe switches 1082(A)-1082(H) (collectively referred to herein as PCIe switches 1082), and / or CPUs 1080(A)-1080(B) (collectively referred to herein as CPUs 1080). The GPUs 1084, the CPUs 1080, and the PCIe switches may be interconnected via high-speed links, such as high-speed interconnects. B. and without limitation, the NVLink interfaces 1088 and / or PCIe connections 1086 developed by NVIDIA. In some examples, the GPUs 1084 are connected via NVLink and / or NVSwitch SoC and the GPUs 1084 and the PCIe switches 1082 are connected via PCIe connections.Although eight GPUs 1084, two CPUs 1080, and two PCIe switches are illustrated, this is not intended to be limiting. Depending on the embodiment, each of the servers 1078 may include any number of GPUs 1084, CPUs 1080, and / or PCIe switches. For example, the servers 1078 may each include eight, sixteen, thirty-two, and / or more GPUs 1084.
[0172] The server(s) 1078 may receive, via the network(s) 1090 and from the vehicles, image data representative of images depicting unexpected or changed road conditions, such as recently commenced roadwork. The server(s) 1078 may transmit, via the network(s) 1090 and to the vehicles, neural networks 1092, updated neural networks 1092, and / or map information 1094, including information about traffic and road conditions. The updates to the map information 1094 may include updates to the HD map 1022, such as information about construction, potholes, detours, flooding, and / or other obstacles.In some examples, the neural networks 1092, the updated neural networks 1092, and / or the map information 1094 may result from new training and / or experience represented in data received from any number of vehicles in the area and / or based on training performed in a data center (e.g., using the server(s) 1078 and / or other servers).
[0173] The server(s) 1078 may be used to train machine learning models (e.g., neural networks) based on training data. The training data may be generated by the vehicles and / or in a simulation (e.g., with a game engine). In some examples, the training data is tagged (e.g., if the neural network benefits from supervised learning) and / or undergoes other preprocessing, while in other examples, the training data is not tagged and / or preprocessed (e.g., if the neural network does not require supervised learning).Training may be performed using one or more classes of machine learning techniques, including, but not limited to, supervised training, semi-supervised training, unsupervised training, self-learning, reinforcement learning, federated learning, transfer learning, feature learning (including principal component and cluster analysis), multilinear subspace learning, manifold learning, representation learning (including dictionary replacement learning), rule-based machine learning, anomaly detection, and any variations or combinations thereof. Once the machine learning models are trained, the machine learning models may be used by the vehicles (e.g., by transmitting them to the vehicles via the network(s) 1090) and / or the machine learning models may be used by the server(s) 1078 to remotely monitor the vehicles.
[0174] In some examples, the servers 1078 may receive data from the vehicles and apply the data to real-time, real-time neural networks for intelligent reasoning. The server(s) 1078 may include deep learning supercomputers and / or dedicated AI computers powered by GPU(s) 1084, such as the DGX and DGX Station machines developed by NVIDIA. However, in some examples, the servers 1078 may also include deep learning infrastructures using only CPU-powered data centers.
[0175] The deep learning infrastructure of server(s) 1078 may be capable of rapid, real-time inferencing and may utilize this capability to evaluate and verify the state of the processors, software, and / or associated hardware in vehicle 1000. For example, the deep learning infrastructure may receive regular updates from vehicle 1000, such as a sequence of images and / or objects that vehicle 1000 has located within that sequence of images (e.g., via computer vision and / or other machine object classification techniques). The deep learning infrastructure may run its own neural network to identify the objects and compare them to the objects identified by vehicle 1000.If the results do not match and the infrastructure concludes that the AI in the vehicle 1000 is malfunctioning, the server 1078 may send a signal to the vehicle 1000 instructing a fail-safe computer of the vehicle 1000 to take control, notify the passengers, and perform a safe parking maneuver.
[0176] For inference building, the server(s) 1078 may include the GPU(s) 1084 and one or more programmable inference accelerators (e.g., NVIDIA's TensorRT). The combination of GPU-driven servers and inference acceleration may enable real-time responsiveness. In other examples, such as when performance is less critical, servers powered by CPUs, FPGAs, and other processors may be used for inference building. EXAMPLE OF A DATA PROCESSING SYSTEM
[0177] Fig. 11 is a block diagram of exemplary computing device(s) 1100 suitable for use in implementing some embodiments of the present disclosure. Computing device 1100 may include an interconnection system 1102 that directly or indirectly interconnects the following devices: memory 1104, one or more central processing units (CPUs) 1106, one or more graphics processing units (GPUs) 1108, a communications interface 1110, input / output (I / O) ports 1112, input / output components 1114, a power supply 1116, one or more presentation components 1118 (e.g., display(s)), and one or more logic units 1120. In at least one embodiment, computing device(s) 1100 may include one or more virtual machines (VMs), and / or each of the components thereof may include virtual components (e.g., virtual hardware components).As non-limiting examples, one or more of the GPUs 1108 may include one or more vGPUs, one or more of the CPUs 1106 may include one or more vCPUs, and / or one or more of the logic units 1120 may include one or more virtual logic units. As such, one or more computing devices 1100 may include discrete components (e.g., an entire GPU dedicated to the computing device 1100), virtual components (e.g., a portion of a GPU dedicated to the computing device 1100), or a combination thereof.
[0178] Although the different blocks in Fig. 11 as being connected to wires via the interconnect system 1102, this is not limiting and is for clarity only. For example, in some embodiments, a presentation component 1118, such as a display device, may be considered an I / O component 1114 (e.g., if the display is a touchscreen). Another example is that the CPUs 1106 and / or GPUs 1108 may include memory (e.g., the memory 1104 may represent a storage device in addition to the memory of the GPUs 1108, the CPUs 1106, and / or other components). In other words, the computing device of Fig. 11 is merely illustrative. No distinction is made between categories such as “workstation”, “server”, “laptop”, “desktop”, “tablet”, “client device”, “mobile device”, “handheld device”, “game console”, “electronic control unit (ECU)”, “virtual reality system” and / or other device or system types, as they all fall within the scope of the computing device of Fig. 11 fall.
[0179] The interconnect system 1102 may represent one or more connections or buses, such as an address bus, a data bus, a control bus, or a combination thereof. The interconnect system 1102 may include one or more bus or connection types, such as an Industry Standard Architecture (ISA) bus, an Extended Industry Standard Architecture (EISA) bus, a Video Electronics Standards Association (VESA) bus, a Peripheral Component Interconnect (PCI) bus, a Peripheral Component Interconnect Express (PCIe) bus, and / or another type of bus or connection. In some embodiments, there are direct connections between the components. For example, the CPU 1106 may be directly connected to the memory 1104. Additionally, the CPU 1106 may be directly connected to the GPU 1108. For direct or point-to-point connections between components, the interconnect system 1102 may include a PCIe connection to establish the connection.In these examples, a PCI bus does not need to be integrated into the computing device.
[0180] The memory 1104 may consist of a variety of computer-readable media. The computer-readable media may be any available media accessible by the computing device 1100. The computer-readable media may include both volatile and non-volatile media, as well as removable and non-removable media. By way of example and without limitation, the computer-readable media may include computer storage media and communication media.
[0181] The computer storage media may include both volatile and non-volatile media and / or removable and non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules, and / or other data types. For example, memory 1104 may store computer-readable instructions (e.g., those representing a program(s) and / or a program element(s), such as an operating system.Computer storage media may include, but is not limited to, RAM, ROM, EEPROM, flash memory or other storage technologies, CD-ROM, Digital Versatile Disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and that can be accessed by the computing device 1100. The term "computer storage medium" as used herein does not per se include signals.
[0182] The computer storage media may embody computer-readable instructions, data structures, program modules, and / or other data types in a modulated data signal, such as a carrier wave or other transport mechanism, and include any information transmission media. The term "modulated data signal" may refer to a signal having one or more of its characteristics adjusted or altered to encode information in the signal. The computer storage media may include, for example, wired media, such as a wired network or a direct cable connection, and wireless media, such as acoustic, RF, infrared, and other wireless media. Combinations of the above media should also be considered computer-readable media.
[0183] The CPU(s) 1106 may be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 1100 to perform one or more of the methods and / or processes described herein. The CPU(s) 1106 may each include one or more cores (e.g., one, two, four, eight, twenty-eight, seventy-two, etc.) capable of concurrently executing a plurality of software threads. The CPU(s) 1106 may include any type of processor and may include different types of processors depending on the type of computing device 1100 implemented (e.g., processors with fewer cores for mobile devices and processors with more cores for servers).Depending on the type of computing device 1100, the processor may be, for example, an Advanced RISC Machines (ARM) processor operating with Reduced Instruction Set Computing (RISC) or an x86 processor operating with Complex Instruction Set Computing (CISC). Computing device 1100 may include one or more CPUs 1106 in addition to one or more microprocessors or additional coprocessors, such as math coprocessors.
[0184] In addition to or alternatively to the CPU(s) 1106, the GPU(s) 1108 may be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 1100 to perform one or more of the methods and / or processes described herein. One or more of the GPU(s) 1108 may be an integrated GPU (e.g., with one or more of the CPU(s) 1106) and / or one or more of the GPU(s) 1108 may be a discrete GPU. In some embodiments, one or more of the GPU(s) 1108 may be a coprocessor of one or more of the CPU(s) 1106. The graphics processor(s) 1108 may be used by the computing device 1100 to render graphics (e.g., 3D graphics) or to perform general-purpose computations. The GPU(s) 1108 may be used, for example, for General-Purpose Computing on GPUs (GPGPU).The GPU(s) 1108 may include hundreds or thousands of cores capable of processing hundreds or thousands of software threads simultaneously. The GPU(s) 1108 may generate pixel data for output images in response to rendering commands (e.g., rendering commands from the CPU(s) 1106 received via a host interface). The GPU(s) 1108 may include graphics memory, such as display memory, for storing pixel data or other suitable data, such as GPGPU data. The display memory may be included as part of the memory 1104. The GPU(s) 1108 may include two or more GPUs operating in parallel (e.g., via an interconnect). The interconnect may connect the GPUs directly (e.g., using NVLINK) or via a switch (e.g., using NVSwitch).When combined, each GPU can generate 1108 pixel data or GPGPU data for different parts of an output or for different outputs (e.g., a first GPU for a first image and a second GPU for a second image). Each GPU can have its own memory or share memory with other GPUs.
[0185] In addition to or alternatively to the CPU(s) 1106 and / or the GPU(s) 1108, the logic unit(s) 1120 may be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 1100 to perform one or more of the methods and / or processes described herein. In embodiments, the CPU(s) 1106, the GPU(s) 1108, and / or the logic unit(s) 1120 may discretely or jointly execute any combination of the methods, processes, and / or portions thereof. One or more of the logic units 1120 may be part of one or more of the CPU(s) 1106 and / or the GPU(s) 1108, and / or one or more of the logic units 1120 may be discrete components or otherwise external to the CPU(s) 1106 and / or the GPU(s) 1108.In embodiments, one or more of the logic units 1120 may be a coprocessor of one or more of the CPU(s) 1106 and / or one or more of the GPU(s) 1108.
[0186] Examples of the logic unit(s) 1120 include one or more compute cores and / or components thereof, such as data processing units (DPUs), tensor cores (TCs), tensor processing units (TPUs), pixel visual cores (PVCs), vision processing units (VPUs), graphics processing clusters (GPCs), texture processing clusters (TPCs), streaming multiprocessors (SMs), tree traversal units (TTUs), artificial intelligence accelerators (AIAs), deep learning accelerators (DLAs), arithmetic logic units (ALUs), application-specific integrated circuits (ASICs), floating point units (FPUs), input / output (I / O) elements, peripheral component interconnect (PCI) or peripheral component interconnect express (PCIe) elements, and / or the like.
[0187] The communication interface 1110 may include one or more receivers, transmitters, and / or transceivers that enable the computing device 1100 to communicate with other computing devices over an electronic communication network, including wired and / or wireless communication. The communication interface 1110 may include components and functions to enable communication over a variety of networks, such as wireless networks (e.g., Wi-Fi, Z-Wave, Bluetooth, Bluetooth LE, ZigBee, etc.), wired networks (e.g., communication over Ethernet or InfiniBand), low-power wide area networks (e.g., LoRaWAN, SigFox, etc.), and / or the Internet.In one or more embodiments, the logic unit(s) 1120 and / or the communication interface 1110 may include one or more data processing units (DPUs) to transfer data received over a network and / or via the interconnect system 1102 directly to one or more GPU(s) 1108 (e.g., a memory).
[0188] The I / O ports 1112 may enable the computing device 1100 to be logically connected to other devices, including the I / O components 1114, the presentation component(s) 1118, and / or other components, some of which may be built into (e.g., integrated) the computing device 1100. Illustrative I / O components 1114 include a microphone, a mouse, a keyboard, a joystick, a gamepad, a game controller, a satellite dish, a scanner, a printer, a wireless device, etc. The I / O components 1114 may provide a natural user interface (NUI) that processes air gestures, speech, or other physiological inputs from a user. In some cases, the inputs may be transmitted to a suitable network element for further processing.An NUI may implement any combination of speech recognition, pen recognition, facial recognition, biometric recognition, both on-screen and off-screen gesture recognition, air gestures, head and eye tracking, and touch detection (as described in more detail below) in conjunction with a display of computing device 1100. Computing device 1100 may include depth cameras, such as stereoscopic camera systems, infrared camera systems, RGB camera systems, touchscreen technology, and combinations thereof, for gesture recognition and capture. In addition, computing device 1100 may include accelerometers or gyroscopes (e.g., as part of an inertial measurement unit (IMU)) that enable the detection of motion. In some examples, the output of the accelerometers or gyroscopes from computing device 1100 may be used to present immersive augmented reality or virtual reality.
[0189] The power supply 1116 may be a hardwired power supply, a battery power supply, or a combination thereof. The power supply 1116 may supply power to the computing device 1100 to enable operation of the components of the computing device 1100.
[0190] The presentation component(s) 1118 may include a display (e.g., a monitor, a touchscreen, a television screen, a heads-up display (HUD), other display types, or a combination thereof), speakers, and / or other presentation components. The presentation component(s) 1118 may receive data from other components (e.g., the GPU(s) 1108, the CPU(s) 1106, DPUs, etc.) and output the data (e.g., as an image, video, audio, etc.). EXAMPLE DATA CENTER
[0191] Fig. Figure 12 shows an example of a data center 1200 that may be used in at least one embodiment of the present disclosure. The data center 1200 may include a data center infrastructure layer 1210, a framework layer 1220, a software layer 1230, and / or an application layer 1240.
[0192] As in Fig. 12, the data center infrastructure layer 1210 may include a resource orchestrator 1212, clustered computing resources 1214, and node computing resources (“node CRs”) 1216(1)-1216(N), where “N” represents any positive integer. In at least one embodiment, the node CRs 1216(1)-1216(N) may include any number of central processing units (CPUs) or other processors (including DPUs, accelerators, field programmable gate arrays (FPGAs), graphics processors or graphics processing units (GPUs), etc.), memory devices (e.g., dynamic read-only memory), storage devices (e.g., solid-state or disk drives), network input / output (NW I / O) devices, network switches, virtual machines (VMs), power modules and / or cooling modules, etc. In some embodiments, one or more node CRs among the node CRs 1216(1)-1216(N) may correspond to a server having one or more of the above-mentioned computing resources. Furthermore, in some embodiments, node CRs 1216(1)-1216(N) may include one or more virtual components, such as vGPUs, vCPUs, and / or the like, and / or one or more of node CRs 1216(1)-1216(N) may correspond to a virtual machine (VM).
[0193] In at least one embodiment, the grouped computing resources 1214 may include separate groupings of node CRs 1216 housed in one or more racks (not shown) or in many racks in data centers in different geographic locations (also not shown). Separate groupings of node CRs 1216 within the grouped computing resources 1214 may include grouped computing, networking, storage, or memory resources that may be configured or allocated to support one or more workloads. In at least one embodiment, multiple node CRs 1216 having CPUs, GPUs, DPUs, and / or other processors may be grouped in one or more racks to provide computing resources to support one or more workloads.The one or more racks may also contain any number of power modules, cooling modules and / or network switches in any combination.
[0194] Resource orchestrator 1212 may configure or otherwise control one or more node CRs 1216(1)-1216(N) and / or clustered computing resources 1214. In at least one embodiment, resource orchestrator 1212 may include a software design infrastructure (SDI) management entity for data center 1200. Resource orchestrator 1212 may include hardware, software, or a combination thereof.
[0195] In at least one embodiment, as in Fig. 12, the framework layer 1220 may include a job scheduler 1233, a configuration manager 1234, a resource manager 1236, and / or a distributed file system 1238. The framework layer 1220 may include a framework to support the software 1232 of the software layer 1230 and / or one or more applications 1242 of the application layer 1240. The software 1232 or application(s) 1242 may include web-based service software or applications such as those provided by Amazon Web Services, Google Cloud, and Microsoft Azure. The framework layer 1220 may be some type of free and open source software web application framework such as Apache Spark™ (hereinafter "Spark"), which may utilize a distributed file system 1238 for processing large amounts of data (e.g., "Big Data"), but is not limited to.In at least one embodiment, the job scheduler 1233 may include a Spark driver to facilitate the scheduling of workloads supported by various layers of the data center 1200. The configuration manager 1234 may be capable of configuring various layers such as the software layer 1230 and the framework layer 1220, including Spark and the distributed file system 1238, to support the processing of large amounts of data. The resource manager 1236 may be capable of managing clustered or grouped computing resources allocated to support the distributed file system 1238 and the job scheduler 1233. In at least one embodiment, the clustered or grouped computing resources may include the clustered computing resources 1214 at the infrastructure layer 1210 of the data center.The resource manager 1236 may coordinate with the resource orchestrator 1212 to manage these allocated or assigned computing resources.
[0196] In at least one embodiment, the software 1232 included in software layer 1230 may include software used by at least portions of node CRs 1216(1)-1216(N), clustered computing resources 1214, and / or distributed file system 1238 of framework layer 1220. One or more types of software may include, but are not limited to, Internet search software, email virus scanning software, database software, and streaming video content software.
[0197] In at least one embodiment, the application(s) 1242 included in application layer 1240 may include one or more types of applications used by at least portions of node CRs 1216(1)-1216(N), clustered computing resources 1214, and / or distributed file system 1238 of framework layer 1220. One or more types of applications may include any number of genomic applications, cognitive computation, and machine learning applications, including, but not limited to, training or inference software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), and / or other machine learning applications used in connection with one or more embodiments.
[0198] In at least one embodiment, configuration manager 1234, resource manager 1236, and resource orchestrator 1212 may implement any number and type of self-modifying actions based on any amount and type of data collected in any technically feasible manner. Self-modifying actions may relieve a data center operator of data center 1200 from potentially making poor configuration decisions and potentially avoid underutilized and / or malfunctioning portions of a data center.
[0199] Data center 1200 may include tools, services, software, or other resources to train one or more machine learning models or to predict or infer information using one or more machine learning models according to one or more embodiments described herein. For example, a machine learning model(s) may be trained by calculating weight parameters according to a neural network architecture using software and / or computing resources described above with respect to data center 1200.In at least one embodiment, trained or deployed machine learning models corresponding to one or more neural networks may be used to infer or predict information using the resources described above with respect to data center 1200 by using weighting parameters calculated by one or more training techniques, such as, but not limited to, those described herein.
[0200] In at least one embodiment, the data center 1200 may utilize CPUs, application-specific integrated circuits (ASICs), GPUs, FPGAs, and / or other hardware (or corresponding virtual computing resources) to perform training and / or inferencing with the resources described above. Furthermore, one or more of the software and / or hardware resources described above may be configured as a service to enable users to train or infer information, such as image recognition, speech recognition, or other artificial intelligence services. EXAMPLE OF NETWORK ENVIRONMENTS
[0201] Network environments suitable for use in implementing embodiments of the disclosure may include one or more client devices, servers, network-attached storage (NAS), other backend devices, and / or other device types. The client devices, servers, and / or other device types (e.g., each device) may be based on one or more instances of the computing device(s) 1100 of Fig. 11 - e.g., each device may include similar components, features, and / or functionality of the computing device(s) 1100. If backend devices (e.g., servers, NAS, etc.) are implemented, the backend devices may also be part of a data center 1200, an example of which is described here with respect to Fig. 12 is described in more detail.
[0202] The components of a network environment can communicate with each other over one or more networks, which can be wired, wireless, or both. The network can comprise multiple networks or a network of networks. For example, the network can comprise one or more wide area networks (WANs), one or more local area networks (LANs), one or more public networks such as the Internet and / or a public switched telephone network (PSTN), and / or one or more private networks. If the network comprises a wireless telecommunications network, components such as a base station, a communications tower, or even access points (as well as other components) can provide wireless connectivity.
[0203] Compatible network environments include one or more peer-to-peer network environments—in which case, a server cannot be included in a network environment—and one or more client-server network environments—in which case, one or more servers can be included in a network environment. In peer-to-peer network environments, the functionality described herein with respect to one or more servers can be implemented on any number of client devices.
[0204] In at least one embodiment, a network environment may include one or more cloud-based network environments, a distributed computing environment, a combination thereof, etc. A cloud-based network environment may include a framework layer, a job scheduler, a resource manager, and a distributed file system implemented on one or more servers, which may include one or more core network servers and / or edge servers. A framework layer may include a framework for supporting software of a software layer and / or one or more applications of an application layer. The software or application(s) may each include web-based service software or applications. In embodiments, one or more client devices may utilize the web-based service software or applications (e.g., by accessing the service software and / or applications via one or more application programming interfaces (APIs)).The framework layer can be some type of free and open source software web application framework that uses, for example, but is not limited to, a distributed file system for processing large amounts of data (e.g., "Big Data").
[0205] A cloud-based network environment may provide cloud computing and / or cloud storage performing any combination of the computing and / or data storage functions described herein (or one or more portions thereof). Each of these various functions may be distributed across multiple locations of central or core servers (e.g., one or more data centers that may be located across a state, region, country, globe, etc.). Where a connection to a user (e.g., a client device) is relatively close to one or more edge servers, a core server may allocate at least some functionality to the edge server(s). A cloud-based network environment may be private (e.g., limited to a single organization), public (e.g., available to many organizations), and / or a combination thereof (e.g., a hybrid cloud environment).
[0206] The client device(s) may include at least some of the components, features, and functions of the device(s) described herein with respect to Fig.11. By way of example and without limitation, a client device may be a personal computer (PC), a laptop, a mobile device, a smartphone, a tablet computer, a smartwatch, a wearable computer, a personal digital assistant (PDA), an MP3 player, a virtual reality headset, a global positioning system (GPS) or device, a video player, a video camera, a surveillance device or system, a vehicle, a boat, a flying vessel, a virtual machine, a computer capable of processing the data stored in the database, or other device, a boat, a flying vessel, a virtual machine, a drone, a robot, a portable communication device, a hospital device, a gaming device or system, an entertainment system, a vehicle computing system, an embedded system controller,a remote control, a device, a consumer electronics device, a workstation, an edge device, any combination of these described devices, or any other suitable device.
[0207] The disclosure may be described in the general context of computer code or machine-usable instructions, including computer-executable instructions such as program modules, that are executed by a computer or other machine, such as a personal data assistant or other handheld device. In general, program modules, including routines, programs, objects, components, data structures, etc., refer to code that performs specific tasks or implements specific abstract data types. The disclosure may be practiced in a variety of system configurations, including handheld devices, consumer electronics, general-purpose computers, more specialized data processing devices, etc. The disclosure may also be applied in distributed computing environments where tasks are performed by remotely controlled devices connected via a communications network.
[0208] When reference is made herein to "and / or" in reference to two or more elements, this should be understood to mean only one element or a combination of elements. For example, "element A, element B, and / or element C" may include only element A, only element B, only element C, element A and element B, element A and element C, element B and element C, or elements A, B, and C. Furthermore, "at least one of elements A or B" may include at least one of elements A, at least one of elements B, or at least one of elements A and at least one of elements B.
[0209] The subject matter of the present disclosure is described herein with a level of particularity consistent with legal requirements. However, the description itself is not intended to limit the scope of this disclosure. Rather, the inventors contemplated that the claimed subject matter could be embodied in other ways to include various steps or combinations of steps similar to those described herein in connection with other present or future technologies. Although the terms "step" and / or "block" may be used herein to refer to various elements of the methods employed, the terms should not be construed to imply any particular ordering among or between the various steps disclosed herein unless the order of each step is expressly described.
[0210] The disclosure of this application also includes the following numbered sections: Clause 1. A method comprising: determining, at least based on image data representing an image, that an image portion is associated with a type of content; generating, at least based on the image portion associated with the type of content, first data indicating the image portion; generating, using an encoder and at least based on the first data, encoded image data by at least encoding the image data such that the image portion is obscured; and causing the encoded image data to be sent to one or more computing devices. Clause 2. The method of clause 1, wherein generating the first data comprises generating the first data representative of a boundary shape indicative of the image portion associated with the type of content. Clause 3. The method of clause 1 or clause 2, wherein generating the first data comprises generating the first data representative of a feature map associated with the image, the feature map indicating the portion of the image associated with the type of content. Clause 4. The method of clause 3, wherein: a first portion of the feature map associated with the image portion includes one or more first values associated with obscuring the image; and a second portion of the feature map associated with a second image portion includes one or more second values associated with non-obscuring the image. Clause 5. The method of any preceding paragraph, wherein generating the encoded image data further comprises encoding the image data using the encoder and based on at least a second image portion that does not include the type of content, such that the second image portion is not obscured. Clause 6. The method of any preceding paragraph, further comprising: determining a quantization parameter associated with obscuring the image portion, wherein generating the encoded image data comprises generating the encoded image data using the encoder and at least based on the first data by encoding at least a portion of the image data using the quantization parameter, the portion of the image data being associated with the image portion. Clause 7. The method of any preceding paragraph, wherein the quantization parameter includes a maximum quantization parameter associated with the encoding of the image data. Clause 8. The method of any preceding paragraph, wherein generating the encoded image data comprises: updating one or more residual coefficient values associated with the image patch using the encoder to be less than or equal to a threshold coefficient value; and after updating the one or more residual coefficient values, generating the encoded image data using the encoder by performing at least one of a transform and a quantization on the image data. Clause 9. The method of any preceding paragraph, wherein determining the image portion associated with the type of content comprises determining, at least based on the image data representing the image, that the image portion represents sensitive information. Clause 10. The method of any preceding paragraph further comprises: receiving the image data generated using one or more image sensors of a machine navigating in an environment; determining, at least based on the image data, one or more operations for the machine to perform in the environment; and causing the machine to perform the one or more operations. Clause 11. A system comprising: one or more processing units configured to: determine, at least based on image data representing an image, that an image portion is associated with a type of content; using an encoder and at least based on the image portion associated with the type of content, generate encoded image data by at least encoding the image data such that the image portion is obscured; and cause the encoded image data to be sent to one or more computing devices. Clause 12. The system of clause 11, wherein the one or more processing units are further operable to: generate first data representative of a boundary shape associated with the image section, wherein the encoded image data is further generated at least based on the first data. Clause 13. The system of Clause 11 or Clause 12, wherein the one or more processing units are further operable to generate a feature map associated with the image data, wherein generating the feature map comprises at least: dividing the feature map into a plurality of blocks; causing a first portion of the plurality of blocks to indicate that the image portion is associated with the type of content; and causing a second portion of the plurality of blocks to indicate that a second image portion is not associated with the type of content, wherein the encoded image data is further generated at least based on the feature map. Clause 14. The system of any one of clauses 11 to 13, wherein generating the encoded image data is further performed by encoding the image data using the encoder and based on at least a second image portion associated with a second type of content, such that the second image portion is not obscured. Clause 15. The system of any of clauses 11-14, wherein the one or more processing units are further operable to: determine a quantization parameter associated with obscuring the image portion, wherein generating the encoded image data comprises generating the encoded image data using the encoder and based on at least the image portion associated with the type of content by at least encoding a portion of the image data using the quantization parameter, the portion of the image data being associated with the image portion. Clause 16. The system of any one of clauses 11 to 15, wherein generating the encoded image data comprises: updating, using the encoder, one or more residual coefficient values associated with the image patch to be less than or equal to a threshold coefficient value; and after updating the one or more residual coefficient values, generating, using the encoder, the encoded image data by performing a transformation or quantization of the image data. Clause 17. The system of any of Clauses 11-16, wherein determining that the image portion is associated with the type of content comprises determining, at least based on the image data representing the image, that the image portion is associated with sensitive information. Clause 18. The system of any of paragraphs 11-17, wherein the system consists of at least one of the following: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing one or more simulation operations; a system for performing one or more digital twin operations; a system for performing light transport simulation; a system for performing collaborative content creation for 3D assets; a system for performing one or more deep learning operations; a system implemented using an edge device; a system implemented using a robot; a system for performing one or more generative AI operations; a system for performing operations using a large language model;a system for performing one or more conversational AI operations; a system for generating synthetic data; a system for presenting virtual reality content, augmented reality content, or mixed reality content; a system including one or more virtual machines (VMs); a system implemented at least in part in a data center; or a system implemented at least in part using cloud computing resources. Clause 19. A processor comprising: one or more processing units for generating coded image data by encoding a first portion of image data to obscure a first image portion and a second portion of the image data to make a second image portion unrecognizable, wherein the encoding is based at least on the first image portion being associated with a first type of content and the second image portion being associated with a second type of content. Clause 20. The processor of Clause 19, wherein the processor is included in at least one of the following systems: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing one or more simulation operations; a system for performing one or more digital twin operations; a system for performing light transport simulation; a system for performing collaborative content creation for 3D assets; a system for performing one or more deep learning operations; a system implemented using an edge device; a system implemented using a robot; a system for performing one or more generative AI operations; a system for performing operations using a large language model;a system for performing one or more conversational AI operations; a system for generating synthetic data; a system for presenting virtual reality content, augmented reality content, or mixed reality content; a system including one or more virtual machines (VMs); a system implemented at least in part in a data center; or a system implemented at least in part using cloud computing resources.
[0211] It is to be understood that the aspects and embodiments described above are exemplary only and that changes in detail may be made within the scope of the claims.
[0212] Each device, method, and feature disclosed in the description and (where appropriate) in the claims and drawings may be provided independently or in any suitable combination.
[0213] The reference numbers contained in the claims are for illustrative purposes only and do not limit the scope of the claims. QUOTES CONTAINED IN THE DESCRIPTION
[0000] This list of documents submitted by the applicant was generated automatically and is included solely for the convenience of the reader. This list is not part of the German patent or utility model application. The DPMA assumes no liability for any errors or omissions. Cited patent literature
[0000] US 16 / 101,232
[0112]
Claims
[1] Procedure comprising: Determining, at least on the basis of image data representative of an image, that an image portion is associated with a type of content; generating first data indicating the image section, at least based on the image section associated with the type of content; Generating coded image data using an encoder and at least based on the first data by encoding at least the image data such that the image section is obscured; and to send the encoded image data to one or more computing devices. [2] The method of claim 1, wherein generating the first data comprises generating the first data representative of a boundary shape indicating the image portion associated with the type of content. [3] The method of claim 1 or claim 2, wherein generating the first data comprises generating the first data representative of a feature map associated with the image, the feature map indicating the image portion associated with the type of content. [4] The method of claim 3, wherein: a first part of the feature map associated with the image section contains one or more first values associated with the blurring of the image; and a second part of the feature map, which is assigned to a second image section, contains one or more second values which are associated with the non-obscuration of the image. [5] A method according to any one of the preceding claims, wherein generating the coded image data further comprises encoding the image data using the encoder and on the basis of at least a second image section that does not contain the type of content, such that the second image section is not obscured. [6] A method according to any one of the preceding claims, further comprising: Determining a quantization parameter associated with the blurring of the image section, wherein generating the encoded image data comprises generating the encoded image data using the encoder and at least based on the first data by encoding at least a portion of the image data using the quantization parameter, the portion of the image data being associated with the image section. [7] The method of claim 6, wherein the quantization parameter includes a maximum quantization parameter associated with the encoding of the image data. [8] A method according to any one of the preceding claims, wherein generating the coded image data comprises: Updating, using the encoder, one or more residual coefficient values associated with the image patch to be less than or equal to a threshold coefficient value; and after updating the one or more residual coefficient values, generating the encoded image data using the encoder by at least one of performing a transformation or quantization of the image data. [9] The method of any preceding claim, wherein determining the image portion associated with the type of content comprises determining, at least based on the image data representing the image, that the image portion represents sensitive information. [10] A method according to any one of the preceding claims, further comprising: Receiving image data generated by one or more image sensors of a machine navigating in an environment; Determining, at least based on the image data, one or more operations that the machine is to perform in the environment; and Causing the machine to perform one or more operations. [11] System comprising: one or more processing units configured to: Determining, at least based on image data representative of an image, that an image portion is associated with a type of content; Generating coded image data using an encoder and based on at least the image section associated with the type of content by encoding at least the image data such that the image section is obscured; and Causing the encoded image data to be sent to one or more computing devices. [12] The system of claim 11, wherein the one or more processing units further serve to: to generate initial data representing a boundary shape assigned to the image section, wherein the coded image data is further generated at least on the basis of the first data. [13] The system of claim 11 or claim 12, wherein the one or more processing units are further operable to: to generate a feature map associated with the image data, wherein generating the feature map comprises at least the following: Dividing the feature map into a plurality of blocks; causing a first portion of the plurality of blocks to indicate that the image section is associated with the type of content; and Causing a second portion of the plurality of blocks to indicate that a second image portion is not associated with the type of content, wherein the coded image data is further generated at least on the basis of the feature map. [14] The system of any one of claims 11 to 13, wherein generating the encoded image data further comprises encoding the image data using the encoder and based on at least a second image portion associated with a second type of content such that the second image portion is not obscured. [15] The system of any one of claims 11 to 14, wherein the one or more processing units further serve to: to determine a quantization parameter associated with the blurring of the image section, wherein generating the encoded image data comprises generating the encoded image data using the encoder and based on at least the image portion associated with the type of content by at least encoding a portion of the image data using the quantization parameter, the portion of the image data being associated with the image portion. [16] A system according to any one of claims 11 to 15, wherein generating the coded image data comprises: Updating, using the encoder, one or more residual coefficient values associated with the image patch to be less than or equal to a threshold coefficient value; and after updating the one or more residual coefficient values, generating the encoded image data using the encoder by at least one of performing a transformation or quantization of the image data. [17] The system of any of claims 11 to 16, wherein determining that the image portion is associated with the type of content comprises determining, at least based on the image data representing the image, that the image portion is associated with sensitive information. [18] System according to one of claims 11 to 17, wherein the system is comprised in at least one of the following elements: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing one or more simulation operations; a system for performing one or more digital twin operations; a system for performing light transport simulations; a collaborative content creation system for 3D assets; a system for performing one or more deep learning operations; a system implemented with an edge device; a system that is implemented with the help of a robot; a system for performing one or more generative AI operations; a system for performing operations using a large language model; a system for performing one or more AI conversational operations; a system for generating synthetic data; a system for displaying virtual reality, augmented reality or mixed reality content; a system that contains one or more virtual machines (VMs); a system that is at least partially implemented in a data center; or a system implemented at least in part using cloud computing resources. [19] Processor comprising: one or more processing units for generating coded image data by encoding a first portion of the image data to obscure a first image portion and a second portion of the image data to make a second image portion unrecognizable, wherein the encoding is based at least on the first image portion being associated with a first type of content and the second image portion being associated with a second type of content. [20] The processor of claim 19, wherein the processor is comprised in at least one of the following elements: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing one or more simulation operations; a system for performing one or more digital twin operations; a system for performing light transport simulations; a collaborative content creation system for 3D assets; a system for performing one or more deep learning operations; a system implemented with an edge device; a system that is implemented with the help of a robot; a system for performing one or more generative AI operations; a system for performing operations using a large language model; a system for performing one or more AI conversational operations; a system for generating synthetic data; a system for displaying virtual reality, augmented reality or mixed reality content; a system that contains one or more virtual machines (VMs); a system that is at least partially implemented in a data center; or a system implemented at least in part using cloud computing resources.
Citation Information
Patent Citations
US-PATENTANMELDUNGNR.16/101,232