Secure execution for multiple processor devices using a trusted execution environment.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- NVIDIA CORP
- Filing Date
- 2022-04-27
- Publication Date
- 2026-08-03
Smart Images

Figure 0007898922000005 
Figure 0007898922000006 
Figure 0007898922000007
Abstract
Description
Technical Field
[0001] The present invention relates to secure execution for multiple processor devices using a trusted execution environment.
Background Art
[0002] Parallel processing units (PPUs) such as graphics processing units (GPUs) have been increasingly improving in performance in recent years. With the improvement of PPU computing performance, users may not be able to fully utilize PPU resources in a single central processing unit (CPU) process. Furthermore, a multi-tenant environment in which computing resources including PPUs and CPUs can be utilized by service platforms and service providers to provide various services has become possible through virtualization.
Summary of the Invention
Problems to be Solved by the Invention
[0003] However, ensuring the security of multiple processing units (e.g., PPUs and CPUs) in a virtualized environment is extremely difficult, especially when multiple tenants utilize the same physical computing resources.
Brief Description of the Drawings
[0004] [Figure 1] A diagram showing an example of a trusted execution environment including multiple accelerators according to at least one embodiment. [Figure 2] A diagram showing an example of a copy operation within a trusted execution environment including multiple accelerators according to at least one embodiment. [Figure 3] A diagram showing an example of a copy operation within a trusted execution environment including multiple accelerators according to at least one embodiment. [Figure 4] A diagram showing an exemplary process for creating a trusted execution environment including multiple accelerators according to at least one embodiment. [Figure 5]This figure shows an exemplary process for terminating a trusted execution environment including multiple accelerators, according to at least one embodiment. [Figure 6] This figure shows an exemplary process for copying data within a trusted execution environment including multiple accelerators, according to at least one embodiment. [Figure 7] This figure shows an exemplary process for copying data within a trusted execution environment including multiple accelerators, according to at least one embodiment. [Figure 8A] This figure shows the inference and / or training logic according to at least one embodiment. [Figure 8B] This figure shows the inference and / or training logic according to at least one embodiment. [Figure 9] This figure shows the training and deployment of a neural network according to at least one embodiment. [Figure 10] This figure shows an exemplary data center system according to at least one embodiment. [Figure 11A] This figure shows an example of an autonomous vehicle, based on at least one embodiment. [Figure 11B] This figure shows an example of the camera location and field of view of the autonomous vehicle shown in Figure 11A, according to at least one embodiment. [Figure 11C] This is a block diagram illustrating an exemplary system architecture of an autonomous vehicle in at least one embodiment, as shown in Figure 11A. [Figure 11D] This figure shows a system for communication between a cloud-based server and the autonomous vehicle shown in Figure 11A, according to at least one embodiment. [Figure 12] A block diagram of a computer system according to at least one embodiment. [Figure 13] A block diagram of a computer system according to at least one embodiment. [Figure 14] This figure shows a computer system according to at least one embodiment. [Figure 15] This figure shows a computer system according to at least one embodiment. [Figure 16A] A diagram showing a computer system according to at least one embodiment. [Figure 16B] A diagram showing a computer system according to at least one embodiment. [Figure 16C] A diagram showing a computer system according to at least one embodiment. [Figure 16D] A diagram showing a computer system according to at least one embodiment. [Figure 16E] A diagram showing a shared programming model according to at least one embodiment. [Figure 16F] A diagram showing a shared programming model according to at least one embodiment. [Figure 17] A diagram showing an exemplary integrated circuit and associated graphics processor according to at least one embodiment. [Figure 18A] A diagram showing an exemplary integrated circuit and associated graphics processor according to at least one embodiment. [Figure 18B] A diagram showing an exemplary integrated circuit and associated graphics processor according to at least one embodiment. [Figure 19A] A diagram showing additional exemplary graphics processor logic according to at least one embodiment. [Figure 19B] A diagram showing additional exemplary graphics processor logic according to at least one embodiment. [Figure 20] A diagram showing a computer system according to at least one embodiment. [Figure 21A] A diagram showing a parallel processor according to at least one embodiment. [Figure 21B] A diagram showing a partition unit according to at least one embodiment. [Figure 21C] A diagram showing a processing cluster according to at least one embodiment. [Figure 21D] A diagram showing a graphics multiprocessor according to at least one embodiment. [Figure 22] A diagram showing a multi - graphics processing unit (GPU) system according to at least one embodiment. [Figure 23] A diagram showing a graphics processor according to at least one embodiment. [Figure 24] A block diagram showing a processor micro - architecture for a processor according to at least one embodiment. [Figure 25] A diagram showing a deep - learning application processor according to at least one embodiment. [Figure 26] A block diagram showing an exemplary neuromorphic processor according to at least one embodiment. [Figure 27] A diagram showing at least a portion of a graphics processor according to one or more embodiments. [Figure 28] A diagram showing at least a portion of a graphics processor according to one or more embodiments. [Figure 29] A diagram showing at least a portion of a graphics processor according to one or more embodiments. [Figure 30] A block diagram of a graphics processing engine of a graphics processor according to at least one embodiment. [Figure 31] A block diagram of at least a portion of a graphics processor core according to at least one embodiment. [Figure 32A] A diagram showing thread execution logic including an array of processing elements of a graphics processor core according to at least one embodiment. [Figure 32B] A diagram showing thread execution logic including an array of processing elements of a graphics processor core according to at least one embodiment. [Figure 33] A diagram showing a parallel processing unit ("PPU") according to at least one embodiment. [Figure 34] A diagram showing a general - purpose processing cluster ("GPC") according to at least one embodiment. [Figure 35] This figure shows a memory partition unit of a parallel processing unit ("PPU") according to at least one embodiment. [Figure 36] This figure shows a streaming multiprocessor according to at least one embodiment. [Figure 37] This is an example data flow diagram for an advanced computing pipeline, based on at least one embodiment. [Figure 38] This is a system diagram of an example system for training, adapting, instantiating, and deploying machine learning models in an advanced computing pipeline, according to at least one embodiment. [Figure 39] This figure includes an example of an advanced computing pipeline 3810A for processing imaging data, according to at least one embodiment. [Figure 40A] This figure includes an example data flow of a virtual instrument supporting an ultrasonic device, according to at least one embodiment. [Figure 40B] This figure includes an example data flow of a virtual device supporting a CT scanner, according to at least one embodiment. [Figure 41A] This is a data flow diagram of a process for training a machine learning model, according to at least one embodiment. [Figure 41B] This figure shows an example of a client-server architecture for extending annotation tools using a pre-trained annotation model, based on at least one embodiment. [Modes for carrying out the invention]
[0005] Embodiments of this disclosure provide novel solutions for executing user code or performing other operations in a virtualized environment, as described in more detail below, by leveraging a parallel processing unit (PPU), such as a graphics processing unit (GPU), to provide a secure execution environment. In the embodiments, the PPU is configured to operate within a trusted execution environment (TEE), which is at least partially implemented by the operation of one or more central processing units (CPUs). In one embodiment, an encrypted virtual machine running within the TEE is provided with access to the PPU by the hypervisor, but the data of applications running on the encrypted virtual machine is inaccessible to the hypervisor and / or a physical attacker. To achieve this, the PPU memory is protected using various techniques, as described in more detail below. In one embodiment, a protected memory region is created, which blocks the compute engine from writing outside the protected memory region when the compute engine of the PPU accesses one or more memory regions within that protected memory region. In another example, a write-protected memory area is created that blocks access to the PPU memory from other computing devices (for example, a CPU accessing the PPU memory via the system bus or other communication channels).
[0006] In various examples, a PPU includes integrated memory, one or more secure microcontrollers, and the private key of a public / private key pair. The public key of the public / private key pair may be provided by the PPU manufacturer and may be used to authenticate the PPU and / or prove information associated with the PPU. Furthermore, these PPUs may be included as hardware in server computer systems in data centers that provide computing resources to users across one or more networks. In such environments, optimizations to leverage the execution streams and computing resources of the PPU may be performed by the compiler. In one example, an execution stream includes a series of operations that are executed sequentially, where different execution streams can be executed simultaneously and in no particular order relative to other execution streams. The use of these execution streams can improve performance by at least duplicating memory copies and kernel execution. In various examples, PPUs available to virtual machines in a multi-tenant environment perform parallel execution of multiple execution streams. Furthermore, in such examples, the execution stream corresponds to either a Compute Unified Device Architecture ("CUDA") execution stream or an OpenCL (Open Computing Language) execution stream.
[0007] As will be explained in more detail below, to secure a TEE (e.g., an encrypted virtual machine or other secure environment) that includes a PPU, the virtual machine running within the TEE and the secure microcontroller of the PPU negotiate a shared key. In such examples, the secure microcontroller acts as the root of trust for the PPU within the TEE. Furthermore, direct memory access between the CPU (e.g., a virtual machine within the TEE) and the PPU can be secured using a shared key and a bounce buffer or similar non-secure memory area to transmit data. In some examples where data is transmitted from a virtual machine to a PPU, once the shared key is negotiated, the virtual machine uses the shared key to encrypt the data and stores the encrypted data in a memory area accessible to the PPU, and then the secure microcontroller retrieves the encrypted data, decrypts it using the shared key, and stores the result in a protected memory area of the PPU. Similarly, in an example where data is sent from a PPU to a virtual machine, a secure microcontroller encrypts the data from the PPU's protected memory area and stores it in a memory area accessible to the CPU. The virtual machine then retrieves the encrypted data, decrypts it using a shared key, and stores the decrypted data (e.g., data in plain text format) in the TEE.
[0008] Figure 1 shows an example of an environment 100 that includes a Trusted Execution Environment (TEE) 106 having access to multiple accelerators, according to at least one embodiment. In various embodiments, the server computer system includes a central processing unit (CPU) 102 and parallel processing units (PPUs), such as a graphics processing unit (GPU) 104. In one embodiment, the server computer system is located in a data center and is utilized to provide computing resources to users of a computing resource service provider. Furthermore, in the embodiment, the CPU 102 is used to implement the TEE 106, within which users can run various applications such as applications 108A-108N. Furthermore, in such an embodiment, the TEE 106 includes a guest operating system 112. In one embodiment, the TEE 106 includes a virtual machine, which can be encrypted and secured using encryption techniques such as Secure Encrypted Virtualization (SEV). In various embodiments, cryptographic material (e.g., an encryption key) is used to encrypt TEE 106 and data within a secure area 116 of system memory. In such embodiments, data stored in system memory 142 without encryption using the encryption key may be considered insecure 118 (e.g., accessible to the hypervisor 114 and / or other components of the service computer system). In various embodiments, the encryption key is used to isolate the guest operating system 112 from the hypervisor 114. Furthermore, in various examples, the encryption key used to isolate the guest operating system 112 from the hypervisor 114 is managed by the CPU 102 and is not exposed to or inaccessible to the hypervisor 114.
[0009] In various embodiments, the server computer system includes system memory 142 and a system bus 120. For example, system memory 142 includes various types of memory, such as volatile memory or non-volatile memory. Furthermore, system memory 142 may include semiconductor memory (e.g., included in the CPU 102) or separate hardware, such as random access memory (RAM). In embodiments, the system bus 120 includes computer hardware that connects or provides communication channels to components of the server computer system, such as the CPU 102 and the GPU 104. In one example, system bus 102 includes a Peripheral Component Interconnect Express (PCIe) bus. Furthermore, in various embodiments, the server computer system runs a hypervisor 114. In one example, the hypervisor enables virtualization and provides access to the underlying hardware of the server computer system, such as the CPU 102, GPU 104, system memory 142, and other components not shown in Figure 1 for clarity, such as network interfaces, storage devices, or other physical hardware. In one example, the hypervisor provides access to various devices included in or accessible to the server. For example, the hypervisor can provide access to the GPU 104 to the TEE 106 by providing virtualization of at least the GPU 104. In various embodiments, the devices include the CPU 102, the GPU 104, storage devices, network devices, or any other devices that can be included in the TEE 106.
[0010] As shown in Figure 1, in various embodiments, components within TEE106 (e.g., applications 108A-108N, guest operating systems 112, or other executable logic) can communicate with GPU104 via system bus 120 using driver 110. In various embodiments, driver 110 includes source code or other executable logic that causes the server computer system to perform various operations, which, as a result of being executed by CPU 102, make the computing resources of GPU104 available to applications 108A-108N or other components within TEE106. In one example, driver 110 enables a machine learning application running within TEE106 to utilize GPU104 to perform inference operations. Furthermore, in various embodiments, the GPU 104 includes a GPU trust boundary 126, one or more computing engines 128A to 128C, a secure processor 132, firmware 140, MPRMPR (memory protection area) 130, memory 134, and protective memory 136.
[0011] In various embodiments, in order to include GPU104 in TEE106 (for example, to allow guest OS112 and / or applications 108A-108N to access GPU104), GPU104 is placed in secure mode (for example, enclave mode). Furthermore, in such embodiments, GPU104 must be idle (for example, no other virtual machines and / or tenants are accessing or utilizing GPU104). For example, when GPU104 is idle, the system blocks access to GPU104 to allow GPU104 to enter secure mode. In embodiments, a secure processor 132 prevents access to GPU104. In various embodiments, the secure processor 132 is a microcontroller containing microcode embedded in GPU104. Furthermore, in various embodiments, the secure processor 132 initializes GPU104 in secure mode. For example, a secure processor 132 creates or causes a protected memory area 136 and / or a GPU trust boundary 126. In various embodiments, the GPU trust boundary 126 includes logical representations of secure components contained in the GPU 104 (e.g., data protected by cryptographic keys, components protected from unauthorized access, etc.).
[0012] In various embodiments, protected memory 136 comprises the majority of the GPU 104 memory, leaving a smaller portion of memory 134 unprotected. In one example, when the GPU 104 enters secure mode, 95 percent of the GPU 104 memory is initialized as protected memory 136, and 5 percent of memory 134 remains unprotected. In various embodiments, protected memory 136 includes model weights for machine learning algorithms, inference results, source data, or any other data that the user wishes to protect. In other embodiments, memory 134 includes internal data structures of the driver 110, semaphores, or other data used by the system to perform the operations described herein. Furthermore, in various embodiments, the driver 110 includes executable code that, as a result of being executed (for example, by a virtual processor in TEE 106), causes data to be stored in the protected memory 136 of the GPU 104. For example, as a result of application 108A calling a specific function of driver 110, encrypted data is transferred to a secure processor 132 via the system bus 120 using buffer 144, decrypted, and stored in protected memory 136.
[0013] In various embodiments, the system bus 120 includes a virtual bus. In one embodiment, the virtual bus includes a physical function (PF) of the system bus 120 (e.g., PCI Express (PCIe)), where the PF is a function of the GPU that supports a Single Root I / O Virtualization (SR-IOV) interface. Furthermore, in such embodiments, the PF includes an SR-IOV extension in the PCIe configuration space, which is used to configure and manage the GPU's SR-IOV capabilities, such as enabling virtualization and exposing PCIe virtual functions (VFs). In various embodiments, the PF is exposed as a virtual GPU in the management operating system of the hypervisor's parent partition.
[0014] In the embodiment, the MPR130 includes a hardware block of the GPU104 that manages the protected memory 136. In one embodiment, the MPR130 works together with the memory management unit of the GPU104 to control access to the protected memory. In various embodiments, when a computer engine (e.g., compute engines 128A-128C) accesses the protected memory 136 (e.g., writes to and / or reads from a memory region associated with the protected memory 136), the MPR130 manages access to the protected memory 136 so that the compute engine is unable to access and / or is prevented from accessing any other memory region, such as system memory 142 or memory 134. For example, when a particular compute engine accesses the protected memory 136, the MPR130 and / or the memory management unit evaluate the memory request from that particular compute engine and output an error if that particular compute engine is attempting to access memory outside of the protected memory 136.
[0015] As shown in Figure 1 with the key symbol, in various embodiments, the driver 110 creates a shared secret (e.g., an encryption key) with the GPU 104. In one embodiment, the GPU 104 contains a private key that is either burned into a fuse by the manufacturer or stored in the device hardware, and a public key corresponding to the private key can be issued by the manufacturer. In the embodiment, the driver 110 generates the shared secret by performing a Security Protocol and Data Model (SPDM) key exchange with the GPU 104 (e.g., via the CPU 102). Furthermore, in various embodiments, the driver 110 generates the shared secret by having a user of the system obtain the public key (e.g., by requesting the public key from the GPU 104 or another entity such as a server operated by the manufacturer) and using that public key (e.g., using the Diffie-Hellman key exchange algorithm). In one embodiment, the shared secret is a symmetric encryption key. Furthermore, in various embodiments, the shared secret is held in the TEE 106 and the secure processor 132. As will be explained in more detail below in relation to Figures 2, 3, 6, and 7, the data exchanged between the CPU 102 and the GPU 104 (for example, using buffer 144) is encrypted or protected by a shared secret.
[0016] In various embodiments, when copying data (for example, data stored in a secure memory area 116), the driver 110 causes the system to retrieve the data from the secure memory area 116, encrypt the data with a shared secret, and store the encrypted data in a buffer 144. In various embodiments, the buffer 144 includes a non-secure memory area 118. For example, the buffer 144 is a memory area accessible to a secure processor 132 via the system bus 120 or its components. Furthermore, in various embodiments, the data in the secure memory area 116 is encrypted and / or protected from access by the CPU 102, but when the data in the secure memory area 116 is retrieved by a component in TEE 106, the data is decrypted. As a result, in such embodiments, the driver 110 or other component in TEE 106 encrypts or protects the data before it is transmitted outside of TEE 106 (for example, via the system bus 102 using buffer 144). Returning to the example above, once the driver 110 encrypts the data and stores the encrypted data in the buffer 144, the secure processor 132 retrieves the encrypted data (for example, via the system bus 120), decrypts the data using a shared secret, and stores the plaintext data in the protected memory 136. In such an embodiment, the data transmitted via the system bus 120 is encrypted and as a result is protected from attacks by intruders, etc.
[0017] Similarly, when sending data from the GPU 104's protected memory 136 to the TEE 106, in various embodiments, the secure processor 132 encrypts the data to generate encrypted data and copies the encrypted data to an insecure memory area 118 of the system memory 142 using buffer 144 (for example, sending the data via the system bus 120). In response, in such embodiments, the driver 110 retrieves the encrypted data from buffer 144, decrypts the encrypted data using a shared key, thereby converting the data into plaintext, which becomes accessible to one or more components within the TEE 106.
[0018] In various embodiments, other components of the secure processor 132 or GPU 104 (such as a memory management unit) prevent the CPU 102 or its components from accessing the protected memory 136. In contrast, in various embodiments, the CPU 102 or its components can access the memory 134. In one example, the CPU 102 and / or driver 110 copy insecure data (e.g., executable code, kernel, data structures, etc.) to the memory 134.
[0019] Figure 2 shows an environment 200 in which a copy operation is performed within a trusted execution environment including multiple accelerators, according to at least one embodiment. As shown in Figure 2, data is copied from GPU memory 204 to CPU memory 202 within a trusted execution environment (TEE), such as the one described above in relation to Figure 1. In the embodiment, GPU memory 204 is used by the GPU, PPU, or other accelerator to store data used by the accelerator during the execution of source code or other executable instructions. In one example, GPU memory 204 is integrated with a GPU 236 operating in a secure execution mode (e.g., enclave mode). In various embodiments, GPU memory includes a non-secure region (e.g., accessible to the CPU or other accelerators) and a secure region (e.g., inaccessible to the CPU or other accelerators).
[0020] In the embodiment, a memory copy between a secure region 220 of GPU memory 204 and CPU memory 202 and / or other system memory is encrypted and transmitted over the bus via a bounce buffer. In one example, memory copy 238 contains data from output buffer 226 in the secure region 220 of GPU memory 204 and is encrypted by a secure processor 232. As described above, in various embodiments, an entity in the TEE (not shown in Figure 2 for clarity) (e.g., a driver, application, guest operating system, etc.) retrieves the encrypted result 210 from the non-secure region 206 of CPU memory 202, decrypts the encrypted result 210, and copies the result 212 (e.g., data in plain text format) to a secure region 208 of CPU memory 202. In one example, the non-secure region 206 of CPU memory 202 contains a memory region that stores data in plain text format or unencrypted by the TEE or an entity associated with the TEE (e.g., a virtual processor of the TEE). In another example, the secure region 208 of CPU memory 202 includes a memory area where data is encrypted before being stored.
[0021] In various embodiments, the non-secure region of the GPU memory 214 includes a driver data structure 216 and a kernel 218. For example, the driver data structure 216 includes data used by a driver loaded into the TEE to enable the use of the GPU 236 by an application running within the TEE. In one example, the kernel 218 includes a CUDA kernel used during processing operations performed by the GPU 236. In various embodiments, the secure processor 232 includes a secure microcontroller or other processor integrated into the GPU 236, such as the secure processor 130 described above in relation to Figure 1. In various embodiments, the secure processor 232 generates an encryption key for encrypting data in the output buffer 226 and generates the encrypted result 210. In an embodiment, the compute engine 234 generates data to be stored in the output buffer 226. For example, the compute engine 234 executes source code or other instructions and puts the result into the output buffer 226. In various embodiments, the compute engine 234 includes a compute engine such as the one described above in relation to Figure 1.
[0022] In some embodiments, the secure processor 232 generates and / or negotiates a cryptographic key with another accelerator (e.g., a CPU) based at least in part on cryptographic material stored in the GPU 236. In one example, the secure processor 232 generates a symmetric key based at least in part on a secret key stored in the GPU 236 using the Diffie-Hellman shared secret generation algorithm. Furthermore, in various embodiments, the secure processor 232 generates proof that the GPU 236 is operating in a secure execution mode. For example, the secure processor 232 retrieves data associated with the GPU 236 and signs the data with a secret key stored in the GPU 236. In yet another embodiment, the secure processor 232 generates information usable by the TEE or its entities to authenticate the GPU 236, ensuring the security of the TEE when adding the GPU 236 to the TEE. In an embodiment, the service provider (for example, a computing resource service provider that provides computing resources to run the TEE) and / or the manufacturer of the GPU236 provides additional information to authenticate and / or certify the GPU236. In one example, the service provider provides a list of GPUs connected to the server running the TEE.
[0023] In some embodiments, cryptographic material (e.g., public and private keys associated with GPU236) is stored in a read-only memory device, such as a fuse block, within GPU236. In various embodiments, the cryptographic material is written to secure write-once memory (e.g., a fuse block), so that once the data is written to secure write-once memory, it cannot be rewritten or modified. In one example, the cryptographic material is stored within GPU236, and therefore the public key is accessible to various components of the server (e.g., the CPU), while the private key is accessible only to the secure processor 232. In other words, in such an example, access to the private key associated with GPU236 is blocked for all entities except the secure processor 232 of GPU236.
[0024] In various embodiments, once an encryption key (e.g., cryptographic material used to encrypt data and generate the encryption result 210) is generated, only the secure processor 232 can access it. For example, the memory management unit of the GPU 236 prevents all entities that are not the secure processor 232 from accessing the memory area where the encryption key is stored. In another embodiment, the secure processor 232 includes memory that is accessible only to the secure processor 232. Furthermore, in various embodiments, the secure processor 232 manages the process of putting the GPU 236 into a secure execution mode (e.g., process 400, which is described in more detail below in relation to Figure 4). In such embodiments, the generation of a shared encryption key is part of the process of incorporating the GPU 236 into the TEE. In such embodiments, the hypervisor writes data to a memory location in a PROM (e.g., non-volatile memory) attached to the GPU 236 via an out-of-band (OOB) channel, indicating that the GPU 236 will operate in a secure execution mode after a reset of the GPU 236. Furthermore, when a secure virtual machine running on GPU236 terminates, the hypervisor writes data to PROM again to indicate that GPU236 will exit secure execution mode at the next reset. In various embodiments, GPU236 includes non-volatile memory for storing data indicating that GPU236 is capable of generating a TEE, data indicating that GPU236 is operating within a TEE, data indicating that GPU236 is exiting a TEE, or other information associated with GPU236.
[0025] Figure 3 shows an environment 300 in which a copy operation is performed between accelerators in a trusted execution environment including multiple accelerators, according to at least one embodiment. As shown in Figure 3, in a trusted execution environment (TEE) such as the one described above in relation to Figure 1, data is copied from CPU memory 302 to GPU memory 304. In embodiments, CPU memory 302 is used by the CPU, PPU, or other accelerators to store data used by accelerators during the execution of source code or other executable instructions. In various embodiments, CPU memory 302 includes a non-secure area 306 (e.g., accessible to the GPU or other accelerators) and a secure area 308 (e.g., inaccessible to the GPU or other accelerators).
[0026] In the embodiment, a memory copy between a secure region 308 of CPU memory 302 and GPU memory 304 is encrypted and transmitted over the bus via a bounce buffer. In one embodiment, memory copy 338 includes the secure region 308 of CPU memory 302 and encrypted model weights 342 obtained from an entity in the TEE (e.g., a guest operating system or its components). In the embodiment shown in Figure 3, model weights 340 for a machine learning algorithm or other artificial intelligence (AI) are used, but arbitrary data may be transferred between CPU memory 302 and GPU memory 304 using the environment 300 and its components. As described above, in various embodiments, an entity in the TEE (not shown in Figure 3 for clarity) (e.g., a driver, application, guest operating system, etc.) encrypts the model weights 342 obtained from the secure region 308 of CPU memory 302 and stores the encrypted model weights 342 in a non-secure region 306 of CPU memory 302 (e.g., an unencrypted memory region accessible to GPU 336). In one example, the non-secure region 306 of CPU memory 302 includes a memory area that stores data in plain text format or in an unencrypted state by the TEE or an entity associated with the TEE (e.g., the TEE's virtual processor). Furthermore, in various embodiments, the non-secure region 306 of CPU memory 302 includes a bounce buffer for transmitting data between the TEE and the GPU 336 via the system bus.
[0027] In various embodiments, an insecure region 306 of CPU memory 302 includes a driver data structure 310 and a kernel 318. For example, the driver data structure 310 includes data used by the compute engine 334 to perform inference using model weights 340. In one example, the kernel 318 includes a CUDA kernel used during processing operations performed by the compute engine 334 of the GPU 336. In embodiments, the driver data structure 310 and kernel 318 are transmitted to GPU memory 304 via the system bus in an unencrypted state. In yet another embodiment, the driver data structure 310 and kernel 318 are encrypted before transmission. Furthermore, as shown in Figure 3, according to at least one embodiment, the driver data structure 310 and kernel 318 are stored in an insecure region 314 of GPU memory 304. The non-secure regions of GPU memory 304 include memory regions of GPU memory 304 that are accessible to other accelerators or are not protected (e.g., not encrypted and / or read / write protected).
[0028] In various embodiments, the secure processor 332 includes a secure microcontroller or other processor integrated with the GPU 336, such as the secure processor 130 described above in relation to Figure 1. As described above, in various embodiments, the secure processor 332 generates an encryption key shared with the TEE to encrypt the model weights 340. Furthermore, in various embodiments, the secure processor 332 retrieves the encrypted model weights 342 from an insecure area 306 of the CPU memory 302 and performs a memory copy 338 by at least copying the encrypted model weights 342 over the system bus, decrypting the encrypted model weights 342, and storing the result (e.g., model weights 340) in a secure area 320 of the GPU memory 304. In various embodiments, one or more methods are used to secure the secure area 320 of the GPU memory 304.
[0029] In one embodiment, when a compute engine 334 accesses a secure region 320 of GPU memory 304, access to the secure region 320 of GPU memory 304 is restricted so that data writing by the compute engine 334 outside the secure region 320 of GPU memory 304 is blocked. In another embodiment, read and write access to the secure region 320 of GPU memory 304 by one or more accelerators (e.g., CPU, GPU, and / or PPU) is blocked or prevented. For example, the memory management unit of GPU 336 prevents access to the secure region 320 of GPU memory 304 by preventing access at least via the system bus. In such an example, the memory management unit prevents access to the secure region 320 of GPU memory 304 based at least in part on a hardware identifier or other identifier of the entity attempting to access the secure region 320 of GPU memory 304. In various embodiments, the memory management unit returns an error or other error if an unauthorized entity attempts to access a secure region 320 of GPU memory 304. In one embodiment, the unauthorized entity includes any entity other than the secure processor 332 and / or a particular computing engine that has access to the secure region 320 of GPU memory 304. In such an embodiment, the memory management unit issues an error if the particular computing engine attempts to access a memory region that is not within the secure region 320 of GPU memory 304.
[0030] Figure 4 shows a process 400 for creating a trusted execution environment including multiple accelerators, according to at least one embodiment. Referring to Figures 4-7, it should be understood that this and other configurations described herein are described merely as examples. Other configurations and elements (e.g., machines, interfaces, functions, sequences, groups of functions, etc.) may be used in addition to or instead of those shown, and some elements may be executed in parallel or omitted collectively. Furthermore, many of the elements described herein are functional entities, and these functional entities may be implemented as separate or distributed components or in combination with other components in any preferred combination and location. The various functions described herein executed by entities may be executed by hardware, firmware, and / or software. For example, the various functions may be executed by a processor that executes instructions stored in memory. Furthermore, the various functions shown in Figures 4-7 may be executed in various orders (e.g., serial or parallel) or omitted collectively.
[0031] Referring here to Figure 4, each block of the method 400 described herein includes a computation process that may be executed using any combination of hardware, firmware, and / or software. For example, various functions may be executed by a processor that executes instructions stored in memory. This method may also be embodied as computer-available instructions stored on a computer storage medium. This method may be provided, to a few examples, by a standalone application, a service or host service (standalone or in combination with another host service), or by plugging into another product. Furthermore, method 400 is described in relation to TEE106 in Figure 1 as an example. However, these methods may be performed additionally or alternatively by any one system, or any combination of systems, including but not limited to those described herein.
[0032] Figure 4 is a flowchart illustrating a process 400 for providing a PPU (such as a GPU) to a TEE to run an application within the TEE, according to some embodiments of the present disclosure. Process 400 includes a step in block 402 to put the GPU into an idle state. As described above, before operating the GPU in secure execution mode, the system running process 400 ensures that no virtual machine or other application is using the GPU. In embodiments, a secure virtual machine requests hypervisor-mediated access to a GPU accessible to the server running the secure virtual machine.
[0033] In block 404, the system running process 400 writes a request to enable the GPU to operate within the TEE. In various embodiments, in order to include the GPU in the TEE, the GPU operates in a secure execution mode (e.g., enclave mode). As described above, the secure execution mode of the GPU enables various data protection features of the GPU by restricting or blocking access to at least one or more memory regions of the GPU memory. In one embodiment, a computing engine accessing a protected memory region of the GPU memory cannot write to any other memory region. In another embodiment, the GPU's memory management unit prevents access to the GPU's protected memory region from entities (e.g., the CPU or other accelerators) via the system bus. In an embodiment, the system running process 400 writes data to the GPU's non-volatile memory to indicate that the GPU will enter a secure execution mode when reset.
[0034] In block 406, the system running process 400 resets the GPU. In one example, the system running process 400 resets the GPU to put it into a secure execution mode and begin the process of handing the GPU over to the TEE. In block 408, once the GPU is reset, the system running process 400 causes the GPU to implement a protected memory region. As described above, the secure processor of the GPU causes the protected memory region to implement in this embodiment. In one example, the GPU's computing engine can access the protected memory region, but the protected memory region is managed by a memory management unit so that after accessing the protected memory region, writing to other memory regions is not possible.
[0035] In block 410, the system running process 400 implements write-protected memory regions. For example, a memory management unit prevents a specific entity from writing data to a memory region that corresponds to a protected memory region. In yet another example, write access via the system bus is blocked. In block 412, the system running process 400 hands over the GPU to a virtual machine. In one example, the hypervisor uses various virtualization techniques to grant the virtual machine access to the GPU. In various embodiments, the virtual machine and / or its components (e.g., application software, guest operating system, drivers, etc.) perform one or more checks to identify that the GPU is operating in a secure execution mode and authenticate the GPU.
[0036] In block 414, the system running process 400 performs a shared key generation operation. In various embodiments, the secure processor of the GPU generates a shared cryptographic key with the TEE. As described above, the shared cryptographic key is used to encrypt data so that it can be transmitted between the CPU and the GPU. In block 416, the system running process 400 runs an application using the GPU at least partially. For example, the system running process 400 runs a machine learning model that performs inference using the GPU.
[0037] Figure 5 shows a process 500 for excluding a GPU from a Trusted Execution Environment (TEE) that includes multiple accelerators, according to at least one embodiment. Process 500 includes a step in block 502 to detect the exit of a virtual machine. In various embodiments, the TEE includes a virtual machine that runs an application for a tenant, and after the application has finished running, or at any point in time, the tenant terminates the execution of the virtual machine. Terminating a virtual machine includes, for example, releasing the physical hardware used to support the virtual machine. In an embodiment, the hypervisor indicates to the GPU that the virtual machine has terminated execution. In another embodiment, the GPU or its components, such as a secure processor, detect that the virtual machine has terminated. For example, when the virtual machine terminates, the TEE is destroyed and the GPU is isolated from the virtual machine and / or the TEE.
[0038] In block 504, the system running process 500 writes a request to disable secure execution mode on the GPU. In one embodiment, writing this request causes the GPU to exit secure execution mode upon the next reset. As described above, secure execution mode protects GPU memory from unauthorized access in various embodiments. In one embodiment, the request (e.g., data indicating the operating mode to the GPU) is recorded in GPU memory (e.g., non-volatile memory (PROM) attached to the GPU). In block 506, the system running process 500 resets the GPU. For example, the secure processor causes the GPU to perform a soft reset. In block 508, after the GPU has been reset, the system running process 500 erases the GPU memory. In one embodiment, all contents of the GPU memory are erased. In another embodiment, the secure processor erases the contents of the GPU memory and the shared encryption key after the reset to ensure that data generated within the TEE is not exposed. In yet another embodiment, the secure processor erases the contents of the GPU's protected memory area.
[0039] Figure 6 shows a process 600 for copying data from CPU memory to GPU memory in a trusted execution environment (TEE) including multiple accelerators, according to at least one embodiment. Process 600 includes the step of encrypting the data with a shared encryption key in block 602. In various embodiments, an application is run by a virtual machine in the TEE, and this application provides data to the GPU for processing (e.g., inference, video processing, audio processing, rendering, or other operations that can be performed by the GPU). Furthermore, in such embodiments, the secure processor of the GPU negotiates a shared key with the virtual machine to encrypt the data being transmitted between the virtual machine and the GPU. For example, the virtual machine or its components (e.g., the driver 110 described above in Figure 1) encrypt the data to be copied to GPU memory.
[0040] In block 604, the system running process 600 stores encrypted data in an insecure memory region accessible to the GPU. For example, the encrypted data is stored in a bounce buffer in the memory of the server running the virtual machine. In such an example, the insecure memory region is accessible to the GPU via the system bus, network interface, or other communication channel. In block 606, the system running process 600 sends the encrypted data to a secure processor on the GPU. In one example, the secure processor copies the encrypted data from memory. In another example, the virtual machine or its components ensure that the encrypted data is copied to the secure processor via the system bus.
[0041] In block 608, the system running process 600 decrypts the encrypted data using a shared encryption key. For example, a secure processor holds a copy of the shared encryption key and uses that key to decrypt the data. In block 610, the system running process 600 stores the plaintext data in a protected memory region within GPU memory. As described above, when the GPU operates in secure execution mode, it creates a protected memory region so that the plaintext data in the protected memory region is inaccessible to unauthorized entities.
[0042] Figure 7 shows a process 700 for copying data from GPU memory to CPU memory in a Trusted Execution Environment (TEE) containing multiple accelerators, according to at least one embodiment. Process 700 includes a step in block 702 in which the GPU's secure processor encrypts the data with a shared encryption key. As described above, in at least one embodiment, during the process of including the GPU in the TEE, the GPU's secure processor generates a shared encryption key with the drivers included in the TEE. For example, the shared encryption key is used to encrypt the data before it is transmitted over a system bus that is otherwise protected during transmission between accelerators included in the TEE. In the embodiment, the data is stored in a protected memory area of the GPU memory.
[0043] In block 704, the system running process 700 stores the encrypted data in an insecure area of memory accessible to the TEE or the CPU running its components. For example, a secure processor encrypts the data and then transmits it via the system bus so that the data is stored in an insecure memory area, such as the insecure memory area 118 described above in relation to Figure 1. In block 706, the system running process 700 decrypts the encrypted data with a shared encryption key. For example, a driver in the TEE stores the encryption key and causes the guest operating system, a driver, or another component of the TEE to open the encrypted data from the insecure memory area and decrypt the encrypted data, at least partially based on the shared encryption key. In block 708, the system running process 700 provides the decrypted data to the TEE. In one example, the encrypted data includes the result of an operation performed by the GPU. Furthermore, in various embodiments, providing the decrypted data to the TEE includes ensuring that the plaintext data is stored in a secure memory area of CPU memory protected by the TEE.
[0044] Logic of reasoning and training Figure 8A shows the inference and / or training logic 815 used to perform inference and / or training operations with respect to one or more embodiments. Further details regarding the inference and / or training logic 815 are provided below in conjunction with Figures 8A and / or 8B.
[0045] In at least one embodiment, the inference and / or training logic 815 may include, without limitation, code and / or data storage 801 for storing forward and / or output weights and / or input / output data and / or other parameters for constituting neurons or layers of a neural network that are trained and / or used to infer in one or more embodiments. In at least one embodiment, the training logic 815 may include, or be coupled to, code and / or data storage 801 for storing graph code or other software for controlling timing and / or sequence, the code and / or data storage 801 being loaded with weight and / or other parameter information to constitute logic including integer and / or floating-point units (collectively, arithmetic logic units (ALUs)). In at least one embodiment, the code, such as graph code, loads weight or other parameter information into the processor ALU based on the architecture of the neural network to which such code corresponds. In at least one embodiment, the code and / or data storage 801 stores the weight parameters and / or input / output data of each layer of the neural network being trained or used in conjunction with one or more embodiments while forward-propagating the input / output data and / or weight parameters during training and / or inference using one or more embodiments. In at least one embodiment, any portion of the code and / or data storage 801 may be included together with other on-chip or off-chip data storage, including L1, L2, or L3 caches of the processor or system memory.
[0046] In at least one embodiment, any portion of the code and / or data storage 801 may be inside or outside one or more processors or other hardware logic devices or circuits. In at least one embodiment, the code and / or code and / or data storage 801 may be cache memory, dynamic randomly addressable memory ("DRAM"), static randomly addressable memory ("SRAM"), non-volatile memory (e.g., flash memory), or other storage. In at least one embodiment, the choice of whether the code and / or code and / or data storage 801 is inside or outside a processor, or whether it includes DRAM, SRAM, flash, or any other type of storage, may depend on the storage available on-chip versus off-chip, the latency requirements of the training and / or inference functions being performed, the batch size of the data used for neural network inference and / or training, or any combination of these factors.
[0047] In at least one embodiment, the inference and / or training logic 815 may include, without limitation, code and / or data storage 805 for storing backpropagated and / or output weights and / or input / output data corresponding to neurons or layers of a neural network used to train and / or infer in one or more embodiments. In at least one embodiment, the code and / or data storage 805 stores the weight parameters and / or input / output data of each layer of the neural network used to train or in conjunction with one or more embodiments while backpropagating the input / output data and / or weight parameters during training and / or inference using one or more embodiments. In at least one embodiment, the training logic 815 may include, or be coupled to, code and / or data storage 805 for storing graph code or other software for controlling timing and / or sequence, the code and / or data storage 805 being loaded with weight and / or other parameter information to constitute logic including integer and / or floating-point units (collectively referred to as arithmetic logic units (ALUs)).
[0048] In at least one embodiment, the code, such as graph code, causes the processor ALU to load weight or other parameter information based on the architecture of the neural network to which the code corresponds. In at least one embodiment, any portion of the code and / or data storage 805 may be included with other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache, or system memory. In at least one embodiment, any portion of the code and / or data storage 805 may be inside or outside one or more processors, or other hardware logic devices or circuits. In at least one embodiment, the code and / or data storage 805 may be cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other storage. In at least one embodiment, the choice of whether the code and / or data storage 805 is, for example, internal or external to the processor, or whether it includes DRAM, SRAM, flash memory, or any other type of storage, may be determined depending on the storage available on-chip versus off-chip, the latency requirements of the training and / or inference functions performed, the batch size of the data used for neural network inference and / or training, or any combination of these factors.
[0049] In at least one embodiment, the code and / or data storage 801 and the code and / or data storage 805 may be separate storage structures. In at least one embodiment, the code and / or data storage 801 and the code and / or data storage 805 may be a combined storage structure. In at least one embodiment, the code and / or data storage 801 and the code and / or data storage 805 may be partially combined and partially separate. In at least one embodiment, any portion of the code and / or data storage 801 and the code and / or data storage 805 may be included together with other on-chip or off-chip data storage, including L1, L2, or L3 caches of the processor or system memory.
[0050] In at least one embodiment, the inference and / or training logic 815 may include, without limitation, one or more arithmetic logic units ("ALUs") 810, including integer and / or floating-point units, for performing logical and / or arithmetic operations that are at least partially based on or shown by training and / or inference code (e.g., graph code), the results of which may produce activations (e.g., output values from layers or neurons in a neural network) stored in activation storage 820, which are functions of input / output and / or weight parameter data stored in code and / or data storage 801 and / or code and / or data storage 805. In at least one embodiment, the activation stored in the activation storage 820 is generated according to linear algebra calculations and / or matrix-based calculations performed by the ALU 810 in response to the execution of an instruction or other code, where the weight values stored in the code and / or data storage 805 and / or data 801 are used as operands along with other values such as bias values, gradient information, momentum values, or other parameters or hyperparameters, any or all of which may be stored in the code and / or data storage 805, or the code and / or data storage 801, or other on-chip or off-chip storage.
[0051] In at least one embodiment, the ALU 810 is contained within one or more processors or other hardware logic devices or circuits, while in another embodiment, the ALU 810 may be outside of the processors or other hardware logic devices or circuits that use them (e.g., a coprocessor). In at least one embodiment, the ALU 810 may be contained within an execution unit of a processor, or otherwise contained within an ALU bank accessible by execution units of a processor, either within the same processor or distributed among different types of processors (e.g., a central processing unit, graphics processing unit, fixed-function unit, etc.). In at least one embodiment, the code and / or data storage 801, code and / or data storage 805, and activation storage 820 may share a processor or other hardware logic devices or circuits, while in another embodiment, they may be in different processors or other hardware logic devices or circuits, or in any combination of the same processor or other hardware logic devices or circuits and different processors or other hardware logic devices or circuits. In at least one embodiment, any portion of the activated storage 820 may be included with other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache, or system memory. Furthermore, the inference and / or training code may be stored with other code accessible to the processor or other hardware logic or circuitry, and may be fetched and / or processed using the processor's fetch, decode, schedule, execute, retire, and / or other logic circuits.
[0052] In at least one embodiment, the activated storage 820 may be cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other storage. In at least one embodiment, the activated storage 820 may be entirely or partially inside or outside one or more processors or other logic circuits. In at least one embodiment, the choice of whether the activated storage 820 is, for example, inside or outside a processor, or whether it includes DRAM, SRAM, flash memory, or any other type of storage, may be determined depending on the available on-chip versus off-chip storage, the latency requirements of the training and / or inference functions being performed, the batch size of the data used for neural network inference and / or training, or any combination of these factors.
[0053] In at least one embodiment, the inference and / or training logic 815 shown in Figure 8A may be used in conjunction with an application-specific integrated circuit ("ASIC") such as a TensorFlow® processing unit from Google, an inference processing unit (IPU) from Graphcore®, or a Nervana® (e.g., "Lake Crest") processor from Intel Corp. In at least one embodiment, the inference and / or training logic 815 shown in Figure 8A may be used in conjunction with other hardware such as a central processing unit ("CPU"), a graphics processing unit ("GPU"), or a field-programmable gate array ("FPGA").
[0054] Figure 8B shows an inference and / or training logic 815 according to at least one embodiment. In at least one embodiment, the inference and / or training logic 815 may include, without limitation, hardware logic in which computational resources are dedicated to, or otherwise used in conjunction with, weight values or other information corresponding to one or more layers of neurons in a neural network. In at least one embodiment, the inference and / or training logic 815 shown in Figure 8B may be used in conjunction with an application-specific integrated circuit (ASIC), such as a TensorFlow® processing unit from Google, an Inference Processing Unit (IPU) from Graphcore®, or a Nervana® (e.g., “Lake Crest”) processor from Intel Corporation. In at least one embodiment, the inference and / or training logic 815 shown in Figure 8B may be used in conjunction with other hardware, such as a central processing unit (CPU) hardware, a graphics processing unit (“GPU”) hardware, or a field-programmable gate array (FPGA). In at least one embodiment, the inference and / or training logic 815 may include, without limitation, code and / or data storage 801 and code and / or data storage 805, which may be used to store other information including code (e.g., graph code), weight values and / or bias values, gradient information, momentum values and / or other parameter or hyperparameter information. In at least one embodiment shown in Figure 8B, each of the code and / or data storage 801 and code and / or data storage 805 is associated with dedicated computing resources such as computing hardware 802 and computing hardware 806, respectively. In at least one embodiment, each of the computing hardware 802 and computing hardware 806 includes one or more ALUs that execute mathematical functions, such as linear algebraic functions, only on the information stored in the code and / or data storage 801 and code and / or data storage 805, respectively, and the results are stored in activation storage 820.
[0055] In at least one embodiment, the code and / or data storage 801 and 805, respectively, and the corresponding computing hardware 802 and 806, each correspond to different layers of a neural network, and the activation resulting from one storage / computation pair 801 / 802 with the code and / or data storage 801 and computing hardware 802 is provided as input to the next storage / computation pair 805 / 806 with the code and / or data storage 805 and computing hardware 806 to reflect the conceptual organization of the neural network. In at least one embodiment, the storage / computation pairs 801 / 802 and 805 / 806 may correspond to two or more layers of a neural network. In at least one embodiment, additional storage / computation pairs (not shown) may be included in the inference and / or training logic 815 after or in parallel with the storage / computation pairs 801 / 802 and 805 / 806.
[0056] Training and implementation of neural networks Figure 9 shows the training and deployment of a deep neural network in at least one embodiment. In at least one embodiment, an untrained neural network 906 is trained using a training dataset 902. In at least one embodiment, the training framework 904 is the PyTorch framework, while in other embodiments, the training framework 904 is TensorFlow, Boost, Caffe, Microsoft Cognitive Toolkit / CNTK, MXNet, Chainer, Keras, Deeplearning4j, or other training frameworks. In at least one embodiment, the training framework 904 trains the untrained neural network 906 and generates a trained neural network 908, enabling it to be trained using the processing resources described herein. In at least one embodiment, the weights may be randomly selected or selected by pre-training using a deep belief network. In at least one embodiment, training may be performed in a supervised, partially supervised, or unsupervised manner.
[0057] In at least one embodiment, the untrained neural network 906 is trained using supervised learning, where the training dataset 902 includes inputs paired with desired outputs for each input, or the training dataset 902 includes inputs with known outputs, and the outputs of the neural network 906 are manually scored. In at least one embodiment, the untrained neural network 906 is trained in a supervised manner, processing inputs from the training dataset 902 and comparing the resulting outputs to a set of expected or desired outputs. In at least one embodiment, the error is then backpropagated through the untrained neural network 906. In at least one embodiment, the training framework 904 adjusts the weights controlling the untrained neural network 906. In at least one embodiment, the training framework 904 includes a tool to monitor how well the untrained neural network 906 is converging toward a model such as a trained neural network 908 that is suitable for generating the correct answer in the result 914, etc., based on input data such as a new dataset 912. In at least one embodiment, the training framework 904 iteratively trains the untrained neural network 906 while adjusting the weights to refine the output of the untrained neural network 906 using a loss function and tuning algorithms such as stochastic gradient descent. In at least one embodiment, the training framework 904 trains the untrained neural network 906 until it reaches a desired accuracy. In at least one embodiment, the trained neural network 908 can then be introduced to implement any number of machine learning operations.
[0058] In at least one embodiment, an untrained neural network 906 is trained using unsupervised learning, where the untrained neural network 906 attempts to train itself using unlabeled data. In at least one embodiment, the training dataset 902 for unsupervised learning includes input data with no associated output data or "ground truth" data. In at least one embodiment, the untrained neural network 906 can learn grouping within the training dataset 902 and determine how individual inputs relate to the untrained dataset 902. In at least one embodiment, unsupervised training can be used within a trained neural network 908 that can perform operations useful for reducing the dimensionality of a novel dataset 912 to generate a self-organizing map. In at least one embodiment, anomaly detection can also be performed using unsupervised training, which allows for the identification of data points in the novel dataset 912 that deviate from the normal pattern of the novel dataset 912.
[0059] In at least one embodiment, semi-supervised learning may be used, which is a technique in which labeled and unlabeled data are mixed in the training dataset 902. In at least one embodiment, incremental learning, such as a transmission learning technique, may be performed using the training framework 904. In at least one embodiment, incremental learning enables the trained neural network 908 to adapt to a new dataset 912 without forgetting the knowledge that was taught to the trained neural network 908 during the initial training.
[0060] In at least one embodiment, the training framework 904 is a framework processed in relation to a software development toolkit such as the OpenVINO (Open Visual Inference and Neural Network Optimization) toolkit. In at least one embodiment, the OpenVINO toolkit is a toolkit such as one developed by Intel Corporation in Santa Clara, California.
[0061] In at least one embodiment, OpenVINO is a toolkit for facilitating the development of applications for a variety of tasks and operations, particularly neural network applications, such as human visual emulation, speech recognition, natural language processing, recommendation systems, and / or variations thereof. In at least one embodiment, OpenVINO supports neural networks such as convolutional neural networks (CNNs), recurrent and / or attention-based neural networks, and / or various other neural network models. In at least one embodiment, OpenVINO supports various software libraries such as OpenCV, OpenCL, and / or variations thereof.
[0062] In at least one embodiment, OpenVINO supports neural network models for a variety of tasks and actions, including classification, segmentation, object detection, face recognition, speech recognition, pose estimation (e.g., human and / or object), monocular depth estimation, image restoration, style transfer, motion recognition, colorization, and / or modified forms thereof.
[0063] In at least one embodiment, OpenVINO includes one or more software tools and / or modules for model optimization, also known as model optimizers. In at least one embodiment, the model optimizer is a command-line tool that facilitates the transition between training and deployment of a neural network model. In at least one embodiment, the model optimizer optimizes a neural network model so that it can run on various devices and / or processing units such as GPUs, CPUs, PPUs, GPGPUs, and / or variations thereof. In at least one embodiment, the model optimizer generates an internal representation of the model and optimizes the model to generate an intermediate representation. In at least one embodiment, the model optimizer reduces the number of layers in the model. In at least one embodiment, the model optimizer removes layers of the model that are used for training. In at least one embodiment, the model optimizer performs various neural network operations, such as modifying the input to the model (e.g., resizing the input to the model), modifying the size of the input to the model (e.g., modifying the batch size of the model), modifying the model structure (e.g., modifying the layers of the model), normalization, standardization, quantization (e.g., converting the model weights from a first representation such as floating-point numbers to a second representation such as integers), and / or variations thereof.
[0064] In at least one embodiment, OpenVINO includes one or more software libraries for inference, also called an inference engine. In at least one embodiment, the inference engine is a C++ library or a library in any preferred programming language. In at least one embodiment, the inference engine is used to infer input data. In at least one embodiment, the inference engine implements various classes for inferring input data and producing one or more results. In at least one embodiment, the inference engine implements one or more API functions for processing intermediate representations, formatting inputs and / or outputs, and / or running the model on one or more devices.
[0065] In at least one embodiment, OpenVINO provides various functions for heterogeneous execution of one or more neural network models. In at least one embodiment, heterogeneous execution or heterogeneous computation means one or more computing processes and / or systems that utilize one or more types of processors and / or cores. In at least one embodiment, OpenVINO provides various software functions for running a program on one or more devices. In at least one embodiment, OpenVINO provides various software functions for running a program and / or parts of a program on different devices. In at least one embodiment, OpenVINO provides various software functions for running, for example, a first part of the code on a CPU and a second part of the code on a GPU and / or FPGA. In at least one embodiment, OpenVINO provides various software functions for running one or more layers of a neural network on one or more devices (for example, running a first set of layers on a first device such as a GPU and running a second set of layers on a second device such as a CPU).
[0066] In at least one embodiment, OpenVINO includes a variety of functions similar to those associated with CUDA programming models, such as various neural network model behaviors associated with frameworks such as TensorFlow, PyTorch, and / or variations thereof. In at least one embodiment, one or more CUDA programming model behaviors are performed using OpenVINO. In at least one embodiment, various systems, methods, and / or techniques described herein are implemented using OpenVINO.
[0067] Data center Figure 10 shows an exemplary data center 1000 in which at least one embodiment may be used. In at least one embodiment, the data center 1000 includes a data center infrastructure layer 1010, a framework layer 1020, a software layer 1030, and an application layer 1040.
[0068] As shown in Figure 10, in at least one embodiment, the data center infrastructure layer 1010 may include a resource orchestrator 1012, grouped computing resources 1014, and node computing resources ("node CRs") 1016(1) to 1016(N), where "N" represents a positive integer (which may be a different integer "N" than that used in other figures). In at least one embodiment, nodes CR1016(1) to 1016(N) may include, but are not limited to, any number of central processing units ("CPUs") or other processors (including accelerators, field-programmable gate arrays (FPGAs), graphics processors, etc.), memory and storage devices 1018(1) to 1018(N) (e.g., dynamic read-only memory, semiconductor storage drives, or disk drives), network input / output ("NW I / O") devices, network switches, virtual machines ("VMs"), power modules, and cooling modules. In at least one embodiment, one or more nodes CR1016(1) to 1016(N) may be servers having one or more of the computing resources described above.
[0069] In at least one embodiment, the grouped computing resources 1014 may include separate groups of node CRs housed in one or more racks (not shown), or a number of racks housed in a data center in various graphical locations (also not shown). In at least one embodiment, separate groups of node CRs within the grouped computing resources 1014 may include grouped compute resources, network resources, memory resources, or storage resources that may be configured or allocated to support one or more workloads. In at least one embodiment, several node CRs, including CPUs or processors, may be grouped in one or more racks to provide compute resources to support one or more workloads. In at least one embodiment, one or more racks may also include any number of power modules, cooling modules, and network switches in any combination.
[0070] In at least one embodiment, the resource orchestrator 1012 may constitute or otherwise control one or more nodes CR1016(1) to 1016(N) and / or grouped computing resources 1014. In at least one embodiment, the resource orchestrator 1012 may include a software design infrastructure ("SDI") management entity for the data center 1000. In at least one embodiment, the resource orchestrator 812 may include hardware, software, or any combination thereof.
[0071] In at least one embodiment shown in Figure 10, the framework layer 1020 includes a job scheduler 1022, a configuration manager 1024, a resource manager 1026, and a distribution file system 1028. In at least one embodiment, the framework layer 1020 may include a framework for supporting software 1032 of the software layer 1030 and / or one or more applications 1042 of the application layer 1040. In at least one embodiment, the software 1032 or application 1042 may each include web-based service software or applications, such as those provided by Amazon Web Services, Google Cloud, and Microsoft Azure. In at least one embodiment, the framework layer 1020 may be, but is not limited to, a type of free, open-source software web application framework, such as Apache Spark® ("Spark"), which can use the distribution file system 1028 for large-scale data processing (e.g., "big data"). In at least one embodiment, the job scheduler 1022 may include a Spark driver to facilitate scheduling of workloads supported by various layers of the data center 1000. In at least one embodiment, the configuration manager 1024 may be capable of configuring different layers, such as the software layer 1030 and the framework layer 1020, which includes Spark and a distribution file system 1028 to support large-scale data processing. In at least one embodiment, the resource manager 1026 may be capable of managing clustered or grouped computing resources that are mapped or allocated to support the distribution file system 1028 and the job scheduler 1022. In at least one embodiment, the clustered or grouped computing resources may include grouped computing resources 1014 located in the data center infrastructure layer 1010.In at least one embodiment, the resource manager 1026 may work in conjunction with the resource orchestrator 1012 to manage these mapped or allocated computing resources.
[0072] In at least one embodiment, the software 1032 included in the software layer 1030 may include software used by at least a portion of the nodes CR1016(1) to 1016(N), the grouped computing resources 1014, and / or the distribution file system 1028 of the framework layer 1020. In at least one embodiment, one or more types of software may include, but are not limited to, internet web page search software, email virus scanning software, database software, and streaming video content software.
[0073] In at least one embodiment, application 1042 included in application layer 1040 may include one or more types of applications used by at least a portion of nodes CR1016(1) to 1016(N), grouped computing resources 1014, and / or distribution file system 1028 of framework layer 1020. In at least one embodiment, one or more types of applications may include, but are not limited to, any number of genomics applications, recognition compute, and training or inference software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), and machine learning applications, or other machine learning applications used in conjunction with one or more embodiments.
[0074] In at least one embodiment, any of the configuration manager 1024, resource manager 1026, and resource orchestrator 1012 may implement any number and type of self-correcting measures based on any amount and type of data obtained in any technically feasible manner. In at least one embodiment, the self-correcting measures may enable the data center operator of data center 1000 to avoid determining potentially faulty configurations and to eliminate underutilized and / or underperforming portions of the data center.
[0075] In at least one embodiment, the data center 1000 may include tools, services, software, or other resources for training one or more machine learning models or for predicting or inferring information using one or more machine learning models according to one or more embodiments described herein. For example, in at least one embodiment, a machine learning model may be trained by computing weight parameters according to a neural network architecture using the software and computing resources described above with respect to the data center 1000. In at least one embodiment, a trained machine learning model corresponding to one or more neural networks may be used to infer or predict information using the resources described above with respect to the data center 1000 by using weight parameters computed by one or more techniques described herein.
[0076] In at least one embodiment, the data center may use a CPU, application-specific integrated circuit (ASIC), GPU, FPGA, or other hardware to perform training and / or inference using the resources described above. Furthermore, one or more of the software and / or hardware resources described above may be configured as a service to enable users to perform training or inference on information such as image recognition, speech recognition, or other artificial intelligence services.
[0077] The inference and / or training logic 815 is used to perform inference and / or training operations related to one or more embodiments. Details relating to the inference and / or training logic 815 are provided herein in conjunction with Figures 8A and / or 8B. In at least one embodiment, the inference and / or training logic 815 may be used in the system of Figure 10 for inference or prediction operations, at least in part, based on weight parameters calculated using the neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0078] Autonomous vehicles Figure 11A shows an example of an autonomous vehicle 1100 according to at least one embodiment. In at least one embodiment, the autonomous vehicle 1100 (or, as referred to herein as "vehicle 1100") may be, without limitation, a passenger vehicle such as a car, truck, bus, and / or another type of vehicle accommodating one or more occupants. In at least one embodiment, vehicle 1100 may be a semi-tractor trailer truck for cargo transport. In at least one embodiment, vehicle 1100 may be an aircraft, a robotic vehicle, or another type of vehicle.
[0079] Autonomous vehicles may also be described in terms of automation levels as defined by the National Highway Traffic Safety Administration ("NHTSA"), a division of the U.S. Department of Transportation, and the Society of Automotive Engineers ("SAE") in their "Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles" (for example, standard No. J3016-201806, published June 15, 2018; standard No. J3016-201609, published September 30, 2016; and previous and newer versions of this standard). In at least one embodiment, vehicle 1100 may be capable of functioning at one or more of the autonomous driving levels from Level 1 to Level 5. For example, in at least one embodiment, vehicle 1100 may be capable of conditional automation (Level 3), high automation (Level 4), and / or full automation (Level 5), depending on the embodiment.
[0080] In at least one embodiment, the vehicle 1100 may include, without limitation, components such as a chassis, a vehicle body, wheels (2, 4, 6, 8, 18, etc.), tires, axles, and other components of the vehicle. In at least one embodiment, the vehicle 1100 may include, without limitation, a propulsion system 1150 such as an internal combustion engine, a hybrid power plant, a fully electric engine, and / or another type of propulsion system. In at least one embodiment, the propulsion system 1150 may be coupled to the drivetrain of the vehicle 1100, and the drivetrain may include, without limitation, a transmission to enable the propulsion of the vehicle 1100. In at least one embodiment, the propulsion system 1150 may be controlled in response to receiving a signal from a throttle / accelerator 1152.
[0081] In at least one embodiment, a steering system 1154, which may include a steering wheel (for example), is used to steer the vehicle 1100 (for example, along a desired path or route) when the propulsion system 1150 is operating (for example, when the vehicle 1100 is moving). In at least one embodiment, the steering system 1154 may receive signals from a steering actuator 1156. In at least one embodiment, the steering wheel may be optional with respect to fully automated (Level 5) functionality. In at least one embodiment, a brake sensor system 1146 may be used to operate the vehicle brakes in response to receiving signals from a brake actuator 1148 and / or a brake sensor.
[0082] In at least one embodiment, the controller 1136, which may include, without limitation, one or more system-on-chip ("SoC") (not shown in Figure 11A) and / or graphics processing unit ("GPU"), provides signals (e.g., representing commands) to one or more components and / or systems of the vehicle 1100. For example, in at least one embodiment, the controller 1136 may transmit signals to operate the vehicle brakes via the brake actuator 1148, signals to operate the steering system 1154 via the steering actuator 1156, and signals to operate the propulsion system 1150 via the throttle / accelerator 1152. In at least one embodiment, the controller 1136 may include one or more onboard (e.g., integrated) computing devices (e.g., supercomputers) that process sensor signals and output operational commands (e.g., signals representing commands) to enable autonomous driving and / or assist a human driver in the driving vehicle 1100. In at least one embodiment, the controller 1136 may include a first controller for autonomous driving functions, a second controller for functional safety functions, a third controller for artificial intelligence functions (e.g., computer vision), a fourth controller for infotainment functions, a fifth controller for emergency redundancy, and / or other controllers. In at least one embodiment, a single controller may address two or more of the above functionalities, two or more controllers may address a single functionality, and / or any combination thereof.
[0083] In at least one embodiment, the controller 1136 provides signals for controlling one or more components and / or systems of the vehicle 1100 in response to sensor data (e.g., sensor inputs) received from one or more sensors. In at least one embodiment, the sensor data may include, for example, a global navigation satellite system ("GNSS") sensor 1158 (e.g., a global positioning system sensor), a radar sensor 1160, an ultrasonic sensor 1162, a lithium-ion radar sensor 1164, and an inertial measurement unit ("IMU"), without limiting them. The unit may receive signals from sensors 1166 (e.g., accelerometer, gyroscope, magnetic compass, magnetometer, etc.), microphone 1196, stereo camera 1168, wide-angle camera 1170 (e.g., fisheye camera), infrared camera 1172, ambient camera 1174 (e.g., 360-degree camera), long-range camera (not shown in Figure 11A), medium-range camera (not shown in Figure 11A), speed sensor 1144 (e.g., for measuring the speed of vehicle 1100), vibration sensor 1142, steering sensor 1140, brake sensor (e.g., as part of brake sensor system 1146), and / or other types of sensors.
[0084] In at least one embodiment, one or more of the controllers 1136 may receive inputs (represented, for example, by input data) from the instrument cluster 1132 of the vehicle 1100 and provide outputs (represented, for example, by output data, display data, etc.) via the human-machine interface ("HMI") display 1134, an audible annunciator, a loudspeaker, and / or other components of the vehicle 1100. In at least one embodiment, the outputs may include information such as vehicle speed, speed, time, map data (e.g., a high-definition map (not shown in Figure 11A)), location data (e.g., the location of the vehicle 1100, such as on a map), direction, location of other vehicles (e.g., an occupied grid), and information about objects and the state of objects sensed by the controller 1136. For example, in at least one embodiment, the HMI display 1134 may display information about the presence of one or more objects (e.g., road signs, warning signs, changes in traffic signals, etc.) and / or information about driving operations that the vehicle has performed, is performing, or will perform (e.g., changing lanes, exiting at Exit 34B 3.22 km (2 miles) ahead, etc.).
[0085] In at least one embodiment, the vehicle 1100 further includes a network interface 1124, which may use a wireless antenna 1126 and / or a modem for communication over one or more networks. For example, in at least one embodiment, the network interface 1124 may be able to communicate over Long-Term Evolution ("LTE"), Wideband Code Division Multiple Access ("WCDMA®"), Universal Mobile Telecommunications System ("UMTS"), Global System for Mobile Communications ("GSM"), IMT-CDMA Multi-Carrier ("CDMA2000") networks, and the like. In at least one embodiment, the wireless antenna 1126 may also enable communication between objects in the environment (e.g., vehicles, mobile devices, etc.) using local area networks such as Bluetooth, Bluetooth Low Energy ("LE"), Z-Wave, ZigBee, and / or low power wide-area networks ("LPWAN") such as LoRaWAN and SigFox.
[0086] The inference and / or training logic 815 is used to perform inference and / or training operations related to one or more embodiments. Details relating to the inference and / or training logic 815 are provided herein in conjunction with Figures 8A and / or 8B. In at least one embodiment, the inference and / or training logic 815 may be used in the system of Figure 11A for inference or prediction operations, at least in part, based on weight parameters calculated using the neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0087] Figure 11B shows an example of camera locations and fields of view for the autonomous vehicle 1100 of Figure 11A according to at least one embodiment. In at least one embodiment, the cameras and their respective fields of view are illustrative and not limiting. For example, in at least one embodiment, additional and / or alternative cameras may be included, and / or the cameras may be positioned at different locations on the vehicle 1100.
[0088] In at least one embodiment, the camera type of the camera may include, but is not limited to, a digital camera which may be adapted for use with components and / or systems of the vehicle 1100. In at least one embodiment, the camera may operate at Automotive Safety Integrity Level ("ASIL") B and / or another ASIL. In at least one embodiment, the camera type may be capable of handling any image capture rate, such as 60 frames per second (fps), 1220 fps, 240 fps, etc., depending on the embodiment. In at least one embodiment, the camera may be capable of using a roll shutter, a global shutter, another type of shutter, or a combination thereof. In at least one embodiment, the color filter array may include a red-clear-clear-clear ("RCCC") color filter array, a red-clear-clear-blue ("RCCB") color filter array, a red-blue-green-clear ("RBGC") color filter array, a Foveon X3 color filter array, a Bayer sensor (RGGB) color filter array, a monochrome sensor color filter array, and / or other types of color filter arrays. In at least one embodiment, a clear pixel camera, such as a camera having RCCC, RCCB, and / or RBGC color filter arrays, may be used to increase light sensitivity.
[0089] In at least one embodiment, one or more of the cameras may be used to perform advanced driver assistance system ("ADAS") functions (for example, as part of a redundant or fail-safe design). For example, in at least one embodiment, a multi-function mono-camera may be installed to provide functions including lane departure warning, traffic sign assistance, and intelligent headlight control. In at least one embodiment, one or more of the cameras (for example, all of the cameras) may simultaneously record and provide image data (for example, video).
[0090] In at least one embodiment, one or more cameras may be mounted on a mounting assembly, such as a custom-designed (three-dimensionally printed) assembly, to eliminate stray light and reflections from inside the vehicle 1100 (e.g., reflections from the dashboard to the windshield) that could interfere with the camera's image data acquisition performance. Referring to a door mirror mounting assembly, in at least one embodiment, the door mirror assembly may be custom 3D printed so that the camera mounting plate conforms to the shape of the door mirror. In at least one embodiment, the camera may be integrated with the door mirror. In at least one embodiment, for a side-view camera, the camera may also be integrated with one of the four pillars at each corner of the cabin.
[0091] In at least one embodiment, a camera having a field of view that includes a portion of the environment in front of the vehicle 1100 (e.g., a front camera) may be used for the surrounding view to facilitate the identification of the path and obstacles ahead, and may be used in conjunction with one or more controllers 1136 and / or control SoCs to assist in providing information essential for generating an occupied grid and / or determining a preferred vehicle path. In at least one embodiment, the front camera may be used to perform many of the ADAS functions similar to LIDAR, including, but not limited to, emergency braking, pedestrian detection, and collision avoidance. In at least one embodiment, the front camera may also be used for ADAS functions and systems, including, but not limited to, other functions such as lane departure warnings ("LDW"), autonomous cruise control ("ACC"), and / or traffic sign recognition.
[0092] In at least one embodiment, various cameras, including a monocular camera platform including, for example, a CMOS (complementary metal oxide semiconductor) color imaging device, may be used in a front configuration. In at least one embodiment, a wide-angle camera 1170 may be used to sense objects entering the view from the surroundings (e.g., pedestrians, cross-traffic, or bicycles). Although only one wide-angle camera 1170 is shown in Figure 11B, in other embodiments, the vehicle 1100 may have any number of wide-angle cameras (including zero). In at least one embodiment, any number of long-range cameras 1198 (e.g., pairs of stereo cameras with long-range views) may be used for depth-based object detection, particularly for objects for which the neural network has not yet been trained. In at least one embodiment, the long-range cameras 1198 may also be used for object detection and classification, as well as basic object tracking.
[0093] In at least one embodiment, any number of stereo cameras 1168 may also be included in a front configuration. In at least one embodiment, one or more stereo cameras 1168 may include an integrated control unit with an expandable processing unit, which may provide a programmable logic ("FPGA") and a multi-core microprocessor having an integrated Controller Area Network ("CAN") or Ethernet® interface on a single chip. In at least one embodiment, such a unit may be used to generate a 3D map of the vehicle 1100's environment, including distance estimation for all points in the image. In at least one embodiment, one or more of the stereo cameras 1168 may include, without limitation, a compact stereo vision sensor, which may include, without limitation, two camera lenses (one on the left and one on the right) and an image processing chip that can measure the distance from the vehicle 1100 to a target object and use the generated information (e.g., metadata) to activate autonomous emergency braking and lane departure warning functions. In at least one embodiment, other types of stereo cameras 1168 may be used in addition to or instead of those described herein.
[0094] In at least one embodiment, a camera having a field of view that includes a portion of the environment to the sides of the vehicle 1100 (e.g., a side-view camera) may be used for the surrounding view to provide information used for creating and updating the occupancy grid and generating side collision warnings. For example, in at least one embodiment, surrounding cameras 1174 (e.g., four surrounding cameras as shown in Figure 11B) may be positioned on the vehicle 1100. In at least one embodiment, the surrounding cameras 1174 may include, without limitation, any number and combination of wide-angle cameras, fisheye cameras, 360-degree cameras and / or similar cameras. For example, in at least one embodiment, four fisheye cameras may be positioned in front of, behind, and to the sides of the vehicle 1100. In at least one embodiment, the vehicle 1100 may use three surrounding cameras 1174 (e.g., left, right, and rear) and utilize one or more other cameras (e.g., a front camera) as a fourth surrounding camera.
[0095] In at least one embodiment, a camera having a field of view that includes a portion of the environment behind the vehicle 1100 (e.g., a rear-view camera) may be used for parking assistance, surrounding view, and rear collision warning to create and update the occupancy grid. In at least one embodiment, a wide variety of cameras may be used, including, but not limited to, cameras also suitable as front cameras as described herein (e.g., long-range camera 1198 and / or medium-range camera 1176, stereo camera 1168, infrared camera 1172, etc.).
[0096] The inference and / or training logic 815 is used to perform inference and / or training operations related to one or more embodiments. Details relating to the inference and / or training logic 815 are provided herein in conjunction with Figures 8A and / or 8B. In at least one embodiment, the inference and / or training logic 815 may be used in the system of Figure 11B for inference or prediction operations, at least in part, based on weight parameters calculated using the neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0097] Figure 11C is a block diagram illustrating an exemplary system architecture of the autonomous vehicle 1100 of Figure 11A according to at least one embodiment. In at least one embodiment, each of the components, features, and systems of the vehicle 1100 in Figure 11C is shown as being connected via a bus 1102. In at least one embodiment, the bus 1102 may include, without limitation, a CAN data interface (or, as referred to herein, the (CAN bus)). In at least one embodiment, the CAN may be an internal network of the vehicle 1100 used to assist in the control of various features and functions of the vehicle 1100, such as brake activation, acceleration, brake control, steering, and windshield wipers. In at least one embodiment, the bus 1102 may be configured to have tens or even hundreds of nodes, each having its own unique identifier (e.g., a CAN ID). In at least one embodiment, the bus 1102 may be read to find steering angle, ground speed, engine revolutions per minute ("RPM"), button position, and / or other vehicle state indicators. In at least one embodiment, bus 1102 may be a CAN bus compliant with ASIL B.
[0098] In at least one embodiment, the FlexRay and / or Ethernet® protocol may be used in addition to or instead of CAN. In at least one embodiment, there may be any number of buses forming bus 1102, which may include, without limitation, zero or more CAN buses, zero or more FlexRay buses, zero or more Ethernet® buses, and / or zero or more other types of buses using other protocols. In at least one embodiment, two or more buses may be used to perform different functions and / or to provide redundancy. For example, a first bus may be used for collision avoidance functions and a second bus may be used for operation control. In at least one embodiment, each bus of bus 1102 may communicate with any of the components of vehicle 1100, and two or more buses of bus 1102 may communicate with corresponding components. In at least one embodiment, each of any number of system-on-chip ("SoC") 1104 (such as SoC1104(A) and SoC1104(B)), each of the controllers 1136, and / or each computer in the vehicle may have access to the same input data (e.g., inputs from sensors in the vehicle 1100) and may be connected to a common bus such as a CAN bus.
[0099] In at least one embodiment, the vehicle 1100 may include one or more controllers 1136, such as those described herein with respect to Figure 11A. In at least one embodiment, the controllers 1136 may be used for a variety of functions. In at least one embodiment, the controllers 1136 may be coupled to any of the various other components and systems of the vehicle 1100 and may be used to control the vehicle 1100, the artificial intelligence of the vehicle 1100, the infotainment and / or other functions of the vehicle 1100.
[0100] In at least one embodiment, the vehicle 1100 may include any number of SoCs 1104. In at least one embodiment, each of the SoCs 1104 may include, without limitation, a central processing unit ("CPU") 1106, a graphics processing unit ("GPU") 1108, a processor 1110, a cache 1112, an accelerator 1114, a data store 1116, and / or other components and features not shown. In at least one embodiment, the SoCs 1104 may be used to control the vehicle 1100 in various platforms and systems. For example, in at least one embodiment, the SoCs 1104 may be incorporated into a system (e.g., the system of the vehicle 1100) having a high-definition ("HD") map 1122 that can be refreshed and / or updated from one or more servers (not shown in Figure 11C) via a network interface 1124.
[0101] In at least one embodiment, the CPU 1106 may include a CPU cluster, or CPU complex (or referred to herein as “CCPLEX”). In at least one embodiment, the CPU 1106 may include multiple cores and / or Level 2 ("L2") caches. For example, in at least one embodiment, the CPU 1106 may include eight cores in a coherent multiprocessor configuration. In at least one embodiment, the CPU 1106 may include four dual-core clusters, each having its own dedicated L2 cache (e.g., a 2 megabyte (MB) L2 cache). In at least one embodiment, the CPU 1106 (e.g., CCPLEX) may be configured to support concurrent cluster operation, which allows any combination of clusters of the CPU 1106 to be activated at any given time.
[0102] In at least one embodiment, one or more of the CPUs 1106 may implement a power management function, which includes, but is not limited to, one or more of the following features: individual hardware blocks can be automatically clock-gated when idle to conserve dynamic power; each core clock can be gated when a core is not actively executing instructions due to the execution of an interrupt-wait ("WFI") / event-wait ("WFE") instruction; each core can be power-gated independently; each core cluster can be clock-gated independently when all cores are clock-gated or power-gated; and / or each core cluster can be power-gated independently when all cores are power-gated. In at least one embodiment, the CPU 1106 may further implement an extended algorithm for managing power states, where an acceptable power state and an expected wake-up time are specified, and the hardware / microcode determines which power state is best for the cores, clusters, and CCPLEX to enter. In at least one embodiment, the processing core may support in software a simple sequence for entering a power state with the work offloaded to microcode.
[0103] In at least one embodiment, the GPU 1108 may include an integrated GPU (or, as referred to herein, an "iGPU"). In at least one embodiment, the GPU 1108 may be programmable and efficient for parallel workloads. In at least one embodiment, the GPU 1108 may use an extended tensor instruction set. In at least one embodiment, the GPU 1108 may include one or more streaming microprocessors, each of which may include a Level 1 ("L1") cache (e.g., an L1 cache with at least 96KB of storage capacity), and two or more streaming microprocessors may share an L2 cache (e.g., an L2 cache with 512KB of storage capacity). In at least one embodiment, the GPU 1108 may include at least eight streaming microprocessors. In at least one embodiment, the GPU 1108 may use a compute application programming interface (API). In at least one embodiment, the GPU1108 may use one or more parallel computing platforms and / or programming modules (for example, NVIDIA's CUDA model).
[0104] In at least one embodiment, one or more of the GPU1108s may be power-optimized to best perform in automotive and embedded use cases. For example, in at least one embodiment, the GPU1108 may be fabricated on a Finn field-effect transistor ("FinFET") circuit. In at least one embodiment, each streaming microprocessor may incorporate a number of mixed-precision processing cores divided into multiple blocks. For example, 64 PF32 cores and 32 PF64 cores may be divided into four processing blocks, without limitation. In at least one embodiment, each processing block may be allocated 16 FP32 cores, 8 FP64 cores, 16 INT32 cores, 2 mixed-precision NVIDIA Tensor cores for deep learning matrix operations, a level zero ("L0") instruction cache, a warp scheduler, a dispatch unit, and / or a 64KB register file. In at least one embodiment, the streaming microprocessor includes independent parallel data paths for integers and floating-point numbers, enabling efficient execution of workloads by combining computer processing and addressing calculations. In at least one embodiment, the streaming microprocessor may include independent thread scheduling capabilities, enabling finer-grained synchronization and coordination between parallel threads. In at least one embodiment, the streaming microprocessor may include a combination of an L1 data cache and a shared memory unit to improve performance while simplifying programming.
[0105] In at least one embodiment, one or more of the GPUs 1108 include high-bandwidth memory ("HBM") and / or a 16GB HBM2 memory subsystem, which in some examples may provide a peak memory bandwidth of approximately 900 GB / s. In at least one embodiment, in addition to or instead of HBM memory, synchronous graphics random-access memory ("SGRAM") such as graphics double data rate type five synchronous random-access memory ("GDDR5") may be used.
[0106] In at least one embodiment, the GPU 1108 may include integrated memory technology. In at least one embodiment, address translation services ("ATS") support may be used to allow the GPU 1108 to directly access the page table of the CPU 1106. In at least one embodiment, when a GPU in the GPU 1108 memory management unit ("MMU") encounters a miss, an address translation request may be sent to the CPU 1106. In at least one embodiment, in response, two of the CPUs in the CPU 1106 may look up the virtual-to-physical address mapping in their own page tables and send the translation back to the GPU 1108. In at least one embodiment, integrated memory technology makes it possible to provide a single, unified virtual address space for the memory of both the CPU 1106 and the GPU 1108, thereby simplifying the programming of the GPU 1108 and the porting of applications to the GPU 1108.
[0107] In at least one embodiment, the GPU 1108 may include any number of access counters that can record how often the GPU 1108 accesses the memory of other processors. In at least one embodiment, the access counters may help ensure that memory pages are moved to the physical memory of the processor that accesses them most frequently, thereby improving the efficiency of memory ranges shared between processors.
[0108] In at least one embodiment, one or more of the SoC1104 may include any number of caches 1112, including those described herein. For example, in at least one embodiment, the cache 1112 may include a Level 3 ("L3") cache that is available to both the CPU 1106 and the GPU 1108 (e.g., connected to both the CPU 1106 and the GPU 1108). In at least one embodiment, the cache 1112 may include a write-back cache that can record line states by using a cache coherence protocol such as MEI, MESI, or MSI. In at least one embodiment, the L3 cache may include 4 MB or more of memory, depending on the embodiment, but a smaller cache size may be used.
[0109] In at least one embodiment, one or more of the SoC1104 may include one or more accelerators 1114 (e.g., hardware accelerators, software accelerators, or a combination thereof). In at least one embodiment, the SoC1104 may include a hardware acceleration cluster that may include optimized hardware accelerators and / or large on-chip memory. In at least one embodiment, the large on-chip memory (e.g., 4 MB of SRAM) may enable the hardware acceleration cluster to accelerate neural networks and other computations. In at least one embodiment, the hardware acceleration cluster may be used to complement the GPU1108 and offload some of the tasks of the GPU1108 (e.g., free up more cycles of the GPU1108 to perform other tasks). In at least one embodiment, accelerator 1114 can be used for target workloads that are stable enough to accept acceleration (e.g., perception, convolutional neural networks ("CNN"), recurrent neural networks ("RNN"), etc.). In at least one embodiment, the CNN may include region-based CNNs, i.e., regional convolutional neural networks ("RCNN"), and fast RCNNs (e.g., used for object detection), or other types of CNNs.
[0110] In at least one embodiment, the accelerator 1114 (e.g., a hardware acceleration cluster) may include one or more deep learning accelerators ("DLAs"). In at least one embodiment, the DLAs may include, without limitation, one or more Tensor processing units ("TPUs"), which may be configured to provide an additional 10 trillion operations per second for deep learning applications and inference. In at least one embodiment, the TPUs may be configured to perform image processing functions (e.g., CNNs, RCNNs, etc.) and may be accelerators optimized for that purpose. In at least one embodiment, the DLAs may be further optimized for specific neural network types and sets of floating-point operations, as well as for inference. In at least one embodiment, the design of the DLA can improve performance per millisecond compared to a typical general-purpose GPU, and typically far exceed the performance of a CPU. In at least one embodiment, the TPU may execute several functions, including, for example, a single-instance convolution function supporting INT8, INT16, and FP16 data types for both features and weights, as well as post-processing functions. In at least one embodiment, the DLA may rapidly and efficiently execute neural networks, particularly CNNs, on processed or raw data for any of a variety of functions, including, for example, a CNN for object recognition and detection using data from camera sensors, a CNN for distance estimation using data from camera sensors, a CNN for emergency vehicle detection and identification and detection using data from microphones, a CNN for face recognition and vehicle owner identification using data from camera sensors, and / or a CNN for security and / or safety-related events.
[0111] In at least one embodiment, the DLA may perform any function of the GPU 1108, and the designer may target either the DLA or the GPU 1108 for any function, for example by using an inference accelerator. For example, in at least one embodiment, the designer may concentrate CNN and floating-point arithmetic processing on the DLA and offload other functions to the GPU 1108 and / or other accelerators 1114.
[0112] In at least one embodiment, the accelerator 1114 may include a programmable vision accelerator ("PVA"), which may be referred to herein as a computer vision accelerator instead. In at least one embodiment, the PVA may be designed and configured to accelerate computer vision algorithms for advanced driver assistance systems ("ADAS") 1, autonomous driving, augmented reality ("AR") applications, and / or virtual reality ("VR") applications. In at least one embodiment, the PVA may provide a balance between performance and flexibility. For example, in at least one embodiment, each PVA may include, for example, any number of reduced instruction set computer ("RISC") cores, direct memory access ("DMA"), and / or any number of vector processors.
[0113] In at least one embodiment, the RISC core may interact with an image sensor (e.g., an image sensor of any camera described herein), an image signal processor, and the like. In at least one embodiment, each RISC core may include any amount of memory. In at least one embodiment, the RISC core may use any of a plurality of protocols, depending on the embodiment. In at least one embodiment, the RISC core may run a real-time operating system ("RTOS"). In at least one embodiment, the RISC core may be implemented using one or more integrated circuit devices, application-specific integrated circuits ("ASICs"), and / or memory devices. For example, in at least one embodiment, the RISC core may include an instruction cache and / or tightly coupled RAM.
[0114] In at least one embodiment, the DMA may allow components of the PVA to access system memory independently of the CPU 1106. In at least one embodiment, the DMA may support any number of features used to provide optimization to the PVA, including but not limited to multidimensional addressing and / or circular addressing. In at least one embodiment, the DMA may support up to six or more addressing dimensions, which may include, without limitation, block width, block height, block depth, horizontal block stepping, vertical block stepping, and / or depth stepping.
[0115] In at least one embodiment, the vector processor may be a programmable processor designed to efficiently and flexibly perform programming for computer vision algorithms and provide signal processing capabilities. In at least one embodiment, the PVA may include a PVA core and two vector processing subsystem partitions. In at least one embodiment, the PVA core may include a processor subsystem, a DMA engine (e.g., two DMA engines), and / or other peripherals. In at least one embodiment, the vector processing subsystem may act as the primary processing engine of the PVA and may include a vector processing unit ("VPU"), an instruction cache, and / or vector memory (e.g., "VMEM"). In at least one embodiment, the VPU may include digital signal processors such as single instruction, multiple data ("SIMD") and very long instruction word ("VLIW") digital signal processors. In at least one embodiment, a combination of SIMD and VLIW may improve throughput and speed.
[0116] In at least one embodiment, each vector processor may include an instruction cache and may be coupled to dedicated memory. As a result, in at least one embodiment, each vector processor may be configured to run independently of other vector processors. In at least one embodiment, the vector processors included in a particular PVA may be configured to use data parallelism. For example, in at least one embodiment, multiple vector processors included in a single PVA may execute a common computer vision algorithm on different regions of an image. In at least one embodiment, the vector processors included in a particular PVA may execute different computer vision algorithms simultaneously on a single image, or furthermore, execute different algorithms on consecutive images or on parts of an image. In at least one embodiment, in particular, any number of PVAs may be included in a hardware acceleration cluster, and any number of vector processors may be included in each PVA. In at least one embodiment, the PVA may include additional error correction code ("ECC") memory to enhance the overall safety of the system.
[0117] In at least one embodiment, accelerator 1114 may include an on-chip computer vision network and static random access memory ("SRAM"), providing high-bandwidth, low-latency SRAM for accelerator 1114. In at least one embodiment, the on-chip memory may include at least 4 MB of SRAM, including, for example, eight field-configurable memory blocks, which may be accessible from both the PVA and DLA. In at least one embodiment, each pair of memory blocks may include an advanced peripheral bus ("APB") interface, configuration circuitry, a controller, and a multiplexer. In at least one embodiment, any type of memory may be used. In at least one embodiment, the PVA and DLA may access the memory via a backbone that provides high-speed access to the memory. In at least one embodiment, the backbone may include an on-chip computer vision network interconnecting the PVA and DLA to the memory (for example, using the APB).
[0118] In at least one embodiment, the on-chip computer vision network may include an interface that determines whether both the PVA and DLA provide ready and enable signals before transmitting any control signals / addresses / data. In at least one embodiment, the interface may provide separate phases and separate channels for transmitting control signals / addresses / data, as well as burst communication for continuous data transfer. In at least one embodiment, the interface may conform to the International Organization for Standardization ("ISO") 26262 or the International Electrotechnical Commission ("IEC") 61508 standard, but other standards and protocols may be used.
[0119] In at least one embodiment, one or more of the SoC1104s may include a real-time ray tracing hardware accelerator. In at least one embodiment, the real-time ray tracing hardware accelerator may be used to quickly and efficiently determine the position and extent of an object (e.g., in a world model) and generate real-time visualization simulations for RADAR signal interpretation, sound propagation synthesis and / or analysis, SONAR system simulation, general waveform propagation simulation, comparison with LIDAR data for localization and / or other functions, and / or other uses.
[0120] In at least one embodiment, accelerator 1114 can have a variety of uses for autonomous driving. In at least one embodiment, PVA can be used in key processing stages of ADAS and autonomous vehicles. In at least one embodiment, the performance of PVA is well suited to algorithmic domains that require low-power and low-latency predictable processing. In other words, PVA performs well even with small datasets for semi-dense or dense regular computations that may require low-latency and low-power predictable run times. In at least one embodiment, PVA can be designed to run conventional computer vision algorithms, such as within a vehicle 1100, because they can be effective for object detection and integer numerical calculations.
[0121] For example, according to at least one embodiment of the technology, computer stereo vision may be performed using PVA. In at least one embodiment, algorithms based on semi-global matching may be used in some examples, but this is not limited to them. In at least one embodiment, applications for Level 3-5 autonomous driving use motion estimation / stereo matching (e.g., structuring from motion, pedestrian recognition, lane detection, etc.) on the fly. In at least one embodiment, PVA may perform computer stereo vision functions for input from two monocular cameras.
[0122] In at least one embodiment, PVA may be used to perform high-density optical flow. For example, in at least one embodiment, PVA can process raw RADAR data (e.g., using a 4D Fast Fourier Transform) to provide processed RADAR data. In at least one embodiment, PVA is used for time-of-flight depth processing, and processed time-of-flight data is provided by processing raw time-of-flight data, for example.
[0123] In at least one embodiment, DLA may be used to run any type of network for enhancing control and driving safety, including, for example, a neural network that outputs a confidence scale for each object detection. In at least one embodiment, confidence may be expressed or interpreted as the probability of each detection compared to other detections, or as providing a relative “weight.” In at least one embodiment, the confidence scale allows the system to make further decisions about which detections should be considered positive rather than false positives. In at least one embodiment, the system may set a threshold for confidence and consider only detections that exceed the threshold as positive. In embodiments where automatic emergency braking (“AEB”) is used, a false positive would cause the vehicle to automatically apply the emergency brakes, which is obviously undesirable. In at least one embodiment, a very reliable detection may be considered a trigger for the AEB. In at least one embodiment, DLA may run a neural network to regress confidence values. In at least one embodiment, the neural network may take as its input at least a subset of parameters such as the dimensions of the bounding box, ground estimation obtained (e.g., from another subsystem), output from the IMU sensor 1166 correlated with the orientation of the vehicle 1100, distance, and 3D location estimation of an object obtained from the neural network and / or other sensors (e.g., the LIDAR sensor 1164 or the RADAR sensor 1160).
[0124] In at least one embodiment, one or more of the SoCs 1104 may include a data store 1116 (e.g., memory). In at least one embodiment, the data store 1116 may be on-chip memory of the SoC 1104, which may store a neural network running on the GPU 1108 and / or DLA. In at least one embodiment, the capacity of the data store 1116 may be large enough to store multiple instances of the neural network for redundancy and safety. In at least one embodiment, the data store 1116 may include an L2 or L3 cache.
[0125] In at least one embodiment, one or more of the SoC1104 may include any number of processors 1110 (e.g., embedded processors). In at least one embodiment, the processors 1110 may include a boot and power management processor, which may be a dedicated processor and subsystem for handling boot power and management functions and associated security enforcement. In at least one embodiment, the boot and power management processor may be part of the boot sequence of the SoC1104 and may provide run-time power management services. In at least one embodiment, the boot power and management processor may provide clock and voltage programming, assistance in transitioning the system to a low-power state, management of thermal and temperature sensors of the SoC1104, and / or management of the power state of the SoC1104. In at least one embodiment, each temperature sensor may be implemented as a ring oscillator whose output frequency is proportional to the temperature, and the SoC1104 may use the ring oscillator to detect the temperature of the CPU 1106, GPU 1108, and / or accelerator 1114. In at least one embodiment, if it is determined that the temperature exceeds a threshold, the boot and power management processor may enter a temperature failure routine, put the SoC 1104 into a low-power state, and / or put the vehicle 1100 into driver-safe stop mode (for example, safely stop the vehicle 1100).
[0126] In at least one embodiment, the processor 1110 may further include a set of embedded processors that can serve as an audio processing engine, which may be an audio subsystem enabling full hardware support for multi-channel audio via multiple interfaces and a wide range of flexible audio I / O interfaces. In at least one embodiment, the audio processing engine is a dedicated processor core having a digital signal processor with dedicated RAM.
[0127] In at least one embodiment, the processor 1110 may further include an always-on processor engine that can provide the hardware features necessary to support low-power sensor management and startup use cases. In at least one embodiment, the always-on processor engine may include, without limitation, a processor core, tightly coupled RAM, supporting peripherals (e.g., timers and interrupt controllers), various I / O controller peripherals, and routing logic.
[0128] In at least one embodiment, the processor 1110 may further include a safety cluster engine, which may include, without limitation, a dedicated processor subsystem for addressing safety management in automotive applications. In at least one embodiment, the safety cluster engine may include, without limitation, two or more processor cores, tightly coupled RAM, supporting peripherals (e.g., timers and interrupt controllers), and / or routing logic. In safety mode, in at least one embodiment, two or more cores operate in lockstep mode and may function as a single core with comparison logic for detecting any differences between their operations. In at least one embodiment, the processor 1110 may further include a real-time camera engine, which may include, without limitation, a dedicated processor subsystem for addressing real-time camera management. In at least one embodiment, the processor 1110 may further include a high dynamic range signal processor, which may include, without limitation, an image signal processor that is a hardware engine that is part of a camera processing pipeline.
[0129] In at least one embodiment, the processor 1110 may include a video image synthesizer, which may be a processing block (e.g., implemented on a microprocessor) that implements video post-processing functions required by a video playback application to produce a final image for the playback device window. In at least one embodiment, the video image synthesizer may perform lens distortion correction on the wide-angle camera 1170, the ambient camera 1174, and / or the in-cabin surveillance camera sensor. In at least one embodiment, the in-cabin surveillance camera sensor is preferably monitored by a neural network running on another instance of the SoC 1104, which is configured to identify events in the cabin and respond thereto accordingly. In at least one embodiment, the in-cabin system may perform lip-reading, without limitation, to activate cellular services, make phone calls, write emails, change the vehicle's destination, activate or change the vehicle's infotainment system and settings, and provide voice-activated web surfing. In at least one embodiment, certain functions are available to the driver when the vehicle is operating in autonomous mode and unavailable at other times.
[0130] In at least one embodiment, the video image synthesizer may include extended temporal noise reduction for both spatial and temporal noise reduction. For example, in at least one embodiment, if motion occurs in the video, the noise reduction appropriately weights the spatial information to reduce the weight of the information provided by adjacent frames. In at least one embodiment, if the image or part of the image does not contain motion, the temporal noise reduction performed by the video image synthesizer may use information from previous images to reduce noise in the current image.
[0131] In at least one embodiment, the video image synthesizer may also be configured to perform stereo parallelization on the input stereo lens frame. In at least one embodiment, the video image synthesizer may further be used to synthesize the user interface when the operating system desktop is in use, so that the GPU 1108 does not need to continuously render new surfaces. In at least one embodiment, when the GPU 1108 is powered on and actively performing 3D rendering, the video image synthesizer may be used to offload the GPU 1108 to improve performance and responsiveness.
[0132] In at least one embodiment, one or more of the SoC1104 SoCs may further include a camera serial interface, a high-speed interface, and / or a video input block which may be used for input functions of the camera and associated pixels of a Mobile Industry Processor Interface ("MIPI") for receiving video and camera inputs. In at least one embodiment, one or more of the SoC1104 SoCs may further include an input / output controller which may be software-controlled and may be used to receive I / O signals that are not tied to a specific role.
[0133] In at least one embodiment, one or more of the SoCs 1104 may further include a broad peripheral interface for enabling communication with peripheral devices, audio encoders / decoders ("codecs"), power management, and / or other devices. In at least one embodiment, the SoC 1104 may be used to process data from cameras (connected, for example, via Gigabit Multimedia Serial Link and Ethernet® channels), data from sensors (e.g., Lidar sensor 1164, Radar sensor 1160, etc., which may be connected via Ethernet® channels), data from bus 1102 (e.g., vehicle speed, steering wheel position, etc.), data from GNSS sensor 1158 (connected, for example, via Ethernet® bus or CAN bus), and the like. In at least one embodiment, one or more of the SoCs 1104 may further include a dedicated high-performance mass storage controller, which may include its own DMA engine and may be used to free the CPU 1106 from routine data management tasks.
[0134] In at least one embodiment, the SoC1104 may be an end-to-end platform with a flexible architecture spanning automation levels 3 to 5, thereby providing a comprehensive functional safety architecture that leverages and efficiently utilizes computer vision and ADAS techniques to achieve diversity and redundancy, and a flexible, reliable driving software stack along with deep learning tools. In at least one embodiment, the SoC1104 is faster, more reliable, and more energy-efficient and space-efficient than conventional systems. For example, in at least one embodiment, the accelerator 1114, when combined with the CPU 1106, GPU 1108, and data store 1116, can realize a fast and efficient platform for Level 3 to 5 autonomous vehicles.
[0135] In at least one embodiment, the computer vision algorithm may run on a CPU, which may be constructed using a high-level programming language such as C, and may execute a variety of processing algorithms across diverse visual data. However, in at least one embodiment, the CPU often fails to meet the performance requirements of many computer vision applications, such as requirements regarding execution time and power consumption. In at least one embodiment, many CPUs are unable to execute complex object detection algorithms used in in-vehicle ADAS applications and in realistic Level 3-5 autonomous vehicles in real time.
[0136] The embodiments described herein allow multiple neural networks to run simultaneously and / or sequentially, and the results can be combined to enable Level 3–5 autonomous driving capabilities. For example, in at least one embodiment, a DLA or a CNN running on a separate GPU (e.g., GPU1120) may include text and word recognition, enabling the neural network to read and understand traffic signs, including signs that the neural network has not been specifically trained to read. In at least one embodiment, the DLA may further include a neural network that can identify and interpret signs, provide a semantic understanding of the signs, and pass that semantic understanding to a route planning module running on a CPU complex.
[0137] In at least one embodiment, multiple neural networks may run simultaneously with respect to Level 3, 4, or 5 driving. For example, in at least one embodiment, a warning sign that displays "Caution: Flashing indicates frozen conditions" along with an electric light may be interpreted separately or collectively by several neural networks. In at least one embodiment, the warning sign itself may be identified as a traffic sign by a first introduced neural network (e.g., a trained neural network), and the words "Flashing indicates frozen conditions" may be interpreted by a second introduced neural network, which, if flashing light is detected, notifies the vehicle's route planning software (preferably running on the CPU complex) that frozen conditions are present. In at least one embodiment, the flashing light may also be identified by running a third introduced neural network over multiple frames, and the presence (or absence) of the flashing light is notified to the vehicle's route planning software. In at least one embodiment, all three neural networks may run simultaneously within the DLA and / or on the GPU1108, etc.
[0138] In at least one embodiment, a CNN for facial recognition and vehicle owner identification may use data from camera sensors to identify the presence of an authorized driver and / or owner of vehicle 1100. In at least one embodiment, an always-on sensor processing engine may be used to unlock the vehicle and turn on the lights when the owner approaches the driver's door, and to disable the vehicle in security mode when the owner leaves the vehicle. Thus, SoC1104 provides security against theft and / or vehicle hijacking.
[0139] In at least one embodiment, the CNN for emergency vehicle detection and identification may use data from microphone 1196 to detect and identify emergency vehicle sirens. In at least one embodiment, SoC 1104 uses the CNN to classify visual data as well as environmental and urban sounds. In at least one embodiment, the CNN running on DLA is trained to identify the relative speed of approaching emergency vehicles (for example, by using the Doppler effect). In at least one embodiment, the CNN may also be trained to identify emergency vehicles specific to the area in which the vehicle is operating, which are identified by GNSS sensor 1158. In at least one embodiment, if operating in Europe, the CNN attempts to detect European sirens, and if in North America, it attempts to identify only North American sirens. In at least one embodiment, when an emergency vehicle is detected, a control program for executing an emergency vehicle safety routine may be used to slow down the vehicle, pull over to the side of the road, stop the vehicle, and / or idle the vehicle using ultrasonic sensor 1162 until the emergency vehicle has passed.
[0140] In at least one embodiment, the vehicle 1100 may include a CPU 1118 (e.g., a separate CPU or dCPU) which may be coupled to the SoC 1104 via a high-speed interconnect (e.g., PCIe). In at least one embodiment, the CPU 1118 may include, for example, an x86 processor. The CPU 1118 may be used to perform any of a variety of functions, including, for example, mediating potentially inconsistent results between ADAS sensors and the SoC 1104, and / or monitoring the status and health of the controller 1136 and / or the infotainment system ("Infotainment SoC") 1130 on the chip.
[0141] In at least one embodiment, the vehicle 1100 may include a GPU 1120 (e.g., a separate GPU or dGPU) which may be coupled to the SoC 1104 via a high-speed interconnect (e.g., NVIDIA's NVLINK channel). In at least one embodiment, the GPU 1120 may provide additional artificial intelligence capabilities, such as by running redundant and / or different neural networks, and may be used to train and / or update neural networks based at least in part on input from the vehicle 1100's sensors (e.g., sensor data).
[0142] In at least one embodiment, the vehicle 1100 may further include a network interface 1124, which may include, but is not limited to, a wireless antenna 1126 (e.g., one or more wireless antennas for different communication protocols, such as a cellular antenna or a Bluetooth antenna). In at least one embodiment, the network interface 1124 may be used to enable wireless connectivity to Internet cloud services (e.g., servers and / or other network devices) with other vehicles and / or computing devices (e.g., occupant client devices). In at least one embodiment, a direct link may be established between the vehicle 110 and other vehicles for communication with other vehicles, and / or an indirect link may be established (e.g., over a network and via the Internet). In at least one embodiment, the direct link may be provided using a vehicle-to-vehicle communication link. In at least one embodiment, the vehicle-to-vehicle communication link may provide the vehicle 1100 with information about nearby vehicles (e.g., vehicles in front of, to the side of, and / or behind the vehicle 1100). In at least one embodiment, the aforementioned functions may be part of the cooperative adaptive cruise control function of the vehicle 1100.
[0143] In at least one embodiment, the network interface 1124 may include an SoC that provides modulation and demodulation functions, enabling the controller 1136 to communicate over a wireless network. In at least one embodiment, the network interface 1124 may include a radio frequency front end for up-conversion from baseband to radio frequency and down-conversion from radio frequency to baseband. In at least one embodiment, frequency conversion may be performed in any technically feasible manner. For example, frequency conversion can be performed by a well-known process and / or using a superheterodyne process. In at least one embodiment, the radio frequency front end functionality may be provided by a separate chip. In at least one embodiment, the network interface may include wireless functionality for communication over LTE, WCDMA, UMTS, GSM, CDMA2000, Bluetooth, Bluetooth LE, Wi-Fi, Z-Wave, ZigBee, LoRaWAN, and / or other wireless protocols.
[0144] In at least one embodiment, the vehicle 1100 may further include a data store 1128, which may include, without limitation, off-chip (e.g., not on the SoC 1104) storage. In at least one embodiment, the data store 1128 may include, without limitation, one or more storage elements, including RAM, SRAM, dynamic random-access memory ("DRAM"), video random-access memory ("VRAM"), flash memory, hard disk, and / or other components and / or devices capable of storing at least one bit of data.
[0145] In at least one embodiment, the vehicle 1100 may further include GNSS sensors 1158 (e.g., GPS and / or auxiliary GPS sensors) to assist in mapping, perception, occupancy grid generation, and / or route planning functions. In at least one embodiment, any number of GNSS sensors 1158, including, for example, GPS, may be used, using a USB connector having an Ethernet® to serial (e.g., RS-232) bridge.
[0146] In at least one embodiment, the vehicle 1100 may further include a RADAR sensor 1160. In at least one embodiment, the RADAR sensor 1160 may be used by the vehicle 1100 to perform long-range vehicle detection even in darkness and / or severe weather conditions. In at least one embodiment, the functional safety level of the RADAR may be ASIL B. In at least one embodiment, the RADAR sensor 1160 may use a CAN bus and / or bus 1102 for control (e.g., to transmit data generated by the RADAR sensor 1160) and to access object tracking data, and in some examples, it may have access to an Ethernet® channel to access raw data. In at least one embodiment, various types of RADAR sensors may be used. For example, without limitation, the RADAR sensor 1160 may be suitable for forward, rear, and side RADAR use. In at least one embodiment, one or more of the sensors of the RADAR sensor 1160 are pulsed Doppler RADAR sensors.
[0147] In at least one embodiment, the RADAR sensor 1160 may include different configurations, such as a narrow-field long-range, a wide-field short-range, and a lateral-covering short-range. In at least one embodiment, the long-range RADAR may be used for adaptive cruise control functionality. In at least one embodiment, the long-range RADAR system may provide a wide field of view, such as within 250 m (meters), achieved by two or more independent scans. In at least one embodiment, the RADAR sensor 1160 may be designed to facilitate the distinction between static and moving objects and may be used by an ADAS system 1138 to provide emergency braking assistance and forward collision warning. In at least one embodiment, the sensor 1160 included in the long-range RADAR system may include, without limitation, multiple (e.g., six or more) fixed RADAR antennas, as well as a monostatic multimode RADAR having high-speed CAN and FlexRay interfaces. In at least one embodiment, if there are six antennas, the four central antennas may generate a concentrated beam pattern designed to record the area around vehicle 1100 at a higher speed with minimal interference from adjacent lanes. In at least one embodiment, the other two antennas may extend the field of view, enabling rapid detection of vehicles entering or leaving the lane of vehicle 1100.
[0148] In at least one embodiment, the medium-range RADAR system may include, for example, a range of up to 160 m (forward) or 80 m (rearward) and a field of view of up to 42 degrees (forward) or 150 degrees (rearward). In at least one embodiment, the short-range RADAR system may include, without limitation, any number of RADAR sensors 1160 designed to be installed at both ends of the rear bumper. When installed at both ends of the rear bumper, in at least one embodiment, the RADAR sensor system may generate two beams that constantly monitor blind spots in the rearward and adjacent to the vehicle. In at least one embodiment, the short-range RADAR system may be used in an ADAS system 1138 to perform blind spot detection and / or lane change assistance.
[0149] In at least one embodiment, the vehicle 1100 may further include an ultrasonic sensor 1162. In at least one embodiment, the ultrasonic sensor 1162 may be positioned in front of, behind, and / or laterally to the vehicle 1100 and may be used for parking assistance and / or to generate and update the occupancy grid. In at least one embodiment, a variety of ultrasonic sensors 1162 may be used, and different ultrasonic sensors 1162 may be used for different detection ranges (e.g., 2.5m, 4m). In at least one embodiment, the ultrasonic sensor 1162 may operate at functional safety level ASIL B.
[0150] In at least one embodiment, the vehicle 1100 may include a LiDAR sensor 1164. In at least one embodiment, the LiDAR sensor 1164 may be used for object and pedestrian detection, emergency braking, collision avoidance, and / or other functions. In at least one embodiment, the LiDAR sensor 1164 may operate at functional safety level ASIL B. In at least one embodiment, the vehicle 1100 may include a plurality of LiDAR sensors 1164 (e.g., two, four, six, etc.), and these sensors may use an Ethernet® channel (e.g., to provide data to a Gigabit Ethernet® switch).
[0151] In at least one embodiment, the LIDAR sensor 1164 may be capable of providing a list of objects and their distances over a 360-degree field of view. In at least one embodiment, a commercially available LIDAR sensor 1164 may have, for example, an advertised range of approximately 100m, an accuracy of 2cm to 3cm, and support a 100Mbps Ethernet® connection. In at least one embodiment, one or more non-protruding LIDAR sensors may be used. In such embodiments, the LIDAR sensor 1164 may include small devices that can be incorporated into the front, rear, side, and / or corner positions of the vehicle 1100. In at least one embodiment, the LIDAR sensor 1164 of such embodiments may provide a horizontal field of view of up to 120 degrees and a vertical field of view of 35 degrees, even for low-reflectivity objects, over a range of 200m. In at least one embodiment, a front-mounted LIDAR sensor 1164 may be configured to provide a horizontal field of view of 45 to 135 degrees.
[0152] In at least one embodiment, LiDAR technology such as 3D flash LiDAR may also be used. In at least one embodiment, the 3D flash LiDAR uses a laser flash as a source to illuminate the area around the vehicle 1100 up to approximately 200 m. In at least one embodiment, the flash LiDAR unit includes, but is not limited to, a receptor which records the transit time of the laser pulse and the reflected light at each pixel, corresponding to the range from the vehicle 1100 to the object. In at least one embodiment, the flash LiDAR enables the generation of a highly accurate and distortion-free ambient image with each laser flash. In at least one embodiment, four flash LiDARs may be introduced, one on each side of the vehicle 1100. In at least one embodiment, the 3D flash LiDAR system includes, but is not limited to, a semiconductor 3D staring array LiDAR camera (e.g., a non-scanning LiDAR device) with no moving parts other than a fan. In at least one embodiment, the flash LiDAR device may use a Class I (eye-safe) laser pulse of 5 nanoseconds per frame to capture the reflected laser light as a 3D range point cloud and position-synchronized (co-registered) intensity data.
[0153] In at least one embodiment, the vehicle 1100 may further include an IMU sensor 1166. In at least one embodiment, the IMU sensor 1166 may be positioned in the center of the rear axle of the vehicle 1100. In at least one embodiment, the IMU sensor 1166 may include, for example, an accelerometer, a magnetometer, a gyroscope, a magnetic compass, multiple magnetic compasses, and / or other types of sensors, without limitation. In at least one embodiment, such as a 6-axis application, the IMU sensor 1166 may include, without limitation, an accelerometer and a gyroscope. In at least one embodiment, such as a 9-axis application, the IMU sensor 1166 may include, without limitation, an accelerometer, a gyroscope, and a magnetometer.
[0154] In at least one embodiment, the IMU sensor 1166 may be implemented as a small, high-performance GPS-Aided Inertial Navigation System ("GPS / INS") that combines a micro-electro-mechanical system ("MEMS") inertial sensor, a high-sensitivity GPS receiver, and an advanced Kalman filtering algorithm to provide estimates of position, velocity, and attitude. In at least one embodiment, the IMU sensor 1166 allows the vehicle 1100 to estimate its bearing without requiring input from magnetic sensors by directly observing velocity changes and correlating them from GPS to the IMU sensor 1166. In at least one embodiment, the IMU sensor 1166 and the GNSS sensor 1158 may be combined into a single integrated unit.
[0155] In at least one embodiment, the vehicle 1100 may include a microphone 1196 installed inside and / or around the vehicle 1100. In at least one embodiment, the microphone 1196 may be used, in particular, for the detection and identification of emergency vehicles.
[0156] In at least one embodiment, the vehicle 1100 may further include any number of camera types, including a stereo camera 1168, a wide-angle camera 1170, an infrared camera 1172, a perimeter camera 1174, a long-range camera 1198, a medium-range camera 1176, and / or other camera types. In at least one embodiment, the cameras may be used to capture image data around the entire perimeter of the vehicle 1100. In at least one embodiment, the type of camera used will vary depending on the vehicle 1100. In at least one embodiment, any combination of camera types may be used to provide the required coverage area around the vehicle 1100. In at least one embodiment, the number of cameras introduced may vary depending on the embodiment. For example, in at least one embodiment, the vehicle 1100 may include six cameras, seven cameras, ten cameras, twelve cameras, or any other number of cameras. In at least one embodiment, the camera may support Gigabit Multimedia Serial Link ("GMSL") and / or Gigabit Ethernet® communication, without being limited to the example. In at least one embodiment, each camera may be as described above in further detail with respect to Figures 11A and 11B.
[0157] In at least one embodiment, the vehicle 1100 may further include a vibration sensor 1142. In at least one embodiment, the vibration sensor 1142 may measure vibrations of components of the vehicle 1100, such as axles. For example, in at least one embodiment, a change in vibration may indicate a change in the road surface. In at least one embodiment, if two or more vibration sensors 1142 are used, the difference in vibration may be used to determine the amount of friction or slip on the road surface (for example, if there is a difference in vibration between a power-driven axle and a freely rotating axle).
[0158] In at least one embodiment, the vehicle 1100 may include an ADAS system 1138. In at least one embodiment, the ADAS system 1138 may include a SoC in some examples, without limitation. In at least one embodiment, the ADAS system 1138 may include, without limitation, any number and any combination of autonomous / adaptive / automatic cruise control ("ACC") systems, cooperative adaptive cruise control ("CACC") systems, forward crash warning ("FCW") systems, automatic emergency braking ("AEB") systems, lane departure warning ("LDW") systems, lane keep assist ("LKA") systems, blind spot warning ("BSW") systems, rear cross-traffic warning ("RCTW") systems, collision warning ("CW") systems, lane centering ("LC") systems, and / or other systems, features, and / or functions.
[0159] In at least one embodiment, the ACC system may use a RADAR sensor 1160, a LIDAR sensor 1164, and / or any number of cameras. In at least one embodiment, the ACC system may include a longitudinal ACC system and / or a lateral ACC system. In at least one embodiment, the longitudinal ACC system monitors and controls the distance of vehicle 1100 to another vehicle directly in front and automatically adjusts the speed of vehicle 1100 to maintain a safe distance from the vehicle ahead. In at least one embodiment, the lateral ACC system performs distance maintenance and notifies vehicle 1100 to change lanes when necessary. In at least one embodiment, the lateral ACC is related to other ADAS applications such as LC and CW.
[0160] In at least one embodiment, the CACC system uses information from other vehicles, which may be received from other vehicles via a network interface 1124 and / or wireless antenna 1126, either via a wireless link or indirectly via a network connection (e.g., via the Internet). In at least one embodiment, the link may be provided directly by a vehicle-to-vehicle ("V2V") communication link, while the link may be provided indirectly by an infrastructure-to-vehicle ("I2V") communication link. Generally, V2V communication provides information about the vehicle immediately ahead (e.g., a vehicle in the same lane immediately ahead of vehicle 1100), and I2V communication provides information about traffic further ahead. In at least one embodiment, the CACC system may include either or both I2V and V2V information sources. In at least one embodiment, having information about the vehicle ahead of vehicle 1100 can further enhance the reliability of the CACC system, potentially leading to smoother traffic flow and reduced congestion on the road.
[0161] In at least one embodiment, the FCW system is designed to alert the driver to hazardous materials so that such a driver can take corrective action. In at least one embodiment, the FCW system uses a front camera and / or a radar sensor 1160, which are coupled to a dedicated processor, DSP, FPGA, and / or ASIC that are electrically coupled to provide feedback to the driver, such as a display, speaker, and / or vibration component. In at least one embodiment, the FCW system may provide warnings in the form of sound, visual warnings, vibration, and / or quick brake pulses.
[0162] In at least one embodiment, the AEB system may detect an imminent head-on collision with another vehicle or other object and automatically apply the brakes if the driver does not take corrective action within a specified time or distance parameter. In at least one embodiment, the AEB system may use a front camera and / or RADAR sensor 1160 coupled to a dedicated processor, DSP, FPGA, and / or ASIC. In at least one embodiment, when the AEB system detects a hazardous object, the AEB system usually first advises the driver to take corrective action to avoid the collision, and if the driver does not take corrective action, the AEB system may automatically apply the brakes to prevent the anticipated collision or at least mitigate its impact. In at least one embodiment, the AEB system may include techniques such as dynamic brake support and / or pre-collision braking.
[0163] In at least one embodiment, the LDW system provides visual, auditory, and / or tactile warnings, such as vibration of the steering wheel or seat, to alert the driver when the vehicle 1100 crosses a lane marker. In at least one embodiment, the LDW system does not activate if the driver indicates an intentional lane departure, such as by activating the turn signal. In at least one embodiment, the LDW system may use a front camera, which is coupled to a dedicated processor, DSP, FPGA, and / or ASIC that can be electrically coupled to provide feedback to the driver, such as a display, speaker, and / or vibration components. In at least one embodiment, the LKA system is a variation of the LDW system. In at least one embodiment, the LKA system provides steering input or brake control to correct the vehicle 1100 if the vehicle 1100 begins to drift out of its lane.
[0164] In at least one embodiment, the BSW system detects vehicles in the vehicle's blind spot and warns the driver. In at least one embodiment, the BSW system may provide visual, auditory, and / or tactile alerts to indicate that merging or changing lanes is unsafe. In at least one embodiment, the BSW system may provide additional warnings when the driver uses the turn signal. In at least one embodiment, the BSW system may use a rear camera and / or radar sensor 1160 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback such as a display, speaker, and / or vibration components.
[0165] In at least one embodiment, the RCTW system may provide visual, auditory, and / or tactile notifications when an object is detected outside the range of the rear camera while the vehicle 1100 is reversing. In at least one embodiment, the RCTW system includes an AEB system to ensure that the vehicle brakes are applied to avoid a collision. In at least one embodiment, the RCTW system may use one or more rear-facing RADAR sensors 1160, which are coupled to a dedicated processor, DSP, FPGA, and / or ASIC that are electrically coupled to provide feedback to the driver, such as a display, speaker, and / or vibration components.
[0166] In at least one embodiment, the conventional ADAS system 1138s may be prone to producing false positives, which can be annoying and distracting to the driver, but is usually not a major issue. This is because the conventional ADAS system 1138s advises the driver, allowing the driver to determine whether a safety-critical condition truly exists and whether appropriate action should be taken. In at least one embodiment, if the results are contradictory, the vehicle 1100 itself determines whether to follow the result from the primary computer (e.g., the first controller of the controller 1136) or the result from the secondary computer (e.g., the second controller of the controller 1136). For example, in at least one embodiment, the ADAS system 1138 may be a backup and / or secondary computer for resisting perceptual information to the rationality module of the backup computer. In at least one embodiment, the rationality monitor of the backup computer may run various software redundancies on hardware components to detect perceptual errors and dynamic driving tasks. In at least one embodiment, the output from the ADAS system 1138 may be provided to a monitoring MCU. In at least one embodiment, if the output from the primary computer and the output from the secondary computer are inconsistent, the monitoring MCU determines how to reconcile the inconsistency to ensure safe operation.
[0167] In at least one embodiment, the primary computer may be configured to provide the monitoring MCU with a reliability score indicating the reliability of the result selected by the primary computer. In at least one embodiment, if the reliability score exceeds a threshold, the monitoring MCU may follow the instructions of the primary computer, regardless of whether the secondary computer provides contradictory or inconsistent results. In at least one embodiment, if the reliability score does not satisfy the threshold and the primary and secondary computers provide different results (e.g., contradictory), the monitoring MCU may mediate between the computers to determine an appropriate result.
[0168] In at least one embodiment, the monitoring MCU may be configured to run a neural network trained and configured to determine, at least in part, the conditions under which a secondary computer provides a false alarm, based on the output from the primary computer and the output from the secondary computer. In at least one embodiment, the neural network of the monitoring MCU may learn when the output from the secondary computer may be trusted and when it may not be trusted. For example, in at least one embodiment, if the secondary computer is a RADAR-based FCW system, the neural network of the monitoring MCU may learn when the FCW system identifies a metallic object that is not actually a hazard, such as a drain grate or manhole cover, which triggers an alarm. In at least one embodiment, if the secondary computer is a camera-based LDW system, the neural network of the monitoring MCU may learn to disable the LDW when there are cyclists or pedestrians and lane departure is actually the safest operation. In at least one embodiment, the monitoring MCU may include at least one of a DLA or GPU suitable for running the neural network with associated memory. In at least one embodiment, the monitoring MCU may comprise and / or be included as a component of the SoC1104.
[0169] In at least one embodiment, the ADAS system 1138 may include a secondary computer that performs ADAS functions using conventional computer vision rules. In at least one embodiment, the secondary computer may use conventional computer vision rules (if-then rules), and reliability, safety, and performance may be improved by the presence of a neural network in the monitoring MCU. For example, in at least one embodiment, diverse implementations and intentional non-identities increase the overall fault tolerance of the system, particularly to errors caused by the functionality of the software (or software-hardware interface). For example, in at least one embodiment, if there is a bug or error in the software running on the primary computer, and non-identical software code running on the secondary computer provides an overall consistent result, the monitoring MCU may have greater confidence that the overall result is correct and that the software or hardware bug on the primary computer has not caused a critical error.
[0170] In at least one embodiment, the output of the ADAS system 1138 may be supplied to the perception block of the primary computer and / or the dynamic driving task block of the primary computer. For example, in at least one embodiment, if the ADAS system 1138 is issuing a head-on collision warning due to an object immediately preceding, the perception block may use this information when identifying the object. In at least one embodiment, the secondary computer may have its own trained, and therefore false-detection-reducing, neural network, as described herein.
[0171] In at least one embodiment, the vehicle 1100 may further include an infotainment SoC 1130 (for example, an in-vehicle infotainment system (IVI)). The infotainment system 1130 is illustrated and described as an SoC, but in at least one embodiment, it does not have to be an SoC and may include, without limitation, two or more separate components. In at least one embodiment, the infotainment SoC 1130 may include, without limitation, a combination of hardware and software that can be used to provide the vehicle 1100 with audio (e.g., music, personal digital assistant, navigation commands, news, radio, etc.), video (e.g., TV, movies, streaming, etc.), telephone (e.g., hands-free calling), network connectivity (e.g., LTE, Wi-Fi, etc.), and / or information services (e.g., a navigation system, rear parking assist, wireless data system, vehicle-related information, such as fuel level, total mileage, brake fuel level, oil level, door open / closed, air filter information, etc.). For example, the infotainment SoC 1130 may include a radio, disc player, navigation system, video player, USB and Bluetooth connectivity, car computer, in-car entertainment, Wi-Fi, steering wheel audio control, hands-free voice control, head-up display ("HUD"), HMI display 1134, telematics device, control panel (for example, for controlling and / or interacting with various components, features, and / or systems), and / or other components. In at least one embodiment, the infotainment SoC 1130 may further be used to provide the user of the vehicle 1100 with (for example, visual and / or auditory) information such as information from the ADAS system 1138, autonomous driving information such as vehicle operation plans and trajectories, ambient information (for example, intersection information, vehicle information, road information, etc.), and / or other information.
[0172] In at least one embodiment, the infotainment SoC 1130 may include any amount and type of GPU functionality. In at least one embodiment, the infotainment SoC 1130 may communicate with other devices, systems, and / or components of the vehicle 1100 via bus 1102. In at least one embodiment, the infotainment SoC 1130 may be coupled to a monitoring MCU so that the GPU of the infotainment system can perform some self-driving functions when the primary controller 1136 (e.g., the primary and / or backup computer of the vehicle 1100) fails. In at least one embodiment, the infotainment SoC 1130 may put the vehicle 1100 into driver-safe stop mode as described herein.
[0173] In at least one embodiment, the vehicle 1100 may further include an instrument cluster 1132 (e.g., a digital dashboard, electronic instrument cluster, digital instrument panel, etc.). In at least one embodiment, the instrument cluster 1132 may include, without limitation, a controller and / or a supercomputer (e.g., a separate controller or supercomputer). In at least one embodiment, the instrument cluster 1132 may include, without limitation, any number and combination of instrument sets such as a speedometer, fuel level, oil pressure, tachometer, odometer, turn signals, shift lever position indicator, seat belt warning light, parking brake warning light, engine fault light, auxiliary restraint system (e.g., airbag) information, light control, safety system control, navigation information, etc. In some examples, information may be displayed and / or shared between the infotainment SoC 1130 and the instrument cluster 1132. In at least one embodiment, the instrument cluster 1132 may be included as part of the infotainment SoC 1130, or vice versa.
[0174] The inference and / or training logic 815 is used to perform inference and / or training operations related to one or more embodiments. Details relating to the inference and / or training logic 815 are provided herein in conjunction with Figures 8A and / or 8B. In at least one embodiment, the inference and / or training logic 815 may be used in the system of Figure 11C for inference or prediction operations, at least in part, based on weight parameters calculated using the neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0175] Figure 11D is a diagram of a system for communication between a cloud-based server and the autonomous vehicle 1100 of Figure 11A, according to at least one embodiment. In at least one embodiment, the system may include, without limitation, a server 1178, a network 1190, and any number and type of vehicles, including the vehicle 1100. In at least one embodiment, the server 1178 may include, without limitation, a plurality of GPUs 1184(A) to 1184(H) (collectively referred to herein as GPU 1184), PCIe switches 1182(A) to 1182(D) (collectively referred to herein as PCIe switches 1182), and / or CPUs 1180(A) to 1180(B) (collectively referred to herein as CPU 1180). In at least one embodiment, the GPU 1184, CPU 1180, and PCIe switch 1182 may be interconnected by high-speed interconnects such as the NVLink interface 1188 and / or PCIe connection 1186 developed by NVIDIA, for example, without limitation. In at least one embodiment, the GPUs 1184 are connected to each other via NVLink and / or NVS switch SoC, and the GPUs 1184 and PCIe switches 1182 are connected via PCIe interconnects. Eight GPUs 1184, two CPUs 1180, and four PCIe switches 1182 are shown, but this is not limited to them. In at least one embodiment, each server 1178 may include any number of GPUs 1184, CPUs 1180, and / or PCIe switches 1182 in any combination, without limitation. For example, in at least one embodiment, each server 1178 may include 8, 16, 32, and / or more GPUs 1184.
[0176] In at least one embodiment, server 1178 may receive image data from vehicles via network 1190 representing images showing unexpected or altered road conditions, such as recently started road construction. In at least one embodiment, server 1178 may transmit map information 1194, including updated or unupdated neural networks 1192 and / or information about traffic and road conditions, without limitation, to vehicles via network 1190. In at least one embodiment, updates to map information 1194 may include, without limitation, updates to HD map 1122, such as information about construction sites, potholes, detours, floods, and / or other obstacles. In at least one embodiment, neural networks 1192 and / or map information 1194 may be derived from new training and / or experience represented in data received from any number of vehicles in the environment, and / or may be derived at least in part from training performed in a data center (for example, using server 1178 and / or other servers).
[0177] In at least one embodiment, a machine learning model (e.g., a neural network) may be trained using server 1178, at least in part, on training data. In at least one embodiment, the training data may be generated by the vehicle and / or generated by simulation (e.g., using a game engine). In at least one embodiment, any amount of training data is tagged and / or otherwise preprocessed (e.g., if the relevant neural network benefits from supervised learning). In at least one embodiment, any amount of training data is not tagged and / or preprocessed (e.g., if the relevant neural network does not require supervised learning). In at least one embodiment, once the machine learning model is trained, the machine learning model may be used by the vehicle (e.g., transmitted to the vehicle via network 1190 and / or the machine learning model may be used by server 1178 to remotely monitor the vehicle).
[0178] In at least one embodiment, server 1178 may receive data from the vehicle and apply the data to a state-of-the-art real-time neural network to enable real-time intelligent reasoning. In at least one embodiment, server 1178 may include a deep learning supercomputer and / or dedicated AI computer powered by a GPU 1184, such as the DGX and DGX Station Machine developed by NVIDIA. However, in at least one embodiment, server 1178 may include a deep learning infrastructure using a CPU-powered data center.
[0179] In at least one embodiment, the deep learning infrastructure of server 1178 may be capable of high-speed real-time inference and may use this capability to evaluate and verify the health of the vehicle's processor, software, and / or associated hardware. For example, in at least one embodiment, the deep learning infrastructure may receive periodic updates from vehicle 1100, such as a series of images and / or objects located in that series of images (e.g., by computer vision and / or other machine learning object classification techniques). In at least one embodiment, the deep learning infrastructure may run its own neural network to identify objects and compare them to objects identified by vehicle 1100. If the results do not match and the deep learning infrastructure concludes that the vehicle's AI is malfunctioning, server 1178 may send a signal to vehicle 1100 instructing the vehicle's fail-safe computer to take control, notify the occupants, and complete a safe stopping operation.
[0180] In at least one embodiment, server 1178 may include a GPU 1184 and one or more programmable inference accelerators (e.g., NVIDIA TensorRT3 devices). In at least one embodiment, a server powered by a GPU and inference acceleration can be combined to enable real-time response. In at least one embodiment, a server powered by a CPU, FPGA, and other processors may be used for inference, for example, when performance is not critical. In at least one embodiment, a hardware structure 815 is used to perform one or more embodiments. Details relating to the hardware structure 815 are provided herein in conjunction with Figures 8A and / or 8B.
[0181] Computer system Figure 12 is a block diagram showing an exemplary computer system, which may be a system having interconnected devices and components, a system-on-a-chip (SoC), or any combination thereof, formed together with a processor which may include an execution unit for executing instructions, in at least one embodiment. In at least one embodiment, computer system 1200 may include, without limitation, components such as processor 1202 for using an execution unit which includes logic for executing algorithms for processing data in accordance with the disclosure, such as in the embodiments described herein. In at least one embodiment, computer system 1200 may include a processor such as the PENTIUM® processor family, Xeon™, Itanium®, XScale™ and / or StrongARM™, Intel® Core®, or Intel® Nervana® microprocessors, available from Intel Corporation in Santa Clara, California, but other systems may be used (including PCs with other microprocessors, engineering workstations, set-top boxes, etc.). In at least one embodiment, the computer system 1200 may run a version of the WINDOWS® operating system available from Microsoft Corporation in Redmond, Washington, but other operating systems (e.g., UNIX® and Linux®), embedded software, and / or graphical user interfaces may be used.
[0182] The embodiments may be used in other devices, such as portable devices and embedded applications. Some examples of portable devices include cellular phones, Internet Protocol devices, digital cameras, personal digital assistants ("PDAs"), and portable PCs. In at least one embodiment, the embedded application may include a microcontroller, a digital signal processor ("DSP"), a system-on-a-chip, a network computer ("NetPC"), a set-top box, a network hub, a wide area network ("WAN") switch, or any other system capable of executing one or more instructions according to at least one embodiment.
[0183] In at least one embodiment, the computer system 1200 may include, without limitation, a processor 1202 which may include, without limitation, one or more execution units 1208 for training and / or inferring machine learning models using the techniques described herein. In at least one embodiment, the computer system 1200 is a single-processor desktop or server system, but in another embodiment, the computer system 1200 may be a multi-processor system. In at least one embodiment, the processor 1202 may include, without limitation, a complex instruction set computer ("CISC") microprocessor, a reduced instruction set computing ("RISC") microprocessor, a very long instruction word ("VLIW") microprocessor, a processor implementing a combination of instruction sets, or any other processor device, such as a digital signal processor. In at least one embodiment, the processor 1202 may be coupled to a processor bus 1210 which may transmit digital signals between the processor 1202 and other components in the computer system 1200.
[0184] In at least one embodiment, the processor 1202 may include, without limitation, a level 1 ("L1") internal cache memory ("cache") 1204. In at least one embodiment, the processor 1202 may have a single internal cache or multiple levels of internal caches. In at least one embodiment, the cache memory may be external to the processor 1202. Other embodiments may include a combination of both internal and external caches, depending on the specific implementation and requirements. In at least one embodiment, the register file 1206 may store different types of data in various registers, including, without limitation, integer registers, floating-point registers, state registers, and instruction pointer registers.
[0185] In at least one embodiment, the processor 1202 also includes an execution unit 1208 which includes, without limitation, logic for performing integer and floating-point arithmetic. In at least one embodiment, the processor 1202 may also include a microcode ("u-code") read-only memory ("ROM") for storing microcode for certain macro instructions. In at least one embodiment, the execution unit 1208 may include logic for handling a packed instruction set 1209. In at least one embodiment, by including the packed instruction set 1209, along with the associated circuitry for executing the instructions, in the instruction set of a general-purpose processor, arithmetic used by many multimedia applications can be performed using the packed data of the processor 1202. In at least one embodiment, by performing arithmetic on the packed data using the full width of the processor's data bus, many multimedia applications can be accelerated and run more efficiently, thereby eliminating the need to transfer smaller units of data across the processor's data bus to perform one or more arithmetic operations on a single data element at a time.
[0186] In at least one embodiment, the execution unit 1208 may also be used in a microcontroller, embedded processor, graphics device, DSP, and other types of logic circuits. In at least one embodiment, the computer system 1200 may include, without limitation, memory 1220. In at least one embodiment, memory 1220 may be a dynamic random access memory ("DRAM") device, a static random access memory ("SRAM") device, a flash memory device, or other memory device. In at least one embodiment, memory 1220 may store instructions 1219 and / or data 1221, which may be represented by data signals executed by the processor 1202.
[0187] In at least one embodiment, a system logic chip may be coupled to a processor bus 1210 and memory 1220. In at least one embodiment, the system logic chip may include, without limitation, a memory controller hub ("MCH") 1216, and the processor 1202 may communicate with the MCH 1216 via the processor bus 1210. In at least one embodiment, the MCH 1216 may provide a high-bandwidth memory path 1218 to memory 1220 for storing instructions and data, and for storing graphics commands, data, and textures. In at least one embodiment, the MCH 1216 may lead data signals between the processor 1202, memory 1220, and other components of the computer system 1200, and may bridge data signals between the processor bus 1210, memory 1220, and system I / O interface 1222. In at least one embodiment, the system logic chip may provide a graphics port for coupling to a graphics controller. In at least one embodiment, the MCH1216 may be coupled to memory 1220 via a high-bandwidth memory path 1218, and the graphics / video card 1212 may be coupled to the MCH1216 via an Accelerated Graphics Port ("AGP") interconnect 1214.
[0188] In at least one embodiment, the computer system 1200 may use the system I / O interface 1222 as a proprietary hub interface bus for coupling the MCH 1216 to the I / O controller hub ("ICH") 1230. In at least one embodiment, the ICH 1230 may provide direct connectivity to several I / O devices via a local I / O bus. In at least one embodiment, the local I / O bus may include, without limitation, a high-speed I / O bus for connecting peripherals to memory 1220, the chipset, and the processor 1202. Examples may include, but are not limited to, an audio controller 1229, a firmware hub ("Flash BIOS") 1228, a wireless transceiver 1226, data storage 1224, a legacy I / O controller 1223 including a user input and keyboard interface 1225, a serial expansion port 1227 such as a Universal Serial Bus ("USB") port, and a network controller 1234. In at least one embodiment, the data storage 1224 may comprise a hard disk drive, a floppy disk drive, a CD-ROM device, a flash memory device, or other mass storage device.
[0189] In at least one embodiment, Figure 12 shows a system including interconnected hardware devices or “chips,” while in other embodiments, Figure 12 may show an exemplary SoC. In at least one embodiment, the devices shown in Figure 12 may be interconnected by proprietary interconnects, standard interconnects (e.g., PCIe), or any combination thereof. In at least one embodiment, one or more components of the computer system 1200 may be interconnected using a compute express link (CXL) interconnect.
[0190] The inference and / or training logic 815 is used to perform inference and / or training operations related to one or more embodiments. Details relating to the inference and / or training logic 815 are provided herein in conjunction with Figures 8A and / or 8B. In at least one embodiment, the inference and / or training logic 815 may be used in the system of Figure 12 for inference or prediction operations, at least in part, based on weight parameters calculated using the neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0191] Figure 13 is a block diagram showing an electronic device 1300 for utilizing a processor 1310, according to at least one embodiment. In at least one embodiment, the electronic device 1300 may be, for example, a notebook, a tower server, a rack server, a blade server, a laptop, a desktop, a tablet, a mobile device, a telephone, an embedded computer, or any other suitable electronic device, without limitation.
[0192] In at least one embodiment, the electronic device 1300 may include, without limitation, a processor 1310 communicatively coupled to any number or type of preferred components, peripherals, modules, or devices. In at least one embodiment, the processor 1310 is I 2The devices are coupled using buses or interfaces such as the C-bus, System Management Bus ("SMBus"), Low Pin Count (LPC) bus, Serial Peripheral Interface ("SPI"), High Definition Audio ("HDA") bus, Serial Advance Technology Attachment ("SATA") bus, Universal Serial Bus ("USB") (versions 1, 2, 3, etc.), or Universal Asynchronous Receiver / Transmitter ("UART") bus. In at least one embodiment, Figure 13 shows a system including interconnected hardware devices or "chips," while in other embodiments, Figure 13 may show an exemplary SoC. In at least one embodiment, the devices shown in Figure 13 may be interconnected by proprietary interconnects, standard interconnects (e.g., PCIe), or any combination thereof. In at least one embodiment, one or more components of Figure 13 may be interconnected using Compute Express Link (CXL) interconnects.
[0193] In at least one embodiment, Figure 13 shows a display 1324, a touch screen 1325, a touch pad 1330, a Near Field Communications unit ("NFC") 1345, a sensor hub 1340, a thermal sensor 1346, an Express Chipset ("EC") 1335, a Trusted Platform Module ("TPM") 1338, a BIOS / firmware / flash memory ("BIOS, FW flash") 1322, a DSP 1360, a drive 1320 such as a Solid State Disk ("SSD") or Hard Disk Drive ("HDD"), a Wireless Local Area Network Unit ("WLAN") 1350, a Bluetooth unit 1352, and a Wireless Wide Area Network Unit ("WWAN"). The components may include a USB 3.0 camera unit 1356, a Global Positioning System (GPS) unit 1355, a camera such as a USB 3.0 camera ("USB 3.0 camera") 1354, and / or a Low Power Double Data Rate ("LPDDR") memory unit ("LPDDR3") 1315, for example, implemented in the LPDDR3 standard. Each of these components may be implemented in any preferred manner.
[0194] In at least one embodiment, other components may be communicatively coupled to the processor 1310 via the components described above. In at least one embodiment, the accelerometer 1341, ambient light sensor ("ALS") 1342, compass 1343, and gyroscope 1344 may be communicatively coupled to the sensor hub 1340. In at least one embodiment, the thermal sensor 1339, fan 1337, keyboard 1336, and touchpad 1330 may be communicatively coupled to the EC 1335. In at least one embodiment, the speaker 1363, headphones 1364, and microphone ("mic") 1365 may be communicatively coupled to an audio unit (audio codec and class D amplifier) 1362, which may be communicatively coupled to the DSP 1360. In at least one embodiment, the audio unit 1362 may include, for example, an audio coder / decoder ("codec") and a class D amplifier. In at least one embodiment, the SIM card ("SIM") 1357 may be communicatively coupled to the WWAN unit 1356. In at least one embodiment, components such as the WLAN unit 1350 and the Bluetooth unit 1352, as well as the WWAN 1356, may be implemented in a Next Generation Form Factor ("NGFF").
[0195] The inference and / or training logic 815 is used to perform inference and / or training operations related to one or more embodiments. Details relating to the inference and / or training logic 815 are provided herein in conjunction with Figures 8A and / or 8B. In at least one embodiment, the inference and / or training logic 815 may be used in the system of Figure 13 for inference or prediction operations, at least in part, based on weight parameters calculated using the neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0196] Figure 14 shows a computer system 1400 according to at least one embodiment. In at least one embodiment, the computer system 1400 is configured to implement various processes and methods described throughout this disclosure.
[0197] In at least one embodiment, the computer system 1400 includes, but is not limited to, at least one central processing unit ("CPU") 1402, which is connected to a communications bus 1410 implemented using any preferred protocol, such as PCI:Peripheral Component Interconnect ("Peripheral Component Interconnect"), Peripheral Component Interconnect Express ("PCI-Express"), AGP:Accelerated Graphics Port ("Accelerated Graphics Port"), Hypertransport, or any other bus or point-to-point communications protocol. In at least one embodiment, the computer system 1400 includes, but is not limited to, main memory 1404 and control logic (implemented, for example, as hardware, software, or a combination thereof), and data is stored in the main memory 1404, which may take the form of random access memory ("RAM"). In at least one embodiment, the network interface subsystem ("Network Interface") 1422 provides an interface with other computing devices and networks for receiving data from other systems having computer system 1400 and transmitting data to other systems having computer system 1400.
[0198] In at least one embodiment, the computer system 1400 includes, in at least one embodiment without limitation, an input device 1408, a parallel processing system 1412, and a display device 1406, which can be implemented using a conventional cathode ray tube ("CRT"), a liquid crystal display ("LCD"), a light-emitting diode ("LED") display, a plasma display, or other suitable display technology. In at least one embodiment, user input is received from the input device 1408, such as a keyboard, mouse, touchpad, or microphone. In at least one embodiment, each module described herein can be placed on a single semiconductor platform to form a processing system.
[0199] The inference and / or training logic 815 is used to perform inference and / or training operations related to one or more embodiments. Details relating to the inference and / or training logic 815 are provided herein in conjunction with Figures 8A and / or 8B. In at least one embodiment, the training logic 815 may be used in the system of Figure 14 for inference or prediction operations, at least in part, based on weight parameters calculated using the neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0200] Figure 15 shows a computer system 1500 according to at least one embodiment. In at least one embodiment, the computer system 1500 may include, without limitation, a computer 1510 and a USB stick 1520. In at least one embodiment, the computer system 1510 may include, without limitation, any number and type of processors (not shown), as well as memory. In at least one embodiment, the computer 1510 may include, without limitation, a server, a cloud instance, a laptop, and a desktop computer.
[0201] In at least one embodiment, the USB stick 1520 includes, but is not limited to, a processing unit 1530, a USB interface 1540, and a USB interface logic 1550. In at least one embodiment, the processing unit 1530 may be any instruction execution system, apparatus, or device capable of executing instructions. In at least one embodiment, the processing unit 1530 may include, but is not limited to, any number and type of processing cores (not shown). In at least one embodiment, the processing unit 1530 comprises an application-specific integrated circuit ("ASIC") optimized to perform any amount and type of operations related to machine learning. For example, in at least one embodiment, the processing unit 1530 is a tensor processing unit ("TPC") optimized to perform machine learning inference operations. In at least one embodiment, the processing unit 1530 is a vision processing unit ("VPU") optimized to perform machine vision and machine learning inference operations.
[0202] In at least one embodiment, the USB interface 1540 may be any type of USB connector or USB socket. For example, in at least one embodiment, the USB interface 1540 is a USB 3.0 Type-C socket for data and power. In at least one embodiment, the USB interface 1540 is a USB 3.0 Type-A connector. In at least one embodiment, the USB interface logic 1550 may include any amount and type of logic that enables the processing unit 1530 to interface with or to a device (e.g., a computer 1510) via the USB connector 1540.
[0203] The inference and / or training logic 815 is used to perform inference and / or training operations related to one or more embodiments. Details relating to the inference and / or training logic 815 are provided herein in conjunction with Figures 8A and / or 8B. In at least one embodiment, the inference and / or training logic 815 may be used in the system of Figure 15 for inference or prediction operations, at least in part, based on weight parameters calculated using the neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0204] Figure 16A shows an exemplary architecture in which multiple GPUs 1610(1)–1610(N) are communicably coupled to multiple multi-core processors 1605(1)–1605(M) via high-speed links 1640(1)–1640(N) (e.g., bus, point-to-point interconnect). In at least one embodiment, the high-speed links 1640(1)–1640(N) support communication throughput of 4 GB / s, 30 GB / s, 80 GB / s, or higher. Various interconnection protocols may be used, including but not limited to PCIe 4.0 or 5.0 and NVLink 2.0. In various figures, "N" and "M" represent positive integers whose values may differ from figure to figure.
[0205] Furthermore, in at least one embodiment, two or more of the GPUs 1610 are interconnected via high-speed links 1629(1) to 1629(2), which may be implemented using similar or different protocols / links as those used for high-speed links 1640(1) to 1640(N). Similarly, two or more of the multi-core processors 1605 may be connected via high-speed link 1628, which can be a symmetric multiprocessor (SMP) bus operating at 20 GB / s, 30 GB / s, 120 GB / s, or higher. Alternatively, all communication between the various system components shown in Figure 16A may be implemented using similar protocols / links (for example, via a common interconnection fabric).
[0206] In at least one embodiment, each multi-core processor 1605 is communicatively coupled to processor memory 1601(1) to 1601(M) via memory interconnects 1626(1) to 1626(M), and each GPU 1610(1) to 1610(N) is communicatively coupled to GPU memory 1620(1) to 1620(N) via GPU memory interconnects 1650(1) to 1650(N). In at least one embodiment, the memory interconnects 1626 and 1650 may utilize similar or different memory access techniques. For example, but not limited to, the processor memories 1601(1) to 1601(M) and the GPU memory 1620 may be volatile memory such as dynamic random access memory (DRAM) (including stacked DRAM), graphics DDR SDRAM (GDDR) (e.g., GDDR5, GDDR6), or high-bandwidth memory (HBM), and / or non-volatile memory such as 3D XPoint or Nano-Ram. In at least one embodiment, some portions of the processor memory 1601 may be volatile memory and other portions may be non-volatile memory (e.g., using a two-level memory (2LM) hierarchy).
[0207] As described herein, various multi-core processors 1605 and GPUs 1610 may be physically coupled to specific memories 1601 and 1620, respectively, and / or an integrated memory architecture may be implemented in which the address space of a virtual system (also called the “effective address” space) is distributed among various physical memories. For example, each processor memory 1601(1) to 1601(M) may have a system memory address space of 64GB, and each GPU memory 1620(1) to 1620(N) may have a system memory address space of 32GB, resulting in a total of 256GB of addressable memory when M=2 and N=4. Other values for N and M are possible.
[0208] Figure 16B shows further details of the interconnection between a multi-core processor 1607 and a graphics acceleration module 1646 in one exemplary embodiment. In at least one embodiment, the graphics acceleration module 1646 may include one or more GPU chips integrated on a line card coupled to the processor 1607 via a high-speed link 1640 (e.g., PCIe bus, NVLink, etc.). In at least one embodiment, or the graphics acceleration module 1646 may be integrated on a package or chip having the processor 1607.
[0209] In at least one embodiment, the processor 1607 includes a plurality of cores 1660A to 1660D, each core having a translation lookaside buffer ("TLB") 1661A to 1661D and one or more caches 1662A to 1662D. In at least one embodiment, cores 1660A to 1660D may include various other components (not shown) for executing instructions and processing data. In at least one embodiment, caches 1662A to 1662D may have level 1 (L1) and level 2 (L2) caches. Furthermore, one or more shared caches 1656 may be included in caches 1662A to 1662D and shared by the set of cores 1660A to 1660D. For example, one embodiment of the processor 1607 includes 24 cores, each core having its own L1 cache, 12 shared L2 caches, and 12 shared L3 caches. In this embodiment, one or more L2 and L3 caches are shared by two adjacent cores. In at least one embodiment, the processor 1607 and the graphics acceleration module 1646 are connected to system memory 1614, which may include processor memories 1601(1) to 1601(M) in Figure 16A.
[0210] In at least one embodiment, coherence is maintained between the data and instructions stored in the various caches 1662A-1662D, 1656, and system memory 1614 by inter-core communication via the coherence bus 1664. In at least one embodiment, for example, each cache may have associated cache coherence logic / circuits to communicate via the coherence bus 1664 in response to detecting a read or write to a particular cache line. In at least one embodiment, a cache snooping protocol is implemented via the coherence bus 1664 to monitor cache access.
[0211] In at least one embodiment, the proxy circuit 1625 communicatively couples the graphics acceleration module 1646 to the coherence bus 1664, enabling the graphics acceleration module 1646 to participate in the cache coherence protocol as a peer of cores 1660A-1660D. In particular, in at least one embodiment, interface 1635 provides a connection to the proxy circuit 1625 via high-speed link 1640, and interface 1637 connects the graphics acceleration module 1646 to high-speed link 1640.
[0212] In at least one embodiment, the accelerator integration circuit 1636 provides cache management, memory access, content management, and interrupt management services on behalf of the multiple graphics processing engines 1631(1) to 1631(N) of the graphics acceleration module 1646. In at least one embodiment, each of the graphics processing engines 1631(1) to 1631(N) may comprise a separate graphics processing unit (GPU). In at least one embodiment, or the graphics processing engines 1631(1) to 1631(N) may comprise different types of graphics processing engines within the GPU, such as a graphics execution unit, a media processing engine (e.g., a video encoder / decoder), a sampler, and a bullet engine. In at least one embodiment, the graphics acceleration module 1646 may be a GPU having a plurality of graphics processing engines 1631(1) to 1631(N), or the graphics processing engines 1631(1) to 1631(N) may be individual GPUs integrated on a common package, line card, or chip.
[0213] In at least one embodiment, the accelerator integration circuit 1636 includes a memory management unit (MMU) 1639 for performing various memory management functions, such as virtual-to-physical memory translation (also known as effective-to-real memory translation), and a memory access protocol for accessing system memory 1614. In at least one embodiment, the MMU 1639 may also include a translation lookaside buffer (TLB) (not shown) for caching virtual / effective-to-physical / real address translations. In at least one embodiment, the cache 1638 can store commands and data so that they can be efficiently accessed by graphics processing engines 1631(1) to 1631(N). In at least one embodiment, data stored in cache 1638 and graphics memory 1633(1)-1633(M) is kept coherent with core caches 1662A-1662D, 1656 and system memory 1614, possibly using a fetch unit 1644. As stated above, this may be achieved via a proxy circuit 1625 instead of cache 1638 and memory 1633(1)-1633(M) (for example, by sending updates regarding cache line modifications / access in processor caches 1662A-1662D, 1656 to cache 1638 and receiving updates from cache 1638).
[0214] In at least one embodiment, a set of registers 1645 stores context data for threads executed by graphics processing engines 1631(1) to 1631(N), and a context management circuit 1648 manages the thread contexts. For example, the context management circuit 1648 may perform save and restore operations to save and restore the contexts of various threads during a context switch (for example, the first thread is saved and the second thread is stored so that the second thread can be executed by the graphics processing engine). For example, during a context switch, the context management circuit 1648 may store the current register values in a designated area of memory (for example, identified by a context pointer). Then, when returning to the context, the context management circuit 1648 may restore the register values. In at least one embodiment, an interrupt management circuit 1647 receives and processes interrupts received from system devices.
[0215] In at least one embodiment, virtual / effective addresses from the graphics processing engine 1631 are translated by the MMU 1639 to real / physical addresses in system memory 1614. In at least one embodiment, one embodiment of the accelerator integration circuit 1636 supports multiple (e.g., 4, 8, or 16) graphics accelerator modules 1646 and / or other accelerator devices. In at least one embodiment, the graphics accelerator module 1646 may be dedicated to a single application running on the processor 1607, or it may be shared among multiple applications. In at least one embodiment, there exists a virtualized graphics execution environment in which the resources of the graphics processing engines 1631(1) to 1631(N) are shared among multiple applications or virtual machines (VMs). In at least one embodiment, the resources may be subdivided into "slices" which are allocated to different VMs and / or applications based on processing requirements and the priority associated with the VMs and / or applications.
[0216] In at least one embodiment, the accelerator integration circuit 1636 functions as a bridge to the system for the graphics acceleration module 1646 and provides address translation and system memory caching services. Furthermore, in at least one embodiment, the accelerator integration circuit 1636 may provide virtualization facilities for the host processor to manage the virtualization, interrupts, and memory management of the graphics processing engines 1631(1) to 1631(N).
[0217] In at least one embodiment, the hardware resources of the graphics processing engines 1631(1) to 1631(N) are explicitly mapped to the real address space seen by the host processor 1607, so that any host processor can directly address these resources using effective address values. In at least one embodiment, one function of the accelerator integration circuit 1636 is to physically isolate the graphics processing engines 1631(1) to 1631(N) so that they appear as independent units to the system.
[0218] In at least one embodiment, one or more graphics memories 1633(1) to 1633(M) are each coupled to each of the graphics processing engines 1631(1) to 1631(N), where N=M. In at least one embodiment, the graphics memories 1633(1) to 1633(M) store instructions and data processed by each of the graphics processing engines 1631(1) to 1631(N). In at least one embodiment, the graphics memories 1633(1) to 1633(M) may be volatile memory such as DRAM (including stacked DRAM), GDDR memory (e.g., GDDR5, GDDR6), or HBM, and / or non-volatile memory such as 3D XPoint or Nano-Ram.
[0219] In at least one embodiment, biasing techniques may be used to reduce data traffic over the high-speed link 1640, such that the data stored in the graphics memory 1633(1)-1633(M) is the data that will be most frequently used by the graphics processing engines 1631(1)-1631(N), and preferably the data that will not be used (or at least not frequently used) by the cores 1660A-1660D. Similarly, in at least one embodiment, the biasing mechanism attempts to keep the data that the cores need (and therefore preferably not needed by the graphics processing engines 1631(1)-1631(N)) in the core caches 1662A-1662D, 1656, and system memory 1614.
[0220] Figure 16C shows another exemplary embodiment in which the accelerator integration circuit 1636 is integrated within the processor 1607. In this embodiment at least, the graphics processing engines 1631(1) to 1631(N) communicate directly with the accelerator integration circuit 1636 via the high-speed link 1640 through interfaces 1637 and 1635 (which can also be any form of bus or interface protocol). In at least one embodiment, the accelerator integration circuit 1636 may perform operations similar to those described with respect to Figure 16B, but may potentially operate at higher throughput given its proximity to the coherence bus 1664 and caches 1662A to 1662D, 1656. In at least one embodiment, the accelerator integration circuit supports different programming models, including a dedicated process programming model (without virtualization of the graphics acceleration module) and a shared programming model (with virtualization), which may include a programming model controlled by the accelerator integration circuit 1636 and a programming model controlled by the graphics acceleration module 1646.
[0221] In at least one embodiment, the graphics processing engines 1631(1) to 1631(N) are dedicated to a single application or process under a single operating system. In at least one embodiment, a single application can achieve virtualization within a VM / partition by directing the requests of other applications to the graphics processing engines 1631(1) to 1631(N).
[0222] In at least one embodiment, the graphics processing engines 1631(1) to 1631(N) may be shared by multiple VM / application partitions. In at least one embodiment, the sharing model may use a system hypervisor to virtualize the graphics processing engines 1631(1) to 1631(N) to allow access by each operating system. In at least one embodiment, in a single-partition system without a hypervisor, the graphics processing engines 1631(1) to 1631(N) are owned by the operating system. In at least one embodiment, the operating system can virtualize the graphics processing engines 1631(1) to 1631(N) to provide access to each process or application.
[0223] In at least one embodiment, the graphics acceleration module 1646 or the individual graphics processing engines 1631(1) to 1631(N) select a process element using a process handle. In at least one embodiment, the process element is stored in system memory 1614 and is addressable using the effective address-to-actual address translation technique described herein. In at least one embodiment, the process handle may be an implementation-specific value provided to the host process when registering the host process context with the graphics processing engines 1631(1) to 1631(N) (i.e., calling system software to add a process element to the process element link list). In at least one embodiment, the lower 16 bits of the process handle may be the offset of the process element in the process element link list.
[0224] Figure 16D shows an exemplary accelerator integration slice 1690. In at least one embodiment, the “slice” comprises a portion of the processing resources of the accelerator integration circuit 1636. In at least one embodiment, the application effective address space 1682 in system memory 1614 stores a process element 1683. In at least one embodiment, the process element 1683 is stored in response to a GPU call 1681 from an application 1680 running on processor 1607. In at least one embodiment, the process element 1683 contains the process state of the corresponding application 1680. In at least one embodiment, the work descriptor (WD) 1684 contained in the process element 1683 may be a single job requested by the application, or it may contain a pointer to a queue of jobs. In at least one embodiment, the WD 1684 is a pointer to a job request queue in the application’s effective address space 1682.
[0225] In at least one embodiment, the graphics acceleration module 1646 and / or individual graphics processing engines 1631(1) to 1631(N) can be shared by all or a subset of processes in the system. In at least one embodiment, infrastructure may be included for setting process states and sending WD1684 to the graphics acceleration module 1646 to start jobs in a virtualized environment.
[0226] In at least one embodiment, the dedicated process programming model is implementation-specific. In at least one embodiment, in this model, a single process owns the graphics acceleration module 1646 or the individual graphics processing engines 1631. In at least one embodiment, when the graphics acceleration module 1646 is owned by a single process, when the graphics acceleration module 1646 is allocated, the hypervisor initializes the accelerator integration circuit 1636 for the owning partition, and the operating system initializes the accelerator integration circuit 1636 for the owning process.
[0227] In at least one embodiment, during operation, a WD fetch unit 1691 in the accelerator integration slice 1690 fetches the next WD 1684 containing a representation of the work to be performed by one or more graphics processing engines of the graphics acceleration module. In at least one embodiment, as illustrated, the data from WD 1684 may be stored in register 1645 and used by the MMU 1639, interrupt management circuit 1647, and / or context management circuit 1648. For example, one embodiment of the MMU 1639 includes a segment / page walk circuit for accessing a segment / page table 1686 in the OS virtual address space 1685. In at least one embodiment, the interrupt management circuit 1647 may process an interrupt event 1692 received from the graphics acceleration module 1646. In at least one embodiment, when performing graphics operations, the effective address 1693 generated by the graphics processing engines 1631(1) to 1631(N) is translated to a real address by the MMU 1639.
[0228] In at least one embodiment, register 1645 may be duplicated for each graphics processing engine 1631(1) to 1631(N) and / or graphics acceleration module 1646 and initialized by the hypervisor or operating system. In at least one embodiment, each of these duplicated registers may be included in the accelerator integration slice 1690. Exemplary registers that may be initialized by the hypervisor are shown in Table 1. [Table 1]
[0229] Table 2 shows exemplary registers that may be initialized by the operating system. [Table 2]
[0230] In at least one embodiment, each WD1684 is specific to a particular graphics acceleration module 1646 and / or graphics processing engine 1631(1) to 1631(N). In at least one embodiment, the WD1684 may contain all the information necessary for the graphics processing engine 1631(1) to 1631(N) to perform the work, or it may be a pointer to a memory location where the application has set up a command queue for the work to be completed.
[0231] Figure 16E provides further details of an exemplary embodiment of the shared model. This embodiment includes a hypervisor real address space 1698 in which the process element list 1699 is stored. In at least one embodiment, the hypervisor real address space 1698 is accessible via a hypervisor 1696 that virtualizes the graphics acceleration module engine of the operating system 1695.
[0232] In at least one embodiment, a shared programming model allows all or a subset of processes from all or a subset of partitions in the system to use the graphics acceleration module 1646. In at least one embodiment, there are two programming models in which the graphics acceleration module 1646 is shared by multiple processes and partitions: time-slice sharing and graphics-directed sharing.
[0233] In at least one embodiment, in this model, the system hypervisor 1696 owns the graphics acceleration module 1646 and makes its functionality available to all operating systems 1695. In at least one embodiment, in order for the graphics acceleration module 1646 to support virtualization by the system hypervisor 1696, the graphics acceleration module 1646 may comply with several requirements, such as 1) application job requests must be autonomous (i.e., no state must be maintained between jobs), or the graphics acceleration module 1646 must provide a mechanism for saving and restoring context; 2) application job requests must be guaranteed by the graphics acceleration module 1646 to be completed within a specified amount of time, including any translation errors, or the graphics acceleration module 1646 must provide a function to preempt job processing; and 3) when operating in a specified shared programming model, the graphics acceleration module 1646 must ensure fairness between processes.
[0234] In at least one embodiment, application 1680 is required to make a system call to operating system 1695 with the graphics acceleration module type, a work descriptor (WD), an authorization mask register (AMR) value, and a context save / restore area pointer (CSRP). In at least one embodiment, the graphics acceleration module type describes the acceleration function intended by the system call. In at least one embodiment, the graphics acceleration module type may be a system-specific value. In at least one embodiment, the WD is specifically formatted for graphics acceleration module 1646 and can be in the form of a command for graphics acceleration module 1646, an effective address pointer to a user-defined structure, an effective address pointer to a queue of commands, or any other data structure for describing the work performed by graphics acceleration module 1646.
[0235] In at least one embodiment, the AMR value is the AMR state for use in the current process. In at least one embodiment, the value passed to the operating system is the same as that of the application setting the AMR. In at least one embodiment, if the implementation of the accelerator integration circuit 1636 (not shown) and the graphics acceleration module 1646 does not support a user privilege mask override register (UAMOR), the operating system may apply the current UAMOR value to the AMR value before passing the AMR to the hypervisor call. In at least one embodiment, the hypervisor 1696 may optionally apply the current privilege mask override register (AMOR) value before placing the AMR into the process element 1683. In at least one embodiment, CSRP is one of the registers 1645 that contains the effective address of an area in the application's effective address space 1682 for the graphics acceleration module 1646 to save and restore context state. In at least one embodiment, this pointer is optional if there is no need to save any state between jobs or when a job is preempted. In at least one embodiment, the context save / restore area may be pinned system memory.
[0236] Upon receiving the system call, the operating system 1695 may verify that application 1680 is registered and authorized to use the graphics acceleration module 1646. In at least one embodiment, the operating system 1695 then calls the hypervisor 1696 with the information shown in Table 3. [Table 3]
[0237] In at least one embodiment, upon receiving a hypervisor call, the hypervisor 1696 verifies that the operating system 1695 is registered and authorized to use the graphics acceleration module 1646. In at least one embodiment, the hypervisor 1696 then places the process element 1683 into a process element link list of the corresponding graphics acceleration module 1646 type. In at least one embodiment, the process element may include the information shown in Table 4. [Table 4]
[0238] In at least one embodiment, the hypervisor initializes the registers 1645 of multiple accelerator integration slices 1690.
[0239] As shown in Figure 16F, in at least one embodiment, integrated memory is used that is addressable via a common virtual memory address space used to access physical processor memories 1601(1) to 1601(N) and GPU memories 1620(1) to 1620(N). In this implementation, operations performed on GPUs 1610(1) to 1610(N) utilize the same virtual / effective memory address space as accessing processor memories 1601(1) to 1601(N), and vice versa, thereby simplifying programmability. In at least one embodiment, a first portion of the virtual / effective address space is allocated to processor memory 1601(1), a second portion to a second processor memory 1601(N), a third portion to GPU memory 1620(1), and so on. In at least one embodiment, the entire virtual / effective memory space (sometimes called the effective address space) is distributed across processor memory 1601 and GPU memory 1620, respectively, so that either processor or GPU can access either physical memory, with virtual addresses mapped to physical memory.
[0240] In at least one embodiment, bias / coherence management circuits 1694A-1694E in one or more of the MMUs 1639A-1639E ensure cache coherence between the cache of one or more host processors (e.g., 1605) and the cache of the GPU 1610, implement bias techniques to indicate physical memory where a particular type of data should be stored. In at least one embodiment, multiple instances of bias / coherence management circuits 1694A-1694E are shown in Figure 16F, but the bias / coherence circuits may be implemented within the MMU of one or more host processors 1605 and / or within the accelerator integration circuit 1636.
[0241] In one embodiment, GPU memory 1620 can be mapped as part of system memory and made accessible using shared virtual memory (SVM) techniques without the performance degradation associated with full system cache coherence. In at least one embodiment, the accessibility of GPU memory 1620 as system memory without cumbersome cache coherence overhead provides a beneficial operating environment for GPU offloading. In at least one embodiment, this configuration allows host processor 1605 software to set operands and access computation results without the overhead of conventional I / O DMA data copying. In at least one embodiment, such conventional copies require driver calls, interrupts, and memory-mapped I / O (MMIO) access, all of which are less efficient than simple memory access. In at least one embodiment, the ability to access GPU memory 1620 without cache coherence overhead may be essential for the execution time of offloaded computations. In at least one embodiment, for example, when there is significant streaming write memory traffic, cache coherence overhead can significantly reduce the effective write bandwidth seen by the GPU1610. In at least one embodiment, the efficiency of operand configuration, the efficiency of accessing results, and the efficiency of GPU computation can be helpful in determining the effectiveness of GPU offloading.
[0242] In at least one embodiment, the selection between GPU bias and host processor bias is determined by a bias tracker data structure. In at least one embodiment, a bias table may be used, for example, which may be a page-granular structure containing 1 or 2 bits per GPU-enabled memory page (for example, controlled by memory page granularity). In at least one embodiment, the bias table may be implemented in a stolen memory range of one or more GPU memories 1620, with or without a bias cache (for example, for caching frequently used / recently used entries in the bias table) present on the GPU 1610. Alternatively, in at least one embodiment, the entire bias table may be maintained within the GPU.
[0243] In at least one embodiment, an entry in the bias table associated with each access to the GPU-biased memory 1620 is accessed before the actual access to the GPU memory, resulting in the following behavior: In at least one embodiment, a local request from GPU 1610 to find its page in the GPU bias is forwarded directly to the corresponding GPU memory 1620. In at least one embodiment, a local request from a GPU to find its page in the host bias is forwarded to processor 1605 (for example, via the high-speed link described above). In at least one embodiment, a request from processor 1605 to find the requested page in the host processor bias completes the request in the same way as a normal memory read. Alternatively, a request directed to a GPU-biased page may be forwarded to GPU 1610. In at least one embodiment, the GPU may then move the page to the host processor bias if the page is not currently in use. In at least one embodiment, the bias state of a page can be changed by either a software-based mechanism, a hardware-assisted software-based mechanism, or, for a limited set of cases, simply by a hardware-based mechanism.
[0244] In at least one embodiment, one mechanism for changing the bias state utilizes an API call (e.g., OpenCL) which calls the GPU's device driver, which sends a message to the GPU (or queues a command descriptor) to change the bias state and, for some transitions, directs the GPU to perform a cache-flushing operation on the host. In at least one embodiment, the cache-flushing operation is used for transitions from a host processor 1605 bias to a GPU bias, but not for transitions in the opposite direction.
[0245] In at least one embodiment, cache coherence is maintained by temporarily rendering GPU-biased pages that cannot be cached by the host processor 1605. In at least one embodiment, to access these pages, the processor 1605 may request access from the GPU 1610, and the GPU 1610 may immediately grant access or not grant it. In at least one embodiment, it is therefore beneficial to have GPU-biased pages requested by the GPU but not by the host processor 1605, or vice versa, in order to reduce communication between the processor 1605 and the GPU 1610.
[0246] Hardware structure 815 is used to carry out one or more embodiments. Further details regarding hardware structure 815 may be provided herein in conjunction with Figures 8A and / or 8B.
[0247] Figure 17 shows exemplary integrated circuits and associated graphics processors that can be fabricated using one or more IP cores according to various embodiments described herein. In addition to those shown, at least one embodiment may include other logic and circuitry, including additional graphics processors / cores, peripheral interface controllers, or general-purpose processor cores.
[0248] FIG. 17 is a block diagram showing an exemplary system-on-chip integrated circuit 1700 that can be fabricated using one or more IP cores according to at least one embodiment. In at least one embodiment, integrated circuit 1700 includes one or more application processors 1705 (e.g., CPUs), at least one graphics processor 1710, and may further include an image processor 1715 and / or a video processor 1720, any of which may be a modular IP core. In at least one embodiment, integrated circuit 1700 includes a peripheral device or bus logic including a USB controller 1725, a UART controller 1730, an SPI / SDIO controller 1735, and an I 2 2S / I 2 2C controller 1740. In at least one embodiment, integrated circuit 1700 may include a display device 1745 coupled to one or more of a high-definition multimedia interface (HDMI™: high-definition multimedia interface™) controller 1750 and a mobile industry processor interface (MIPI) display interface 1755. In at least one embodiment, storage may be provided by a flash memory subsystem 1760 including a flash memory and a flash memory controller. In at least one embodiment, a memory interface may be provided via a memory controller 1765 to access a SDRAM or SRAM memory device. In at least one embodiment, some integrated circuits further include an embedded security engine 1770.
[0249] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 815 is used. Details regarding the inference and / or training logic 815 are provided herein in conjunction with FIGS. 8A and / or 8B. In at least one embodiment, the inference and / or training logic 815 may be used in the integrated circuit 1700 for inference or prediction operations, based at least in part on weight parameters calculated using the training operations, functions and / or architectures of the neural networks described herein, or the use cases of the neural networks.
[0250] FIGS. 18A-18B illustrate an exemplary integrated circuit and associated graphics processor that can be fabricated using one or more IP cores, according to various embodiments described herein. In addition to what is shown, in at least one embodiment, other logic and circuits may be included, including additional graphics processors / cores, peripheral device interface controllers, or general purpose processor cores.
[0251] FIGS. 18A-18B are block diagrams illustrating exemplary graphics processors for use within a SoC, according to embodiments described herein. FIG. 18A illustrates an exemplary graphics processor 1810 of a system-on-chip integrated circuit that can be fabricated using one or more IP cores, according to at least one embodiment. FIG. 18B illustrates a further exemplary graphics processor 1840 of a system-on-chip integrated circuit that can be fabricated using one or more IP cores, according to at least one embodiment. In at least one embodiment, the graphics processor 1810 of FIG. 18A is a low power graphics processor core. In at least one embodiment, the graphics processor 1840 of FIG. 18B is a high performance graphics processor core. In at least one embodiment, each of the graphics processors 1810, 1840 can be a variation of the graphics processor 1710 of FIG. 17.
[0252] In at least one embodiment, the graphics processor 1810 includes a vertex processor 1805 and one or more fragment processors 1815A to 1815N (e.g., 1815A, 1815B, 1815C, 1815D to 1815N-1, and 1815N). In at least one embodiment, the graphics processor 1810 can execute different shader programs via separate logic, thereby optimizing the vertex processor 1805 to perform operations for a vertex shader program, while one or more fragment processors 1815A to 1815N perform fragment (e.g., pixel) shading operations for a fragment or pixel shader program. In at least one embodiment, the vertex processor 1805 executes the vertex processing stage of the 3D graphics pipeline and generates primitive and vertex data. In at least one embodiment, the fragment processors 1815A to 1815N use primitive and vertex data generated by the vertex processor 1805 to generate a frame buffer for display on a display device. In at least one embodiment, the fragment processors 1815A to 1815N are optimized to execute fragment shader programs provided in the OpenGL API, and the OpenGL API may be used to perform similar operations to pixel shader programs provided in the Direct 3D API.
[0253] In at least one embodiment, the graphics processor 1810 further includes one or more memory management units (MMUs) 1820A-1820B, caches 1825A-1825B, and circuit interconnects 1830A-1830B. In at least one embodiment, one or more MMUs 1820A-1820B include vertex processors 1805 and / or fragment processors 1815A-1815N, providing virtual-to-physical address mappings for the graphics processor 1810, which may reference vertex or image / text data stored in memory, in addition to vertex or image / text data stored in one or more caches 1825A-1825B. In at least one embodiment, one or more MMUs 1820A-1820B may be synchronized with other MMUs in the system, including one or more MMUs associated with one or more application processors 1705, image processor 1715, and / or video processor 1720 in Figure 17, so that each processor 1705-1720 can participate in a shared or integrated virtual memory system. In at least one embodiment, one or more circuit interconnects 1830A-1830B allow the graphics processor 1810 to interface with other IP cores in the SoC via the SoC's internal bus or via a direct connection.
[0254] In at least one embodiment, as shown in Figure 18B, the graphics processor 1840 includes one or more shader cores 1855A-1855N (for example, 1855A, 1855B, 1855C, 1855D, 1855E, 1855F-1855N-1, and 1855N), which provide an integrated shader core architecture in which a single core, or type, or core can execute all types of programmable shader code, including shader program code for implementing vertex shaders, fragment shaders, and / or compute shaders. In at least one embodiment, the number of shader cores can be varied. In at least one embodiment, the graphics processor 1840 includes an intercore task manager 1845 that acts as a thread dispatcher for dispatching execution threads to one or more shader cores 1855A-1855N, and a tiling unit 1858 for accelerating tiling operations for tile-based rendering, where the rendering operation of a scene is subdivided in image space, for example, to take advantage of local space coherence in a scene or to optimize the use of an internal cache.
[0255] The inference and / or training logic 815 is used to perform inference and / or training operations related to one or more embodiments. Details relating to the inference and / or training logic 815 are provided herein in conjunction with Figures 8A and / or 8B. In at least one embodiment, the inference and / or training logic 815 may be used in integrated circuits 18A and / or 18B for inference or prediction operations, at least in part, based on weight parameters calculated using the neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0256] Figures 19A and 19B illustrate further exemplary graphics processor logic according to the embodiments described herein. Figure 19A shows a graphics core 1900, which in at least one embodiment may be included in the graphics processor 1710 of Figure 17, and in at least one embodiment may be an integrated shader core 1855A to 1855N, as shown in Figure 18B. Figure 19B shows a highly parallel general-purpose graphics processing unit ("GPGPU") 1930 suitable for deployment in a multi-chip module in at least one embodiment.
[0257] In at least one embodiment, the graphics core 1900 includes a shared instruction cache 1902, a texture unit 1918, and a cache / shared memory 1920, which are common to the execution resources within the graphics core 1900. In at least one embodiment, the graphics core 1900 may include multiple slices 1901A-1901N, or per-core partitions, and the graphics processor may include multiple instances of the graphics core 1900. In at least one embodiment, slices 1901A-1901N may include support logic including local instruction caches 1904A-1904N, thread schedulers 1906A-1906N, thread dispatchers 1908A-1908N, and register sets 1910A-1910N. In at least one embodiment, slices 1901A to 1901N may include a set of additional function units (AFU1912A to 1912N), floating-point units (FPU1914A to 1914N), integer arithmetic logic units (ALU1916 to 1916N), address calculation units (ACU1913A to 1913N), double-precision floating-point units (DPFPU1915A to 1915N), and matrix processing units (MPU1917A to 1917N).
[0258] In at least one embodiment, the FPU1914A-1914N can perform single-precision (32-bit) and half-precision (16-bit) floating-point operations, and the DPFPU1915A-1915N can perform double-precision (64-bit) floating-point operations. In at least one embodiment, the ALU1916A-1916N can perform variable-precision integer operations with 8-bit, 16-bit, and 32-bit precision, and can be configured to perform mixed-precision operations. In at least one embodiment, the MPU1917A-1917N can also be configured to perform mixed-precision matrix operations, including half-precision floating-point and 8-bit integer operations. In at least one embodiment, the MPU1917A-1917N can perform various matrix operations to accelerate machine learning application frameworks, including supporting General-Purpose Matrix Multiplication (GEMM) acceleration. In at least one embodiment, AFU1912A~1912N can perform additional logical operations not supported by the floating-point unit or integer unit, including trigonometric function operations (e.g., sine, cosine, etc.).
[0259] The inference and / or training logic 815 is used to perform inference and / or training operations related to one or more embodiments. Details relating to the inference and / or training logic 815 are provided herein in conjunction with Figures 8A and / or 8B. In at least one embodiment, the inference and / or training logic 815 may be used in the graphics core 1900 for inference or prediction operations, at least in part, based on weight parameters calculated using the neural network training operations, neural network functions and / or architecture, or neural network use cases described herein.
[0260] Figure 19B shows a General Purpose Processing Unit (GPGPU) 1930, which can be configured in at least one embodiment to enable highly parallel computational operations by an array of graphics processing units. In at least one embodiment, the GPGPU 1930 can be directly linked to other instances of the GPGPU 1930 to generate multiple GPU clusters to improve the training speed of deep neural networks. In at least one embodiment, the GPGPU 1930 includes a host interface 1932 for enabling connectivity with a host processor. In at least one embodiment, the host interface 1932 is a PCI Express interface. In at least one embodiment, the host interface 1932 can be a vendor-specific communication interface or communication fabric. In at least one embodiment, the GPGPU 1930 receives commands from the host processor and uses a global scheduler 1934 to distribute the execution threads associated with these commands to a set of compute clusters 1936A-1936H. In at least one embodiment, compute clusters 1936A to 1936H share a cache memory 1938. In at least one embodiment, the cache memory 1938 can act as a high-level cache for the cache memory within compute clusters 1936A to 1936H.
[0261] In at least one embodiment, the GPGPU 1930 includes memories 1944A to 1944B coupled to compute clusters 1936A to 1936H via a set of memory controllers 1942A to 1942B. In at least one embodiment, the memories 1944A to 1944B may include various types of memory devices, including dynamic random access memory (DRAM) or graphics random access memory, such as synchronous graphics random access memory (SGRAM) including graphics double data rate (GDDR) memory.
[0262] In at least one embodiment, each of the compute clusters 1936A to 1936H includes a set of graphics cores, such as the graphics core 1900 in Figure 19A, which may include multiple types of integer and floating-point logic units capable of performing computational operations at varying precisions, including those suitable for machine learning computations. For example, in at least one embodiment, at least a subset of the floating-point units in each of the compute clusters 1936A to 1936H may be configured to perform 16-bit or 32-bit floating-point operations, while another subset of the floating-point units may be configured to perform 64-bit floating-point operations.
[0263] In at least one embodiment, multiple instances of GPGPU 1930 can be configured to operate as a compute cluster. In at least one embodiment, the communication used for synchronization and data exchange by compute clusters 1936A-1936H differs across embodiments. In at least one embodiment, multiple instances of GPGPU 1930 communicate via a host interface 1932. In at least one embodiment, GPGPU 1930 includes an I / O hub 1939, which couples GPGPU 1930 to a GPU link 1940, enabling direct connections to other instances of GPGPU 1930. In at least one embodiment, the GPU link 1940 is coupled to a dedicated GPU-to-GPU bridge, enabling communication and synchronization between multiple instances of GPGPU 1930. In at least one embodiment, the GPU link 1940 is coupled to a high-speed interconnect for sending and receiving data to and from other GPGPUs or parallel processors. In at least one embodiment, multiple instances of the GPGPU 1930 are located on separate data processing systems and communicate via a network device accessible through the host interface 1932. In at least one embodiment, the GPU link 1940 can be configured to enable connection to a host processor in addition to, or instead of, the host interface 1932.
[0264] In at least one embodiment, the GPGPU 1930 can be configured to train a neural network. In at least one embodiment, the GPGPU 1930 can be used within an inference platform. In at least one embodiment where the GPGPU 1930 is used for inference, the GPGPU 1930 may include fewer compute clusters 1936A-1936H than when the GPGPU 1930 is used to train a neural network. In at least one embodiment, the memory technology associated with memory 1944A-1944B may differ between the inference configuration and the training configuration, with high-bandwidth memory technology being used in the training configuration. In at least one embodiment, the inference configuration of the GPGPU 1930 can support inference-specific instructions. For example, in at least one embodiment, the inference configuration can support one or more 8-bit integer dot product instructions, which may be used during the inference operation of a deployed neural network.
[0265] The inference and / or training logic 815 is used to perform inference and / or training operations related to one or more embodiments. Details relating to the inference and / or training logic 815 are provided herein in conjunction with Figures 8A and / or 8B. In at least one embodiment, the inference and / or training logic 815 may be used in the GPGPU 1930 for inference or prediction operations, at least in part, based on weight parameters calculated using the neural network training operations, neural network functions and / or architecture, or neural network use cases described herein.
[0266] Figure 20 is a block diagram of a computing system 2000 according to at least one embodiment. In at least one embodiment, the computing system 2000 includes a processing subsystem 2001 having one or more processors 2002 and system memory 2004 communicating via an interconnection path which may include a memory hub 2005. In at least one embodiment, the memory hub 2005 may be a separate component within a chipset component or may be integrated within one or more processors 2002. In at least one embodiment, the memory hub 2005 is coupled to an I / O subsystem 2011 via a communication link 2006. In at least one embodiment, the I / O subsystem 2011 includes an I / O hub 2007 which can enable the computing system 2000 to receive input from one or more input devices 2008. In at least one embodiment, the I / O hub 2007 may enable a display controller, which may be included in one or more processors 2002 and provide output to one or more display devices 2010A. In at least one embodiment, one or more display devices 2010A coupled to the I / O hub 2007 may include local, internal, or embedded display devices.
[0267] In at least one embodiment, the processing subsystem 2001 includes one or more parallel processors 2012 coupled to a memory hub 2005 via a bus or other communication link 2013. In at least one embodiment, the communication link 2013 may use one of any number of communication link technologies or protocols based on standards such as PCI Express, or it may be a vendor-specific communication interface or communication fabric. In at least one embodiment, one or more parallel processors 2012 may form a computationally intensive parallel or vector processing system that may include a large number of processing cores and / or processing clusters, such as a many integrated core (MIC) processor. In at least one embodiment, some or all of the parallel processors 2012 may form a graphics processing subsystem that can output pixels to one of one or more display devices 2010A coupled via an I / O hub 2007. In at least one embodiment, the parallel processors 2012 may also include a display controller and display interface (not shown) that enables direct connection to one or more display devices 2010B.
[0268] In at least one embodiment, the system storage unit 2014 can be connected to the I / O hub 2007 to provide storage functionality for the computing system 2000. In at least one embodiment, an I / O switch 2016 can be used to provide an interface mechanism for enabling communication between the I / O hub 2007 and other components such as a network adapter 2018 and / or a wireless network adapter 2019, which may be integrated into the platform, as well as various other devices that can be added via one or more add-in devices 2020. In at least one embodiment, the network adapter 2018 may be an Ethernet® adapter or another wired network adapter. In at least one embodiment, the wireless network adapter 2019 may include one or more other network devices, including Wi-Fi, Bluetooth, Near Field Communication (NFC), or one or more wireless radios.
[0269] In at least one embodiment, the computing system 2000 may include other components not explicitly shown, such as USB or other port connections, optical storage drives, and video capture devices, which may also be connected to the I / O hub 2007. In at least one embodiment, the communication paths interconnecting the various components of Figure 20 may be implemented using any preferred protocol, such as a PCI (Peripheral Component Interconnect) based protocol (e.g., PCI-Express), or other bus or point-to-point communication interfaces, such as NV-Link High-Speed Interconnect, or other interconnection protocols.
[0270] In at least one embodiment, parallel processor 2012 incorporates circuitry optimized for graphics and video processing, such as a video output circuit, and constitutes a graphics processing unit (GPU). In at least one embodiment, parallel processor 2012 incorporates circuitry optimized for general-purpose processing. In at least one embodiment, the components of computing system 2000 may be integrated with one or more other system elements on a single integrated circuit. For example, in at least one embodiment, parallel processor 2012, memory hub 2005, processor 2002, and I / O hub 2007 can be integrated into a system-on-chip (SoC) integrated circuit. In at least one embodiment, the components of computing system 2000 may be integrated into a single package to form a system-in-package (SIP) configuration. In at least one embodiment, at least a portion of the components of computing system 2000 can be integrated into a multi-chip module (MCM), and this module can be interconnected with other multi-chip modules to form a modular computing system.
[0271] Inference and / or training logic 815 is used to perform inference and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 815 are provided herein in conjunction with FIGS. 8A and / or 8B. In at least one embodiment, inference and / or training logic 815 may be used in the system of FIG. 2000 for inference or prediction operations, based at least in part on weight parameters calculated using the training operations of the neural networks described herein, the functions and / or architectures of the neural networks, or the use cases of the neural networks.
[0272] Processor Figure 21A shows a parallel processor 2100 according to at least one embodiment. In at least one embodiment, various components of the parallel processor 2100 may be implemented using one or more integrated circuit devices such as a programmable processor, an application-specific integrated circuit (ASIC), or a field-programmable gate array (FPGA). In at least one embodiment, the illustrated parallel processor 2100 is a variation of one or more parallel processors 2012 shown in Figure 20 according to an exemplary embodiment.
[0273] In at least one embodiment, the parallel processor 2100 includes a parallel processing unit 2102. In at least one embodiment, the parallel processing unit 2102 includes an I / O unit 2104 that enables communication with other devices, including other instances of the parallel processing unit 2102. In at least one embodiment, the I / O unit 2104 may be directly connected to the other devices. In at least one embodiment, the I / O unit 2104 is connected to the other devices via the use of a hub or switch interface, such as a memory hub 2105. In at least one embodiment, the connection between the memory hub 2105 and the I / O unit 2104 forms a communication link 2113. In at least one embodiment, the I / O unit 2104 is connected to a host interface 2106 and a memory crossbar 2116, where the host interface 2106 receives commands targeting the execution of processing operations and the memory crossbar 2116 receives commands targeting the execution of memory operations.
[0274] In at least one embodiment, when the host interface 2106 receives a command buffer via the I / O unit 2104, the host interface 2106 can direct work operations to the front end 2108 to execute these commands. In at least one embodiment, the front end 2108 is coupled to a scheduler 2110, which is configured to distribute commands or other work items to the processing cluster array 2112. In at least one embodiment, the scheduler 2110 ensures that the processing cluster array 2112 is properly configured and enabled before tasks are distributed to the clusters of the processing cluster array 2112. In at least one embodiment, the scheduler 2110 is implemented via firmware logic running on a microcontroller. In at least one embodiment, the microcontroller-implemented scheduler 2110 can be configured to perform complex scheduling and work distribution operations with both coarse and fine granularity, enabling rapid preemption and context switching of threads running on the processing array 2112. In at least one embodiment, host software can assign scheduling workloads on the processing cluster array 2112 via one of several graphics processing paths. In at least one embodiment, the workload can then be automatically distributed across the processing cluster array 2112 by scheduler 2110 logic within a microcontroller, including scheduler 2110.
[0275] In at least one embodiment, the processing cluster array 2112 may contain up to "N" processing clusters (e.g., cluster 2114A, cluster 2114B to cluster 2114N), where "N" represents a positive integer (which may be a different integer "N" than that used in other diagrams). In at least one embodiment, each cluster 2114A to 2114N of the processing cluster array 2112 may execute a large number of concurrent threads. In at least one embodiment, the scheduler 2110 may distribute work to clusters 2114A to 2114N of the processing cluster array 2112 using various scheduling and / or work distribution algorithms, which may differ depending on the workload arising for each type of program or computation. In at least one embodiment, scheduling may be handled dynamically by the scheduler 2110 or partially assisted by compiler logic during the compilation of program logic configured to be executed by the processing cluster array 2112. In at least one embodiment, different clusters 2114A to 2114N of the processing cluster array 2112 can be allocated to process different types of programs or to perform different types of calculations.
[0276] In at least one embodiment, the processing cluster array 2112 can be configured to perform various types of parallel processing operations. In at least one embodiment, the processing cluster array 2112 is configured to perform general-purpose parallel compute operations. For example, in at least one embodiment, the processing cluster array 2112 may include logic for performing processing tasks, including filtering video and / or audio data, performing modeling operations including physical operations, and performing data transformations.
[0277] In at least one embodiment, the processing cluster array 2112 is configured to perform parallel graphics processing operations. In at least one embodiment, the processing cluster array 2112 may include additional logic to support the execution of such graphics processing operations, including but not limited to texture sampling logic for performing texture operations, as well as mosaic logic and other vertex processing logic. In at least one embodiment, the processing cluster array 2112 can be configured to execute graphics processing-related shader programs, including but not limited to vertex shaders, mosaic shaders, geometry shaders, and pixel shaders. In at least one embodiment, the parallel processing unit 2102 can transfer data from system memory via I / O unit 2104 for processing. In at least one embodiment, during processing, the transferred data can be stored in on-chip memory (e.g., parallel processor memory 2122) during processing and then written back to system memory.
[0278] In at least one embodiment, when graphics processing is performed using the parallel processing unit 2102, the scheduler 2110 can be configured to divide the processing workload into tasks of roughly equal size in order to better distribute the graphics processing operations among the multiple clusters 2114A to 2114N of the processing cluster array 2112. In at least one embodiment, parts of the processing cluster array 2112 can be configured to perform different types of processing. For example, in at least one embodiment, to generate and display a rendered image, a first part may be configured to perform vertex shading and topology generation, a second part may be configured to perform mosaic and geometry shading, and a third part may be configured to perform pixel shading or other screen-space operations. In at least one embodiment, intermediate data generated by one or more of the clusters 2114A to 2114N may be stored in a buffer so that the intermediate data can be transmitted between the clusters 2114A to 2114N for further processing.
[0279] In at least one embodiment, the processing cluster array 2112 may receive processing tasks to be executed via a scheduler 2110, which receives commands defining the processing tasks from a front-end 2108. In at least one embodiment, the processing task may include an index of the data to be processed, such as surface (patch) data, primitive data, vertex data, and / or pixel data, as well as state parameters, and commands defining how the data should be processed (e.g., which program to run). In at least one embodiment, the scheduler 2110 may be configured to fetch the index corresponding to the task, or to receive the index from the front-end 2108. In at least one embodiment, the front-end 2108 may be configured to ensure that the processing cluster array 2112 is configured to be in a valid state before the workload specified by an incoming command buffer (e.g., a batch buffer, a push buffer, etc.) is initiated.
[0280] In at least one embodiment, each of one or more instances of the parallel processing unit 2102 can be coupled to a parallel processor memory 2122. In at least one embodiment, the parallel processor memory 2122 can be accessed via a memory crossbar 2116, which can receive memory requests from the processing cluster array 2112 and the I / O unit 2104. In at least one embodiment, the memory crossbar 2116 can access the parallel processor memory 2122 via a memory interface 2118. In at least one embodiment, the memory interface 2118 may include a plurality of partition units (for example, partition unit 2120A, partition units 2120B to 2120N), each of which can be coupled to a portion of the parallel processor memory 2122 (for example, a memory unit). In at least one embodiment, the number of partition units 2120A to 2120N is configured to be equal to the number of memory units, so that the first partition unit 2120A has a corresponding first memory unit 2124A, the second partition unit 2120B has a corresponding memory unit 2124B, and the Nth partition unit 2120N has a corresponding Nth memory unit 2124N. In at least one embodiment, the number of partition units 2120A to 2120N does not have to be equal to the number of memory devices.
[0281] In at least one embodiment, memory units 2124A to 2124N may include various types of memory devices, including dynamic random access memory (DRAM) or graphics random access memory, such as synchronous graphics random access memory (SGRAM) including graphics double data rate (GDDR) memory. In at least one embodiment, memory units 2124A to 2124N may also include, but not limited to, high-bandwidth memory (HBM) and 3D stacked memory. In at least one embodiment, to efficiently use the available bandwidth of parallel processor memory 2122, render targets such as frame buffers or texture maps may be stored across memory units 2124A to 2124N, allowing partition units 2120A to 2120N to write each portion of the render target in parallel. In at least one embodiment, local instances of parallel processor memory 2122 may be excluded to favor an integrated memory design using system memory and local cache memory together.
[0282] In at least one embodiment, any one of the clusters 2114A to 2114N of the processing cluster array 2112 can process data that will be written to any of the memory units 2124A to 2124N in the parallel processor memory 2122. In at least one embodiment, the memory crossbar 2116 can be configured to forward the output of each cluster 2114A to 2114N to any partition unit 2120A to 2120N, or another cluster 2114A to 2114N, which can perform further processing operations on the output. In at least one embodiment, each cluster 2114A to 2114N can communicate with the memory interface 2118 through the memory crossbar 2116 to read from or write to various external memory devices. In at least one embodiment, the memory crossbar 2116 has a connection to a memory interface 2118 for communicating with the I / O unit 2104, and a connection to a local instance of the parallel processor memory 2122, enabling processing units in different processing clusters 2114A to 2114N to communicate with system memory or other memory not local to the parallel processing unit 2102. In at least one embodiment, the memory crossbar 2116 can use virtual channels to separate traffic streams between the clusters 2114A to 2114N and the partition units 2120A to 2120N.
[0283] In at least one embodiment, multiple instances of the parallel processing unit 2102 may be provided on a single add-in card, or multiple add-in cards may be interconnected. In at least one embodiment, different instances of the parallel processing unit 2102 can be configured to interact with each other, even if different instances have different numbers of processing cores, different amounts of local parallel processor memory, and / or other different configurations. For example, in at least one embodiment, some instances of the parallel processing unit 2102 may include higher-precision floating-point units than other instances. In at least one embodiment, a system incorporating one or more instances of the parallel processing unit 2102 or parallel processor 2100 can be implemented in a variety of configurations and form factors, including, but not limited to, desktop, laptop, or portable personal computers, servers, workstations, game consoles, and / or embedded systems.
[0284] Figure 21B is a block diagram of partition unit 2120 according to at least one embodiment. In at least one embodiment, partition unit 2120 is an instance of one of partition units 2120A to 2120N in Figure 21A. In at least one embodiment, partition unit 2120 includes an L2 cache 2121, a frame buffer interface 2125, and a ROP (raster operations unit) 2126. In at least one embodiment, the L2 cache 2121 is a read / write cache configured to perform load and store operations received from the memory crossbar 2116 and the ROP 2126. In at least one embodiment, read misses and urgent write-back requests are output by the L2 cache 2121 to the frame buffer interface 2125 for processing. In at least one embodiment, updates are also sent to the frame via the frame buffer interface 2125 for processing. In at least one embodiment, the frame buffer interface 2125 interfaces with one of the memory units of the parallel processor memory, such as memory units 2124A to 2124N (for example, in the parallel processor memory 2122) in Figure 21.
[0285] In at least one embodiment, ROP2126 is a processing unit that performs raster operations such as stenciling, z-testing, and blending. In at least one embodiment, ROP2126 then outputs the processed graphics data stored in graphics memory. In at least one embodiment, ROP2126 includes compression logic for compressing depth or color data written to memory and for decompressing depth or color data read from memory. In at least one embodiment, the compression logic may be lossless compression logic that utilizes one or more of a plurality of compression algorithms. In at least one embodiment, the type of compression performed by ROP2126 may be modified based on the statistical characteristics of the data being compressed. For example, in at least one embodiment, delta color compression is performed on a per-tile basis for depth and color data.
[0286] In at least one embodiment, ROP2126 is located within each processing cluster (for example, clusters 2114A to 2114N in Figure 21A) rather than within the partition unit 2120. In at least one embodiment, read and write requests for pixel data, rather than pixel fragment data, are transmitted via the memory crossbar 2116. In at least one embodiment, the processed graphics data may be displayed on a display device, such as one of the one or more display devices 2010 in Figure 20, routed for further processing by a processor 2002, or routed for further processing by one of the processing entities in the parallel processor 2100 in Figure 21A.
[0287] Figure 21C is a block diagram of a processing cluster 2114 within a parallel processing unit according to at least one embodiment. In at least one embodiment, the processing cluster is an instance of one of the processing clusters 2114A to 2114N in Figure 21A. In at least one embodiment, the processing cluster 2114 may be configured to run a large number of threads in parallel, where “thread” refers to an instance of a particular program running on a particular set of input data. In at least one embodiment, a single-instruction, multiple-data (SIMD) instruction issuing technique is used to support the parallel execution of a large number of threads without providing multiple independent instruction units. In at least one embodiment, a single-instruction, multiple-thread (SIMT) technique is used to support the parallel execution of a large number of threads in a globally synchronized manner, using a common instruction unit configured to issue instructions to a set of processing engines within each processing cluster.
[0288] In at least one embodiment, the operation of the processing cluster 2114 can be controlled via a pipeline manager 2132 that distributes processing tasks to SIMT parallel processors. In at least one embodiment, the pipeline manager 2132 receives instructions from the scheduler 2110 in Figure 21A and manages the execution of these instructions via the graphics multiprocessor 2134 and / or texture unit 2136. In at least one embodiment, the graphics multiprocessor 2134 is an exemplary instance of a SIMT parallel processor. However, in at least one embodiment, various types of SIMT parallel processors with different architectures may be included within the processing cluster 2114. In at least one embodiment, one or more instances of the graphics multiprocessor 2134 may be included within the processing cluster 2114. In at least one embodiment, the graphics multiprocessor 2134 can process data, and a data crossbar 2140 may be used to distribute the processed data to one of several possible destinations, including other shader units. In at least one embodiment, the pipeline manager 2132 can facilitate the distribution of processed data by specifying the destination of the processed data to be distributed through the data crossbar 2140.
[0289] In at least one embodiment, each graphics multiprocessor 2134 within the processing cluster 2114 may include an identical set of function execution logic (e.g., arithmetic logic units, load / store units, etc.). In at least one embodiment, the function execution logic can be configured in a pipelined manner, allowing new instructions to be issued before the previous instruction is completed. In at least one embodiment, the function execution logic supports a variety of operations, including integer and floating-point arithmetic, comparison operations, Boolean operations, bit shifts, and calculations of various algebraic functions. In at least one embodiment, different operations can be performed by leveraging the hardware of the same function units, and any combination of function units may exist.
[0290] In at least one embodiment, instructions sent to the processing cluster 2114 constitute a thread. In at least one embodiment, a set of threads running across a set of parallel processing engines is a thread group. In at least one embodiment, a thread group executes a common program for different input data. In at least one embodiment, each thread in a thread group can be assigned to a different processing engine in the graphics multiprocessor 2134. In at least one embodiment, a thread group may contain fewer threads than the number of processing engines in the graphics multiprocessor 2134. In at least one embodiment, if a thread group contains fewer threads than the number of processing engines, one or more of the processing engines may be idle during the cycle in which the thread group is being processed. In at least one embodiment, a thread group may also contain more threads than the number of processing engines in the graphics multiprocessor 2134. In at least one embodiment, if a thread group contains more threads than the number of processing engines in the graphics multiprocessor 2134, processing can be performed over consecutive clock cycles. In at least one embodiment, multiple thread groups can run simultaneously on the graphics multiprocessor 2134.
[0291] In at least one embodiment, the graphics multiprocessor 2134 includes internal cache memory for performing load and store operations. In at least one embodiment, the graphics multiprocessor 2134 can abandon its internal cache and use cache memory within the processing cluster 2114 (e.g., L1 cache 2148). In at least one embodiment, each graphics multiprocessor 2134 may also access an L2 cache within a partition unit (e.g., partition units 2120A-2120N in Figure 21A), and these caches may be shared among all processing clusters 2114 and used to transfer data between threads. In at least one embodiment, the graphics multiprocessor 2134 may also access off-chip global memory, which may include one or more of the local parallel processor memory and / or system memory. In at least one embodiment, any memory outside the parallel processing unit 2102 may be used as global memory. In at least one embodiment, the processing cluster 2114 includes multiple instances of a graphics multiprocessor 2134, which can share common instructions and data, which may be stored in an L1 cache 2148.
[0292] In at least one embodiment, each processing cluster 2114 may include an MMU 2145 (Memory Management Unit) configured to map virtual addresses to physical addresses. In at least one embodiment, one or more instances of the MMU 2145 may be located within the memory interface 2118 in Figure 21A. In at least one embodiment, the MMU 2145 includes a set of page table entries (PTEs) used to map virtual addresses to physical addresses of tiles and optionally cache line indices. In at least one embodiment, the MMU 2145 may include an address translation lookaside buffer (TLB) or cache, which may be located within the graphics multiprocessor 2134 or L1 2148 cache, or within the processing cluster 2114. In at least one embodiment, physical addresses are processed to distribute surface data access locally, enabling efficient interleaving of requests between partition units. In at least one embodiment, a cache line index may be used to determine whether a cache line request is a hit or a miss.
[0293] In at least one embodiment, each graphics multiprocessor 2134 may be coupled to a texture unit 2136 to configure a processing cluster 2114 so that texture mapping operations, such as determining texture sample locations, reading texture data, and filtering texture data, are performed. In at least one embodiment, texture data is read from an internal texture L1 cache (not shown) or from an L1 cache within the graphics multiprocessor 2134 and, if necessary, fetched from an L2 cache, local parallel processor memory, or system memory. In at least one embodiment, each graphics multiprocessor 2134 outputs processed tasks to a data crossbar 2140 to provide processed tasks to another processing cluster 2114 for further processing, or stores processed tasks in an L2 cache, local parallel processor memory, or system memory via a memory crossbar 2116. In at least one embodiment, a pre-ROP 2142 (pre-raster arithmetic unit) is configured to receive data from a graphics multiprocessor 2134 and direct the data to an ROP unit, which may be located within a partition unit (for example, partition units 2120A-2120N in Figure 21A) as described herein. In at least one embodiment, the pre-ROP 2142 unit can perform color blending optimization, organize pixel color data, and perform address translation.
[0294] The inference and / or training logic 815 is used to perform inference and / or training operations related to one or more embodiments. Details relating to the inference and / or training logic 815 are provided herein in conjunction with Figures 8A and / or 8B. In at least one embodiment, the inference and / or training logic 815 may be used in the graphics processing cluster 2114 for inference or prediction operations, at least in part, based on weight parameters calculated using the neural network training operations, neural network functions and / or architecture, or neural network use cases described herein.
[0295] Figure 21D shows a graphics multiprocessor 2134 according to at least one embodiment. In at least one embodiment, the graphics multiprocessor 2134 is coupled with a pipeline manager 2132 of a processing cluster 2114. In at least one embodiment, the graphics multiprocessor 2134 has an execution pipeline that includes, but is not limited to, an instruction cache 2152, an instruction unit 2154, an address mapping unit 2156, a register file 2158, one or more general-purpose graphics processing unit (GPGPU) cores 2162, and one or more load / store units 2166. In at least one embodiment, the GPGPU cores 2162 and the load / store units 2166 are coupled to a cache memory 2172 and a shared memory 2170 via a memory and cache interconnect 2168.
[0296] In at least one embodiment, the instruction cache 2152 receives a stream of instructions to be executed from the pipeline manager 2132. In at least one embodiment, the instructions are cached in the instruction cache 2152 and dispatched to be executed by the instruction unit 2154. In at least one embodiment, the instruction unit 2154 can dispatch the instructions as a thread group (e.g., a warp), and each thread in the thread group is assigned to a different execution unit within the GPGPU core 2162. In at least one embodiment, the instructions can access either the local, shared, or global address space by specifying an address in the unified address space. In at least one embodiment, an address mapping unit 2156 can be used to translate an address in the unified address space to a separate memory address accessible by the load / store unit 2166.
[0297] In at least one embodiment, the register file 2158 provides a set of registers to the functional units of the graphics multiprocessor 2134. In at least one embodiment, the register file 2158 provides temporary storage for operands connected to the data paths of the functional units of the graphics multiprocessor 2134 (e.g., GPGPU core 2162, load / store unit 2166). In at least one embodiment, the register file 2158 is divided among the functional units such that each functional unit is allocated a dedicated portion of the register file 2158. In one embodiment, the register file 2158 is divided among different warps being executed by the graphics multiprocessor 2134.
[0298] In at least one embodiment, each GPGPU core 2162 may include a floating-point unit (FPU) and / or an integer arithmetic logic unit (ALU) used to execute instructions of the graphics multiprocessor 2134. In at least one embodiment, the GPGPU cores 2162 may have similar architectures or different architectures. In at least one embodiment, a first part of the GPGPU core 2162 includes a single-precision FPU and an integer ALU, and a second part of the GPGPU core includes a double-precision FPU. In at least one embodiment, the FPU may implement IEEE 754-2008 standard floating-point arithmetic or enable variable-precision floating-point arithmetic. In at least one embodiment, the graphics multiprocessor 2134 may further include one or more fixed-function units or special-function units for performing specific functions such as rectangular copying or pixel blending operations. In at least one embodiment, one or more GPGPU cores 2162 may also include fixed or special-function logic.
[0299] In at least one embodiment, the GPGPU core 2162 includes SIMD logic that can execute a single instruction for multiple datasets. In at least one embodiment, the GPGPU core 2162 can physically execute SIMD4, SIMD8, and SIMD16 instructions and logically execute SIMD1, SIMD2, and SIMD32 instructions. In at least one embodiment, SIMD instructions for the GPGPU core may be generated at compile time by a shader compiler, or they may be automatically generated when running a program written and compiled for a single program multiple data (SPMD) or SIMT architecture. In at least one embodiment, multiple threads of a program configured for a SIMT execution model can be executed via a single SIMD instruction. For example, in at least one embodiment, eight SIMT threads performing the same or similar operations can be executed in parallel via a single SIMD8 logic unit.
[0300] In at least one embodiment, the memory and cache interconnect 2168 is an interconnect network connecting each functional unit of the graphics multiprocessor 2134 to the register file 2158 and shared memory 2170. In at least one embodiment, the memory and cache interconnect 2168 is a crossbar interconnect that allows the load / store unit 2166 to implement load and store operations between the shared memory 2170 and the register file 2158. In at least one embodiment, the register file 2158 can operate at the same frequency as the GPGPU core 2162, and therefore data transfer between the GPGPU core 2162 and the register file 2158 can have very low latency. In at least one embodiment, the shared memory 2170 can be used to enable communication between threads running in the functional units within the graphics multiprocessor 2134. In at least one embodiment, the cache memory 2172 can be used, for example, as a data cache to cache texture data communicated between the functional units and the texture unit 2136. In at least one embodiment, the shared memory 2170 can also be used as a program-managed cache. In at least one embodiment, a thread running on the GPGPU core 2162 can programmatically store data in the shared memory in addition to the automatically cached data stored in the cache memory 2172.
[0301] In at least one embodiment, the parallel processor or GPGPU described herein is communicatively coupled to a host / processor core to accelerate graphics operations, machine learning operations, pattern analysis operations, and various general-purpose GPU (GPGPU) functions. In at least one embodiment, the GPU may be communicatively coupled to the host processor / core via a bus or other interconnect (e.g., a high-speed interconnect such as PCIe or NVLink). In at least one embodiment, the GPU may be integrated as a core in a package or chip and communicatively coupled to the core via an internal processor bus / interconnection within the package or chip. In at least one embodiment, regardless of how the GPU is connected, the processor core may allocate work to such GPUs in the form of a sequence of commands / instructions contained in a work descriptor. In at least one embodiment, the GPU then uses dedicated circuitry / logic to efficiently process these commands / instructions.
[0302] The inference and / or training logic 815 is used to perform inference and / or training operations related to one or more embodiments. Details relating to the inference and / or training logic 815 are provided herein in conjunction with Figures 8A and / or 8B. In at least one embodiment, the inference and / or training logic 815 may be used in the graphics multiprocessor 2134 for inference or prediction operations, at least in part, based on weight parameters calculated using the neural network training operations, neural network functions and / or architecture, or neural network use cases described herein.
[0303] Figure 22 shows a multi-GPU computing system 2200 according to at least one embodiment. In at least one embodiment, the multi-GPU computing system 2200 may include a processor 2202 coupled to a plurality of general-purpose graphics processing units (GPGPUs) 2206A-D via a host interface switch 2204. In at least one embodiment, the host interface switch 2204 is a PCI Express switch device that couples the processor 2202 to a PCI Express bus, through which the processor 2202 can communicate with the GPGPUs 2206A-D. In at least one embodiment, the GPGPUs 2206A-D can be interconnected via a set of high-speed point-to-point GPU-to-GPU links 2216. In at least one embodiment, the GPU-to-GPU links 2216 are connected to each of the GPGPUs 2206A-D via dedicated GPU links. In at least one embodiment, the P2P GPU link 2216 enables direct communication between each of the GPGPUs 2206A-D without requiring communication via the host interface bus 2204 to which the processor 2202 is connected. In at least one embodiment, when there is GPU-to-GPU traffic directed to the P2P GPU link 2216, the host interface bus 2204 is kept available to allow access to system memory or to communicate with other instances of the multi-GPU computing system 2200, for example, via one or more network devices. In at least one embodiment, the GPGPUs 2206A-D are connected to the processor 2202 via the host interface switch 2204, and in at least one embodiment, the processor 2202 can be directly connected to the GPGPUs 2206A-D, including direct support for the P2P GPU link 2216.
[0304] The inference and / or training logic 815 is used to perform inference and / or training operations related to one or more embodiments. Details relating to the inference and / or training logic 815 are provided herein in conjunction with Figures 8A and / or 8B. In at least one embodiment, the inference and / or training logic 815 may be used in a multi-GPU computing system 2200 for inference or prediction operations, at least in part, based on weight parameters calculated using the neural network training operations, neural network functions and / or architecture, or neural network use cases described herein.
[0305] Figure 23 is a block diagram of a graphics processor 2300 according to at least one embodiment. In at least one embodiment, the graphics processor 2300 includes a ring interconnect 2302, a pipeline front end 2304, a media engine 2337, and graphics cores 2380A to 2380N. In at least one embodiment, the ring interconnect 2302 connects the graphics processor 2300 to other graphics processors or other processing units including one or more general-purpose processor cores. In at least one embodiment, the graphics processor 2300 is one of a number of processors integrated within a multi-core processing system.
[0306] In at least one embodiment, the graphics processor 2300 receives batches of commands via a ring interconnect 2302. In at least one embodiment, incoming commands are interpreted by a command streamer 2303 on a pipeline front end 2304. In at least one embodiment, the graphics processor 2300 includes scalable execution logic for performing 3D geometry processing and media processing via graphics cores 2380A-2380N. In at least one embodiment, for 3D geometry processing commands, the command streamer 2303 feeds the commands to a geometry pipeline 2336. In at least one embodiment, for at least some media processing commands, the command streamer 2303 feeds the commands to a video front end 2334, which is coupled to a media engine 2337. In at least one embodiment, the media engine 2337 includes a Video Quality Engine (VQE) 2330 for post-processing of video and images, and a multi-format encoding / decoding (MFX) 2333 engine for encoding and decoding hardware-accelerated media data. In at least one embodiment, the geometry pipeline 2336 and the media engine 2337 each generate execution threads for thread execution resources provided by at least one graphics core 2380.
[0307] In at least one embodiment, the graphics processor 2300 includes a scalable thread execution resource featuring graphics cores 2380A-2380N (which may be modular and sometimes called core slices), each graphics core 2380A-2380N having multiple sub-cores 2350A-50N, 2360A-2360N (sometimes called core sub-slices). In at least one embodiment, the graphics processor 2300 may have any number of graphics cores 2380A. In at least one embodiment, the graphics processor 2300 includes a graphics core 2380A having at least a first sub-core 2350A and a second sub-core 2360A. In at least one embodiment, the graphics processor 2300 is a low-power processor having a single sub-core (e.g., 2350A). In at least one embodiment, the graphics processor 2300 includes a plurality of graphics cores 2380A to 2380N, each of which includes a first set of subcores 2350A to 2350N and a second set of subcores 2360A to 2360N. In at least one embodiment, each of the first subcores 2350A to 2350N includes at least a first set of execution units 2352A to 2352N and media / texture samplers 2354A to 2354N. In at least one embodiment, each of the second subcores 2360A to 2360N includes at least a second set of execution units 2362A to 2362N and samplers 2364A to 2364N. In at least one embodiment, each sub-core 2350A-2350N, 2360A-2360N shares a set of shared resources 2370A-2370N. In at least one embodiment, the shared resources include shared cache memory and pixel operation logic.
[0308] The inference and / or training logic 815 is used to perform inference and / or training operations related to one or more embodiments. Details relating to the inference and / or training logic 815 are provided herein in conjunction with Figures 8A and / or 8B. In at least one embodiment, the inference and / or training logic 815 may be used in the graphics processor 2300 for inference or prediction operations, at least in part, based on weight parameters calculated using the neural network training operations, neural network functions and / or architecture, or neural network use cases described herein.
[0309] Figure 24 is a block diagram showing the microarchitecture of a processor 2400, which may include logic circuits for executing instructions, according to at least one embodiment. In at least one embodiment, the processor 2400 may execute instructions including x86 instructions, AMR instructions, and special instructions for application-specific integrated circuits (ASICs). In at least one embodiment, the processor 2400 may include registers for storing packed data, such as 64-bit wide MMX™ registers in a microprocessor enabled by MMX technology, as provided by Intel Corporation in Santa Clara, California. In at least one embodiment, MMX registers, available in both integer and floating-point formats, may operate with packed data elements accompanied by Single Instruction Multiple Data ("SIMD") and Streaming SIMD Extensions ("SSE") instructions. In at least one embodiment, 128-bit wide XMM registers relating to SSE2, SSE3, SSE4, AVX, or higher technologies (collectively referred to as "SSEx") may hold operands of such packed data. In at least one embodiment, the processor 2400 may execute instructions to accelerate machine learning or deep learning algorithms, training, or inference.
[0310] In at least one embodiment, the processor 2400 includes an in-order front-end ("front-end") 2401 that fetches instructions to be executed and prepares instructions for later use in the processor pipeline. In at least one embodiment, the front-end 2401 may include several units. In at least one embodiment, an instruction prefetcher 2426 fetches instructions from memory and supplies them to an instruction decoder 2428, which decodes or interprets the instructions. For example, in at least one embodiment, the instruction decoder 2428 decodes the received instruction into one or more operations called "microinstructions" or "microoperations" (also called "microops" or "uops") that the machine can execute. In at least one embodiment, the instruction decoder 2428 may parse the instruction into opcodes and corresponding data, as well as control fields, which are used by the microarchitecture to perform the operation according to at least one embodiment. In at least one embodiment, the trace cache 2430 may assemble the decoded uops into a program-order sequence or trace in the uop queue 2434 so that they can be executed. In at least one embodiment, when the trace cache 2430 encounters a complex instruction, the microcode ROM 2432 provides the uops necessary to complete the operation.
[0311] In at least one embodiment, some instructions can be converted into a single micro-ops, while others require several micro-ops to complete the entire operation. In at least one embodiment, if five or more micro-ops are required to complete an instruction, the instruction decoder 2428 may access the microcode ROM 2432 to execute the instruction. In at least one embodiment, an instruction may be decoded into a small number of micro-ops so that it can be processed by the instruction decoder 2428. In at least one embodiment, if many micro-ops are required to complete such an operation, the instruction may be stored in the microcode ROM 2432. In at least one embodiment, the trace cache 2430 refers to an entry-point programmable logic array ("PLA") to determine the correct microinstruction pointer for reading a microcode sequence in order to complete one or more instructions from the microcode ROM 2432 according to at least one embodiment. In at least one embodiment, after the microcode ROM 2432 has finished sequencing microops for instructions, the machine's front-end 2401 may resume fetching microops from the trace cache 2430.
[0312] In at least one embodiment, the out-of-order execution engine ("out-of-order engine") 2403 may prepare instructions for execution. In at least one embodiment, the out-of-order execution logic has a large number of buffers to smooth the flow of instructions and change their order, optimizing performance as instructions are pipelined and scheduled for execution. In at least one embodiment, the out-of-order execution engine 2403 includes, without limitation, an allocator / register renamer 2440, a memory uop queue 2442, an integer / floating-point uop queue 2444, a memory scheduler 2446, a fast scheduler 2402, a slow / general-purpose floating-point scheduler ("slow / general-purpose FP: floating-point scheduler") 2404, and a simple floating-point scheduler ("simple FP scheduler") 2406. In at least one embodiment, the fast scheduler 2402, the slow / general-purpose floating-point scheduler 2404, and the simple floating-point scheduler 2406 are collectively referred to herein as “uop schedulers 2402, 2404, and 2406.” In at least one embodiment, the allocator / register renamer 2440 allocates the machine buffers and resources required by each uop to run. In at least one embodiment, the allocator / register renamer 2440 renames logical registers upon entry into the register file. In at least one embodiment, the allocator / register renamer 2440 also allocates the entries of each uop to one of two uop queues, namely the memory uop queue 2442 for memory operations and the integer / floating-point uop queue 2444 for non-memory operations, prior to the memory scheduler 2446 and the uop schedulers 2402, 2404, 2406. In at least one embodiment, the uop schedulers 2402, 2404, 2406 determine when uops are ready to execute based on whether the sources for their dependent input register operands are ready and whether the execution resources required by the uop to complete their operations are available.In at least one embodiment, the high-speed scheduler 2402 may schedule every half of the main clock cycle, while the slow / general-purpose floating-point scheduler 2404 and the simple floating-point scheduler 2406 may schedule once per main processor clock cycle. In at least one embodiment, the uop schedulers 2402, 2404, and 2406 arbitrate dispatch ports to schedule uops so that they can be executed.
[0313] In at least one embodiment, the execution block 2411 includes, without limitation, an integer register file / bypass network 2408, a floating-point register file / bypass network ("FP register file / bypass network") 2410, address generation units ("AGUs") 2412 and 2414, fast arithmetic logic units (ALUs) ("fast ALUs") 2416 and 2418, a slow arithmetic logic unit ("slow ALU") 2420, a floating-point ALU ("FP") 2422, and a floating-point move unit ("FP move") 2424. In at least one embodiment, the integer register file / bypass network 2408 and the floating-point register file / bypass network 2410 are also referred to herein as "register files 2408, 2410". In at least one embodiment, AGU2412 and 2414, high-speed ALU2416 and 2418, low-speed ALU2420, floating-point ALU2422, and floating-point movement unit 2424 are also referred to herein as “execution units 2412, 2414, 2416, 2418, 2420, 2422, and 2424”. In at least one embodiment, execution block 2411 may include, without limitation, any number and type of register files (including zero), bypass networks, address generation units, and execution units in any combination.
[0314] In at least one embodiment, register networks 2408, 2410 may be located between the uop schedulers 2402, 2404, 2406 and the execution units 2412, 2414, 2416, 2418, 2420, 2422, and 2424. In at least one embodiment, the integer register file / bypass network 2408 performs integer arithmetic. In at least one embodiment, the floating-point register file / bypass network 2410 performs floating-point arithmetic. In at least one embodiment, each of the register networks 2408, 2410 may include, without limitation, a bypass network which may bypass or transfer newly completed results that have not yet been written to the register file to new dependent uops. In at least one embodiment, the register networks 2408, 2410 may communicate data with each other. In at least one embodiment, the integer register file / bypass network 2408 may include, without limitation, two separate register files: one register file for low-order 32-bit data and a second register file for high-order 32-bit data. In at least one embodiment, since floating-point instructions typically have operands of 64 to 128 bits in width, the floating-point register file / bypass network 2410 may include, without limitation, 128-bit wide entries.
[0315] In at least one embodiment, execution units 2412, 2414, 2416, 2418, 2420, 2422, and 2424 may execute instructions. In at least one embodiment, register networks 2408 and 2410 store operand values of integer and floating-point data that microinstructions need to execute. In at least one embodiment, processor 2400 may include, without limitation, any number and combinations of execution units 2412, 2414, 2416, 2418, 2420, 2422, and 2424. In at least one embodiment, floating-point ALU 2422 and floating-point movement unit 2424 may perform floating-point, MMX, SIMD, AVX, and SEE, or other operations including special machine learning instructions. In at least one embodiment, the floating-point ALU2422 may include, without limitation, 64-bit floating-point dividers and perform division, square root, and the remaining micro-operations. In at least one embodiment, instructions involving floating-point values may be handled by floating-point hardware. In at least one embodiment, ALU operations may be passed to high-speed ALU2416, 2418. In at least one embodiment, high-speed ALU2416, 2418 may perform high-speed operations with an effective latency of half a clock cycle. In at least one embodiment, the slow ALU2420 may include, without limitation, integer execution hardware for long-latency types of operations such as multipliers, shifts, flag logic, and branching, so that most complex integer operations are passed to the slow ALU2420. In at least one embodiment, memory load / store operations may be performed by AGUS2412, 2414. In at least one embodiment, the high-speed ALU2416, high-speed ALU2418, and low-speed ALU2420 may perform integer arithmetic with 64-bit data operands. In at least one embodiment, the high-speed ALU2416, high-speed ALU2418, and low-speed ALU2420 may be implemented to support various data bit sizes, including 16, 32, 128, 256, and so on.In at least one embodiment, the floating-point ALU 2422 and the floating-point movement unit 2424 may be implemented to support a wide range of operands with various bit widths, such as 128-bit wide packed data operands in combination with SIMD and multimedia instructions.
[0316] In at least one embodiment, the uop schedulers 2402, 2404, and 2406 dispatch dependent operations before the parent load finishes execution. In at least one embodiment, uops may be speculatively scheduled and executed in processor 2400, so processor 2400 may also include logic for handling memory misses. In at least one embodiment, if a data load misses in the data cache, there may be ongoing dependent operations in the pipeline that have passed the scheduler with temporarily inaccurate data. In at least one embodiment, a replay mechanism tracks and redelivers instructions that use inaccurate data. In at least one embodiment, dependent operations may need to be replayed, while independent operations may be allowed to complete. In at least one embodiment, the scheduler and replay mechanism of at least one embodiment of the processor may also be designed to capture instruction sequences for text string comparison operations.
[0317] In at least one embodiment, “register” may refer to an onboard processor storage location that can be used as part of an instruction to identify an operand. In at least one embodiment, a register may be accessible from outside the processor (from the programmer’s perspective). In at least one embodiment, a register may not be limited to a particular type of circuit. Rather, in at least one embodiment, a register may store data, provide data, and perform the functions described herein. In at least one embodiment, the registers described herein may be implemented by circuitry within the processor using any number of different techniques, such as dedicated physical registers, dynamically allocated physical registers using register renaming, or a combination of dedicated and dynamically allocated physical registers. In at least one embodiment, an integer register stores 32-bit integer data. The register file in at least one embodiment also includes eight multimedia SIMD registers for packed data.
[0318] The inference and / or training logic 815 is used to perform inference and / or training operations related to one or more embodiments. Details relating to the inference and / or training logic 815 are provided herein in conjunction with Figures 8A and / or 8B. In at least one embodiment, some or all of the inference and / or training logic 815 may be incorporated into the execution block 2411 and other memories or registers, whether illustrated or not. For example, in at least one embodiment, the training and / or inference techniques described herein may use one or more of the ALUs shown in the execution block 2411. Furthermore, weight parameters may be stored in on-chip or off-chip memories and / or registers (illustrated or not illustrated) that constitute the ALUs of the execution block 2411 for performing one or more machine learning algorithms, neural network architectures, use cases, or training techniques described herein.
[0319] Figure 25 shows a deep learning application processor 2500 according to at least one embodiment. In at least one embodiment, the deep learning application processor 2500 uses instructions that cause the deep learning application processor 2500 to perform some or all of the processes and techniques described throughout this disclosure when executed by the deep learning application processor 2500. In at least one embodiment, the deep learning application processor 2500 is an application-specific integrated circuit (ASIC). In at least one embodiment, the application processor 2500 performs a matrix multiplication operation which is hardwired to hardware as a result of executing one or more instructions or both. In at least one embodiment, the deep learning application processor 2500 includes, without limitation, processing clusters 2510(1) to 2510(12), inter-chip links ("ICL") 2520(1) to 2520(12), inter-chip controllers ("ICC") 2530(1) to 2530(2), high-bandwidth memory second generation ("HBM2") 2540(1) to 2540(4), memory controllers ("Mem Ctrlrs") 2542(1) to 2542(4), and high-bandwidth memory physical layer ("HBM"). This includes the PHYs (2544(1)-2544(4)), the Management-Controller Central Processing Unit ("Management-Controller CPU") (2550), serial peripheral interfaces, inter-integrated and general-purpose input / output blocks ("SPI, I2C, GPIO") (2560), peripheral component interconnect express controllers and direct memory access blocks ("PCIe controllers and DMA") (2570), and a 16-lane peripheral component interconnect express port ("PCI Express x16") (2580).
[0320] In at least one embodiment, the processing cluster 2510 may perform deep learning operations, including inference or prediction operations, based on weight parameters computed using one or more training techniques, including the techniques described herein. In at least one embodiment, each processing cluster 2510 may include any number and type of processors, without limitation. In at least one embodiment, the deep learning application processor 2500 may include any number and type of processing clusters 2500. In at least one embodiment, the inter-chip link 2520 is bidirectional. In at least one embodiment, the inter-chip link 2520 and the inter-chip controller 2530 enable multiple deep learning application processors 2500 to exchange information, including activation information obtained as a result of executing one or more machine learning algorithms embodied in one or more neural networks. In at least one embodiment, the deep learning application processor 2500 may include any number and type of ICL2520 and ICC2530 (including zero).
[0321] In at least one embodiment, the HBM2 2540 provides a total of 32 gigabytes (GB) of memory. In at least one embodiment, the HBM2 2540(i) is associated with both the memory controller 2542(i) and the HBM PHY2544(i), where "i" is any integer. In at least one embodiment, any number of HBM2 2540s may provide high-bandwidth memory of any type and total amount, and may be associated with any number and type of memory controllers 2542 and HBM PHY2544 (including zero). In at least one embodiment, SPI, I2C, GPIO2560, PCIe controller and DMA2570, and / or PCIe2580 may be replaced with any number and type of blocks enabling any number and type of communication standards in any technically viable way.
[0322] The inference and / or training logic 815 is used to perform inference and / or training operations related to one or more embodiments. Details relating to the inference and / or training logic 815 are provided herein in conjunction with Figures 8A and / or 8B. In at least one embodiment, the deep learning application processor is used to train a machine learning model, such as a neural network, to predict or infer information provided to the deep learning application processor 2500. In at least one embodiment, the deep learning application processor 2500 is used to infer or predict information based on a trained machine learning model (e.g., a neural network) that has been trained by another processor or system, or by the deep learning application processor 2500. In at least one embodiment, the processor 2500 may be used to perform one or more neural network use cases described herein.
[0323] Figure 26 is a block diagram of a neuromorphic processor 2600 according to at least one embodiment. In at least one embodiment, the neuromorphic processor 2600 receives one or more inputs from an external source. In at least one embodiment, these inputs may be transmitted to one or more neurons 2602 within the neuromorphic processor 2600. In at least one embodiment, the neurons 2602 and their components may be implemented using circuits or logic including one or more arithmetic logic units (ALUs). In at least one embodiment, the neuromorphic processor 2600 may include thousands or millions of instances of neurons 2602, but any preferred number of neurons 2602 may be used. In at least one embodiment, each instance of neuron 2602 may include a neuron input 2604 and a neuron output 2606. In at least one embodiment, neuron 2602 may produce an output which may be transmitted to the input of other instances of neuron 2602. For example, in at least one embodiment, the neuron input 2604 and the neuron output 2606 may be interconnected via a synapse 2608.
[0324] In at least one embodiment, neuron 2602 and synapse 2608 may be interconnected so that the neuromorphic processor 2600 can operate to process or analyze information received by the neuromorphic processor 2600. In at least one embodiment, neuron 2602 may transmit an output pulse (or “fire” or “spike”) when the input received via neuron input 2604 exceeds a threshold. In at least one embodiment, neuron 2602 may sum or integrate the signals received at neuron input 2604. For example, in at least one embodiment, neuron 2602 may be implemented as a leaky integrate-and-fire neuron, where neuron 2602 may generate an output (or “fire”) using a transfer function such as a sigmoid function or a threshold function when the sum (called “membrane potential”) exceeds a threshold. In at least one embodiment, the leakage integral firing neuron may sum the signals received at neuron input 2604 to obtain the membrane potential, or it may apply a collapse factor (or leak) to reduce the membrane potential. In at least one embodiment, the leakage integral firing neuron may fire if multiple input signals are received at neuron input 2604 quickly enough to exceed a threshold (i.e., before the membrane potential collapse is too small to fire). In at least one embodiment, neuron 2602 may be implemented using a circuit or logic that receives the input, integrates the input to obtain the membrane potential, and collapses the membrane potential. In at least one embodiment, the input may be averaged, or any other suitable transfer function may be used. Furthermore, in at least one embodiment, neuron 2602 may include, without limitation, a comparator circuit or logic that generates an output spike at neuron 2606 when the result of applying the transfer function to neuron 2604 exceeds a threshold. In at least one embodiment, once neuron 2602 fires, it may ignore previously received input information, for example, by resetting the membrane potential to 0 or another suitable default value.In at least one embodiment, once the membrane potential is reset to 0, neuron 2602 may resume normal operation after a suitable period (or refractory period).
[0325] In at least one embodiment, neurons 2602 may be interconnected through synapses 2608. In at least one embodiment, synapses 2608 may operate to transmit a signal from the output of a first neuron 2602 to the input of a second neuron 2602. In at least one embodiment, neuron 2602 may transmit information through two or more instances of synapses 2608. In at least one embodiment, one or more instances of neuron output 2606 may be connected to an instance of neuron input 2604 of the same neuron 2602 via an instance of synapse 2608. In at least one embodiment, an instance of neuron 2602 that generates an output to be transmitted through an instance of synapse 2608 may be called a “presynaptic neuron” with respect to that instance of synapse 2608. In at least one embodiment, an instance of neuron 2602 that receives an input to be transmitted through an instance of synapse 2608 may be called a “postsynaptic neuron” with respect to that instance of synapse 2608. In at least one embodiment, an instance of neuron 2602 may receive input from one or more instances of synapse 2608 and transmit output through one or more instances of synapse 2608, so that a single instance of neuron 2602 may be both a "presynaptic neuron" and a "postsynaptic neuron" with respect to various instances of synapse 2608.
[0326] In at least one embodiment, the neurons 2602 may be organized into one or more layers. In at least one embodiment, each instance of neuron 2602 may have one neuron output 2606 that can fan out to one or more neuron inputs 2604 through one or more synapses 2608. In at least one embodiment, the neuron output 2606 of neuron 2602 in the first layer 2610 may be connected to the neuron input 2604 of neuron 2602 in the second layer 2612. In at least one embodiment, layer 2610 may be called a “feedforward” layer. In at least one embodiment, each instance of neuron 2602 in an instance of the first layer 2610 may fan out to each instance of neuron 2602 in the second layer 2612. In at least one embodiment, the first layer 2610 may be called a “fully connected feedforward layer”. In at least one embodiment, each instance of neuron 2602 in an instance of the second layer 2612 may be fanned out to fewer instances of neuron 2602 in the third layer 2614 than the total number of instances of neuron 2602 in the third layer 2614. In at least one embodiment, the second layer 2612 may be called a “loosely connected feedforward layer”. In at least one embodiment, neurons 2602 in the second layer 2612 may be fanned out to neurons 2602 in multiple other layers, including neurons 2602 in the second layer 2612. In at least one embodiment, the second layer 2612 may be called a “regression layer”. In at least one embodiment, the neuromorphic processor 2600 may include, without limitation, any preferred combination of regression layers and feedforward layers, including, without limitation, both loosely connected feedforward layers and fully connected feedforward layers.
[0327] In at least one embodiment, the neuromorphic processor 2600 may include, without limitation, a reconfigurable interconnect architecture or dedicated hardwired interconnect for connecting synapses 2608 to neurons 2602. In at least one embodiment, the neuromorphic processor 2600 may include, without limitation, circuits or logic that, based on the neural network topology and fan-in / fan-out of neurons, allow synapses to be allocated to different neurons 2602 as needed. For example, in at least one embodiment, synapse 2608 may be connected to neuron 2602 using an interconnect fabric such as a network-on-a-chip or using a dedicated connection. In at least one embodiment, synaptic interconnects and their components may be implemented using circuits or logic.
[0328] Figure 27 is a block diagram of a processing system according to at least one embodiment. In at least one embodiment, system 2700 includes one or more processors 2702 and one or more graphics processors 2708, and may be a single-processor desktop system, a multi-processor workstation system, or a server system having a large number of processors 2702 or processor cores 2707. In at least one embodiment, system 2700 is a processing platform embedded in a system-on-a-chip (SoC) integrated circuit for use in a mobile device, portable device, or embedded device.
[0329] In at least one embodiment, system 2700 may include, or be incorporated into, a server-based gaming platform, a game console including a game and media console, a mobile gaming console, a portable game console, or an online game console. In at least one embodiment, system 2700 is a mobile phone, a smartphone, a tablet computing device, or a mobile internet device. In at least one embodiment, processing system 2700 may also include, be coupled to, or be integrated into wearable devices such as a smartwatch wearable device, a smart eyewear device, an augmented reality device, or a virtual reality device. In at least one embodiment, processing system 2700 is a television or set-top box device having one or more processors 2702 and a graphical interface produced by one or more graphics processors 2708.
[0330] In at least one embodiment, each of the one or more processors 2702 includes one or more processor cores 2707 for processing instructions that, when executed, perform actions for the system and user software. In at least one embodiment, each of the one or more processor cores 2707 is configured to process a particular instruction sequence 2709. In at least one embodiment, the instruction sequence 2709 may facilitate computing via composite instruction set computing (CISC), reduced instruction set computing (RISC), or very long instruction words (VLIW). In at least one embodiment, each of the processor cores 2707 may process a different instruction sequence 2709, which may include instructions that facilitate the emulation of other instruction sequences. In at least one embodiment, the processor cores 2707 may also include other processing devices, such as a digital signal processor (DSP).
[0331] In at least one embodiment, the processor 2702 includes a cache memory 2704. In at least one embodiment, the processor 2702 may have a single internal cache or multiple levels of internal caches. In at least one embodiment, the cache memory is shared among various components of the processor 2702. In at least one embodiment, the processor 2702 also uses an external cache (e.g., a Level 3 (L3) cache or a Last Level Cache (LLC)) (not shown), which may be shared among processor cores 2707 using known cache coherence techniques. In at least one embodiment, the processor 2702 further includes a register file 2706, which may contain different types of registers for storing different types of data (e.g., integer registers, floating-point registers, state registers, and instruction pointer registers). In at least one embodiment, the register file 2706 may contain general-purpose registers or other registers.
[0332] In at least one embodiment, one or more processors 2702 are coupled to one or more interface buses 2710 to transmit communication signals, such as addresses, data, or control signals, between the processors 2702 and other components in the system 2700. In at least one embodiment, the interface bus 2710 may be a processor bus, such as a version of a Direct Media Interface (DMI) bus. In at least one embodiment...
Claims
1. A non-temporary computer-readable storage medium, which, as a result of being executed by one or more processors of a computer system, Sending a request to the Parallel Processing Unit (PPU) to enter a secure execution mode, In response to the aforementioned request, the PPU is reset, and when the entity performs an operation within the first portion of the memory, the first portion of the memory is protected so as to prevent the entity from performing an operation outside of the first portion of the memory. This involves having the hypervisor provide the aforementioned PPU to virtual machines running within a trusted execution environment, The virtual machine and the PPU are made to generate an encryption key used to encrypt data transmission between the virtual machine and the PPU, To enable the virtual machine to run applications using the PPU and A non-temporary computer-readable storage medium that stores executable instructions for causing the computer system to perform the aforementioned action.
2. The instruction is executed by one or more processors as a result of the execution of the instruction. The virtual machine is instructed to encrypt the data using the aforementioned encryption key and generate encrypted data. The virtual machine is instructed to write the encrypted data to a memory area outside the trusted execution environment that is accessible to the PPU. The encrypted data is to be acquired by the PPU, The PPU is instructed to write the encrypted data to the first portion of the memory by decrypting it at least using the encryption key. A non-temporary computer-readable storage medium according to claim 1, further comprising an instruction to cause the computer system to perform the following.
3. The instruction is executed by one or more processors as a result of the execution of the instruction. The secure processor of the PPU is made to encrypt the data using the aforementioned encryption key and generate encrypted data. The secure processor is instructed to write the encrypted data to a memory area outside the trusted execution environment that is accessible to the virtual machine. The virtual machine is instructed to obtain the data by at least decrypting the encrypted data using the aforementioned encryption key. A non-temporary computer-readable storage medium according to claim 1, further comprising an instruction to cause the computer system to perform the following.
4. The non-temporary computer-readable storage medium according to claim 1, wherein the instruction further includes an instruction causing the computer system to cause the virtual machine to authenticate the PPU based at least in part on a proof generated by the PPU as a result of execution by the one or more processors.
5. The non-transient computer-readable storage medium according to claim 1, wherein the application is a Compute Unified Device Architecture (CUDA) application.
6. The non-temporary computer-readable storage medium according to claim 1, wherein the instruction further includes an instruction causing the computer system to prevent the virtual machine and the hypervisor from writing to the first portion of the memory as a result of being executed by the one or more processors.
7. The non-temporary computer-readable storage medium according to claim 6, wherein the instruction causing the system to cause the virtual machine and the PPU to generate the encryption key further includes an instruction causing the computer system to generate the encryption key based at least in part on the cryptographic material associated with the PPU as a result of being executed by the one or more processors.
8. The non-temporary computer-readable storage medium according to claim 1, wherein the encryption key is a secret key.
9. It is a system, The method involves causing a parallel processing unit (PPU) to create a protected memory area, wherein the PPU's computing engine, when accessing the protected memory area, is prevented from accessing memory outside the protected memory area. The hypervisor is instructed to provide the aforementioned PPU to virtual machines within the trusted execution environment. It comprises one or more processors for performing the following: The aforementioned one or more processors further, The virtual machine is made to encrypt data using a shared encryption key and generate encrypted data, The encrypted data is stored in a buffer accessible to the PPU, The PPU is instructed to store the encrypted data in the protected memory area by at least decrypting the encrypted data with the shared encryption key. A system designed to perform a specific task.
10. The system according to claim 9, wherein the data transmitted between the PPU and the virtual machine is encrypted with an encryption key that is inaccessible to the hypervisor.
11. The system according to claim 10, wherein the cryptographic key is generated at least in part based on a secret key held by the secure processor of the PPU.
12. The system according to claim 9, wherein the one or more processors are further for preventing the virtual machine from accessing the protected memory area.
13. The system according to claim 9, wherein the one or more processors are further configured to indicate the memory region corresponding to the protected memory region to the memory management unit of the PPU.
14. The aforementioned one or more processors further, Determining that the aforementioned virtual machine has terminated, Resetting the aforementioned PPU, To erase the data held in the aforementioned protective memory area. The system according to claim 9, which is for performing the following.
15. The system according to claim 9, wherein the PPU is a graphics processing unit (GPU).
16. A method performed by one or more processors of a computer system, Sending a request to the Parallel Processing Unit (PPU) to enter a secure execution mode, In response to the aforementioned request, the PPU is reset, and when the entity performs an operation within the first portion of the memory, the first portion of the memory is protected so as to prevent the entity from performing an operation outside of the first portion of the memory. This involves having the hypervisor provide the aforementioned PPU to virtual machines running within a trusted execution environment, The virtual machine and the PPU are made to generate an encryption key used to encrypt data transmission between the virtual machine and the PPU, A method for enabling the virtual machine to run an application using the PPU.
17. A method performed by one or more processors, The method involves causing a parallel processing unit (PPU) to create a protected memory area, wherein the PPU's computing engine, when accessing the protected memory area, is prevented from accessing memory outside the protected memory area. The hypervisor is instructed to provide the aforementioned PPU to virtual machines within the trusted execution environment. moreover, The virtual machine is made to encrypt data using a shared encryption key and generate encrypted data, The encrypted data is stored in a buffer accessible to the PPU, A method comprising causing the PPU to store the encrypted data in the protected memory area by at least decrypting the encrypted data with the shared encryption key.