Techniques for parallel execution

By identifying opportunities for parallel instruction execution through a deep learning compiler, generating memory allocation plans and stream schedulers, and speculatively executing instructions, the problem of high resource consumption in neural network training and inference is solved, thereby improving efficiency and processor utilization.

CN115237551BActive Publication Date: 2026-04-10NVIDIA CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NVIDIA CORP
Filing Date
2022-04-19
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Training and using neural networks consumes significant amounts of memory, time, and computing resources, leading to inefficiency.

Method used

The deep learning compiler identifies opportunities for parallel instruction execution, generates memory allocation plans and stream schedulers, speculatively executes instructions to optimize resource utilization, and uses recurrent neural networks for training and inference.

Benefits of technology

It improves the efficiency of neural network training and inference, reduces memory and time consumption, and enhances processor utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115237551B_ABST
    Figure CN115237551B_ABST
Patent Text Reader

Abstract

Techniques directed to parallel execution are involved. Apparatuses, systems, and techniques to identify instructions for advanced execution are involved. In at least one embodiment, a processor executes one or more instructions that have been identified by a compiler as to be speculatively executed in parallel.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] At least one embodiment relates to processing resources for performing and facilitating artificial intelligence. For example, at least one embodiment relates to a processor or computing system for performing training and / or inference using neural networks in accordance with the various novel techniques described herein. BACKGROUND

[0002] Training neural networks and / or inference using neural networks can use significant memory, time, or computational resources. The amount of memory, time, or computational resources used for training neural networks and / or inference using neural networks can be improved. BRIEF DESCRIPTION OF DRAWINGS

[0003] Figure 1 is a block diagram illustrating a system for identifying instructions that can be speculatively executed, in accordance with at least one embodiment;

[0004] Figure 2 is a block diagram illustrating a system for speculatively executing instructions by launching from a host to a device, in accordance with at least one embodiment;

[0005] Figure 3 is a flow diagram of a technique for generating instructions, in accordance with at least one embodiment;

[0006] Figure 4 is a flow diagram of a technique for identifying possible speculative instructions, in accordance with at least one embodiment;

[0007] Figure 5 is a flow diagram of a technique for speculatively launching instructions, in accordance with at least one embodiment;

[0008] Figure 6 shows a comparison of inference operations performed over time, in accordance with at least one embodiment;

[0009] Figure 7A shows inference and / or training logic, in accordance with at least one embodiment;

[0010] Figure 7B shows inference and / or training logic, in accordance with at least one embodiment;

[0011] Figure 8 shows training and deployment of a neural network, in accordance with at least one embodiment;

[0012] Figure 9 shows an example data center system, in accordance with at least one embodiment;

[0013] Figure 10A shows an example of an autonomous vehicle, in accordance with at least one embodiment;

[0014] Figure 10B An example of a camera position and field of view of an autonomous vehicle is shown in accordance with at least one embodiment Figure 10A An example of a camera position and field of view of an autonomous vehicle is shown in accordance with at least one embodiment

[0015] Figure 10C A block diagram illustrating an example system architecture of an autonomous vehicle is shown in accordance with at least one embodiment Figure 10A A block diagram illustrating an example system architecture of an autonomous vehicle is shown in accordance with at least one embodiment

[0016] Figure 10D A diagram illustrating a system for communication between one or more cloud-based servers and an autonomous vehicle is shown in accordance with at least one embodiment Figure 10A A diagram illustrating a system for communication between one or more cloud-based servers and an autonomous vehicle is shown in accordance with at least one embodiment

[0017] Figure 11 A block diagram illustrating a computer system is shown in accordance with at least one embodiment

[0018] Figure 12 A block diagram illustrating a computer system is shown in accordance with at least one embodiment

[0019] Figure 13 A computer system is shown in accordance with at least one embodiment

[0020] Figure 14 A computer system is shown in accordance with at least one embodiment

[0021] Figure 15A A computer system is shown in accordance with at least one embodiment

[0022] Figure 15B A computer system is shown in accordance with at least one embodiment

[0023] Figure 15C A computer system is shown in accordance with at least one embodiment

[0024] Figure 15D A computer system is shown in accordance with at least one embodiment

[0025] Figure 15E A computer system is shown in accordance with at least one embodiment Figure 15F A computer system is shown in accordance with at least one embodiment

[0026] Figure 16 An exemplary integrated circuit and associated graphics processor are shown in accordance with at least one embodiment

[0027] Figure 17A An exemplary integrated circuit and associated graphics processor are shown in accordance with at least one embodiment Figure 17B An exemplary integrated circuit and associated graphics processor are shown in accordance with at least one embodiment

[0028] Figure 18A An exemplary integrated circuit and associated graphics processor are shown in accordance with at least one embodiment Figure 18BAn additional exemplary graphics processor logic is shown in accordance with at least one embodiment;

[0029] Figure 19 A computer system is shown in accordance with at least one embodiment;

[0030] Figure 20A A parallel processor is shown in accordance with at least one embodiment;

[0031] Figure 20B A partition unit is shown in accordance with at least one embodiment;

[0032] Figure 20C A processing cluster is shown in accordance with at least one embodiment;

[0033] Figure 20D A graphics multiprocessor is shown in accordance with at least one embodiment;

[0034] Figure 21 A multi-GPU system is shown in accordance with at least one embodiment;

[0035] Figure 22 A graphics processor is shown in accordance with at least one embodiment;

[0036] Figure 23 is a block diagram showing a processor micro-architecture for a processor in accordance with at least one embodiment;

[0037] Figure 24 A deep learning application processor is shown in accordance with at least one embodiment;

[0038] Figure 25 is a block diagram showing an example neuromorphic processor in accordance with at least one embodiment;

[0039] Figure 26 At least portions of a graphics processor are shown in accordance with one or more embodiments;

[0040] Figure 27 At least portions of a graphics processor are shown in accordance with one or more embodiments;

[0041] Figure 28 At least portions of a graphics processor are shown in accordance with one or more embodiments;

[0042] Figure 29 is a block diagram showing a graphics processing engine of a graphics processor in accordance with at least one embodiment;

[0043] Figure 30 is a block diagram showing at least portions of a graphics processor core in accordance with at least one embodiment;

[0044] Figure 31A and Figure 31B Thread execution logic is shown that includes an array of processing elements of a graphics processor core, in accordance with at least one embodiment.

[0045] Figure 32 A parallel processing unit (“PPU”) is shown, in accordance with at least one embodiment.

[0046] Figure 33 A general processing cluster (“GPC”) is shown, in accordance with at least one embodiment.

[0047] Figure 34 A memory partition unit of a parallel processing unit (“PPU”) is shown, in accordance with at least one embodiment.

[0048] Figure 35 A streaming multiprocessor is shown, in accordance with at least one embodiment.

[0049] Figure 36 An example dataflow graph of an advanced computing pipeline, in accordance with at least one embodiment.

[0050] Figure 37 A system diagram of an example system for training, adapting, instantiating, and deploying machine learning models in an advanced computing pipeline, in accordance with at least one embodiment.

[0051] Figure 38 An example illustration of an advanced computing pipeline 3710A for processing imaging data, in accordance with at least one embodiment.

[0052] Figure 39A An example dataflow graph of a virtual instrument that supports an ultrasound device, in accordance with at least one embodiment.

[0053] Figure 39B An example dataflow graph of a virtual instrument that supports a CT scanner, in accordance with at least one embodiment.

[0054] Figure 40A A dataflow graph of a process for training a machine learning model is shown, in accordance with at least one embodiment; and

[0055] Figure 40B An example illustration of a client-server architecture that leverages a pre-trained annotation model to augment an annotation tool, in accordance with at least one embodiment. DETAILED DESCRIPTION

[0056] Figure 1is a block diagram illustrating a system 100 for identifying instructions that can be speculatively executed in parallel, in accordance with at least one embodiment. In at least one embodiment, a deep learning (DL) compiler 102 uses a representation 104 of a computer program to generate a modified representation 106 of the computer program that indicates at least one operation that can be executed speculatively. In at least one embodiment, DL compiler 102 is a computer program that runs on a processor (e.g., a CPU) and is accessible via an application programming interface (API). In at least one embodiment, representation 104 of a computer program includes instructions to be launched by a host (e.g., a computer system having a CPU) on a device (e.g., a parallel processing unit (PPU), such as a graphics processing unit (GPU)). In at least one embodiment, representation 104 of a computer program includes operations using a neural network, such as a recurrent neural network (RNN). In at least one embodiment, representation 104 of a computer program is a graph representation.

[0057] In at least one embodiment, modified representation 106 of a computer program includes instructions to be launched by a host (e.g., a computer system having a CPU) on a device (e.g., a PPU, a GPU, or other suitable acceleration device) and indicates instructions that can be launched speculatively by the host on the device. In at least one embodiment, modified representation 106 of a computer program is a modified graph representation. In at least one embodiment, modified representation 106 of a computer program includes representation 104 of a computer program and a list and / or other data structure that indicates operations in representation 104 of a computer program that can be safely executed speculatively. In at least one embodiment, modified representation 106 of a computer program is a tagged and / or annotated version of representation 104 of a computer program.

[0058] In at least one embodiment, deep learning compiler 102 also generates a memory allocation plan 108 based at least in part on modified representation 106 of a computer program. In at least one embodiment, memory allocation plan 108 includes extended live ranges of variables and / or values (e.g., tensors) used in instructions (e.g., operations) that have indications that they can be executed speculatively. In at least one embodiment, deep learning compiler 102 generates extended live ranges and stores indications of extended live ranges in modified representation 106 of a computer program or some other data structure instead of or in addition to memory allocation plan 108.

[0059] In at least one embodiment, a stream scheduler 110 of deep learning compiler 102 generates modified representation 106 of computer program. In at least one embodiment, stream scheduler 110 is a computer program that runs on a processor (e.g., CPU) and is accessible via an API. In at least one embodiment, stream scheduler 106 identifies operations in representation 104 of computer program that can be speculatively executed, and modified representation 106 of computer program includes at least one indication of operations that can be speculatively executed. In at least one embodiment, stream scheduler 110 performs stream scheduling only within a basic block (e.g., no tasks from different basic blocks will be overlapped with each other). In at least one embodiment, stream scheduler 110 performs stream scheduling within and across basic blocks. In at least one embodiment, instead of or in addition to stream scheduler 110, some other part of deep learning compiler 102 generates at least part of modified representation 106 of computer program.

[0060] In at least one embodiment, a memory allocator 112 of deep learning compiler 102 generates memory allocation plan 108 and / or other indication of extended live periods. In at least one embodiment, memory allocator 112 is a computer program that runs on a processor (e.g., CPU) and is accessible via an API. In at least one embodiment, instead of or in addition to memory allocator 112, some other part of deep learning compiler 102 generates memory allocation plan 108 and / or other indication of extended live periods.

[0061] In at least one embodiment, DL compiler 102 finds high-level execution opportunities (e.g., with stream scheduler 110) based at least in part on representation 104 of computer program (e.g., a graph), finds operations that can be safely executed during those opportunities, and allocates memory addresses accordingly (e.g., with memory allocator 112). In at least one embodiment, high-level execution and / or speculatively executing an instruction or operation refers to launching an instruction (e.g., one or more operations in a kernel) from a host to a device before receiving an indication that the instruction needs to be launched (e.g., a value of a branch condition received during a device-to-host copy operation). In at least one embodiment, at least one action of DL compiler 102 can be described with respect to the following pseudo code:

[0062]

[0063] In at least one embodiment, DL compiler 102 generates a list of speculation opportunities for each by finding device-to-host copy operations in representation 104 of computer program. In at least one embodiment, DL compiler 102 generates a list of speculation opportunities based on at least one other type of operation instead of or in addition to finding device-to-host copy operations. In at least one embodiment, list of speculation opportunities is assigned to variable spec_ops, as indicated by pseudo code above. In at least one embodiment, DL compiler 102 traverses nodes of representation 104 of computer program based at least in part on identified speculation opportunities. In at least one embodiment, DL compiler 102 selects one branch from among multiple branches that follow a conditional branch associated with a speculation opportunity (e.g., selects a branch that follows a conditional value of True or a branch that follows a conditional value of False). In at least one embodiment, DL compiler 102 uses a heuristic technique to select a branch from among multiple branches (e.g., selects a branch that returns to a beginning of a loop when a loop typically iterates multiple times). In at least one embodiment, DL compiler 102 traverses nodes in a selected branch (e.g., nodes in subsequent_operations (op)) to identify whether an operation can be safely performed. In at least one embodiment, DL compiler 102 initially assumes that an operation can be safely performed, but flags it as unsafe if it is a particular operation (e.g., changes a random state, overwrites an output, uses a scan input, uses a signal instruction, uses a wait instruction, uses a different stream, and / or has a parent node that is not flagged as safe) because performing these types of operations can lead to undesirable side effects if performed speculatively. In at least one embodiment, DL compiler 102 extends a variable’s live range for operations flagged as safe to perform speculatively.

[0064] In at least one embodiment, deep learning compiler 102, although referred to as a compiler, generates a modified representation 106 of a computer program (e.g., with stream scheduler 110) and a memory allocation plan 108 (e.g., with memory allocator 112), but does not generate runtime code sufficient to execute a computer program corresponding to representation 104 of a computer program. In at least one embodiment, representation 104 of a computer program is generated by a deep learning framework (e.g., TensorFlow or PyTorch). In at least one embodiment, representation 104 of a computer program is a graph. In at least one embodiment, stream scheduler 110 and memory allocator 112 operate via a common API. In at least one embodiment, compiler and / or interpreter 114 generates runtime code 116 based at least in part on modified representation 106 of a computer program and memory allocation plan 108. In at least one embodiment, memory allocation plan 108 is included as part of modified representation 106 of a computer program. In at least one embodiment, in addition to being based on modified representation 106 of a computer program and memory allocation plan 108, compiler / interpreter 114 generates runtime code 116 based at least in part on other input 118 (e.g., portions of a computer program not represented by a graph representation of a neural network). In at least one embodiment, DL compiler 102 generates runtime code 116 (e.g., by integrating compiler / interpreter 114 in DL compiler 102). In at least one embodiment, runtime code 116 is stored for later use (e.g., in memory and / or persistent storage). In at least one embodiment, runtime code 116 is used soon after being generated (e.g., for just-in-time compilation). In at least one embodiment, stream scheduler 110, memory allocator 112, and compiler / interpreter 114 (e.g., as a compiler) are integrated into a combined compiler that performs operations described with respect to stream scheduler 110, memory allocator 112, and compiler / interpreter 114 to generate runtime code 116 at compile time. In at least one embodiment, combined compiler is accessible via an API.

[0065] In at least one embodiment, the representation 104 of the computer program is structured data (e.g., data according to a predetermined format and / or syntax) that represents an entire computer program. In at least one embodiment, the representation 104 of the computer program is structured data that represents a portion of a computer program rather than an entire computer program, where the representation can define a directed acyclic graph (DAG) to indicate use of tensor data in a deep learning neural network. In at least one embodiment, each node of the DAG represents an operation that produces some tensor output, and each edge represents a tensor producer-consumer relationship. In at least one embodiment, a client using system 100 (e.g., an application using system 100 to compile and / or run deep learning neural network training and / or inference techniques) launches instructions to be executed speculatively based at least in part on the runtime code 116, the modified representation 106 of the program, and / or the memory allocation plan 108.

[0066] In at least one embodiment, the DL compiler 102 generates the modified representation 106 of the computer program based at least in part on adding one or more indicators to the representation 104 of the computer program. In at least one embodiment, the indicators are referred to as annotations. In at least one embodiment, instead of or in addition to adding indicators to the representation 104 of the computer program, the DL compiler 102 generates a data structure that includes a list of instructions in the representation 104 of the computer program that can be safely executed speculatively. In at least one embodiment, the representation 104 of the computer program includes a graph that performs inference using a neural network (e.g., a recurrent neural network (RNN)). In at least one embodiment, the representation 104 of the computer program includes a graph that performs training using a neural network (e.g., a RNN). In at least one embodiment, the representation 104 of the computer program is for an image processing application.

[0067] Figure 2 FIG. 1 is a block diagram illustrating a system 100 to speculatively execute instructions by launching from a host 102 to a device 104, according to at least one embodiment. In at least one embodiment, the host 102 is a computer system including a processor 106 (e.g., a CPU) and a memory 108. In at least one embodiment, the device 104 is an accelerator including a processor 110 (e.g., one or more parallel processors) and a memory 112. In at least one embodiment, the device 104 is a PPU or GPU. In at least one embodiment, the host 102 is a computer system including a processor 106 (e.g., a CPU) and a memory 108. In at least one embodiment, the device 104 is an accelerator including a processor 110 (e.g., one or more parallel processors) and a memory 112. In at least one embodiment, the device 104 is a PPU or GPU. In at least one embodiment, the host 102 is a computer system including a processor 106 (e.g., a CPU) and a memory 108. In at least one embodiment, the device 104 is an accelerator including a processor 110 (e.g., one or more parallel processors) and a memory 112. In at least one embodiment, the device 104 is a PPU or GPU. Figure 1 The DL compiler 102 of FIG. 1 runs on the host 102.

[0068] In at least one embodiment, host 202 initiates operations and / or instructions to be executed on device 204 (e.g., by initiating parallel processing framework instructions, such as a Compute Unified Device Architecture (CUDA) kernel). In at least one embodiment, host 202 bases its operations at least in part on one or more indications that can be speculatively performed (e.g., Figure 1 The modified representation of the computer program 106 (commented instructions and / or operations and / or runtime code 116) initiates one or more instructions to be presumably executed by the device 204. In at least one embodiment, the host 202 and / or device 204 are at least partially based on the instructions and / or operations and / or runtime code 116 of the computer program 106. Figure 1 The DL compiler 102 recognizes extended active periods to extend the active period of variables.

[0069] In at least one embodiment, an executor (not shown for clarity) running on host 202 is at least partially based on a modified representation of a computer program (e.g., Figure 1 The modified representation 106 and / or runtime code 116 of the computer program initiates instructions speculatively. In at least one embodiment, the executor runs on a CPU (e.g., processor 206) and initiates instructions (e.g., as a kernel) on a parallel processing unit (e.g., GPU). In at least one embodiment, the executor is a virtual machine running on processor 206 (e.g., CPU).

[0070] In at least one embodiment, in the compiled graph (e.g., Figure 1 During the execution of the modified representation 106 or runtime code 116, when an instruction with a comment indicating that an operation can be performed speculatively is encountered, at least one aspect of the execution process is described by the following pseudocode:

[0071]

[0072]

[0073] In at least one embodiment, host 202 uses a kernel booted by host 202 and executed by device 204 to initiate operations that will be speculatively executed on device 204. In at least one embodiment, high-level execution and / or speculative execution of instructions or operations refers to starting the kernel before receiving an indication that the kernel needs to be started (e.g., a value of a branch condition received during a device-to-host copy operation). In at least one embodiment, starting the kernel in this manner provides performance advantages, but results in the execution of kernels that may not be necessary in some cases.

[0074] In at least one embodiment, the processor 210 of device 204 includes one or more circuits for executing one or more instructions, which have been prepared by a compiler (e.g.,Figure 1 The DL compiler 102) identifies that the instructions are to be executed speculatively in parallel. In at least one embodiment, the system 200 includes one or more memories (e.g., the memory 208 before the kernel launch instruction and the memory 212 after the kernel launch instruction while the device 204 is executing the instructions) for storing instructions that are identified to be executed speculatively in parallel. In at least one embodiment, the identification of instructions to be executed speculatively in parallel (e.g., on a PPU or GPU) means that the instructions include an identifier (e.g., a tag, a comment, or other suitable identifier associated with the instructions) that a host (e.g., the host 202) can launch the instructions on the device (e.g., using a kernel launch operation) before certain instructions will be needed (e.g., before a value indicating a branch condition is received by the host in a device-to-host copy operation).

[0075] In at least one embodiment, the instructions have been identified by a compiler to be executed speculatively in parallel based at least in part on identifying copy operations, and one or more circuits of the processor 210 are to execute one or more instructions based at least in part on receiving a command from another processor (e.g., the processor 206 of the host 202). In at least one embodiment, the instructions have been identified by a compiler to be executed speculatively in parallel based at least in part on identifying copy operations between a parallel processing unit (e.g., the device 204) and a host computer system (e.g., the host 202) and marking safe operations after one or more identified copy operations. In at least one embodiment, it is understood that the copy operations are identified from general device-to-host copy operations (e.g., in the representation 104 of the computer program) but are performed with respect to a particular device (e.g., the device 204) and a particular host (e.g., the host 202) when executed. In at least one embodiment, the instructions include an extended lifetime of a variable used by an operation associated with the instructions identified to be executed speculatively in parallel. In at least one embodiment, the instructions identified to be executed speculatively in parallel are instructions that can be executed speculatively, not necessarily that they will be executed speculatively. In at least one embodiment, the host 202 stops speculatively launching instructions after a branch condition after the host 202 receives a value indicating the branch condition, even if the host 202 has not launched all possible instructions after the branch condition that have been identified by the compiler to be launched speculatively.

[0076] In at least one embodiment, processor 210 is part of a PPU, and one or more circuits of processor 210 are configured to execute one or more instructions identified for high-level execution after receiving a kernel boot command from a host computer system (e.g., host 202). In at least one embodiment, the instructions identified for high-level execution are part of a while loop. In at least one embodiment, the instructions identified for high-level execution are part of some other type of loop (e.g., a counting loop), or a different type of code segment following a branch condition. In at least one embodiment, the instructions use a recurrent neural network to implement part of the inference operation. In at least one embodiment, the instructions have been compiled by a compiler based at least in part on a representation of a computer program using a neural network (e.g., Figure 1 The computer program representation 104 identifies one or more conditional branches. In at least one embodiment, one or more processors of host 202 (e.g., processor 206) speculatively initiate one or more instructions to be executed by one or more processors of device 204 (e.g., processor 210), and one or more processors of host 202 stop speculatively initiating instructions in response to receiving a value of a condition prior to receiving one or more speculatively executed instructions in the computer program representation via a copy operation (e.g., from device 204).

[0077] Figure 3 A flowchart illustrating a technology 300 for generating instructions according to at least one embodiment is provided. In at least one embodiment, technology 300 is executed by at least one circuit, at least one system, at least one processor, at least one graphics processing unit, at least one parallel processor, and / or at least some other processors or components thereof, as described and / or illustrated herein. In at least one embodiment, at least one aspect of technology 300 is... Figure 1 The DL compiler 102 is executed.

[0078] In at least one embodiment, at block 302, technique 300 includes identifying speculative opportunities (e.g., in...). Figure 1In at least one embodiment, identifying a speculation opportunity comprises finding a device-to-host copy operation. In at least one embodiment, the operation that is allowed to have high-level execution of other operations is an asynchronous device-to-host copy. In at least one embodiment, if pageable memory is used for asynchronous device-to-host copy, such that even for asynchronous copy runs synchronously, another form of memory is used during execution of speculative operations (e.g., pinned memory), so operations can be executed speculatively. In at least one embodiment, identifying a speculation opportunity comprises finding an operation that corresponds to a device-to-host copy operation that can be safely initiated before receiving data that is copied. In at least one embodiment, identifying a speculation opportunity comprises searching for device-to-host copies in a representation (e.g., a graph) of a computer program (e.g., during flow scheduling). In at least one embodiment, device-to-host copies found by searching the graph are treated as speculation points.

[0079] In at least one embodiment, at block 304, technique 300 comprises marking safe operations. In at least one embodiment, marking safe operations comprises identifying safe operations. In at least one embodiment, operations are considered safe if they do not have side effects (e.g., change a random state), do not break any data dependencies, and do not overwrite memory of any tensors that are valid during a copy operation. In at least one embodiment, for each speculation point (e.g., identified by a device-to-host copy operation found at block 302), consecutive operations are iterated in execution order. In at least one embodiment, iteration is interrupted at the end of a basic block, a wait on a default stream, or a speculation point, which can be the same speculation point. In at least one embodiment, when an operation has side effects (e.g., can change something other than inputs and outputs, such as a random or certain custom operations), uses scanned inputs / outputs, uses a non-default stream, is a signal or wait instruction, or depends on another unsafe operation, the operation is marked as unsafe. In at least one embodiment, scanned inputs / outputs are considered safe instead of unsafe if additional checks on memory boundaries are satisfied. In at least one embodiment, all other operations are considered safe. In at least one embodiment, operations are marked as safe or unsafe (e.g., with a data structure, an annotation, a tag, or in a separate data structure that associates operations with corresponding safety indicators).

[0080] In at least one embodiment, identifying a speculation opportunity at block 302 and marking safe operations at block 304 are further explained with reference to the following pseudo-code example:

[0081] Tensors a, b, and f are live into BB1.

[0082] BB1:

[0083] 1. c = add a b

[0084] 2. d = rand a / / unsafe

[0085] 3. e = conv c f

[0086] 4. g = add d e / / unsafe, depends on d

[0087] 5. d2h copy g

[0088] 6. cjmp g [BB1, BB2]

[0089] BB2…

[0090] In at least one embodiment, with respect to the above pseudocode, first search in the basic block (e.g., by Figure 1 ’s deep learning compiler 102) device to host (d2h) copy, which in this case is operation (op) 5. In at least one embodiment, the technique continues iterating over the operations starting at op 5 (e.g., where the d2h is found), the interruption condition of the iteration is the end of the basic block, the stream 0 wait, and other speculation points, where special handling of cjmp loops back to the start of the basic block. In at least one embodiment, during the iteration, at op 6, the iterator moves to op 1. In at least one embodiment, op 1 is marked as safe. In at least one embodiment, op 2 is a random op, which causes side effects, so it is marked as unsafe. In at least one embodiment, op 3 is marked as safe. In at least one embodiment, op 4 depends on d from op 2, so it is marked as unsafe. In at least one embodiment, the iteration stops at op 5, where op 1 and op 3 were previously marked as safe.

[0091] In at least one embodiment, at block 306, technique 300 includes expanding live intervals of variables. In at least one embodiment, expanding live intervals of variables includes scanning the graph (e.g., modified representation 106 of a computer program of Figure 1 ’s deep learning compiler 102) for speculation points, and then for each operation marked as safe for speculation, the live periods of tensors used by the operation are expanded to ensure they are live during the speculation point. In at least one embodiment, expanding live periods ensures that no violation of anti-dependences and / or output dependences occurs after resource allocation. In at least one embodiment, the expanded live intervals of variables at block 306 are used in memory allocation using an assumption that operations can execute out of order.

[0092] In at least one embodiment, at block 308, technique 300 includes generating instructions (e.g., modified representation 106 of computer program, memory allocation plan 108, and / or runtime code 116) that move the copy operation. In at least one embodiment, at block 310, technique 300 includes performing other actions. In at least one embodiment, performing other actions at block 310 includes moving the copy operation in generated instructions and / or modified representation of computer program, such as by moving the device-to-host copy operation to just after defining a device variable, or to just before using a host variable if the copy operation is not already located in such a location. Figure 1

[0093] In at least one embodiment, technique 300 is performed at least in part by execution of a set of instructions (e.g., from a non-transitory machine-readable medium) using one or more processors (e.g., of host 202) or any other suitable processor as shown or described herein. In at least one embodiment, technique 300 includes identifying one or more instructions to be executed speculatively in parallel (e.g., using DL compiler 102 of host 202). In at least one embodiment, technique 300 includes identifying instructions to be executed speculatively based at least in part on identifying a copy operation between a parallel processing unit and a host computer system in a representation of a computer program (e.g., at block 302). In at least one embodiment, technique 300 identifies an operation that is safe to execute after the copy operation (e.g., at block 304). In at least one embodiment, technique 300 includes marking the operation that is safe to execute speculatively (e.g., at block 304). In at least one embodiment, technique 300 includes extending a lifetime of a variable associated with the operation that is marked as safe to execute speculatively (e.g., at block 306). In at least one embodiment, technique 300 includes searching a representation of a computer program for a copy operation between a GPU and a host computer system, and identifying a copy operation that is safe to execute speculatively after the copy operation. In at least one embodiment, technique 300 includes extending a lifetime of a variable associated with the operation that is identified to be executed speculatively in parallel. In at least one embodiment, technique 300 includes, based at least in part on identifying a copy operation in a representation of a computer program, finding a conditional branch in the representation of the computer program; selecting one path from a plurality of paths after the conditional branch; and identifying instructions that are safe to execute speculatively in the selected path. Figure 2 Figure 1

[0094] Figure 4 ​​​A flow diagram illustrating techniques 400 of identifying possible speculative instructions, in accordance with at least one embodiment, is shown. In at least one embodiment, techniques 400 are performed by at least one circuit, at least one system, at least one processor, at least one graphics processing unit, at least one parallel processor, and / or at least some other processor or component described and / or illustrated herein. In at least one embodiment, one or more aspects of techniques 400 are performed by a DL compiler 102 executing a computer program. Figure 1

[0095] In at least one embodiment, at block 402, techniques 400 include identifying a representation of a set of instructions (e.g., a graph of a computer program’s representation 104). In at least one embodiment, identifying a representation of a set of instructions includes receiving an API function call that includes a representation of a set of instructions (e.g., a graph), a pointer to a representation of a set of instructions, a link to a representation of a set of instructions, or any other suitable manner of identifying a representation of a set of instructions. Figure 1

[0096] In at least one embodiment, at block 404, techniques 400 include finding a device-to-host copy operation (e.g., in a computer program’s representation 104 of a computer program). In at least one embodiment, finding a device-to-host copy operation includes traversing a representation of a set of instructions identified at block 402 to search for and find a device-to-host copy operation. In at least one embodiment, at block 404, techniques 400 include finding a conditional branch instead of finding a device-to-host copy operation, in an alternative or additional manner. In at least one embodiment, finding a device-to-host copy operation includes searching for an asynchronous device-to-host copy operation. Figure 1

[0097] In at least one embodiment, at block 406, techniques 400 include selecting a branch path. In at least one embodiment, selecting a branch path at block 406 includes selecting one or more branch paths from a plurality of branch paths. In at least one embodiment, selecting a branch path at block 406 includes using a heuristic to select a branch path (e.g., when a loop is typically executed a number of times, selecting a path that executes additional iteration instructions in a loop instead of instructions after a loop terminates). In at least one embodiment, at block 408, techniques 400 include performing other actions. In at least one embodiment, performing other actions at block 408 includes returning to block 404 to identify additional copy operations.

[0098] Figure 5 ​​​A flowchart of a technique 500 to speculatively launch instructions is shown, in accordance with at least one embodiment. In at least one embodiment, the technique 500 is performed by at least one circuit, at least one system, at least one processor, at least one graphics processing unit, at least one parallel processor, and / or at least some other processor or component thereof, described and / or shown herein. In at least one embodiment, at block 502, the technique 500 includes identifying a speculation opportunity. In at least one embodiment, identifying a speculation opportunity at block 502 includes identifying a point in a deep learning network at which a host must wait for a device to be able to select a branch to take control flow (e.g., replicating an operation for a branch condition). In at least one embodiment, identifying a speculation opportunity at block 502 includes identifying a device-to-host replication operation. In at least one embodiment, identifying a speculation opportunity at block 502 includes identifying an operation previously identified (e.g., by an annotation, tag, or any other suitable identifier) as safe for execution speculatively (e.g., by Figure 1 the DL compiler 102).

[0099] In at least one embodiment, at block 504, the technique 500 includes speculatively launching an instruction. In at least one embodiment, the speculatively launching instruction of block 504 is performed by the host 202. Figure 2 In at least one embodiment, the speculatively launching instruction of block 504 is performed by an executor (e.g., a virtual machine running on the host). In at least one embodiment, the speculatively launching instruction at block 504 includes launching a kernel after a branch condition before some kernel would have been needed.

[0100] In at least one embodiment, at block 506, the technique 500 includes performing other actions. In at least one embodiment, performing other actions at block 506 includes returning to block 502 to identify a next speculation opportunity. In at least one embodiment, performing other actions at block 506 includes overriding a device-to-host replication command, including replicating to a fixed memory, creating an event, logging the event, launching a safe operation (e.g., at block 504) while querying the event, and destroying the event. In at least one embodiment, overriding by a replication implementation is described by the following pseudo code:

[0101]

[0102]

[0103] In at least one embodiment, as shown by the above pseudo code, overriding includes the ability of one instruction to launch other instructions and a way to skip over instructions during a main execution context.

[0104] In at least one embodiment, performing other actions at block 506 includes (e.g., by Figure 2by the compiler (e.g., the DL compiler 102 of FIG. 1) to be speculatively and Figure 1 In at least one embodiment, the instruction has been identified by the compiler to be speculatively and concurrently executed based at least in part on identifying a conditional branch and selecting a path from a plurality of paths after the conditional branch. In at least one embodiment, the instruction has been identified by the compiler to be speculatively and concurrently executed based at least in part on identifying a copy operation. In at least one embodiment, the instruction includes an expansion live range of a variable used in an operation that is executed speculatively. In at least one embodiment, the instruction implements a portion of an inference operation using a neural network.

[0105] Figure 6 A comparison 600 of inference operations performed over time is shown, in accordance with at least one embodiment. In at least one embodiment, a first simplified representation 602 shows inference operations of a neural network performed over time based at least in part on not using advanced execution. In at least one embodiment, a second simplified representation 604 shows inference operations performed over time using the same neural network as about the first simplified representation 602 but using advanced execution (e.g., as described with respect to one or more of Figure 1-5

[0106] In at least one embodiment, the first simplified representation 602 includes a top portion 606 showing operations performed on a host computer system and a bottom portion 608 showing operations performed on a device (e.g., a PPU or GPU). In at least one embodiment, the second simplified representation 604 includes a top portion 610 showing operations performed on a host computer system and a bottom portion 612 showing operations performed on a device (e.g., a PPU or GPU).

[0107] In at least one embodiment, to infer but not perform a speculative operation, an end of a first iteration (e.g., of a time step loop of an RNN) is marked by line 614 and an end of a second iteration is marked by line 616. In at least one embodiment, for inference including performing a speculative operation (e.g., as described with respect to one or more of Figure 1-5 In at least one embodiment, for inference including performing a speculative operation (e.g., as described with respect to one or more of

[0108] ​In at least one embodiment, gap 622 exists in second iteration of operations performed on device indicated in bottom portion 608 prior to first operation 614 being performed on device for second iteration. In at least one embodiment, gap 622 is caused by requirement of host system to wait for result of condition returned from device (e.g., value of variable indicating whether loop is complete) before initiating additional operations on device. In at least one embodiment, host system waits for synchronization operation 624 (e.g., stream synchronization operation between host system and device) to complete before initiating additional instructions when execution of speculative operations is not supported. In at least one embodiment, for inferences including execution of speculative operations, no gap similar to gap 622 exists prior to first operation 626 being performed on device. In at least one embodiment, when execution of speculative operations is supported, host system initiates kernel for next iteration (e.g., kernel containing instructions tagged safe for advanced execution) without performing a synchronization operation and / or before host system has received an indication that kernel will be needed for next iteration. In at least one embodiment, initiating kernel for next iteration in this manner results in device having work available for faster execution (e.g., no gap similar to gap 622 prior to first operation 626). In at least one embodiment, this results in performance advantages, such as greater utilization of device and / or shorter time to complete iterations, as compared to systems that do not support execution of speculative operations.

[0109] In at least one embodiment, comparison 600 illustrates a simplified representation of operations of a decoder (e.g., natural language model, such as Tacotron2 decoder, or any other suitable decoder) over time (e.g., one iteration between line 614 and line 616 for a system that does not support speculative operations, and one iteration between line 618 and line 620 for a system that supports speculative operations). In at least one embodiment, comparison 600 illustrates a simplified representation of what a system tracks (e.g., as generated by a system performance analysis tool, such as NVIDIA Nsight Systems, or any other suitable performance analysis tool).

[0110] In at least one embodiment, comparison 600 illustrates that performing speculative operations (e.g., as described with respect to one or more of Figure 1-5 In at least one embodiment, performing speculative operations provides performance advantages over traditional methods in which host system can only initiate work up to a branch, and must wait for device to complete before selecting a path to take. In at least one embodiment, performing speculative operations (e.g., as described with respect to one or more of Figure 1-5one or more of the Figures) results in performance advantages for inference and / or training with RNNs, where the output of a node is fed back as input, and the RNN behaves like a loop of nodes across a time series. In at least one embodiment, performing speculative operations increases utilization of a device (e.g., device 204 of FIG. 2) and provides a performance improvement of approximately 15% compared to conventional techniques that do not perform operations speculatively (e.g., do not launch a kernel from a host to a device until it is determined that the kernel will be needed). Figure 2 In at least one embodiment, performing speculative operations (e.g., as described with respect to one or more of the Figures) results in performance advantages for inference and / or training with RNNs, where the output of a node is fed back as input, and the RNN behaves like a loop of nodes across a time series. In at least one embodiment, performing speculative operations increases utilization of a device (e.g., device 204 of FIG. 2) and provides a performance improvement of approximately 15% compared to conventional techniques that do not perform operations speculatively (e.g., do not launch a kernel from a host to a device until it is determined that the kernel will be needed). Figure 1-5 In at least one embodiment, performing speculative operations (e.g., as described with respect to one or more of the Figures) results in performance advantages for inference and / or training with RNNs, where the output of a node is fed back as input, and the RNN behaves like a loop of nodes across a time series. In at least one embodiment, performing speculative operations increases utilization of a device (e.g., device 204 of FIG. 2) and provides a performance improvement of approximately 15% compared to conventional techniques that do not perform operations speculatively (e.g., do not launch a kernel from a host to a device until it is determined that the kernel will be needed). Figure 1-5 In at least one embodiment, performing speculative operations (e.g., as described with respect to one or more of the Figures) results in performance advantages for inference and / or training with RNNs, where the output of a node is fed back as input, and the RNN behaves like a loop of nodes across a time series. In at least one embodiment, performing speculative operations increases utilization of a device (e.g., device 204 of FIG. 2) and provides a performance improvement of approximately 15% compared to conventional techniques that do not perform operations speculatively (e.g., do not launch a kernel from a host to a device until it is determined that the kernel will be needed).

[0111]

[0112] In at least one embodiment, instructions in the body of the while loop of the above pseudocode are launched by the host and computed on the device. In at least one embodiment, after all instructions have been launched, the host needs to wait for the final memory copy to complete before it can launch more instructions (e.g., without advanced execution). In at least one embodiment, with advanced execution (e.g., as described with respect to one or more of the Figures), the host can continue launching instructions for the next iteration before evaluating the loop condition. Figure 1-5 In at least one embodiment, instructions in the body of the while loop of the above pseudocode are launched by the host and computed on the device. In at least one embodiment, after all instructions have been launched, the host needs to wait for the final memory copy to complete before it can launch more instructions (e.g., without advanced execution). In at least one embodiment, with advanced execution (e.g., as described with respect to one or more of the Figures), the host can continue launching instructions for the next iteration before evaluating the loop condition.

[0113] Inference and Training Logic

[0114] Figure 7A Inference and / or training logic 715 is shown performing inference and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 715 are provided below in conjunction with FIGS. 7A-7D. Figure 7A and / or Figure 7B Details regarding inference and / or training logic 715 are provided below in conjunction with FIGS. 7A-7D.

[0115] In at least one embodiment, inference and / or training logic 715 can include, without limitation, code and / or data storage 701 for storing forward and / or output weights and / or input / output data, and / or other parameters of neurons or layers of a neural network configured in aspects of one or more embodiments that are trained and / or used for inferencing. In at least one embodiment, training logic 715 can include or be coupled to code and / or data storage 701 for storing graph code or other software to control timing and / or order, where weight and / or other parameter information is loaded to configure logic, including integer and / or floating point units (collectively, arithmetic logic units (ALUs)). In at least one embodiment, code, such as graph code, loads weight or other parameter information into processor ALUs based on an architecture of a neural network to which that code corresponds. In at least one embodiment, code and / or data storage 701 stores weight parameters and / or input / output data of each layer of a neural network trained or used in conjunction with one or more embodiments during forward propagation of input / output data and / or weight parameters during training and / or inferencing using aspects of one or more embodiments. In at least one embodiment, any portion of code and / or data storage 701 can be included with other on-chip or off-chip data storage, including a processor’s LI, L2, or L3 cache or system memory.

[0116] In at least one embodiment, any portion of code and / or data storage 701 can be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, code and / or data storage 701 can be cache memory, dynamic random addressable memory (“DRAM”), static random addressable memory (“SRAM”), non-volatile memory (such as flash memory), or other storage. In at least one embodiment, a choice of whether code and / or data storage 701 is internal or external to a processor, e.g., or made up of DRAM, SRAM, flash or some other storage type, can depend on available storage space on-chip or off-chip, latency requirements of training and / or inferencing functions being performed, batch size of data used in inferencing and / or training of a neural network, or some combination of these factors.

[0117] In at least one embodiment, inference and / or training logic 715 can include, without limitation, code and / or data storage 705 to store backward and / or output weight and / or input / output data for neurons or layers of a neural network trained and / or used for inferencing in aspects of one or more embodiments. In at least one embodiment, code and / or data storage 705 stores weight parameters and / or input / output data for each layer of a neural network trained or used in conjunction with one or more embodiments during backward propagation of input / output data and / or weight parameters during training and / or inferencing using aspects of one or more embodiments. In at least one embodiment, training logic 715 can include or be coupled to code and / or data storage 705 to store graph code or other software to control timing and / or order, wherein weight and / or other parameter information is loaded to configure logic, including integer and / or floating point units (collectively, arithmetic logic units (ALUs)).

[0118] In at least one embodiment, code such as graph code causes weight or other parameter information to be loaded into processor ALUs based on an architecture of a neural network to which the code corresponds. In at least one embodiment, any portion of code and / or data storage 705 can be included with other on-chip or off-chip data storage, including a processor’s LI, L2, or L3 cache or system memory. In at least one embodiment, any portion of code and / or data storage 705 can be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, code and / or data storage 705 can be cache memory, DRAM, SRAM, non-volatile memory (e.g., Flash) or other storage. In at least one embodiment, whether code and / or data storage 705 is internal or external to a processor, e.g., whether made up of DRAM, SRAM, Flash, or some other storage type, is a choice dictated by design parameters including, without limitation, whether available storage is on-chip or off-chip, latency requirements of training and / or inferencing functions being performed, data batch size used in inferencing and / or training of a neural network, or some combination of these factors.

[0119] In at least one embodiment, code and / or data storage 701 and code and / or data storage 705 can be separate storage structures. In at least one embodiment, code and / or data storage 701 and code and / or data storage 705 can be the same storage structures. In at least one embodiment, code and / or data storage 701 and code and / or data storage 705 can be partially combined and partially separate. In at least one embodiment, any portion of code and / or data storage 701 and code and / or data storage 705 can be included with other on-chip or off-chip data storage, including a processor’s LI, L2, or L3 cache or system memory.

[0120] In at least one embodiment, inference and / or training logic 715 can include, without limitation, one or more arithmetic logic units (“ALUs”), including integer and / or floating-point units, for performing logical and / or mathematical operations based, at least in part, on training and / or inference code (e.g., graph code) or instructions therefrom. Results of such operations can result in activations (e.g., output values from layers or neurons within a neural network) stored in activation storage 720, which are functions of input / output and / or weight parameter data stored in code and / or data storage 701 and / or code and / or data storage 705. In at least one embodiment, activations stored in activation storage 720 are generated by ALU 710 in response to executing instructions or other code, linear algebra and / or matrix-based mathematics performed by ALU 710, where weight values stored in code and / or data storage 705 and / or code and / or data storage 701 are used as operands along with other values, such as bias values, gradient information, momentum values, or other parameters or hyperparameters, any or all of which can be stored in code and / or data storage 705 or code and / or data storage 701 or other on-chip or off-chip storage.

[0121] In at least one embodiment, one or more processors or other hardware logic devices or circuits include one or more ALUs 710, while in another embodiment, one or more ALUs 710 may be located outside the processor or other hardware logic device or the circuitry that uses them (e.g., a coprocessor). In at least one embodiment, one or more ALUs 710 may be included within an execution unit of a processor, or otherwise included in a group of ALUs accessible by the execution unit of the processor, which may be within the same processor or distributed among different processors of different types (e.g., a central processing unit, a graphics processing unit, a fixed-function unit, etc.). In at least one embodiment, code and / or data storage 701, code and / or data storage 705, and activation storage 720 may share a processor or other hardware logic device or circuitry, while in another embodiment, they may be located in different processors or other hardware logic devices or circuitry, or in some combination of the same and different processors or other hardware logic devices or circuitry. In at least one embodiment, any portion of activation storage 720 may be included together with other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory. Furthermore, inference and / or training code may be stored together with other code accessible to the processor or other hardware logic or circuitry, and may be retrieved and / or processed using the processor’s fetch, decode, schedule, execute, exit, and / or other logic circuitry.

[0122] In at least one embodiment, the active memory 720 may be a cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other memory. In at least one embodiment, the active memory 720 may be wholly or partially located inside or outside one or more processors or other logic circuits. In at least one embodiment, the choice of whether the active memory 720 is internal to or external to the processor may depend on the available on-chip or off-chip storage, the latency requirements for training and / or inference functions, the batch size of data used in inference and / or training the neural network, or some combination of these factors. For example, it may include DRAM, SRAM, flash memory, or other memory types.

[0123] In at least one embodiment, Figure 7A The inference and / or training logic 715 shown can be used in conjunction with an application-specific integrated circuit (“ASIC”), such as the Tensor from Google. Processing unit, from Graphcore TM Inference processing units (IPUs) or from Intel Corp. (e.g., "Lake Crest") processor. In at least one embodiment, Figure 7AThe illustrated inference and / or training logic 715 can be used in combination with central processing unit (“CPU”) hardware, graphics processing unit (“GPU”) hardware, or other hardware (e.g., field programmable gate arrays (“FPGAs”)).

[0124] Figure 7B Inference and / or training logic 715 are illustrated as a portion of the system 700 in the example of FIG. 7. In other examples, the inference and / or training logic 715 can be used in combination with other systems or devices, including other systems or devices illustrated in FIG. 7. Figure 7B The inference and / or training logic 715 illustrated in FIG. 7 can be used in combination with application-specific integrated circuit (ASIC) hardware, such as Tensor Processing Units from Google, an inference processing unit (IPU) from Graphcore TM AI100 from Alibaba, or a Lake Crest) processor. In at least one embodiment, the inference and / or training logic 715 illustrated in FIG. 7 can be used in combination with central processing unit (CPU) hardware, graphics processing unit (GPU) hardware, or other hardware (e.g., field programmable gate arrays (FPGAs)). Figure 7B Inference and / or training logic 715 are illustrated as a portion of the system 700 in the example of FIG. 7. In other examples, the inference and / or training logic 715 can be used in combination with other systems or devices, including other systems or devices illustrated in FIG. 7. Figure 7B In at least one embodiment, each of code and / or data storage 701 and code and / or data storage 705 are respectively associated with dedicated computing resources (e.g., computing hardware 702 and computing hardware 706). In at least one embodiment, each of computing hardware 702 and computing hardware 706 includes one or more ALUs that only perform mathematical functions (e.g., linear algebraic functions) on information stored in code and / or data storage 701 and code and / or data storage 705, respectively, results of performing the functions being stored in activation storage 720.

[0125] In at least one embodiment, each of code and / or data stores 701 and 705 and corresponding compute hardware 702 and 706, respectively, correspond to different layers of a neural network, such that activations resulting from one “storage / compute pair 701 / 702” of code and / or data store 701 and compute hardware 702 provide input to next “storage / compute pair 705 / 706” of code and / or data store 705 and compute hardware 706 in order to reflect a conceptual organization of a neural network. In at least one embodiment, each storage / compute pair 701 / 702 and 705 / 706 can correspond to more than one neural network layer. In at least one embodiment, additional storage / compute pairs (not shown) can be included in inference and / or training logic 715 after or in parallel with storage / compute pairs 701 / 702 and 705 / 706.

[0126] Neural network training and deployment

[0127] Figure 8 Training and deployment of a deep neural network is shown, in accordance with at least one embodiment. In at least one embodiment, an untrained neural network 806 is trained using a training dataset 802. In at least one embodiment, training framework 804 is a PyTorch framework, while in other embodiments, training framework 804 is TensorFlow, Boost, Caffe, Microsoft Cognitive Toolkit / CNTK, MXNet, Chainer, Keras, Deeplearning4j, or other training framework. In at least one embodiment, training framework 804 trains untrained neural network 806 and enables it to train using processing resources described herein to generate a trained neural network 808. In at least one embodiment, weights can be chosen randomly or by pre-training using a deep belief network. In at least one embodiment, training can be performed in a supervised, partially supervised, or unsupervised manner.

[0128] In at least one embodiment, an untrained neural network 806 is trained using supervised learning, where a training dataset 802 includes inputs paired with desired outputs for inputs, or where a training dataset 802 includes inputs with known outputs and the neural network 806 is manually graded outputs. In at least one embodiment, an untrained neural network 806 is trained in a supervised manner and processes inputs from a training dataset 802 and compares resulting outputs to a set of expected or desired outputs. In at least one embodiment, errors are then propagated back through the untrained neural network 806. In at least one embodiment, a training framework 804 adjusts weights that control the untrained neural network 806. In at least one embodiment, a training framework 804 includes tools for monitoring how well an untrained neural network 806 is converging towards a model (e.g., a trained neural network 808) that is suitable for generating correct answers (e.g., results 814) based on input data (e.g., new datasets 812). In at least one embodiment, a training framework 804 trains an untrained neural network 806 repeatedly while adjusting weights to improve outputs of the untrained neural network 806 using a loss function and an adjustment algorithm (e.g., stochastic gradient descent). In at least one embodiment, a training framework 804 trains an untrained neural network 806 until the untrained neural network 806 reaches a desired accuracy. In at least one embodiment, a trained neural network 808 can then be deployed to implement any number of machine learning operations.

[0129] In at least one embodiment, an untrained neural network 806 is trained using unsupervised learning, where an untrained neural network 806 attempts to train itself using unlabeled data. In at least one embodiment, an unsupervised learning training dataset 802 will include input data without any associated output data or “ground truth” data. In at least one embodiment, an untrained neural network 806 can learn groupings within a training dataset 802 and can determine how individual inputs relate to an untrained dataset 802. In at least one embodiment, unsupervised training can be used to generate a self-organizing map in a trained neural network 808 that is capable of performing operations useful for reducing dimensions of a new dataset 812. In at least one embodiment, unsupervised training can also be used to perform anomaly detection, which allows for identification of data points in a new dataset 812 that deviate from a normal pattern of the new dataset 812.

[0130] In at least one embodiment, semi-supervised learning can be used, which is a technique where a mix of labeled and unlabeled data is included in training dataset 802. In at least one embodiment, training framework 804 can be used to perform incremental learning, for example, through a passing-learning technique. In at least one embodiment, incremental learning enables trained neural network 808 to adapt to new dataset 812 without forgetting knowledge that was imprinted into trained neural network 808 during initial training.

[0131] data center

[0132] Figure 9 An example data center 900 that can use at least one embodiment is shown. In at least one embodiment, data center 900 includes a data center infrastructure layer 910, a framework layer 920, a software layer 930, and an application layer 940.

[0133] In at least one embodiment, as shown Figure 9 In at least one embodiment, data center infrastructure layer 910 can include a resource orchestrator 912, grouped computing resources 914, and node computing resources (“node C.R.s”) 916(1)-916(N), where “N” represents a positive integer (which can be a different integer “N” than used in other figures). In at least one embodiment, node C.R.s 916(1)-916(N) can include, but are not limited to, any number of central processing units (“CPUs” or “processors”) or other processors (including accelerators, field programmable gate arrays (FPGAs), graphics processors, etc.), memory storage devices 918(1)-918(N) (such as dynamic read-only memory, solid-state drives, or disk drives), network input / output (“NW I / O”) devices, network switches, virtual machines (“VMs”), power modules, and cooling modules, etc. In at least one embodiment, one or more of node C.R.s 916(1)-916(N) can be a server having one or more of above-described computing resources.

[0134] In at least one embodiment, the grouped computing resource 914 may include individual groups (not shown) of node CRs housed within one or more racks, or a plurality of racks (also not shown) housed within data centers in various geographical locations. In at least one embodiment, the individual groups of node CRs within the grouped computing resource 914 may include computing, networking, memory, or storage resources that can be configured or allocated to support groups of one or more workloads. In at least one embodiment, several node CRs, including CPUs or processors, may be grouped within one or more racks to provide computing resources to support one or more workloads. In at least one embodiment, the one or more racks may also include any number of power modules, cooling modules, and network switches, in any combination.

[0135] In at least one embodiment, resource coordinator 912 may configure or otherwise control one or more nodes CR916(1)-916(N) and / or grouped computing resources 914. In at least one embodiment, resource coordinator 912 may include a Software Design Infrastructure (“SDI”) management entity for data center 900. In at least one embodiment, resource coordinator 912 may include hardware, software, or some combination thereof.

[0136] In at least one embodiment, such as Figure 9 As shown, framework layer 920 includes job scheduler 922, configuration manager 924, resource manager 926, and distributed file system 928. In at least one embodiment, framework layer 920 may include a framework of software 932 supporting software layer 930 and / or one or more applications 942 supporting application layer 940. In at least one embodiment, software 932 or application 942 may respectively include web-based service software or applications, such as services or applications provided by Amazon Web Services, Google Cloud, and Microsoft Azure. In at least one embodiment, framework layer 920 may be, but is not limited to, a free and open-source software web application framework, such as Apache Spark, which can utilize distributed file system 928 for large-scale data processing (e.g., "big data"). TM(“Spark”). In at least one embodiment, a job scheduler 932 can include a Spark driver to facilitate scheduling of workloads supported by various layers of data center 900. In at least one embodiment, a configuration manager 924 can be capable of configuring different layers, such as software layer 930 and framework layer 920 including Spark and a distributed file system 928 for supporting large-scale data processing. In at least one embodiment, a resource manager 926 can be capable of managing clustered or grouped computing resources mapped to or allocated for supporting distributed file system 928 and job scheduler 922. In at least one embodiment, clustered or grouped computing resources can include grouped computing resources 914 on data center infrastructure layer 910. In at least one embodiment, resource manager 926 can coordinate with resource orchestrator 912 to manage these mapped or allocated computing resources.

[0137] In at least one embodiment, software 932 included in software layer 930 can include software used by at least a portion of node C.R.s 916(1)-916(N), grouped computing resources 914, and / or distributed file system 928 of framework layer 920. In at least one embodiment, one or more types of software can include, but are not limited to, Internet web page search software, email virus scanning software, database software, and streaming video content software.

[0138] In at least one embodiment, one or more applications 942 included in application layer 940 can include one or more types of applications used by at least a portion of node C.R.s 916(1)-916(N), grouped computing resources 914, and / or distributed file system 928 of framework layer 920. In at least one embodiment, one or more types of applications can include, but are not limited to, any number of genomics applications, cognitive computing, applications, and machine learning applications, including training or inferencing software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), or other machine learning applications used in conjunction with one or more embodiments.

[0139] In at least one embodiment, any of configuration manager 924, resource manager 926, and resource orchestrator 912 can implement any number and type of self-modifying actions based on any number and type of data acquired in any technically feasible manner. In at least one embodiment, self-modifying actions can relieve data center operators of data center 900 from making possibly poor configuration decisions and can avoid underutilized and / or poorly performing portions of a data center.

[0140] In at least one embodiment, data center 900 may include tools, services, software, or other resources to train one or more machine learning models or to use one or more machine learning models to predict or infer information according to one or more embodiments described herein. For example, in at least one embodiment, a machine learning model can be trained by calculating weight parameters based on a neural network architecture using the software and computing resources described above with respect to data center 900. In at least one embodiment, information can be inferred or predicted using trained machine learning models corresponding to one or more neural networks using the resources described above with respect to data center 900 by using weight parameters calculated through one or more training techniques described herein.

[0141] In at least one embodiment, the data center may use a CPU, application-specific integrated circuit (ASIC), GPU, FPGA, or other hardware to utilize the aforementioned resources to perform training and / or inference. Furthermore, one or more of the aforementioned software and / or hardware resources may be configured as a service to allow a user to train or perform information inference, such as image recognition, speech recognition, or other artificial intelligence services.

[0142] Inference and / or training logic 715 is used to perform inference and / or training operations associated with one or more embodiments. This document combines... Figure 7A and / or Figure 7B Details regarding inference and / or training logic 715 are provided. In at least one embodiment, inference and / or training logic 715 may be used in System Figure 9 for inference or prediction operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.

[0143] In at least one embodiment, relative to Figure 9 At least one component shown or described is used to achieve the combination Figure 1-6 The described techniques and / or functions. In at least one embodiment, the inference and / or training logic 715 includes and / or operates regarding... Figure 1 At least one aspect described (e.g., deep learning compiler 102, stream scheduler 110, memory allocator 112). In at least one embodiment, inference and / or training logic 715 uses a representation of a computer program to train at least one untrained or partially trained neural network, the representation of which may be as described regarding Figure 1-6 One or more of the operations and / or instructions described in the description are speculatively performed. In at least one embodiment, the inference and / or training logic uses a representation of a computer program to perform at least one inference operation, the representation of which may indicate, as per [the description of the inference and / or training logic]... Figure 1-6one or more of the operations and / or instructions described in the specification.

[0144] Autonomous vehicle

[0145] Figure 10A An example of an autonomous vehicle 1000 is shown in accordance with at least one embodiment. In at least one embodiment, autonomous vehicle 1000 (alternatively referred to herein as “vehicle 1000”) can be, but is not limited to, a passenger vehicle such as a car, truck, bus, and / or another type of vehicle that can accommodate one or more passengers. In at least one embodiment, vehicle 1000 can be a semi-trailer tractor for hauling cargo. In at least one embodiment, vehicle 1000 can be an airplane, a robotic vehicle, or another type of vehicle.

[0146] Autonomous vehicles can be described in terms of levels of automation defined by the National Highway Traffic Safety Administration (“NHTSA”), a division of the U.S. Department of Transportation, and the Society of Automotive Engineers (“SAE”) “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles” (e.g., Standard No. J3016-201806 published on June 15, 2018, Standard No. J3016-201609 published on September 30, 2016, and previous and future versions of this standard). In at least one embodiment, vehicle 1000 can be capable of functioning in accordance with one or more of Levels 1 through Level 5 of the levels of automation. For example, in at least one embodiment, vehicle 1000 can be capable of conditional automation (Level 3), high automation (Level 4), and / or full automation (Level 5), depending on the embodiment.

[0147] In at least one embodiment, vehicle 1000 can include, but is not limited to, components such as a chassis, a body, wheels (e.g., 2, 4, 6, 8, 18, etc.), tires, axles, and other components of a vehicle. In at least one embodiment, vehicle 1000 can include, but is not limited to, a propulsion system 1050 such as a combustion engine, a hybrid electric device, a fully electric motor, and / or another type of propulsion system. In at least one embodiment, propulsion system 1050 can be connected to a drivetrain of vehicle 1000, which can include, but is not limited to, a transmission to enable propulsion of vehicle 1000. In at least one embodiment, propulsion system 1050 can be controlled in response to receiving a signal from a throttle / accelerator 1052.

[0148] In at least one embodiment, when the propulsion system 1050 is operating (e.g., when the vehicle 1000 is traveling), the steering system 1054 (which may include, but is not limited to, a steering wheel) is used to steer the vehicle 1000 (e.g., along a desired path or route). In at least one embodiment, the steering system 1054 may receive signals from the steering actuator 1056. In at least one embodiment, the steering wheel may be optional for fully automated (Level 5) functionality. In at least one embodiment, the brake sensor system 1046 may be used to operate the vehicle brakes in response to signals received from the brake actuator 1048 and / or brake sensors.

[0149] In at least one embodiment, the controller 1036 may include, but is not limited to, one or more system-on-chips (“SoCs”). Figure 10A A controller 1036 (not shown) and / or a graphics processing unit (“GPU”) provides signals (e.g., representing commands) to one or more components and / or systems of vehicle 1000. For example, in at least one embodiment, controller 1036 may send signals to operate vehicle braking via brake actuator 1048, to operate steering system 1054 via one or more steering actuators 1056, and to operate propulsion system 1050 via one or more throttles / accelerators 1052. In at least one embodiment, one or more controllers 1036 may include one or more onboard (e.g., integrated) computing devices that process sensor signals and output operating commands (e.g., signals representing commands) to enable autonomous driving and / or assist a driver in driving vehicle 1000. In at least one embodiment, one or more controllers 1036 may include a first controller for autonomous driving functions, a second controller for functional safety functions, a third controller for artificial intelligence functions (e.g., computer vision), a fourth controller for infotainment functions, a fifth controller for redundancy in emergency situations, and / or other controllers. In at least one embodiment, a single controller may handle two or more of the functions described above, and two or more controllers may handle a single function and / or any combination thereof.

[0150] In at least one embodiment, one or more controllers 1036 provide signals for controlling one or more components and / or systems of vehicle 1000 in response to sensor data received from one or more sensors (e.g., sensor inputs). In at least one embodiment, sensor data can be received from sensors, including but not limited to one or more Global Navigation Satellite System (“GNSS”) sensors 1058 (e.g., one or more Global Positioning System sensors), one or more RADAR sensors 1060, one or more ultrasonic sensors 1062, one or more LiDAR sensors 1064, one or more Inertial Measurement Unit (IMU) sensors 1066 (e.g., one or more accelerometers, one or more gyroscopes, one or more magnetic compasses, one or more magnetometers, etc.), one or more microphones 1096, one or more stereo cameras 1068, one or more wide-angle cameras 1070 (e.g., fisheye cameras), one or more infrared cameras 1072, one or more surround cameras 1074 (e.g., 360-degree cameras), and remote cameras (…). Figure 10A The system receives a mid-range camera (not shown in Figure 10A), one or more speed sensors 1044 (e.g., for measuring the speed of vehicle 1000), one or more vibration sensors 1042, one or more steering sensors 1040, one or more brake sensors (e.g., as part of brake sensor system 1046) and / or other sensor types.

[0151] In at least one embodiment, one or more controllers 1036 may receive input (e.g., represented by input data) from the dashboard 1032 of the vehicle 1000 and provide output (e.g., represented by output data, display data, etc.) via a human-machine interface (“HMI”) display 1034, a voice signaler, a speaker, and / or other components of the vehicle 1000. In at least one embodiment, the output may include information such as vehicle speed, velocity, time, map data (e.g., high-definition map). Figure 10A The HMI display 1034 may display information such as (not shown in the image), location data (e.g., the location of vehicle 1000, for example on a map), direction, the location of other vehicles (e.g., occupancy raster), information about objects, and the state of objects sensed by one or more controllers 1036. For example, in at least one embodiment, the HMI display 1034 may display information about the presence of one or more objects (e.g., road signs, warning signs, traffic light changes, etc.) and / or information about driving operations that the vehicle has already made, is making, or will make (e.g., changing lanes now, exiting exit 34B within two miles, etc.).

[0152] In at least one embodiment, the vehicle 1000 also includes a network interface 1024, which can communicate over one or more networks using one or more wireless antennas 1026 and / or one or more modems. For example, in at least one embodiment, the network interface 1024 may be able to communicate over Long Term Evolution (“LTE”), Wideband Code Division Multiple Access (“WCDMA”), Universal Mobile Telecommunications System (“UMTS”), Global System for Mobile Communications (“GSM”), IMT-CDMA Multicarrier (“CDMA2000”) networks, etc. In at least one embodiment, one or more wireless antennas 1026 may also enable communication between objects in the environment (e.g., vehicles, mobile devices) using one or more local area networks (e.g., Bluetooth, Bluetooth Low Energy (LE), Z-Wave, ZigBee, etc.) and / or one or more low-power wide area networks (hereinafter referred to as “LPWAN”) (e.g., LoRaWAN, SigFox, etc. protocols).

[0153] Inference and / or training logic 715 is used to perform inference and / or training operations associated with one or more embodiments. This document combines... Figure 7A and / or Figure 7B Details regarding the inference and / or training logic 715 are provided. In at least one embodiment, the inference and / or training logic 715 may be used in System Figure 10A to infer or predict operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.

[0154] In at least one embodiment, relative to Figure 10A At least one component shown or described is used to achieve the combination Figure 1-6 The described technologies and / or functions. In at least one embodiment, the inference and / or training logic 715 of the vehicle 1000 (regarding...) Figure 10C (Shown as part of CPU 1006 and GPU 1008) includes and / or runs about Figure 1 At least one aspect described (e.g., deep learning compiler 102, stream scheduler 110, memory allocator 112). In at least one embodiment, inference and / or training logic 715 uses a representation of a computer program to train at least one untrained or partially trained neural network, the representation of which may be as described regarding Figure 1-6 One or more of the operations and / or instructions described in the description are speculatively performed. In at least one embodiment, the inference and / or training logic 715 uses a representation of a computer program to perform at least one inference operation, the representation of which may indicate, as per [the description of the inference operation]. Figure 1-6One or more operations and / or instructions described as speculatively performed. In at least one embodiment, vehicle 1000 includes a computer vision system comprising one or more processors for performing one or more reasoning operations, at least in part, based on a representation of a computer program to identify one or more trajectories of corresponding one or more objects, the computer program including representations already compiled by a compiler (e.g., Figure 1 The DL compiler 102) identifies one or more instructions to be executed speculatively in parallel. In at least one embodiment, the vehicle 1000 includes one or more of a propulsion system, a steering control system, and a vehicle operator notification system to perform one or more actions (e.g., acceleration, braking, steering, warning signals) based at least in part on the identified one or more trajectories.

[0155] Figure 10B The illustration shows an embodiment according to at least one of the embodiments. Figure 10A Examples of camera positions and fields of view for an autonomous vehicle 1000. In at least one embodiment, the camera and its respective field of view are an example embodiment and are not intended to be limiting. For example, in at least one embodiment, additional and / or alternative cameras may be included and / or the cameras may be located at different positions on the vehicle 1000.

[0156] In at least one embodiment, the camera type used for the camera may include, but is not limited to, a digital camera suitable for use with components and / or systems of vehicle 1000. In at least one embodiment, one or more cameras may operate at Automotive Safety Integrity Level (“ASIL”) B and / or other ASILs. In at least one embodiment, the camera type may have any image capture rate, such as 60 frames per second (fps), 1220 fps, 240 fps, etc. In at least one embodiment, the camera may be able to use a rolling shutter, a global shutter, another type of shutter, or a combination thereof. In at least one embodiment, the color filter array may include a red-to-clear (“RCCC”) color filter array, a red-to-clear-blue (“RCCB”) color filter array, a red-blue-green (“RBGC”) color filter array, a Foveon X3 color filter array, a Bayer sensor (“RGGB”) color filter array, a monochrome sensor color filter array, and / or other types of color filter arrays. In at least one embodiment, a transparent pixel camera, such as a camera with an array of RCCC, RCCB and / or RBGC color filters, may be used to improve photosensitivity.

[0157] In at least one embodiment, one or more cameras can be used to perform advanced driver assistance system (“ADAS”) functions (e.g., as part of a redundant or fail-safe design). For example, in at least one embodiment, a multi-functional mono camera can be installed to provide functions including lane departure warning, traffic sign assist, and intelligent headlamp control. In at least one embodiment, one or more cameras (e.g., all cameras) can simultaneously record and provide image data (e.g., video).

[0158] In at least one embodiment, one or more cameras can be mounted in mounting assemblies, such as custom designed (three-dimensional (“3D”) printed) assemblies, in order to cut out stray light and reflections from within vehicle 1000 (e.g., dashboard reflections in windshield mirror) that can interfere with camera’s image data capture capabilities. With regard to rearview mirror mounting assemblies, in at least one embodiment, rearview mirror assemblies can be 3D printed custom designed such that camera mounting plates match a shape of a rearview mirror. In at least one embodiment, one or more cameras can be integrated into a rearview mirror. In at least one embodiment, for side view cameras, one or more cameras can also be integrated within four pillars at each corner of a cabin.

[0159] In at least one embodiment, cameras with a field of view that includes a portion of an environment in front of vehicle 1000 (e.g., front-facing cameras) can be used for surround view, as well as to help identify a path and obstacles ahead with the help of one or more controllers 1036 and / or control SoCs, thereby providing information that is critical to generating an occupancy grid and / or determining a preferred vehicle path. In at least one embodiment, front-facing cameras can be used to perform many ADAS functions similar to LIDAR, including but not limited to emergency braking, pedestrian detection, and collision avoidance. In at least one embodiment, front-facing cameras can also be used for ADAS functions and systems including, but not limited to, lane departure warning (“LDW”), automatic cruise control (“ACC”), and / or other functions (such as traffic sign recognition).

[0160] In at least one embodiment, various cameras can be used in a front-facing configuration, including, for example, a monocular camera platform including a CMOS (“complementary metal-oxide semiconductor”) color imager. In at least one embodiment, a wide-view camera 1070 can be used to perceive objects (e.g., pedestrians, crossing or bicycles) entering from the periphery. Although in Figure 10BOnly one wide-angle camera 1070 is shown; however, in other embodiments, the vehicle 1000 may have any number (including zero) of wide-angle cameras. In at least one embodiment, any number of remote cameras 1098 (e.g., a pair of remote stereo cameras) can be used for depth-based object detection, especially for objects for which a neural network has not yet been trained. In at least one embodiment, the remote camera 1098 can also be used for object detection and classification, as well as basic object tracking.

[0161] In at least one embodiment, any number of stereo cameras 1068 may also be included in a forward configuration. In at least one embodiment, one or more stereo cameras 1068 may include an integrated control unit comprising a scalable processing unit that may provide programmable logic (“FPGA”) and a multi-core microprocessor with a controller area network (“CAN”) or Ethernet interface integrated on a single chip. In at least one embodiment, such a unit may be used to generate a 3D map of the environment of the vehicle 1000, including distance estimates for all points in the image. In at least one embodiment, one or more stereo cameras 1068 may include, but are not limited to, a compact stereo vision sensor, which may include, but is not limited to, two camera lenses (one on the left and one on the right) and an image processing chip that can measure the distance from the vehicle 1000 to a target object and use the generated information (e.g., metadata) to activate autonomous emergency braking and lane departure warning functions. In at least one embodiment, other types of stereo cameras 1068 may also be used in addition to those described herein.

[0162] In at least one embodiment, a camera (e.g., a side-view camera) having a field of view including a portion of the environment on the sides of the vehicle 1000 can be used for surround viewing, thereby providing information for creating and updating the occupied grid, and generating a side collision warning. For example, in at least one embodiment, a surround camera 1074 (e.g., such as...) Figure 10B The four surround cameras shown can be positioned on vehicle 1000. In at least one embodiment, one or more surround cameras 1074 can include, but are not limited to, any number and combination of wide-angle cameras, one or more fisheye lenses, one or more 360-degree cameras, and / or similar cameras. For example, in at least one embodiment, four fisheye lens cameras can be located at the front, rear, and sides of vehicle 1000. In at least one embodiment, vehicle 1000 can use three surround cameras 1074 (e.g., left, right, and rear) and can utilize one or more other cameras (e.g., forward-facing cameras) as a fourth surround-view camera.

[0163] In at least one embodiment, a camera (e.g., a rear-view camera) having a field of view including a portion of the environment behind the vehicle 1000 can be used for parking assistance, surround view, rear collision warning, and creating and updating occupancy raster. In at least one embodiment, a wide variety of cameras can be used, including but not limited to cameras that are also suitable as one or more forward-facing cameras (e.g., long-range camera 1098 and / or one or more mid-range cameras 1076, one or more stereo cameras 1068, one or more infrared cameras 1072, etc.), as described herein.

[0164] The inference and / or training logic 715 is used to perform inference and / or training operations associated with one or more embodiments. Figure 7A and / or Figure 7B This document provides details regarding inference and / or training logic 715. In at least one embodiment, inference and / or training logic 715 can be... Figure 10B Used in systems for reasoning or predicting operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.

[0165] In at least one embodiment, relative to Figure 10B At least one component shown or described is used to achieve the combination Figure 1-6 The described technologies and / or functions. In at least one embodiment, the inference and / or training logic 715 of the vehicle 1000 (regarding...) Figure 10C (Shown as part of CPU 1006 and GPU 1008) includes and / or runs about Figure 1 At least one aspect described (e.g., deep learning compiler 102, stream scheduler 110, memory allocator 112). In at least one embodiment, inference and / or training logic 715 uses a representation of a computer program to train at least one untrained or partially trained neural network, the representation of which may be as described regarding Figure 1-6 One or more of the operations and / or instructions described in the description are speculatively performed. In at least one embodiment, the inference and / or training logic uses a representation of a computer program to perform at least one inference operation, the representation of which may indicate, as per [the description of the inference and / or training logic]... Figure 1-6 One or more of the operations and / or instructions described as being performed speculatively.

[0166] Figure 10C The illustration shows an embodiment according to at least one of the embodiments. Figure 10A A block diagram of an example system architecture for an autonomous vehicle 1000. In at least one embodiment, Figure 10CEach of the one or more components, one or more features, and one or more systems of vehicle 1000 in FIG. 10 are shown to be connected via bus 1002. In at least one embodiment, bus 1002 can include, without limitation, a CAN data interface (alternatively referred to herein as a “CAN bus”). In at least one embodiment, CAN can be a network within vehicle 1000 that is used to help control various features and functions of vehicle 1000, such as actuation of brakes, acceleration, braking, steering, wipers, etc. In one embodiment, bus 1002 can be configured to have tens or even hundreds of nodes, each with its own unique identifier (e.g., CAN ID). In at least one embodiment, bus 1002 can be read to find steering wheel angle, ground speed, engine revolutions per minute (“RPM”), button positions, and / or other vehicle status indicators. In at least one embodiment, bus 1002 can be an ASIL B compliant CAN bus.

[0167] In at least one embodiment, in addition to or instead of CAN, FlexRay and / or Ethernet protocols can be used. In at least one embodiment, there can be any number of buses 1002, which can include, without limitation, zero or more CAN buses, zero or more FlexRay buses, zero or more Ethernet buses, and / or zero or more other types of buses that use other protocols. In at least one embodiment, two or more buses can be used to perform different functions, and / or can be used for redundancy. For example, a first bus can be used for collision avoidance functions, and a second bus can be used for actuation control. In at least one embodiment, each bus of buses 1002 can communicate with any component of vehicle 1000, and two or more of buses 1002 can communicate with respective components. In at least one embodiment, each of any number of system on chips (“SoCs”) 1004 (e.g., SoC 1004(A) and SoC 1004(B)), each of one or more controllers 1036, and / or each computer within a vehicle can have access to the same input data (e.g., inputs from sensors of vehicle 1000), and can be connected to a common bus, such as a CAN bus.

[0168] In at least one embodiment, vehicle 1000 can include one or more controllers 1036, such as described herein with respect to FIG. 9, that are configured to communicate with each other and / or with other components of vehicle 1000 via bus 1002. In at least one embodiment, one or more controllers 1036 can be configured to communicate with each other and / or with other components of vehicle 1000 via bus 1002. Figure 10AThose described. In at least one embodiment, controller 1036 can be used for a variety of functions. In at least one embodiment, controller 1036 can be coupled to any of various other components and systems of vehicle 1000, and can be used to control vehicle 1000, artificial intelligence of vehicle 1000, infotainment of vehicle 1000, and / or other functions.

[0169] In at least one embodiment, vehicle 1000 can include any number of SoCs 1004. In at least one embodiment, each of SoCs 1004 can include, without limitation, central processing units (“one or more CPUs”) 1006, graphics processing units (“one or more GPUs”) 1008, one or more processors 1010, one or more caches 1012, one or more accelerators 1014, one or more data stores 1016, and / or other non- shown components and features. In at least one embodiment, one or more SoCs 1004 can be used to control vehicle 1000 in various platforms and systems. For example, in at least one embodiment, one or more SoCs 1004 can be combined with a high-definition (“HD”) map 1022 in a system (e.g., a system of vehicle 1000) that can obtain map refreshes and / or updates from one or more servers (not shown in FIG. 10) via network interface 1024. Figure 10C

[0170] In at least one embodiment, one or more CPUs 1006 can include a CPU cluster or CPU complex (alternatively referred to herein as a “CCPLEX”). In at least one embodiment, one or more CPUs 1006 can include multiple cores and / or level two (“L2”) caches. For example, in at least one embodiment, one or more CPUs 1006 can include eight cores in a multiprocessor configuration coupled to one another. In at least one embodiment, one or more CPUs 1006 can include four dual-core clusters with each cluster having a dedicated L2 cache (e.g., a 2 MB L2 cache). In at least one embodiment, one or more CPUs 1006 (e.g., a CCPLEX) can be configured to support simultaneous cluster operation, such that any combination of clusters of one or more CPUs 1006 can be active at any given time.

[0171] ​In at least one embodiment, one or more CPU(s) 1006 can implement power management features including, but not limited to, one or more of the following: individual hardware blocks can be automatically clock-gated at idle to conserve dynamic power; each core can be clock-gated when it is not actively executing instructions due to execution of Wait for Interrupt (“WFI”) / Wait for Event (“WFE”) instructions; each core can be independently powered; each cluster of cores can be independently clock-gated when all cores in the cluster are clock-gated or power-gated; and / or each cluster of cores can be independently power-gated when all cores in the cluster are power-gated. In at least one embodiment, one or more CPU(s) 1006 can further implement enhanced algorithms for managing power states where allowed power states and expected wake-up times are specified and hardware / microcode determines optimal power states for core, cluster, and CCPLEX inputs. In at least one embodiment, processing cores can support a simplified power state input sequence in software where work is offloaded to microcode.

[0172] In at least one embodiment, one or more GPU(s) 1008 can include an integrated GPU (also referred to herein as an “iGPU”). In at least one embodiment, one or more GPU(s) 1008 can be programmable and efficient for parallel workloads. In at least one embodiment, one or more GPU(s) 1008 can use enhanced tensor sets of instructions. In at least one embodiment, one or more GPU(s) 1008 can include one or more streaming microprocessors, where each streaming microprocessor can include a level one (“Ll”) cache (e.g., with at least 96 KB of storage capacity), and two or more streaming microprocessors can share an L2 cache (e.g., with 512 KB storage capacity). In at least one embodiment, one or more GPU(s) 1008 can include at least eight streaming microprocessors. In at least one embodiment, one or more GPU(s) 1008 can use a compute Application Programing Interface (API). In at least one embodiment, one or more GPU(s) 1008 can use one or more parallel computing platforms and / or programming models (e.g., NVIDIA’s CUDA model).

[0173] In at least one embodiment, one or more GPU(s) 1008 can be power-optimized to achieve best performance in automotive and embedded use cases. For example, in one embodiment, one or more GPU(s) 1008 can be fabricated on Fin Field Effect Transistor (“FinFET”) circuitry. In at least one embodiment, each streaming microprocessor can include a number of mixed-precision processing cores divided into a number of blocks. For example, and without limitation, 64 PF32 cores and 32 PF64 cores can be divided into four processing blocks. In at least one embodiment, each processing block can be allocated 16 FP32 cores, 8 FP64 cores, 16 INT32 cores, two mixed-precision NVIDIA Tensor Cores for deep learning matrix arithmetic, a Level-Zero (“L0”) instruction cache, a thread bundle scheduler, a dispatch unit, and / or a 64 KB register file. In at least one embodiment, a streaming microprocessor can include separate parallel integer and floating point datapaths to provide efficient execution of compute and address bound workloads. In at least one embodiment, a streaming microprocessor can include independent thread scheduling capabilities to enable finer grain synchronization and cooperation between parallel threads. In at least one embodiment, a streaming microprocessor can include a combined LI data cache and shared memory unit to improve performance while simplifying programming.

[0174] In at least one embodiment, one or more GPU(s) 1008 can include a high bandwidth memory (“HBM”) and / or a 16 GB HBM2 memory subsystem to provide, in some examples, a peak memory bandwidth of about 900 GB / s. In at least one embodiment, in addition to, or instead of, HBM memory, a synchronous graphics random access memory (“SGRAM”) can be used, such as a Graphics Double Data Rate type five synchronous random-access memory (“GDDR5”).

[0175] In at least one embodiment, GPU(s) 1008 can include unified memory technology. In at least one embodiment, address translation services (“ATS”) support can be used to allow GPU(s) 1008 to directly access CPU(s) 1006 page tables. In at least one embodiment, when a memory management unit (“MMU”) of a GPU of GPU(s) 1008 experiences a miss, an address translation request can be sent to CPU(s) 1006. In response, in at least one embodiment, a CPU of CPU(s) 1006 can look up a virtual-physical mapping for an address in its page tables and transmit a translation back to GPU(s) 1008. In at least one embodiment, unified memory technology can allow a single unified virtual address space to be used for memory of both CPU(s) 1006 and GPU(s) 1008, simplifying programming of GPU(s) 1008 and porting of applications to GPU(s) 1008.

[0176] In at least one embodiment, GPU(s) 1008 can include any number of access counters that can track how frequently GPU(s) 1008 are accessing memory of other processors. In at least one embodiment, access counter(s) can help ensure that memory pages are moved into physical memory of a processor that most frequently accesses the page, improving efficiency of memory ranges shared between processors.

[0177] In at least one embodiment, SoC(s) 1004 can include any number of caches 1012, including those described herein. For example, in at least one embodiment, cache(s) 1012 can include a level three (“L3”) cache that can be used by CPU(s) 1006 and GPU(s) 1008 (e.g., connected to CPU(s) 1006 and GPU(s) 1008). In at least one embodiment, cache(s) 1012 can include a write-back cache that can track state of lines, for example, by using a cache coherence protocol (e.g., MESI, MSI, etc.). In at least one embodiment, although smaller cache sizes can be used, L3 cache can include 4 MB of memory or more, according to embodiments.

[0178] In at least one embodiment, one or more SoC(s) 1004 can include one or more accelerator(s) 1014 (e.g., hardware accelerators, software accelerators, or a combination thereof). In at least one embodiment, one or more SoC(s) 1004 can include a hardware acceleration cluster that can include optimized hardware accelerators and / or large on-chip memory. In at least one embodiment, large on-chip memory (e.g., 4 MB of SRAM) can enable hardware acceleration cluster to accelerate neural networks and other computations. In at least one embodiment, hardware acceleration cluster can be used to supplement one or more GPU(s) 1008 and offload some tasks of one or more GPU(s) 1008 (e.g., freeing up more cycles of one or more GPU(s) 1008 to perform other tasks). In at least one embodiment, one or more accelerator(s) 1014 can be used for targeted workloads (e.g., perception, convolutional neural networks (“CNNs”), recurrent neural networks (“RNNs”), etc.) that are stable enough to be amenable to acceleration. In at least one embodiment, CNNs can include region-based or region with convolutional neural networks (“RCNNs”) and fast RCNNs (e.g., as used for object detection) or other types of CNNs.

[0179] In at least one embodiment, one or more accelerators 1014 (e.g., hardware acceleration clusters) can include one or more deep learning accelerators (“DLAs”). In at least one embodiment, one or more DLAs can include, without limitation, one or more Tensor Processing Units (“TPUs”) that can be configured to provide an additional 100 trillion operations per second for deep learning applications and inferencing. In at least one embodiment, a TPU can be an accelerator configured and optimized for performing image processing functions (e.g., for CNNs, RCNNs, etc.). In at least one embodiment, one or more DLAs can be further optimized for a particular set of neural network types and floating point operations and inferencing. In at least one embodiment, design of one or more DLAs can provide higher performance per mm than a typical general purpose GPU, and often significantly outperform CPUs. In at least one embodiment, one or more TPUs can perform several functions including support for INT8, INT16, and FP16 data types for features and weights, single instance convolution functionality, and post-processor functionality, for example. In at least one embodiment, one or more DLAs can quickly and efficiently execute neural networks, especially CNNs, on processed or unprocessed data for any of a variety of functions including, for example and without limitation: CNNs for object recognition and detection using data from camera sensors; CNNs for distance estimation using data from camera sensors; CNNs for emergency vehicle detection, as well as recognition and detection, using data from microphones; CNNs for facial recognition and vehicle owner identification using data from camera sensors; and / or CNNs for safety and / or safety related events.

[0180] In at least one embodiment, a DLA can perform any of functions of GPU(s) 1008, and by using an inferencing accelerator, for example, a designer can target one or more DLAs or GPU(s) 1008 for any function. For example, in at least one embodiment, a designer can concentrate processing and floating point operations for CNNs on one or more DLAs, and leave other functions to GPU(s) 1008 and / or accelerator(s) 1014.

[0181] In at least one embodiment, one or more accelerators 1014 can include a programmable vision accelerator (“PVA”), which can be alternatively referred to herein as a computer vision accelerator. In at least one embodiment, one or more PVAs can be designed and configured to accelerate computer vision algorithms used for advanced driver assistance systems (“ADAS”) 1038, autonomous driving, augmented reality (“AR”) applications, and / or virtual reality (“VR”) applications. In at least one embodiment, one or more PVAs can strike a balance between performance and flexibility. For example, in at least one embodiment, each of one or more PVAs can include, for example and without limitation, any number of reduced instruction set computer (“RISC”) cores, direct memory access (“DMA”), and / or any number of vector processors.

[0182] In at least one embodiment, RISC cores can interact with image sensors (e.g., image sensors of any of cameras described herein), image signal processors, etc. In at least one embodiment, each RISC core can include any number of memories. In at least one embodiment, RISC cores can use any of a number of protocols, depending on embodiment. In at least one embodiment, RISC cores can execute a real-time operating system (“RTOS”). In at least one embodiment, RISC cores can be implemented using one or more integrated circuit devices, application specific integrated circuits (“ASICs”), and / or memory devices. For example, in at least one embodiment, RISC cores can include instruction caches and / or tightly coupled RAM.

[0183] In at least one embodiment, DMA can enable components of a PVA to access system memory independently of one or more CPUs 1006. In at least one embodiment, DMA can support any number of features for providing optimizations to a PVA, including but not limited to, support for multi-dimensional addressing and / or circular addressing. In at least one embodiment, DMA can support up to six or more dimensions of addressing, which can include, but are not limited to, block width, block height, block depth, horizontal block stride, vertical block stride, and / or depth stride.

[0184] In at least one embodiment, vector processors can be programmable processors that can be designed to efficiently and flexibly execute programming for computer vision algorithms and provide signal processing capabilities. In at least one embodiment, a PVA can include a PVA core and two vector processing subsystem partitions. In at least one embodiment, a PVA core can include a processor subsystem, DMA engines (e.g., two DMA engines), and / or other peripherals. In at least one embodiment, a vector processing subsystem can function as a primary processing engine for a PVA and can include a vector processing unit (“VPU”), an instruction cache, and / or a vector memory (e.g., “VMEM”). In at least one embodiment, a VPU core can include a digital signal processor, such as a single instruction multiple data (“SIMD”), very long instruction word (“VLIW”) digital signal processor. In at least one embodiment, a combination of SIMD and VLIW can improve throughput and speed.

[0185] In at least one embodiment, each vector processor can include an instruction cache and can be coupled to a dedicated memory. As a result, in at least one embodiment, each vector processor can be configured to execute independently of other vector processors. In at least one embodiment, vector processors included in a particular PVA can be configured to employ data parallelism. For example, in at least one embodiment, multiple vector processors included in a single PVA can execute a general purpose computer vision algorithm, except on different regions of an image. In at least one embodiment, vector processors included in a particular PVA can execute different computer vision algorithms on one image at a time, or even different algorithms on a sequence of images or portions of images. In at least one embodiment, any number of PVAs can be included in a hardware acceleration cluster, and any number of vector processors can be included in each PVA, among other things. In at least one embodiment, a PVA can include additional error correcting code (“ECC”) memory to enhance overall system security.

[0186] In at least one embodiment, one or more accelerators 1014 can include on-chip computer vision networks and static random access memory (“SRAM”) for providing high bandwidth, low latency SRAM for one or more accelerators 1014. In at least one embodiment, on-chip memory can include at least 4 MB of SRAM that includes, for example and without limitation, eight field-programmable memory blocks that are accessible by both PVA and DLA. In at least one embodiment, each pair of memory blocks can include an advanced peripheral bus (“APB”) interface, configuration circuitry, a controller, and a multiplexer. In at least one embodiment, any type of memory can be used. In at least one embodiment, PVA and DLA can access memory via a backbone that provides PVA and DLA with high-speed access to memory. In at least one embodiment, a backbone can include on-chip computer vision networks that interconnect PVA and DLA to memory (e.g., using APB).

[0187] In at least one embodiment, on-chip computer vision networks can include an interface that determines that both PVA and DLA provide ready and valid signals before transmitting any control signals / addresses / data. In at least one embodiment, an interface can provide separate phases and separate channels for transmitting control signals / addresses / data, as well as burst-type communication for continuous data transmission. In at least one embodiment, although other standards and protocols can be used, an interface can comply with International Organization for Standardization (“ISO”) 26262 or International Electrotechnical Commission (“IEC”) 61508 standards.

[0188] In at least one embodiment, one or more SoC 1004 can include real-time line-of-sight tracking hardware accelerators. In at least one embodiment, real-time line-of-sight tracking hardware accelerators can be used to quickly and efficiently determine locations and ranges of objects (e.g., within a world model) to generate real-time visualizations simulations for RADAR signal interpretation, for sound propagation synthesis and / or analysis, for simulations of SONAR systems, for general wave propagation simulations, for comparison with LIDAR data for positioning and / or other functions, and / or for other uses.

[0189] In at least one embodiment, one or more accelerators 1014 have broad use for autonomous driving. In at least one embodiment, PVAs can be used for key processing stages in ADAS and autonomous vehicles. In at least one embodiment, capabilities of PVAs at low power and low latency match well with algorithmic domains that require predictable processing. In other words, PVAs excel at semi-dense or dense regular computations, even on small data sets that can require predictable runtimes with low latency and low power. In at least one embodiment, PVAs can be designed to run classic computer vision algorithms, such as in vehicle 1000, as they can be efficient at object detection and integer math operations.

[0190] For example, in accordance with at least one embodiment of technology, a PVA is used to perform computer stereo vision. In at least one embodiment, a semi-global matching based algorithm can be used in some examples, although this is not meant to be limiting. In at least one embodiment, applications for level 3-5 autonomous driving use dynamic estimation / stereo matching in run-time (e.g., structure from motion, pedestrian recognition, lane detection, etc.). In at least one embodiment, a PVA can perform computer stereo vision functions on inputs from two monocular cameras.

[0191] In at least one embodiment, a PVA can be used to perform dense optical flow. For example, in at least one embodiment, a PVA can process raw RADAR data (e.g., using a 4D fast Fourier transform) to provide processed RADAR data. In at least one embodiment, a PVA is used for time-of-flight depth processing, e.g., by processing raw time-of-flight data to provide processed time-of-flight data.

[0192] In at least one embodiment, DLA can be used to run any type of network to enhance control and driving safety, including, for example and without limitation, a neural network that outputs a confidence level for each object detection. In at least one embodiment, confidence level can be represented or interpreted as a probability, or as providing a relative “weight” of each detection relative to other detections. In at least one embodiment, a confidence measurement enables system to make further decisions as to which detections should be considered as true positive detections and not false positive detections. In at least one embodiment, system can set a threshold for confidence level, and only consider detections that exceed threshold as true positive detections. In embodiments using automatic emergency braking (“AEB”) systems, false positive detections would result in vehicle automatically performing emergency braking, which is obviously undesirable. In at least one embodiment, highly confident detections can be considered as triggers for AEB. In at least one embodiment, DLA can run a neural network for regression of confidence values. In at least one embodiment, neural network can take as its input at least some subset of parameters, such as bounding box size, ground plane estimates obtained (e.g., from another subsystem), outputs of one or more IMU sensors 1066 related to object’s vehicle 1000 direction, distance, 3D position estimates obtained from neural network and / or other sensors (e.g., one or more LIDAR sensors 1064 or one or more RADAR sensors 1060), etc.

[0193] In at least one embodiment, one or more SoC(s) 1004 can include one or more data storage(s) 1016 (e.g., memory). In at least one embodiment, one or more data storage(s) 1016 can be on-chip memory of one or more SoC(s) 1004 that can store neural networks to be executed on one or more GPU(s) 1008 and / or DLA. In at least one embodiment, one or more data storage(s) 1016 can have sufficient capacity to store multiple instances of a neural network for redundancy and safety. In at least one embodiment, one or more data storage(s) 1016 can include L2 or L3 cache.

[0194] In at least one embodiment, one or more SoC(s) 1004 can include any number of processor(s) 1010 (e.g., embedded processors). In at least one embodiment, one or more processor(s) 1010 can include a boot and power management processor that can be a dedicated processor and subsystem to handle boot power and management functions and related security enforcement. In at least one embodiment, a boot and power management processor can be part of a one or more SoC(s) 1004 boot sequence and can provide run-time power management services. In at least one embodiment, a boot power and management processor can provide clock and voltage programming, assist system low power state transitions, one or more SoC(s) 1004 thermal and temperature sensor management, and / or one or more SoC(s) 1004 power state management. In at least one embodiment, each temperature sensor can be implemented as a ring oscillator whose output frequency is proportional to temperature, and one or more SoC(s) 1004 can use ring oscillators to detect temperature of one or more CPU(s) 1006, one or more GPU(s) 1008, and / or one or more accelerator(s) 1014. In at least one embodiment, if a temperature is determined to exceed a threshold, a boot and power management processor can enter a temperature fault routine and put one or more SoC(s) 1004 into a lower power state and / or put vehicle 1000 into a safe park pattern for the driver (e.g., cause vehicle 1000 to safely park).

[0195] In at least one embodiment, one or more processor(s) 1010 can also include a set of embedded processors that can function as an audio processing engine that can be an audio subsystem that can provide full hardware support for multi-channel audio to hardware through a number of interfaces as well as a broad and flexible range of audio I / O interfaces. In at least one embodiment, an audio processing engine is a dedicated processor core with a digital signal processor with dedicated RAM.

[0196] In at least one embodiment, one or more processor(s) 1010 can also include an always-on processor engine that can provide necessary hardware features to support low-power sensor management and wake-up use cases. In at least one embodiment, a processor on an always-on processor engine can include, but is not limited to, a processor core, tightly coupled RAM, support peripherals (e.g., timers and interrupt controllers), various I / O controller peripherals, and routing logic.

[0197] In at least one embodiment, one or more processors 1010 can further include a safety cluster engine that includes, without limitation, a dedicated processor subsystem for handling safety management for automotive applications. In at least one embodiment, safety cluster engine can include, without limitation, two or more processor cores, tightly coupled RAM, supporting peripherals (e.g., timers, interrupt controllers, etc.), and / or routing logic. In a safety mode, in at least one embodiment, two or more cores can operate in a lockstep mode and can function as a single core with comparison logic to detect any differences between their operations. In at least one embodiment, one or more processors 1010 can further include a real-time camera engine that can include, without limitation, a dedicated processor subsystem for handling real-time camera management. In at least one embodiment, one or more processors 1010 can further include a high dynamic range signal processor that can include, without limitation, an image signal processor that is a hardware engine that is part of a camera processing pipeline.

[0198] In at least one embodiment, one or more processors 1010 can include a video image compositor that can be a processing block (e.g., implemented on a microprocessor) that implements video post-processing functions needed by a video playback application to produce a final video to produce a final image for a player window. In at least one embodiment, video image compositor can perform lens distortion correction on one or more wide-view cameras 1070, one or more surround cameras 1074, and / or one or more in-cabin monitoring camera sensors. In at least one embodiment, preferably, in-cabin monitoring camera sensors are monitored by a neural network running on another instance of SoC 1004 that is configured to identify cabin events and respond accordingly. In at least one embodiment, in-cabin systems can perform, without limitation, lip reading to activate cellular service and place a phone call, dictate an email, change a destination of a vehicle, activate or change a vehicle’s infotainment system and settings, or provide voice-activated web surfing. In at least one embodiment, certain functionality is available to a driver when a vehicle is operating in an autonomous mode, otherwise it is disabled.

[0199] In at least one embodiment, video image compositor can include enhanced temporal noise reduction for simultaneous spatial and temporal noise reduction. For example, in at least one embodiment, where motion occurs in a video, noise reduction appropriately weights spatial information, reducing a weight of information provided by adjacent frames. In at least one embodiment, where an image or portion of an image does not include motion, temporal noise reduction performed by video image compositor can use information from a previous image to reduce noise in a current image.

[0200] In at least one embodiment, video image compositor can also be configured to perform stereo correction on input stereoscopic lens frames. In at least one embodiment, when using an operating system desktop, video image compositor can also be used for user interface composition and one or more GPUs 1008 are not required to continuously render new surfaces. In at least one embodiment, when one or more GPUs 1008 are powered and active for 3D rendering, video image compositor can be used to offload one or more GPUs 1008 to improve performance and responsiveness.

[0201] In at least one embodiment, one or more SoCs in SoC(s) 1004 can also include mobile industry processor interface (“MIPI”) camera serial interfaces for receiving video and input from cameras, high-speed interfaces, and / or video input blocks that can be used for camera and related pixel input functionality. In at least one embodiment, one or more SoCs 1004 can also include an input / output controller that can be controlled by software and can be used to receive I / O signals that are not committed to a specific role.

[0202] In at least one embodiment, one or more SoCs in SoC(s) 1004 can also include a wide range of peripheral interfaces to enable communication with peripherals, audio encoders / decoders (“codecs”), power management, and / or other devices. In at least one embodiment, one or more SoCs 1004 can be used to process data from cameras (e.g., connected over Gigabit Multimedia Serial Link and Ethernet channels), sensors (e.g., one or more LIDAR sensors 1064, one or more RADAR sensors 1060, etc., which can be connected over Ethernet channels), data from bus 1002 (e.g., speed of vehicle 1000, steering wheel position, etc.), data from one or more GNSS sensors 1058 (e.g., connected over Ethernet bus or CAN bus), etc. In at least one embodiment, one or more SoCs in SoC(s) 1004 can also include dedicated high-performance mass storage controllers that can include their own DMA engines and can be used to free one or more CPUs 1006 from regular data management tasks.

[0203] In at least one embodiment, SoC(s) 1004 can be an end-to-end platform with a flexible architecture that spans automation levels 3-5, providing a comprehensive functional safety architecture that leverages and efficiently uses computer vision and ADAS technology for diversity and redundancy, which provides a platform that can provide a flexible, reliable driving software stack, as well as deep learning tools. In at least one embodiment, SoC(s) 1004 can be faster, more reliable, and even more energy and spatial efficient than conventional systems. For example, in at least one embodiment, accelerator(s) 1014, when combined with CPU(s) 1006, GPU(s) 1008, and data storage device(s) 1016, can provide a fast, efficient platform for level 3-5 autonomous vehicles.

[0204] In at least one embodiment, computer vision algorithms can be executed on CPUs, which can be configured using high-level programming languages (e.g., C) to perform a variety of processing algorithms on a variety of visual data. However, in at least one embodiment, CPUs often cannot meet performance requirements of many computer vision applications, such as performance requirements related to execution time and power consumption. In at least one embodiment, many CPUs cannot execute complex object detection algorithms in real-time, which are used in on-board ADAS applications and actual level 3-5 autonomous vehicles.

[0205] Embodiments described herein allow for simultaneous and / or sequential execution of multiple neural networks, and allow for combining results together to enable level 3-5 autonomous driving functionality. For example, in at least one embodiment, CNNs executed on DLAs or discrete GPUs (e.g., GPU(s) 1020) can include text and word recognition, allowing a supercomputer to read and understand traffic signs, including signs that a neural network has not been specifically trained for. In at least one embodiment, DLAs can also include neural networks capable of recognizing, interpreting, and providing semantic understanding of symbols, and passing that semantic understanding to a path planning module running on a CPU Complex.

[0206] In at least one embodiment, multiple neural networks can be run simultaneously for a level 3, 4, or 5 drive. For example, in at least one embodiment, a warning sign consisting of a “Caution: flashing lights indicate icy conditions” sign with flashing lights can be interpreted independently or collectively by multiple neural networks. In at least one embodiment, the warning sign itself can be recognized as a traffic sign by a first deployed neural network (e.g., a neural network that has already been trained), the text “flashing lights indicate icy conditions” can be interpreted by a second deployed neural network that informs vehicle’s path planning software (preferably executing on a CPU Complex) that icy conditions exist when flashing lights are detected. In at least one embodiment, flashing lights can be recognized by a third deployed neural network operating over multiple frames, informing vehicle’s path planning software that flashing lights exist (or do not exist). In at least one embodiment, all three neural networks can be run simultaneously, for example within a DLA and / or on one or more GPU(s) 1008.

[0207] In at least one embodiment, a CNN for facial recognition and vehicle owner identification can use data from a camera sensor to identify presence of an authorized driver and / or owner of vehicle 1000. In at least one embodiment, when an owner approaches a driver door and opens a light, a normally open sensor processor engine can be used to unlock the vehicle, and, in a safe mode, when the owner leaves the vehicle, can be used to disable the vehicle. In this way, one or more SoC(s) 1004 provide safeguards against theft and / or carjacking.

[0208] In at least one embodiment, a CNN for emergency vehicle detection and identification can use data from microphones 1096 to detect and identify emergency vehicle sirens. In at least one embodiment, one or more SoCs 1004 use a CNN to classify ambient and urban sounds, as well as to classify visual data. In at least one embodiment, a CNN running on a DLA is trained to identify relative proximity of emergency vehicles (e.g., by using Doppler effect). In at least one embodiment, a CNN can also be trained to identify emergency vehicles for regions in which a vehicle is operating, as identified by one or more GNSS sensors 1058. In at least one embodiment, when operating in Europe, a CNN will seek to detect European sirens, while in North America, a CNN will seek to identify only North American sirens. In at least one embodiment, once an emergency vehicle is detected, a control program can be used to execute emergency vehicle safety routines, slow vehicle down, pull vehicle to side of road, stop, and / or idle vehicle until emergency vehicle passes, with assistance of one or more ultrasonic sensors 1062.

[0209] In at least one embodiment, vehicle 1000 can include one or more CPUs 1018 (e.g., one or more discrete CPUs or one or more dCPUs) that can be coupled to one or more SoCs 1004 via a high-speed interconnect (e.g., PCIe). In at least one embodiment, one or more CPUs 1018 can include an X86 processor, such as one or more CPUs 1018 can be used to perform any of a variety of functions, such as including arbitrating inconsistent results between ADAS sensors and one or more SoCs 1004, and / or one or more supervisory controllers 1036 state and health and / or an on-chip information system (“information SoC”) 1030.

[0210] In at least one embodiment, vehicle 1000 can include one or more GPUs 1020 (e.g., one or more discrete GPUs or one or more dGPUs) that can be coupled to one or more SoCs 1004 via a high-speed interconnect (e.g., NVIDIA’s NVLINK channel). In at least one embodiment, one or more GPUs 1020 can provide additional artificial intelligence functionality, such as by executing redundant and / or different neural networks, and can be used to train and / or update neural networks based at least in part on input (e.g., sensor data) from sensors of vehicle 1000.

[0211] In at least one embodiment, vehicle 1000 can also include network interface 1024, which can include, without limitation, one or more wireless antennas 1026 (e.g., one or more wireless antennas for different communication protocols, such as cellular antennas, Bluetooth antennas, etc.). In at least one embodiment, network interface 1024 can be used to enable wireless connectivity through Internet cloud services (e.g., with servers and / or other network equipment) with other vehicles and / or computing devices (e.g., client devices of passengers). In at least one embodiment, to communicate with other vehicles, a direct link can be established between vehicle 1000 and another vehicle and / or an indirect link can be established (e.g., through a network and the Internet). In at least one embodiment, a vehicle-to-vehicle communication link can be used to provide a direct link. In at least one embodiment, a vehicle-to-vehicle communication link can provide vehicle 1000 with information about vehicles in a vicinity of vehicle 1000 (e.g., vehicles in front of, to the side of, and / or behind vehicle 1000). In at least one embodiment, this aforementioned functionality can be part of a cooperative adaptive cruise control functionality of vehicle 1000.

[0212] In at least one embodiment, network interface 1024 can include a SoC that provides modulation and demodulation functionality and enables one or more controllers 1036 to communicate over wireless networks. In at least one embodiment, network interface 1024 can include a radio frequency front end for upconversion from baseband to radio frequency and downconversion from radio frequency to baseband. In at least one embodiment, frequency conversion can be performed in any technically feasible way. For example, frequency conversion can be performed through well-known processes and / or using a superheterodyne process. In at least one embodiment, radio frequency front end functionality can be provided by a separate chip. In at least one embodiment, a network interface can include wireless functionality to communicate over LTE, WCDMA, UMTS, GSM, CDMA2000, Bluetooth, Bluetooth LE, Wi-Fi, Z-Wave, Zigbee, LoRaWAN, and / or other wireless protocols.

[0213] In at least one embodiment, vehicle 1000 can also include one or more data stores 1028, which can include, without limitation, off-chip (e.g., of SoC(s) 1004) storage. In at least one embodiment, one or more data stores 1028 can include, without limitation, one or more storage elements, including RAM, SRAM, dynamic random access memory (“DRAM”), video random access memory (“VRAM”), flash memory, hard disks, and / or other components and / or devices that can store at least one bit of data.

[0214] In at least one embodiment, vehicle 1000 can also include one or more GNSS sensors 1058 (e.g., GPS and / or assisted GPS sensors) to assist in mapping, perception, occupancy grid generation, and / or path planning functions. In at least one embodiment, any number of GNSS sensors 1058 can be used including, for example and without limitation, a GPS using a USB connector with Ethernet connected to a serial interface (e.g., RS-232) bridge.

[0215] In at least one embodiment, vehicle 1000 can also include one or more RADAR sensors 1060. In at least one embodiment, one or more RADAR sensors 1060 can be used by vehicle 1000 for long range vehicle detection, even in darkness and / or adverse weather conditions. In at least one embodiment, a RADAR functional safety level can be ASIL B. In at least one embodiment, one or more RADAR sensors 1060 can use CAN bus and / or bus 1002 (e.g., to transmit data generated by one or more RADAR sensors 1060) for control and access to object tracking data, in some examples can access an Ethernet channel for access to raw data. In at least one embodiment, a wide variety of RADAR sensor types can be used. For example and without limitation, one or more of RADAR sensors 1060 can be suitable for front, rear, and side RADAR use. In at least one embodiment, one or more RADAR sensors 1060 are pulse Doppler RADAR sensors.

[0216] In at least one embodiment, RADAR sensor(s) 1060 can include different configurations, such as long-range with narrow field of view, short-range with wide field of view, short-range side coverage, etc. In at least one embodiment, long-range RADAR can be used for adaptive cruise control functionality. In at least one embodiment, long-range RADAR systems can provide a wide field of view achieved through two or more independent scans (e.g., within 250 m range). In at least one embodiment, RADAR sensor(s) 1060 can help distinguish between static and moving objects and can be used by ADAS system 1038 for emergency brake assist and forward collision warning. In at least one embodiment, sensor(s) 1060 included in a long-range RADAR system can include, without limitation, a monostatic multi-mode RADAR with multiple (e.g., six or more) fixed RADAR antennas and a high-speed CAN and FlexRay interface. In at least one embodiment, with six antennas, a central four antennas can create a focused beam pattern designed to record the environment around vehicle 1000 at higher speeds with minimal traffic interference from adjacent lanes. In at least one embodiment, other two antennas can expand the field of view so that vehicles 1000 entering or leaving a lane can be quickly detected.

[0217] In at least one embodiment, as an example, a mid-range RADAR system can include, for example, a range of up to 160 m (front) or 80 m (rear), and a field of view of up to 42 degrees (front) or 150 degrees (rear). In at least one embodiment, a short-range RADAR system can include, without limitation, any number of RADAR sensors 1060 designed to be mounted at either end of a rear bumper. When mounted at either end of a rear bumper, in at least one embodiment, a RADAR sensor system can produce two beams that constantly monitor the vehicle’s rearward direction and a blind spot close by. In at least one embodiment, a short-range RADAR system can be used in ADAS system 1038 for blind spot detection and / or lane change assist.

[0218] In at least one embodiment, vehicle 1000 can also include ultrasonic sensor(s) 1062. In at least one embodiment, ultrasonic sensor(s) 1062, which can be positioned in front, rear, and / or side locations of vehicle 1000, can be used for parking assist and / or to create and update occupancy grids. In at least one embodiment, a wide variety of ultrasonic sensors 1062 can be used, and different ultrasonic sensors 1062 can be used for different detection ranges (e.g., 2.5 m, 4 m). In at least one embodiment, ultrasonic sensors 1062 can operate at a functional safety level of ASIL B.

[0219] In at least one embodiment, vehicle 1000 can include one or more LIDAR sensors 1064. In at least one embodiment, one or more LIDAR sensors 1064 can be used for object and pedestrian detection, emergency braking, collision avoidance, and / or other functions. In at least one embodiment, one or more LIDAR sensors 1064 can operate at a functional safety level of ASIL B. In at least one embodiment, vehicle 1000 can include multiple (e.g., two, four, six, etc.) LIDAR sensors 1064 that can use Ethernet channels (e.g., provide data to a Gigabit Ethernet switch).

[0220] In at least one embodiment, one or more LIDAR sensors 1064 can be capable of providing a list of objects and their distances for a 360 degree field of view. In at least one embodiment, one or more LIDAR sensors 1064 that are commercially available can have an advertised range of approximately 100 m, have an accuracy of 2 cm - 3 cm, and support a 100 Mbps Ethernet connection, for example. In at least one embodiment, one or more non-protruding LIDAR sensors can be used. In such embodiments, one or more LIDAR sensors 1064 can include small devices that can be embedded into front, rear, side, and / or corner locations of vehicle 1000. In at least one embodiment, one or more LIDAR sensors 1064, in such embodiments, can provide a horizontal field of view of up to 120 degrees and a vertical field of view of 35 degrees, with a range of 200 m, even for low reflectivity objects. In at least one embodiment, a forward-facing one or more LIDAR sensors 1064 can be configured for a horizontal field of view between 45 degrees and 105 degrees.

[0221] In at least one embodiment, LIDAR technology such as 3D Flash LIDAR can also be used. In at least one embodiment, 3D Flash LIDAR uses a laser flash as a transmission source to illuminate approximately 200 m around vehicle 1000. In at least one embodiment, a flash LIDAR unit includes, without limitation, a receiver that records laser pulse travel time and reflected light on each pixel, which in turn corresponds to a range from vehicle 1000 to an object. In at least one embodiment, flash LIDAR can allow for highly accurate and distortion-free images of surrounding environment to be generated with each laser flash. In at least one embodiment, four flash LIDAR sensors can be deployed, one on each side of vehicle 1000. In at least one embodiment, a 3D flash LIDAR system includes, without limitation, a solid-state 3D line-of-sight array LIDAR camera with no moving parts other than a fan (e.g., a non-scanning LIDAR device). In at least one embodiment, a flash LIDAR device can use 5 nanosecond Class I (eye-safe) laser pulses per frame, and can capture reflected laser light as a 3D ranging point cloud and co-registered intensity data.

[0222] In at least one embodiment, vehicle 1000 can also include one or more IMU sensors 1066. In at least one embodiment, one or more IMU sensors 1066 can be located at a center of a rear axle of vehicle 1000. In at least one embodiment, one or more IMU sensors 1066 can include, for example and without limitation, one or more accelerometers, one or more magnetometers, one or more gyroscopes, one magnetic compass, multiple magnetic compasses, and / or other sensor types. In at least one embodiment, such as in a six-axis application, one or more IMU sensors 1066 can include, without limitation, an accelerometer and a gyroscope. In at least one embodiment, such as in a nine-axis application, one or more IMU sensors 1066 can include, without limitation, an accelerometer, a gyroscope, and a magnetometer.

[0223] In at least one embodiment, one or more IMU sensors 1066 can be implemented as a miniature, high-performance GPS-aided inertial navigation system (“GPS / INS”) that combines micro-electro-mechanical systems (“MEMS”) inertial sensors, high-sensitivity GPS receiver, and advanced Kalman filtering algorithms to provide estimates of position, velocity, and attitude; in at least one embodiment, one or more IMU sensors 1066 can enable vehicle 1000 to estimate heading without requiring input from a magnetic sensor by directly observing and correlating changes in velocity from GPS to one or more IMU sensors 1066. In at least one embodiment, one or more IMU sensors 1066 and one or more GNSS sensors 1058 can be combined in a single integrated unit.

[0224] In at least one embodiment, vehicle 1000 can include one or more microphones 1096 placed within and / or around vehicle 1000. In at least one embodiment, additionally, one or more microphones 1096 can be used for emergency vehicle detection and identification.

[0225] In at least one embodiment, vehicle 1000 can also include any number of camera types, including one or more stereo cameras 1068, one or more wide-view cameras 1070, one or more infrared cameras 1072, one or more surround-view cameras 1074, one or more long-range cameras 1098, one or more mid-range cameras 1076, and / or other camera types. In at least one embodiment, cameras can be used to capture image data around entire periphery of vehicle 1000. In at least one embodiment, type of cameras used depends on vehicle 1000. In at least one embodiment, any combination of camera types can be used to provide necessary coverage around vehicle 1000. In at least one embodiment, number of cameras deployed can vary from embodiment to embodiment. For example, in at least one embodiment, vehicle 1000 can include six cameras, seven cameras, ten cameras, twelve cameras, or other number of cameras. In at least one embodiment, cameras can support Gigabit Multimedia Serial Link (“GMSL”) and / or Gigabit Ethernet communications, by way of example and without limitation. In at least one embodiment, cameras can be implemented as described in greater detail herein previously with reference to FIG. 6A. Figure 10A and Figure 10B Each camera can be described in greater detail.

[0226] In at least one embodiment, vehicle 1000 can also include one or more vibration sensors 1042. In at least one embodiment, one or more vibration sensors 1042 can measure vibrations of components of vehicle 1000 (e.g., axles). For example, in at least one embodiment, changes in vibration can be indicative of changes in a road surface. In at least one embodiment, when two or more vibration sensors 1042 are used, differences between vibrations can be used to determine friction or slippage of a road surface (e.g., when there is a difference in vibration between a power driven axle and a freely rotating axle).

[0227] In at least one embodiment, vehicle 1000 can include an ADAS system 1038. In at least one embodiment, ADAS system 1038 can include, without limitation, a SoC. In at least one embodiment, ADAS system 1038 can include, without limitation, any number of adaptive / autonomous / automatic cruise control (“ACC”) systems, cooperative adaptive cruise control (“CACC”) systems, forward collision warning (“FCW”) systems, automatic emergency braking (“AEB”) systems, lane departure warning (“LDW”) systems, lane keep assist (“LKA”) systems, blind spot warning (“BSW”) systems, rear cross-traffic warning (“RCTW”) systems, collision warning (“CW”) systems, lane centering (“LC”) systems, and / or other systems, features, and / or functionality, and combinations thereof.

[0228] In at least one embodiment, ACC systems can use one or more RADAR sensors 1060, one or more LIDAR sensors 1064, and / or any number of cameras. In at least one embodiment, ACC systems can include longitudinal ACC systems and / or lateral ACC systems. In at least one embodiment, longitudinal ACC systems monitor and control distance to another vehicle immediately in front of vehicle 1000 and automatically adjust speed of vehicle 1000 to maintain a safe distance from the vehicle in front. In at least one embodiment, lateral ACC systems perform distance keeping and suggest lane changes for vehicle 1000 when needed. In at least one embodiment, lateral ACC is relevant to other ADAS applications, such as LC and CW.

[0229] In at least one embodiment, a CACC system uses information from other vehicles, which can be received from other vehicles via a wireless link or indirectly via a network connection (e.g., via the Internet) via network interface 1024 and / or one or more wireless antennas 1026. In at least one embodiment, a direct link can be provided by a vehicle-to-vehicle (“V2V”) communication link, while an indirect link can be provided by an infrastructure-to-vehicle (“I2V”) communication link. In general, V2V communications provide information about immediately preceding vehicles (e.g., vehicles immediately ahead of and in the same lane as vehicle 1000), while I2V communications provide information about traffic further ahead. In at least one embodiment, a CACC system can include one or both of I2V and V2V information sources. In at least one embodiment, a CACC system can be more reliable with information about vehicles ahead of vehicle 1000, and has potential to improve smoothness of traffic flow and reduce road congestion.

[0230] In at least one embodiment, an FCW system is designed to warn a driver of a hazard so that the driver can take corrective action. In at least one embodiment, an FCW system uses a forward-facing camera and / or one or more RADAR sensors 1060 coupled to a dedicated processor, DSP, FPGA, and / or ASIC that is electrically coupled to provide driver feedback such as a display, a speaker, and / or a vibrating component. In at least one embodiment, an FCW system can provide a warning, for example, in the form of a sound, a visual warning, a vibration, and / or a quick brake pulse.

[0231] In at least one embodiment, an AEB system detects an impending forward collision with another vehicle or other object and can automatically apply brakes if a driver does not take corrective action within specified time or distance parameters. In at least one embodiment, an AEB system can use one or more forward-facing cameras and / or one or more RADAR sensors 1060 coupled to a dedicated processor, DSP, FPGA, and / or ASIC. In at least one embodiment, when an AEB system detects a hazard, it typically first warns a driver to take corrective action to avoid a collision, and if that driver does not take corrective action, the AEB system can automatically apply brakes in an attempt to prevent or at least mitigate the effects of a predicted collision. In at least one embodiment, an AEB system can include techniques such as dynamic brake support and / or crash imminent braking.

[0232] In at least one embodiment, LDW system provides visual, audible, and / or tactile warnings, such as steering wheel or seat vibrations, when vehicle 1000 crosses lane markings to warn driver. In at least one embodiment, LDW system is not active when driver indicates intentional lane departure, such as by activating turn signals. In at least one embodiment, LDW system can use a front-facing camera coupled to a dedicated processor, DSP, FPGA, and / or ASIC that is electrically coupled to provide driver feedback such as a display, speakers, and / or vibrating components. In at least one embodiment, LKA system is a variation of LDW system. In at least one embodiment, if vehicle 1000 begins to deviate from a lane, LKA system provides steering input or braking to correct vehicle 1000.

[0233] In at least one embodiment, BSW system detects and warns vehicle drivers of vehicles in a car’s blind spot. In at least one embodiment, BSW system can provide visual, audible, and / or tactile alerts to indicate that merging or changing lanes is unsafe. In at least one embodiment, BSW system can provide additional warnings when a driver uses turn signals. In at least one embodiment, BSW system can use one or more rear-facing cameras and / or one or more RADAR sensors 1060 coupled to a dedicated processor, DSP, FPGA, and / or ASIC that is electrically coupled to driver feedback such as a display, speakers, and / or vibrating components.

[0234] In at least one embodiment, RCTW system can provide visual, audible, and / or tactile notifications when an object is detected outside of a rear camera range while vehicle 1000 is backing up. In at least one embodiment, RCTW system includes AEB system to ensure application of vehicle brakes to avoid a collision. In at least one embodiment, RCTW system can use one or more rear-facing RADAR sensors 1060 coupled to a dedicated processor, DSP, FPGA, and / or ASIC that is electrically coupled to provide driver feedback such as a display, speakers, and / or vibrating components.

[0235] In at least one embodiment, conventional ADAS systems can be prone to false positives, which can annoy and distract drivers, but are typically not catastrophic because conventional ADAS systems alert the driver and allow the driver to decide whether a safety condition is truly present and take appropriate action. In at least one embodiment, in the event of a result conflict, vehicle 1000 itself decides whether to follow the results of the primary computer or the secondary computer (e.g., a first controller or a second controller of controller 1036). For example, in at least one embodiment, ADAS system 1038 can be a backup and / or secondary computer for providing perception information to a backup computer plausibility module. In at least one embodiment, a backup computer plausibility monitor can run redundant varieties of software on hardware components to detect faults in perception and dynamic driving tasks. In at least one embodiment, output from ADAS system 1038 can be provided to a supervisory MCU. In at least one embodiment, if output from a primary computer and output from a secondary computer conflict, the supervisory MCU decides how to reconcile the conflict to ensure safe operation.

[0236] In at least one embodiment, a primary computer can be configured to provide a confidence score to a supervisory MCU to indicate a confidence of the primary computer in a selected result. In at least one embodiment, if the confidence score exceeds a threshold, the supervisory MCU can follow the indication of the primary computer regardless of whether the secondary computer provides a conflicting or inconsistent result. In at least one embodiment, in the event that the confidence score does not satisfy the threshold, and in the event that the primary computer and the secondary computer indicate different results (e.g., conflict), the supervisory MCU can arbitrate between the computers to determine an appropriate result.

[0237] In at least one embodiment, a supervisory MCU can be configured to run a neural network trained and configured to determine conditions under which an auxiliary computer provides false alarms based at least in part on outputs from a host computer and outputs from an auxiliary computer. In at least one embodiment, a neural network in a supervisory MCU can learn when to trust outputs of an auxiliary computer, and when not to. For example, in at least one embodiment, when the auxiliary computer is a RADAR-based FCW system, a neural network in a supervisory MCU can learn when the FCW system identifies metal objects that are not actually dangerous, such as drain grates or manhole covers that would trigger an alert. In at least one embodiment, when the auxiliary computer is a camera-based LDW system, a neural network in a supervisory MCU can learn to override LDW when there is a bicyclist or pedestrian present and it is actually safest to lane depart. In at least one embodiment, a supervisory MCU can include at least one of a DLA or GPU suitable for running a neural network with associated memory. In at least one embodiment, a supervisory MCU can include and / or be included as a component of one or more SoCs 1004.

[0238] In at least one embodiment, ADAS system 1038 can include an auxiliary computer that performs ADAS functions using traditional computer vision rules. In at least one embodiment, the auxiliary computer can use classic computer vision rules (if-then), and presence of a neural network in a supervisory MCU can improve reliability, safety, and performance. For example, in at least one embodiment, diverse implementations and intentional non-identity make the overall system more fault-tolerant, especially to faults caused by software (or software-hardware interface) functionality. For example, in at least one embodiment, if there is a software bug or error in software running on a host computer, and non-identical software code running on an auxiliary computer provides consistent overall results, then a supervisory MCU can be more confident that overall results are correct, and that the bug in software or hardware on the host computer will not cause a significant error.

[0239] In at least one embodiment, outputs of ADAS system 1038 can be input into a perception module of a host computer and / or a dynamic driving task module of a host computer. For example, in at least one embodiment, if ADAS system 1038 indicates a forward collision warning due to an object directly in front, then a perception block can use that information in identifying the object. In at least one embodiment, as described herein, an auxiliary computer can have its own neural network trained such that risk of false positives is reduced.

[0240] In at least one embodiment, vehicle 1000 can also include infotainment SoC 1030 (e.g., an in-vehicle infotainment system (IVI)). Although illustrated and described as an SoC, in at least one embodiment, infotainment system SoC 1030 can not be an SoC and can include, without limitation, two or more discrete components. In at least one embodiment, infotainment SoC 1030 can include, without limitation, a combination of hardware and software that can be used to provide audio (e.g., music, personal digital assistant, navigation instructions, news, radio, etc.), video (e.g., television, movies, streaming media, etc.), telephony (e.g., hands-free calling), network connectivity (e.g., LTE, WiFi, etc.), and / or information services (e.g., navigation systems, rear park assist, radio data system, vehicle related information such as fuel level, total range, brake fuel level, oil level, doors open / close, air filter information, etc.) to vehicle 1000. For example, infotainment SoC 1030 can include a radio, disc player, navigation system, video player, USB and Bluetooth connectivity, car, car entertainment system, WiFi, steering wheel audio controls, hands-free voice controls, heads-up display (“HUD”), HMI display 1034, telematics equipment, control panel (e.g., for controlling and / or interacting with various components, features, and / or systems), and / or other components. In at least one embodiment, infotainment SoC 1030 can be further used to provide information (e.g., visual and / or audible) to a user of vehicle 1000, such as information from ADAS system 1038, autonomous driving information (such as planned vehicle maneuvers), trajectory, surrounding environment information (e.g., intersection information, vehicle information, road information, etc.), and / or other information.

[0241] In at least one embodiment, infotainment SoC 1030 can include any number and type of GPU functionality. In at least one embodiment, infotainment SoC 1030 can communicate with other devices, systems, and / or components of vehicle 1000 over bus 1002. In at least one embodiment, infotainment SoC 1030 can be coupled to a supervisory MCU such that a GPU of the infotainment system can perform some autonomous driving functionality in the event of a failure of primary controller 1036 (e.g., a host computer and / or backup computer of vehicle 1000). In at least one embodiment, infotainment SoC 1030 can cause vehicle 1000 to enter a driver-to-safe-stop mode, as described herein.

[0242] In at least one embodiment, vehicle 1000 can also include an instrument cluster 1032 (e.g., a digital instrument cluster, an electronic instrument cluster, a digital instrument panel, etc.). In at least one embodiment, instrument cluster 1032 can include, without limitation, a controller and / or supercomputer (e.g., a discrete controller or supercomputer). In at least one embodiment, instrument cluster 1032 can include, without limitation, any number and combination of gauges such as a speedometer, fuel level, oil pressure, tachometer, odometer, turn indicator, shift position indicator, one or more seatbelt warning lights, one or more parking brake warning lights, one or more engine malfunction lights, auxiliary restraint system (e.g., airbag) information, lighting controls, safety system controls, navigation information, etc. In some examples, information can be displayed and / or shared between infotainment SoC 1030 and instrument cluster 1032. In at least one embodiment, instrument cluster 1032 can be included as part of infotainment SoC 1030, and vice versa.

[0243] Inference and / or training logic 715 are used to perform inferencing and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 715 are provided below in conjunction with FIGS. 7A-7C, 8A, 8B, and / or 10C. Figure 7A and / or Figure 7B Details regarding inference and / or training logic 715 are provided below in conjunction with FIGS. 7A-7C, 8A, 8B, and / or 10C.

[0244] In at least one embodiment, at least one component shown or described as being part of inference and / or training logic 715 is implemented as part of one or more components of system 1000. Figure 10C In at least one embodiment, at least one component shown or described as being part of inference and / or training logic 715 is implemented as part of one or more components of system 1000. Figure 1-6 In at least one embodiment, inference and / or training logic 715 includes and / or is used in conjunction with one or more aspects of the technology and / or functionality described in Figure 1 In at least one embodiment, inference and / or training logic 715 includes and / or is used in conjunction with one or more aspects of the technology and / or functionality described in Figure 1-6 In at least one embodiment, inference and / or training logic 715 includes and / or is used in conjunction with one or more aspects of the technology and / or functionality described in Figure 1-6 In at least one embodiment, inference and / or training logic 715 includes and / or is used in conjunction with one or more aspects of the technology and / or functionality described in

[0245] Figure 10Dis a diagram of a system 1076 that communicates between a cloud-based server and Figure 10A a self-driving vehicle 1000, in accordance with at least one embodiment. In at least one embodiment, system 1076 can include, without limitation, one or more servers 1078, one or more networks 1090, and any number and type of vehicles, including vehicle 1000. In at least one embodiment, one or more servers 1078 can include, without limitation, a plurality of GPUs 1084(A)-1084(H) (collectively referred to herein as GPUs 1084), PCIe switches 1082(A)-1082(D) (collectively referred to herein as PCIe switches 1082), and / or CPUs 1080(A)-1080(B) (collectively referred to herein as CPUs 1080), which GPUs 1084, CPUs 1080, and PCIe switches 1082 can be interconnected with high-speed connection lines such as, without limitation, NVLink interfaces 1088 developed by NVIDIA and / or PCIe connections 1086. In at least one embodiment, GPUs 1084 are connected by NVLink and / or NVSwitch SoC connections, and GPUs 1084 and PCIe switches 1082 are connected by PCIe interconnects. Although eight GPUs 1084, two CPUs 1080, and four PCIe switches 1082 are shown, this is not intended to be limiting. In at least one embodiment, each of one or more servers 1078 can include, without limitation, any number of GPUs 1084, CPUs 1080, and / or PCIe switches 1082 in any combination. For example, in at least one embodiment, one or more servers 1078 can each include eight, sixteen, thirty-two, and / or more GPUs 1084.

[0246] In at least one embodiment, one or more servers 1078 can receive image data representative of images from vehicles over one or more networks 1090, which show unexpected or changing road conditions, such as road work that has recently started. In at least one embodiment, one or more servers 1078 can transmit updated equalization neural networks 1092 and / or map information 1094, including but not limited to information about traffic and road conditions, to vehicles over one or more networks 1090. In at least one embodiment, updates to map information 1094 can include, but are not limited to, updates to HD map 1022, such as information about construction sites, potholes, detours, flooding, and / or other obstacles. In at least one embodiment, neural networks 1092 and / or map information 1094 can be a result of new training and / or experience represented in data received from any number of vehicles in an environment, and / or based at least on training performed at a data center (e.g., using one or more servers 1078 and / or other servers).

[0247] In at least one embodiment, one or more servers 1078 can be used to train machine learning models (e.g., neural networks) based at least in part on training data. In at least one embodiment, training data can be generated by vehicles, and / or can be generated in simulations (e.g., using a game engine). In at least one embodiment, any amount of training data is labeled (e.g., where a related neural network benefits from supervised learning) and / or undergoes other pre-processing. In at least one embodiment, any amount of training data is not labeled and / or pre-processed (e.g., where an associated neural network does not require supervised learning). In at least one embodiment, once a machine learning model is trained, a machine learning model can be used by vehicles (e.g., transmitted to vehicles over one or more networks 1090, and / or a machine learning model can be used by one or more servers 1078 to remotely monitor vehicles.

[0248] In at least one embodiment, one or more servers 1078 can receive data from vehicles and apply the data to up-to-date real-time neural networks for real-time intelligent inference. In at least one embodiment, one or more servers 1078 can include deep-learning supercomputers and / or specialized Al computers powered by one or more GPUs 1084, such as DGX and DGX Station machines developed by NVIDIA. However, in at least one embodiment, one or more servers 1078 can include deep-learning infrastructure of a data center using CPU power.

[0249] In at least one embodiment, deep learning infrastructure of server(s) 1078 can be capable of fast, real-time inferencing and can use this capability to assess and validate the health of processors, software, and / or related hardware in vehicle 1000. For example, in at least one embodiment, deep learning infrastructure can receive periodic updates from vehicle 1000, such as sequences of images and / or objects located by vehicle 1000 in that sequence of images (e.g., through computer vision and / or other machine learning object classification techniques). In at least one embodiment, deep learning infrastructure can run its own neural network to identify objects and compare them to objects identified by vehicle 1000, and if results do not match and deep learning infrastructure concludes that AI in vehicle 1000 is malfunctioning, server(s) 1078 can send a signal to vehicle 1000 instructing a failsafe computer of vehicle 1000 to take control, notify passengers, and complete a safe parking operation.

[0250] In at least one embodiment, server(s) 1078 can include GPU(s) 1084 and programmable inference accelerator(s) such as NVIDIA’s TensorRT 3 devices. In at least one embodiment, a combination of GPU-driven servers and inference-accelerated servers can make real-time responses possible. In at least one embodiment, CPU-, FPGA-, and other processor-driven servers can be used for inference, for example, in cases where performance is less critical. In at least one embodiment, hardware structure 715 is used to perform one or more embodiments. Described herein in connection with Figure 7A and / or Figure 7B Details are provided regarding hardware structure 715.

[0251] Computer system

[0252] Figure 11 is a block diagram illustrating an exemplary computer system, which can be a system with interconnected devices and components, a system on a chip (SOC), or some combination thereof formed with a processor that can include execution units to execute an instruction, according to at least one embodiment. In at least one embodiment, according to the present disclosure, such as embodiments described herein, computer system 1100 can include, without limitation, components such as processor 1102 that includes execution units to perform logic to execute an algorithm for process data. In at least one embodiment, computer system 1100 can include a processor such as a Intel® Core® i7 Processor family, Xeon® TM , XScale TM and / or StrongARM TM , Core TM or Nervana TM microprocessor, although other systems (including PCs, engineering workstations, set-top boxes, etc. having other microprocessors) can also be used. In at least one embodiment, computer system 1100 can execute a version of the WINDOWS operating system available from Microsoft Corporation of Redmond, Wash., although other operating systems (UNIX and Linux, for example), embedded software, and / or graphical user interfaces can also be used.

[0253] Embodiments can be used in other devices such as handheld devices and embedded applications. Some examples of handheld devices include cellular phones, Internet Protocol devices, digital cameras, personal digital assistants (“PDAs”), and handheld PCs. In at least one embodiment, embedded applications can include a microcontroller, a digital signal processor (“DSP”), a system on a chip, a network computer (“NetPC”), a set-top box, a network hub, a wide area network (“WAN”) switch, or any other system that can perform one or more instructions in accordance with at least one embodiment.

[0254] In at least one embodiment, computer system 1100 can include, but is not limited to, a processor 1102, which can include, but is not limited to, one or more execution units 1108 to perform machine learning model training and / or inferencing according to techniques described herein. In at least one embodiment, computer system 1100 is a single processor desktop or server system, but in another embodiment, computer system 1100 can be a multiprocessor system. In at least one embodiment, processor 1102 can include, but is not limited to, a complex instruction set computer (“CISC”) microprocessor, a reduced instruction set computing (“RISC”) microprocessor, a very long instruction word (“VLIW”) microprocessor, a processor implementing a combination of instruction sets, or any other processor device, such as a digital signal processor. In at least one embodiment, processor 1102 can be coupled to a processor bus 1110 that can transmit data signals between processor 1102 and other components in computer system 1100.

[0255] In at least one embodiment, processor 1102 can include, without limitation, level 1 (“L1”) internal cache memory (“cache”) 1104. In at least one embodiment, processor 1102 can have a single -level cache or multi-level cache. In at least one embodiment, cache memory can reside in the processor 1102’s external. Other embodiments can include a combination of internal and external caches based on specific implementation and requirements. In at least one embodiment, register file 1106 can store different types of data, including, without limitation, integer registers, floating point registers, status registers, and instruction pointer registers, in various registers.

[0256] In at least one embodiment, execution unit 1108, including, without limitation, logic to perform integer and floating point operations, also resides in processor 1102. In at least one embodiment, processor 1102 can also include microcode (“ucode”) read-only memory (“ROM”) that stores microcode for certain macro instructions. In at least one embodiment, execution unit 1108 can include logic to

[0257] In at least one embodiment, execution unit 1108 can also be used in a microcontroller, embedded processor, graphics device, DSP, and other types of logic circuits. In at least one embodiment, computer system 1100 can include, without limitation, memory 1120. In at least one embodiment, memory 1120 can be a Dynamic Random Access Memory (“DRAM”) device, a Static Random Access Memory (“SRAM”) device, a flash memory device, or another memory device. In at least one embodiment, memory 1120 can store data signals expressed as instructions 1119 and / or data 1121 to be executed by processor 1102.

[0258] In at least one embodiment, a system logic chip can be coupled to processor bus 1110 and memory 1120. In at least one embodiment, system logic chip can include, without limitation, a memory controller hub (“MCH”) 1116 and processor 1102 can communicate with MCH 1116 via processor bus 1110. In at least one embodiment, MCH 1116 can provide a high bandwidth memory path 1118 to memory 1120 for instruction and data storage and for storage of graphics commands, data, and textures. In at least one embodiment, MCH 1116 can initiate data signals at processor bus 1110, memory 1120, and other components in computer system 1100, and can bridge data signals between processor bus 1110, memory 1120, and system I / O interface 1122. In at least one embodiment, system logic chip can provide a graphics port used to couple MCH 1116 to a graphics controller or card. In at least one embodiment, MCH 1116 can be coupled to memory 1120 through high bandwidth memory path 1118, and a graphics / video card 1112 can be coupled to MCH 1116 through an Accelerated Graphics Port (“AGP”) interconnect 1114.

[0259] In at least one embodiment, computer system 1100 can use system I / O interface 1122 as a proprietary hub interface bus to couple MCH 1116 to I / O controller hub (“ICH”) 1130. In at least one embodiment, ICH 1130 can provide a direct connection to some I / O devices and can be connected to other devices through a local I / O bus. In at least one embodiment, the local I / O bus can include, without limitation, a high-speed I / O bus for connecting peripheral devices to memory 1120, chipset, and processor 1102. Examples can include, without limitation, audio controller 1129, firmware hub (“Flash BIOS”) 1128, wireless transceiver 1126, data storage 1124, legacy I / O controller 1123 containing user input and keyboard interfaces, serial expansion port 1127 (e.g., Universal Serial Bus (“USB”) port), and network controller 1134. In at least one embodiment, data storage 1124 can include a hard disk drive, a floppy disk drive, a CD-ROM device, a flash memory device, or other mass storage device.

[0260] In at least one embodiment, Figure 11 A system is shown that includes interconnected hardware devices or “chips,” while in other embodiments, Figure 11A System-on-a-Chip (SoC) may be shown. In at least one embodiment, the device shown in FIG11 may be interconnected with a proprietary interconnect, a standardized interconnect (e.g., PCIe), or some combination thereof. In at least one embodiment, one or more components of the computer system 1100 are interconnected using a Compute Fast Link (CXL) interconnect.

[0261] The inference and / or training logic 715 is used to perform inference and / or training operations associated with one or more embodiments. This document combines... Figure 7A and / or Figure 7B Details are provided regarding the inference and / or training logic 715. In at least one embodiment, the inference and / or training logic 715 may be... Figure 11 Used in systems for reasoning or predicting operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures or neural network use cases described herein.

[0262] In at least one embodiment, regarding Figure 11 At least one component shown or described is used to achieve the combination Figure 1-6 The described techniques and / or functions. In at least one embodiment, the inference and / or training logic 715 includes and / or operates regarding... Figure 1 At least one aspect described (e.g., deep learning compiler 102, stream scheduler 110, memory allocator 112). In at least one embodiment, inference and / or training logic 715 uses a representation of a computer program to train at least one untrained or partially trained neural network, the representation of which may be as described regarding Figure 1-6 One or more of the operations and / or instructions described in the description are speculatively performed. In at least one embodiment, the inference and / or training logic uses a representation of a computer program to perform at least one inference operation, the representation of which may indicate, as per [the description of the inference and / or training logic]... Figure 1-6 One or more of the operations and / or instructions described as being performed speculatively. In at least one embodiment, utilizing Figure 11 The computer system 1100 uses the processor 1102 and / or other components to achieve the combination Figure 1-6 The described technologies and / or functions.

[0263] Figure 12 This is a block diagram illustrating an electronic device 1200 for utilizing a processor 1210 according to at least one embodiment. In at least one embodiment, the electronic device 1200 may be, for example, but not limited to, a laptop computer, tower server, rack server, blade server, desktop computer, tablet computer, mobile device, telephone, embedded computer, or any other suitable electronic device.

[0264] In at least one embodiment, electronic device 1200 can include, without limitation, a processor 1210 communicatively coupled to any suitable number or kind of components, peripherals, modules, or devices. In at least one embodiment, processor 1210 is coupled using a bus or interface such as an Industry Standard 2 C bus, System Management Bus (“SMBus”), Low Pin Count (LPC) bus, Serial Peripheral Interface (“SPI”), High Definition Audio (“HDA”) bus, Serial Advance Technology Attachment (“SATA”) bus, Universal Serial Bus (“USB”) (versions 1, 2, 3, etc.), or Universal Asynchronous Receiver / Transmitter (“UART”) bus. In at least one embodiment, processor 1210 can include, without limitation, a complex instruction set computer (“CISC”) microprocessor, a reduced instruction set computing (“RISC”) microprocessor, or an Figure 12 Systems can be shown that include interconnected hardware devices or “chips,” while in other embodiments, Figure 12 An exemplary SoC can be shown. In at least one embodiment, Figure 12 Devices shown in FIG. 13 can be interconnected with proprietary interconnects, standardized interconnects (e.g., PCIe), or some combination thereof. In at least one embodiment, Figure 12 One or more components of FIG. 13 are interconnected using Compute Express Link (CXL) interconnects.

[0265] In at least one embodiment, Figure 12 may include a display 1224, a touchscreen 1225, a touchpad 1230, a near-field communication unit (“NFC”) 1245, a sensor hub 1240, a thermal sensor 1246, an Express Chipset (“EC”) 1235, a Trusted Platform Module (“TPM”) 1238, a BIOS / firmware / flash memory (“BIOS, FW Flash”) 1222, a DSP 1260, a drive 1220 (such as a solid state disk (“SSD”) or a hard disk drive (“HDD”)), a wireless local area network unit (“WLAN”) 1250, a Bluetooth unit 1252, a wireless wide area network unit (“WWAN”) 1256, a Global Positioning System (“GPS”) unit 1255, a camera (“USB 3.0 camera”) 1254 (such as a USB 3.0 camera), and / or a Low Power Double Data Rate (“LPDDR”) memory unit (“LPDDR3”) 1215 implemented in, for example, LPDDR3 standard. These components can each be implemented in any suitable manner.

[0266] In at least one embodiment, other components can be communicatively coupled to processor 1210 by components described herein. In at least one embodiment, an accelerometer 1241, an ambient light sensor (“ALS”) 1242, a compass 1243, and a gyroscope 1244 can be communicatively coupled to a sensor hub 1240. In at least one embodiment, a thermal sensor 1239, a fan 1237, a keyboard 1236, and a touchpad 1230 can be communicatively coupled to an EC 1235. In at least one embodiment, a speaker 1263, a headphone 1264, and a microphone (“mic”) 1265 can be communicatively coupled to an audio unit (“audio codec and class D amplifier”) 1262, which in turn can be communicatively coupled to a DSP 1260. In at least one embodiment, audio unit 1262 can include, for example and without limitation, an audio coder / decoder (“codec”) and a class D amplifier. In at least one embodiment, a SIM card (“SIM”) 1257 can be communicatively coupled to a WWAN unit 1256. In at least one embodiment, components such as WLAN unit 1250 and Bluetooth unit 1252, and WWAN unit 1256 can be implemented as a next generation form factor (NGFF).

[0267] Inference and / or training logic 715 are used to perform inferencing and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 715 are provided below in conjunction with FIGS. 7L and / or 7M. Figure 7A and / or Figure 7B Details regarding inference and / or training logic 715 are provided below in conjunction with FIGS. 7L and / or 7M.

[0268] In at least one embodiment, at least one component shown or described as being co-located can be Figure 12 implemented in a distributed format. In at least one embodiment, at least one component shown or described as being co-located can be a separate and individual component. Figure 1-6 In at least one embodiment, inference and / or training logic 715 includes and / or runs at least one aspect described as being performed by Figure 1 In at least one embodiment, inference and / or training logic 715 uses a representation of a computer program to train at least one untrained or partially trained neural network, the representation of the computer program indicating operations and / or instructions that can be speculatively executed as described with respect to one or more of FIGS. 7A-7M. Figure 1-6 In at least one embodiment, inference and / or training logic uses a representation of a computer program to perform at least one inferencing operation, the representation of the computer program indicating operations and / or instructions that can be speculatively executed as described with respect to one or more of FIGS. 7A-7M. Figure 1-6one or more of the operations and / or instructions described herein that are performed by the apparatus 700 in the manner described above. In at least one embodiment, the apparatus 700 is configured to perform a method for wireless communication by performing the above-described operations and / or instructions. Figure 12 The system 1200 and / or the processor 1210 are used to implement techniques and / or functions described throughout this disclosure. Figure 1-6 The system 1200 and / or the processor 1210 are used to implement techniques and / or functions described throughout this disclosure.

[0269] Figure 13 A computer system 1300 according to at least one embodiment is shown. In at least one embodiment, computer system 1300 is configured to implement various processes and methods described throughout this disclosure.

[0270] In at least one embodiment, computer system 1300 includes, without limitation, at least one central processing unit (“CPU”) 1302 that is connected to a communication bus 1310 implemented using any suitable protocol, such as PCI (“Peripheral Component Interconnect”), peripheral component interconnect express (“PCI-Express”), AGP (“Accelerated Graphics Port”), HyperTransport, or any other bus or point-to-point communication protocol. In at least one embodiment, computer system 1300 includes, without limitation, a main memory 1304 and control logic (e.g., implemented in hardware, software, or a combination thereof) and data can be stored in the main memory 1304 in the form of random-access memory (“RAM”).

[0271] In at least one embodiment, computer system 1300 includes, without limitation, an input device 1308, parallel processing system 1312, and display device 1306, which can be implemented using a conventional cathode ray tube (“CRT”), liquid crystal display (“LCD”), light emitting diode (“LED”), plasma display, or other suitable display technologies in at least one embodiment. In at least one embodiment, user input is received from input device 1308 such as keyboard, mouse, touchpad, microphone, etc. In at least one embodiment, each of the modules described herein can be located on a single semiconductor platform.

[0272] Inference and / or training logic 715 are used to perform inferencing and / or training operations associated with one or more embodiments. In at least one embodiment, inference and / or training logic 715 can be used in place of, or to complement, one or more of the operations and / or instructions described herein that are performed by the apparatus 700 in the manner described above. In at least one embodiment, the apparatus 700 is configured to perform a method for wireless communication by performing the above-described operations and / or instructions. Figure 7A Inference and / or training logic 715 are used to perform inferencing and / or training operations associated with one or more embodiments. In at least one embodiment, inference and / or training logic 715 can be used in place of, or to complement, one or more of the operations and / or instructions described herein that are performed by the apparatus 700 in the manner described above. In at least one embodiment, the apparatus 700 is configured to perform a method for wireless communication by performing the above-described operations and / or instructions. Figure 7BDetails regarding the inference and / or training logic 715 are provided. In at least one embodiment, the inference and / or training logic 715 may be used in System Figure 13 to perform inference or prediction operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architecture, or neural network use cases described herein.

[0273] In at least one embodiment, relative to Figure 13 At least one component shown or described is used to achieve the combination Figure 1-6 The described techniques and / or functions. In at least one embodiment, the inference and / or training logic 715 includes and / or operates regarding... Figure 1 At least one aspect described (e.g., deep learning compiler 102, stream scheduler 110, memory allocator 112). In at least one embodiment, inference and / or training logic 715 uses a representation of a computer program to train at least one untrained or partially trained neural network, the representation of which may be as described regarding Figure 1-6 One or more of the operations and / or instructions described in the description are speculatively performed. In at least one embodiment, the inference and / or training logic uses a representation of a computer program to perform at least one inference operation, the representation of which may indicate, as per [the description of the inference and / or training logic]... Figure 1-6 One or more of the operations and / or instructions described as being performed speculatively. In at least one embodiment, utilizing Figure 13 The computer system 1300 and / or at least one PPU 1314 are used to achieve the combination. Figure 1-6 The described technologies and / or functions.

[0274] Figure 14 A computer system 1400 according to at least one embodiment is illustrated. In at least one embodiment, the computer system 1400 includes, but is not limited to, a computer 1410 and a USB flash drive 1420. In at least one embodiment, the computer 1410 may include, but is not limited to, any number and type of processors (not shown) and memory (not shown). In at least one embodiment, the computer 1410 includes, but is not limited to, a server, a cloud instance, a laptop computer, and a desktop computer.

[0275] In at least one embodiment, USB stick 1420 includes, without limitation, processing unit 1430, USB interface 1440, and USB interface logic 1450. In at least one embodiment, processing unit 1430 can be any instruction execution system, apparatus, or device capable of executing instructions. In at least one embodiment, processing unit 1430 can include, without limitation, any number and type of processing core (not shown). In at least one embodiment, processing unit 1430 includes an application-specific integrated circuit (“ASIC”) optimized to perform any number and type of operations associated with machine learning. For example, in at least one embodiment, processing unit 1430 is a tensor processing unit (“TPC”) optimized to perform machine learning inference operations. In at least one embodiment, processing unit 1430 is a vision processing unit (“VPU”) optimized to perform machine vision and machine learning inference operations.

[0276] In at least one embodiment, USB interface 1440 can be any type of USB connector or USB receptacle. For example, in at least one embodiment, USB interface 1440 is a USB 3.0 Type-C receptacle for data and power. In at least one embodiment, USB interface 1440 is a USB 3.0 Type-A connector. In at least one embodiment, USB interface logic 1450 can include any number and type of logic that enables processing unit 1430 to interface with a device (e.g., computer 1410) via USB connector 1440.

[0277] Inference and / or training logic 715 are used to perform inferencing and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 715 are provided below in conjunction with FIGS. 6 and 7. Figure 7A and / or Figure 7B Details regarding inference and / or training logic 715 are provided below in conjunction with FIGS. 6 and 7. In at least one embodiment, inference and / or training logic 715 can be used in system FIG. 14 for inferencing or predicting operations based, at least in part, on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.

[0278] In at least one embodiment, at least one component shown or described as being implemented via computer executable instructions can be implemented using one or more ASICs designed to perform the functions recited by the instructions without Figure 14 At least one component shown or described as being implemented via computer executable instructions can be implemented using one or more ASICs designed to perform the functions recited by the instructions without Figure 1-6 In at least one embodiment, inference and / or training logic 715 include and / or run one or more of components shown and described in FIGS. 6 and 7. Figure 1The described at least one aspect (e.g., deep learning compiler 102, stream scheduler 110, memory allocator 112). In at least one embodiment, inference and / or training logic 715 uses a representation of a computer program to train at least one untrained or partially trained neural network, which representation of a computer program indicates operations and / or instructions that can be executed speculatively as described with respect to one or more of Figure 1-6 In at least one embodiment, inference and / or training logic uses a representation of a computer program to perform at least one inferencing operation, which representation of a computer program indicates operations and / or instructions that can be executed speculatively as described with respect to one or more of Figure 1-6 In at least one embodiment, techniques and / or functions described in conjunction with Figure 14 are implemented with processing unit 1430 of Figure 1-6

[0279] Figure 15A An exemplary architecture is shown in which a plurality of GPUs 1510(1)- 1510(N) are communicatively coupled to a plurality of multi-core processors 1505(1)- 1505(M) over high-speed links 1540(1)-1540(N) (e.g., buses, point-to-point interconnects, etc.). In at least one embodiment, high-speed links 1540(1)-1540(N) support a communication throughput of 4GB / s, 30GB / s, 80GB / s or higher. In at least one embodiment, various interconnect protocols can be used including, but not limited to, PCIe 4.0 or 5.0 and NVLink 2.0. In various figures, “N” and “M” represent positive integers, values of which can vary from figure to figure.

[0280] Moreover, in one embodiment, two or more of GPUs 1510 are interconnected by a high-speed link 1529(1)-1529(2), which can be implemented using similar or different protocols / links than those used for high-speed links 1540(1)-1540(N). Similarly, two or more of multi-core processors 1505 can be connected by an interconnection 1528, which can be an SMP bus that runs at 20 GB / s, 30 GB / s, 120 GB / s or higher. Alternatively, all communication between various system components shown in Figure 15A may be accomplished using similar protocols / links (e.g., over common interconnection

[0281] ​In at least one embodiment, each multi-core processor 1505 is communicatively coupled to processor memories 1501(1)-1501(M) via memory interconnects 1526(1)-1526(M), respectively, and each GPU 1510(1)-1510(N) is communicatively coupled to GPU memories 1520(1)-1520(N) by GPU memory interconnects 1550(1)-1550(N), respectively. In at least one embodiment, memory interconnects 1526 and 1550 can utilize similar or different memory access technologies. By way of non-limiting example, processor memories 1501(1)-1501(M) and GPU memories 1520 can be volatile memories such as dynamic random access memory (DRAM) including stacked DRAM, graphics DDR SDRAM (GDDR) such as GDDR5, GDDR6, or high-bandwidth memory (HBM), and / or can be non-volatile memories such as 3D XPoint or Nano-Ram. In at least one embodiment, certain portions of processor memories 1501 can be volatile memory while another portion can be non-volatile memory (e.g., using a two-level memory (2LM) hierarchy).

[0282] As described herein, although various multi-core processors 1505 and GPUs 1510 can be physically coupled to particular memories 1501, 1520, respectively, and / or can implement a unified memory architecture in which a virtual system address space (also referred to as an “effective address” space) is distributed among various physical memories. For example, processor memories 1501(1)-1501(M) can each contain 64 GB of system memory address space, and GPU memories 1520(1)-1520(N) can each contain 32 GB of system memory address space, resulting in a total of 256 GB of addressable memory size when M=2 and N=4. N and M can also be other values.

[0283] Figure 15B Additional details for interconnect between multi-core processor 1507 and graphics acceleration module 1546 are shown, according to one exemplary embodiment. In at least one embodiment, graphics acceleration module 1546 can include one or more GPU chips integrated on a line card that is coupled via a high-speed link 1540 (e.g., a PCIe bus, NVLink, etc.) to processor 1507. In at least one embodiment, graphics acceleration module 1546 can alternatively be integrated on a package or chip with processor 1507.

[0284] In at least one embodiment, processor 1507 includes multiple cores 1560A-1560D each with a translation lookaside buffer (“TLB”) 1561A-1561D and one or more caches 1562A-1562D. In at least one embodiment, cores 1560A-1560D can include various other components not shown for purposes of this description, to execute instructions and process data. In at least one embodiment, caches 1562A-1562D can include Level 1 (“Ll”) and Level 2 (“L2”) caches. In addition, one or more shared caches 1556 can be included in caches 1562A-1562D and shared by groups of cores 1560A-1560D. For example, one embodiment of processor 1507 includes 24 cores, each with its own Ll cache, twelve shared L2 caches, and twelve shared L3 caches. In that embodiment, two adjacent cores share one or more L2 and L3 caches. In at least one embodiment, processor 1507 and graphics acceleration module 1546 are connected with system memory 1514, which can include processor memories 1501(1)-1501(M) in Figure 15A

[0285] In at least one embodiment, consistency of data and instructions stored in respective caches 1562A-1562D, 1556, and system memory 1514 is maintained through inter-core communications via coherence bus 1564. In at least one embodiment, for example, each cache can have cache coherency logic / circuitry associated therewith to communicate through coherence bus 1564 in response to detecting a read or write to a particular cache line. In at least one embodiment, a cache snoop protocol is implemented through coherence bus 1564 to snoop cache accesses.

[0286] In at least one embodiment, agent circuitry 1525 communicatively couples graphics acceleration module 1546 to coherence bus 1564, allowing graphics acceleration module 1546 to participate in cache coherency protocol as a peer to cores 1560A-1560D. In particular, in at least one embodiment, interface 1535 provides connectivity from graphics acceleration module 1546 to agent circuitry 1525 over high-speed link 1540, and interface 1537 connects graphics acceleration module 1546 to high-speed link 1540.

[0287] ​In at least one embodiment, accelerator integration circuit 1536 provides cache management, memory access, context management, and interrupt management services on behalf of graphics processing engines 1531(1)-1531(N). In at least one embodiment, graphics processing engines 1531(1)-1531(N) each comprise a separate graphics processing unit (GPU). In at least one embodiment, graphics processing engines 1531(1)-1531(N) alternatively can comprise different types of graphics processing engines within a GPU, such as graphics execution units, media processing engines, samplers, and blit engines. In at least one embodiment, graphics acceleration module 1546 can be a GPU with a plurality of graphics processing engines 1531(1)-1531(N) or graphics processing engines 1531(1)-1531(N) can be individual GPUs integrated on a common package, line card, or chip.

[0288] In at least one embodiment, accelerator integration circuit 1536 includes a memory management unit (MMU) 1539 to provide for translation of virtual addresses into physical addresses, as is known to those skilled in the art. In at least one embodiment, MMU 1539 can include memory protection facilities that enable different privilege levels in the system. In at least one embodiment, MMU 1539 can provide translation of virtual addresses into physical addresses, including handling of interrupts and exceptions, such as “page faults” that occur if an access is attempted when no mapping exists at the provided address; MMU 1539 can handle translation of virtual addresses into physical addresses, including handling of interrupts and exceptions, such as “page faults” that occur if an access is attempted when no mapping exists at the provided address. In at least one embodiment, MMU 1539 can include a translation lookaside buffer (TLB) to improve translation speeds into physical addresses; TLB can include a cache of recently translated addresses, so that the same address can be translated more quickly; in at least one embodiment, MMU 1539 can include a translation lookaside buffer (TLB) to improve translation speeds into physical addresses; TLB can include a cache of recently translated addresses, so that the same address can be translated more quickly; in at least one embodiment, MMU 1539 can include a translation lookaside buffer (TLB) to improve translation speeds into physical addresses; TLB can include a cache of recently translated addresses, so that the same address can be translated more quickly. In at least one embodiment, cache 1538 can store commands and data used by graphics processing engines 1531(1)-1531(N) for more efficient processing. In at least one embodiment, fetch unit 1544 can be used to maintain coherency between data stored in cache 1538 and graphics memory 1533(1)-1533(M) with core caches 1562A-1562D, 1556, and system memory 1514. As described previously, this can be accomplished via proxy circuitry 1525 representing cache 1538 and graphics memory 1533(1)-1533(M) (e.g., sending updates related to modifications / accesses made to a cache line on processor caches 1562A-1562D, 1556 to cache 1538 and receiving updates from cache 1538).

[0289] In at least one embodiment, a set of registers 1545 store context data for threads executed by graphics processing engines 1531(1)-1531(N), and context management circuit 1548 manages thread contexts. For example, context management circuit 1548 can perform save and restore operations to save and restore context of individual threads during context switches (e.g., where a first thread is saved and a second thread is stored so that it can be executed by a graphics processing engine). For example, context management circuit 1548 can store current register values to a designated area in memory (e.g., identified by a context pointer) at a context switch. Register values can then be restored when returning to a context. In at least one embodiment, interrupt management circuit 1547 receives and processes interrupts received from system devices.

[0290] In at least one embodiment, MMU 1539 translates virtual / effective addresses from graphics processing engines 1531 to real / physical addresses in system memory 1514. In at least one embodiment, accelerator integration circuit 1536 supports multiple (e.g., 4, 8, 16) graphics processor modules 1546 and / or other accelerator devices. In at least one embodiment, graphics processor modules 1546 can be dedicated to a single application executing on processor 1507 or can be shared between multiple applications. In at least one embodiment, a virtualized graphics execution environment is presented in which resources of graphics processing engines 1531(1)-1531(N) are shared between multiple applications or virtual machines (VMs). In at least one embodiment, resources can be subdivided into “slices” that are assigned to different VMs and / or applications based on processing requirements and priority levels associated with VMs and / or applications.

[0291] In at least one embodiment, accelerator integration circuit 1536 performs as a bridge to system for graphics acceleration module 1546, and provides address translation and system memory caching services. Additionally, in at least one embodiment, accelerator integration circuit 1536 can provide virtualization facilities for a host processor to manage virtualization of graphics processing engines 1531(1)-1531(N), interrupts, and memory management.

[0292] In at least one embodiment, because hardware resources of graphics processing engines 1531(1)-1531(N) are explicitly mapped to real address space seen by host processor 1507, any host processor can directly address these resources using effective address values. In at least one embodiment, one function of accelerator integration circuit 1536 is to physically separate graphics processing engines 1531(1)-1531(N) so that they appear as independent units to a system.

[0293] In at least one embodiment, one or more graphics memory 1533(1)-1533(M) are coupled to each graphics processing engines 1531(1)-1531(N) respectively, and N=M. In at least one embodiment, graphics memory 1533(1)-1533(M) stores instructions and data used by each graphics processing engines 1531(1)-1531(N). In at least one embodiment, graphics memory 1533(1)-1533(M) can be a volatile memory, such as DRAM (including stacked DRAM), GDDR memory (e.g., GDDR5, GDDR6), or HBM, and / or can be non-volatile memory, such as 3DXPoint or Nano-Ram.

[0294] In at least one embodiment, to reduce data traffic on high-speed link 1540, a bias technique can be used to ensure that data stored in graphics memory 1533(1)-1533(M) is data that is most frequently used by graphics processing engines 1531(1)-1531(N), and data that is least used (at least frequently) by cores 1560A-1560D. Similarly, in at least one embodiment, a bias mechanism attempts to keep data needed by cores (and preferably not by graphics processing engines 1531(-1)-1531(N)) in caches 1562A-1562D, 1556, and system memory 1514.

[0295] Figure 15C Another example embodiment is shown in which accelerator integration circuit 1536 is integrated within processor 1507. In this embodiment, graphics processing engines 1531(1)-1531(N) communicate directly with accelerator integration circuit 1536 via interface 1537 and interface 1535 (which can also be a bus or can be an interface of some form) over high-speed link 1540. In at least one embodiment, accelerator integration circuit 1536 can perform similar operations to those described with respect to Figure 15B operations described with respect to accelerator integration circuit 1536, but can have higher throughput due to its close proximity to coherence bus 1564 and caches 1562A-1562D, 1556. In at least one embodiment, accelerator integration circuit supports different programming models, including a dedicated process programming model (no graphics acceleration module virtualization) and a shared programming model (with virtualization), which can include programming models controlled by accelerator integration circuit 1536 and programming models controlled by graphics acceleration module 1546.

[0296] In at least one embodiment, graphics processing engines 1531(1)-1531(N) are dedicated to a single application or process under a single operating system. In at least one embodiment, a single application can funnel other application requests to graphics processing engines 1531(1)-1531(N), providing virtualization within a VM / partition.

[0297] In at least one embodiment, graphics processing engines 1531(1)-1531(N) can be shared by multiple VMs / application partitions. In at least one embodiment, a shared model can use a hypervisor to virtualize graphics processing engines 1531(1)-1531(N) to allow access by each operating system. In at least one embodiment, for a single partition system without hypervisor, an operating system owns graphics processing engines 1531(1)-1531(N). In at least one embodiment, an operating system can virtualize graphics processing engines 1531(1)-1531(N) to provide access to each process or application.

[0298] In at least one embodiment, graphics acceleration module 1546 or individual graphics processing engines 1531(1)-1531(N) use a process handle to select a process element. In at least one embodiment, a process element is stored in system memory 1514 and can be addressed using effective to real address translation techniques described herein. In at least one embodiment, a process handle can be an implementation specific value provided to a host process when registering its context with a graphics processing engine 1531(1)-1531(N) (i.e., calling system software to add a process element to a process element linked list). In at least one embodiment, a lower 16 bits of a process handle can be an offset into a process element linked list for a process element.

[0299] Figure 15DAn exemplary accelerator integration slice 1590 is shown. In at least one embodiment, a “slice” comprises a specified portion of processing resources of accelerator integration circuit 1536. In at least one embodiment, an application is an effective address space 1582 in system memory 1514 that stores process elements 1583. In at least one embodiment, process elements 1583 are stored in response to GPU invocations 1581 from an application 1580 executing on processor 1507. In at least one embodiment, process elements 1583 contain process state for respective application 1580. In at least one embodiment, a work descriptor (WD) 1584 contained in process element 1583 can be a single job requested by an application or can contain pointers to a queue of jobs. In at least one embodiment, WD 1584 is a pointer to a job request queue in an application’s effective address space 1582.

[0300] In at least one embodiment, graphics acceleration module 1546 and / or individual graphics processing engines 1531(1)-1531(N) can be shared by all processes or a subset of processes in a system. In at least one embodiment, infrastructure can be included for setting process state and sending WDs 1584 to the graphics acceleration module 1546 to initiate work in a virtualized environment.

[0301] In at least one embodiment, a dedicated process programming model is implementation specific. In at least one embodiment, in this model, a single process owns a graphics acceleration module 1546 or individual graphics processing engines 1531. In at least one embodiment, when a graphics acceleration module 1546 is owned by a single process, a hypervisor initializes the accelerator integration circuit for the owned partition and an operating system initializes the accelerator integration circuit 1536 for the owned process when a graphics acceleration module 1546 is assigned.

[0302] In at least one embodiment, in operation, a WD fetch unit 1591 in accelerator integration slice 1590 fetches a next WD 1584, which includes an indication of work to be done by one or more graphics processing engines of graphics acceleration module 1546. In at least one embodiment, data from WD 1584 can be stored in registers 1545 and used by MMU 1539, interrupt management circuit 1547, and / or context management circuit 1548, as shown. For example, one embodiment of MMU 1539 includes segment / page walk circuitry to access segment / page tables 1586 within an OS virtual address space 1585. In at least one embodiment, interrupt management circuit 1547 can handle interrupt events 1592 received from graphics acceleration module 1546. In at least one embodiment, effective addresses 1593 generated by graphics processing engines 1531(1)-1531(N) are translated to real addresses by MMU 1539 when performing graphics operations.

[0303] In at least one embodiment, registers 1545 are replicated for each graphics processing engine 1531(1)-1531(N) and / or graphics acceleration module 1546, and can be initialized by a hypervisor or operating system. In at least one embodiment, each of these replicated registers can be included in accelerator integration slice 1590. Exemplary registers that can be initialized by a hypervisor are shown in Table 1.

[0304]

[0305]

[0306] Exemplary registers that can be initialized by an operating system are shown in Table 2.

[0307]

[0308] In at least one embodiment, each WD 1584 is specific to a particular graphics acceleration module 1546 and / or graphics processing engines 1531(1)-1531(N). In at least one embodiment, it contains all information needed for a graphics processing engine 1531(1)-1531(N) to finish the work, or it can be a pointer to a memory location where an application has set up a command queue of work to be done.

[0309] Figure 15EAdditional details are shown for one exemplary embodiment of a shared model. This embodiment includes a hypervisor real address space 1598 in which a list of process elements 1599 is stored. In at least one embodiment, the hypervisor real address space 1598 is accessible via a hypervisor 1596 that virtualizes a graphics acceleration module engine for an operating system 1595.

[0310] In at least one embodiment, a shared programming model allows all processes or a subset of processes from all partitions or a subset of partitions in a system to use a graphics acceleration module 1546. In at least one embodiment, there are two programming models in which a graphics acceleration module 1546 is shared by multiple processes and partitions, namely, time-sliced sharing and graphics-directed sharing.

[0311] In at least one embodiment, in this model, a system hypervisor 1596 owns the graphics acceleration module 1546 and makes its functionality available to all operating systems 1595. In at least one embodiment, for a graphics acceleration module 1546 to support virtualization by a system hypervisor 1596, the graphics acceleration module 1546 can adhere to certain requirements, such as (1) an application’s job request must be autonomous (i.e., no state needs to be kept between jobs), or the graphics acceleration module 1546 must provide a context save and restore mechanism, (2) the graphics acceleration module 1546 guarantees that an application’s job request completes within a specified amount of time, including any translation faults, or the graphics acceleration module 1546 provides the ability to preempt job processing, and (3) fairness between graphics acceleration module 1546 processes must be ensured when operating in a directed sharing programming model.

[0312] In at least one embodiment, an application 1580 is required to use a graphics acceleration module type, a work descriptor (WD), an authority mask register (AMR) value, and a context save / restore area pointer (CSRP) for an operating system 1595 system call. In at least one embodiment, the graphics acceleration module type describes a target acceleration function for the system call. In at least one embodiment, the graphics acceleration module type can be a system-specific value. In at least one embodiment, the WD is formatted specifically for the graphics acceleration module 1546 and can take the form of a graphics acceleration module 1546 command, a valid address pointer to a user-defined structure, a valid address pointer to a command queue, or any other data structure that describes work to be done by the graphics acceleration module 1546.

[0313] In at least one embodiment, the AMR value is the AMR state for the current process. In at least one embodiment, the value passed to operating system is similar to the application that set the AMR. In at least one embodiment, if the implementation of accelerator integration circuit 1536 (not shown) and graphics acceleration module 1546 does not support a user access mask override register (UAMOR), then operating system can apply current UAMOR value to AMR value before passing AMR in a hypervisor call. In at least one embodiment, hypervisor 1596 can selectively apply current access mask override register (AMOR) value before placing AMR in process element 1583. In at least one embodiment, CSRP is one of registers 1545 that contains an effective address of a region in application’s effective address space 1582 for graphics acceleration module 1546 to save and restore context state. In at least one embodiment, this pointer is optional if there is no need to save state between jobs or when a job is preempted. In at least one embodiment, context save / restore region can be a fixed system memory.

[0314] Upon receiving the system call, operating system 1595 can verify that application 1580 is registered and has been granted permission to use graphics acceleration module 1546. Operating system 1595 then, in at least one embodiment, uses information shown in Table 3 to call hypervisor 1596.

[0315]

[0316]

[0317] Upon receiving the hypervisor call, hypervisor 1596 verifies, in at least one embodiment, that operating system 1595 is registered and has been granted permission to use graphics acceleration module 1546. Hypervisor 1596 then, in at least one embodiment, places process element 1583 in a corresponding graphics acceleration module 1546 type of process element linked list. In at least one embodiment, process element can include information shown in Table 4.

[0318]

[0319] In at least one embodiment, hypervisor initializes a number of accelerator integration slice 1590 registers 1545.

[0320] As Figure 15FAs shown, in at least one embodiment, a unified memory is used that can be addressed via a common virtual memory address space for accessing physical processor memory 1501(1)-1501(N) and GPU memory 1520(1)-1520(N). In this implementation, operations performed on GPU(s) 1510(1)-1510(N) utilize the same virtual / effective memory address space to access processor memory 1501(1)-1501(M), and vice versa, simplifying programmability. In at least one embodiment, a first portion of the virtual / effective address space is allocated to processor memory 1501(1), a second portion to second processor memory 1501(N), a third portion to GPU memory 1520(1), and so on. In at least one embodiment, the entire virtual / effective memory space (sometimes referred to as an effective address space) is thus distributed among processor memory 1501 and GPU memory 1520, allowing any processor or GPU to access a memory with a virtual address that maps to that memory.

[0321] In at least one embodiment, bias / coherence management circuitry 1594A-1594E within one or more MMU(s) 1539A-1539E ensure cache coherency between one or more host processor(s) (e.g., 1505) and caches of GPU 1510, and implement bias techniques that dictate a physical memory in which certain types of data should be stored. In at least one embodiment, while multiple instances of bias / coherence management circuitry 1594A-1594E are shown in FIG. 15B, bias / coherence circuitry can be implemented within MMU(s) of one or more host processor(s) 1505 and / or within accelerator integration circuit 1536. Figure 15F In at least one embodiment, while multiple instances of bias / coherence management circuitry 1594A-1594E are shown in FIG. 15B, bias / coherence circuitry can be implemented within MMU(s) of one or more host processor(s) 1505 and / or within accelerator integration circuit 1536.

[0322] One embodiment allows GPU memory 1520 to be mapped as part of system memory and accessed using shared virtual memory (SVM) techniques, but without suffering the performance penalties associated with full system cache coherency. In at least one embodiment, the ability to access GPU memory 1520 as system memory without the heavy cache coherency overhead provides a favorable operating environment for GPU offload. In at least one embodiment, this arrangement allows software of host processor 1505 to set operands and access computation results without the overhead of traditional I / O DMA data copies. In at least one embodiment, such traditional copies include driver calls, interrupts, and memory-mapped I / O (MMIO) accesses, which are all less efficient than simple memory accesses. In at least one embodiment, the ability to access GPU memory 1520 without cache coherency overhead can be critical to the execution time of offloaded computations. In at least one embodiment, for example, in cases with a large amount of streaming write memory traffic, cache coherency overhead can significantly reduce the effective write bandwidth seen by GPU 1510. In at least one embodiment, the efficiency of operand setup, the efficiency of result access, and the efficiency of GPU computation can all play a role in determining the effectiveness of GPU offload.

[0323] In at least one embodiment, the selection of GPU bias and host processor bias is driven by a bias tracker data structure. In at least one embodiment, for example, a bias table can be used, which can be a page-granularity structure (e.g., controlled at the granularity of a memory page) that includes a memory page 1 or 2 bits per GPU attachment. In at least one embodiment, with or without a bias cache in GPU 1510 (e.g., to cache frequently / recently used entries of the bias table), the bias table can be implemented in the stolen memory range of one or more GPU memories 1520. Alternatively, in at least one embodiment, the entire bias table can be maintained within the GPU.

[0324] In at least one embodiment, prior to actually accessing GPU memory, a bias table entry associated with each access to GPU-attached memory 1520 is accessed, causing the following operations. In at least one embodiment, local requests from GPU 1510 that find their pages in GPU bias are forwarded directly to corresponding GPU memory 1520. In at least one embodiment, local requests from GPU that find their pages in host bias are forwarded to processor 1505 (e.g., over a high-speed link as described herein). In at least one embodiment, requests from processor 1505 that find requested pages in host processor bias complete requests similar to normal memory reads. Alternatively, requests that point to GPU-biased pages can be forwarded to GPU 1510. In at least one embodiment, if GPU is not currently using a page, GPU can then migrate the page to host processor bias. In at least one embodiment, bias state of a page can be changed by software-based mechanisms, hardware-assisted software-based mechanisms, or in limited cases purely hardware-based mechanisms.

[0325] In at least one embodiment, a mechanism for changing bias state employs an API call (e.g., OpenCL) that then invokes a device driver of a GPU, which then sends a message (or causes a command descriptor to be enqueued) to the GPU, directing the GPU to change bias state, and in certain migrations, to perform a cache flush operation in the host. In at least one embodiment, a cache flush operation is used for a migration from host processor 1505 bias to GPU bias, but not for the reverse migration.

[0326] In at least one embodiment, cache coherency is maintained by temporarily rendering GPU-biased pages that host processor 1505 cannot cache. In at least one embodiment, to access these pages, processor 1505 can request access from GPU 1510, which can or can not grant access immediately. Thus, in at least one embodiment, to reduce communication between processor 1505 and GPU 1510, it is beneficial to ensure that GPU-biased pages are pages that are needed by GPU and not by host processor 1505, and vice versa.

[0327] One or more hardware structures 715 are used to perform one or more embodiments. Details regarding one or more hardware structures 715 can be found in this document in connection with Figure 7A and / or Figure 7B Details regarding one or more hardware structures 715 are provided.

[0328] Figure 16Exemplary integrated circuits and associated graphics processors according to various embodiments described herein are illustrated, which may be fabricated using one or more IP cores. In addition to the illustrations, at least one embodiment may include other logic and circuitry, including additional graphics processor / cores, peripheral interface controllers, or general-purpose processor cores.

[0329] Figure 16 This is a block diagram illustrating an exemplary system on a chip integrated circuit 1600 that can be fabricated using one or more IP cores according to at least one embodiment. In at least one embodiment, the integrated circuit 1600 includes one or more application processors 1605 (e.g., CPU), at least one graphics processor 1610, and may additionally include an image processor 1615 and / or a video processor 1620, any of which may be a modular IP core. In at least one embodiment, the integrated circuit 1600 includes peripheral or bus logic, which includes a USB controller 1625, a UART controller 1630, an SPI / SDIO controller 1635, and an I... 2 2S / I 2 2C controller 1640. In at least one embodiment, integrated circuit 1600 may include display device 1645 coupled to one or more of High Definition Multimedia Interface (HDMI) controller 1650 and Mobile Industrial Processor Interface (MIPI) display interface 1655. In at least one embodiment, storage may be provided by flash memory subsystem 1660, including flash memory and flash memory controller. In at least one embodiment, a memory interface may be provided via memory controller 1665 for accessing SDRAM or SRAM memory devices. In at least one embodiment, some integrated circuits also include embedded security engine 1670.

[0330] Inference and / or training logic 715 is used to perform inference and / or training operations associated with one or more embodiments. This document combines... Figure 7A and / or Figure 7B Details are provided regarding the inference and / or training logic 715. In at least one embodiment, the inference and / or training logic 715 may be used in the integrated circuit 1600 to infer or predict operations based at least in part on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.

[0331] In at least one embodiment, relative to Figure 16 At least one component shown or described is used to achieve the combination Figure 1-6 The described techniques and / or functions. In at least one embodiment, the inference and / or training logic 715 includes and / or operates regarding... Figure 1At least one aspect described (e.g., deep learning compiler 102, stream scheduler 110, memory allocator 112). In at least one embodiment, inference and / or training logic 715 uses a representation of a computer program to train at least one untrained or partially trained neural network, which representation of a computer program indicates operations and / or instructions that can be speculatively executed as described with respect to one or more of Figure 1-6 At least one aspect described (e.g., deep learning compiler 102, stream scheduler 110, memory allocator 112). In at least one embodiment, inference and / or training logic 715 uses a representation of a computer program to train at least one untrained or partially trained neural network, which representation of a computer program indicates operations and / or instructions that can be speculatively executed as described with respect to one or more of Figure 1-6 At least one aspect described (e.g., deep learning compiler 102, stream scheduler 110, memory allocator 112). In at least one embodiment, inference and / or training logic 715 uses a representation of a computer program to train at least one untrained or partially trained neural network, which representation of a computer program indicates operations and / or instructions that can be speculatively executed as described with respect to one or more of Figure 16 Integrated circuit 1600 of FIG. 16A is used to implement techniques and / or functionality described in association with Figure 1-6 FIG. 16A.

[0332] Figure 17A-17B Exemplary integrated circuits and associated graphics processors, which can be manufactured using one or more IP cores, are shown in accordance with various embodiments described herein. In addition to the illustrated, other logic and circuitry can be included, including additional graphics processors / cores, peripheral interface controllers or general-purpose processor cores.

[0333] Figure 17A-17B is a block diagram illustrating an exemplary graphics processor used within a SoC in accordance with the embodiments described herein. Figure 17A An exemplary graphics processor 1710 of a system on a chip integrated circuit, which can be manufactured using one or more IP cores, is shown in accordance with at least one embodiment. Figure 17B Another exemplary graphics processor 1740 of a system on a chip integrated circuit, which can be manufactured using one or more IP cores, is shown in accordance with at least one embodiment. In at least one embodiment, Figure 17A Graphics processor 1710 of FIG. 17A is a low power graphics processor core. In at least one embodiment, Figure 17B Graphics processor 1740 of FIG. 17B is a higher performance graphics processor core. In at least one embodiment, each graphics processor 1710, 1740 can be a variant of graphics processor 1610 of FIG. 16A. Figure 16 Graphics processor 1710 of FIG. 17A is a low power graphics processor core. In at least one embodiment,

[0334] In at least one embodiment, graphics processor 1710 includes a vertex processor 1705 and one or more fragment processor(s) 1715A-1715N (e.g., 1715A, 1715B, 1715C, 1715D through 1715N-1 and 1715N). In at least one embodiment, graphics processor 1710 can execute different shader programs via separate logic for vertex processing and / or for fragment / pixel processing. In at least one embodiment, vertex processor 1705 is optimized to execute operations for vertex shader programs, while one or more fragment processor(s) 1715A-1715N may be optimized to execute fragment or pixel shader programs. In at least one embodiment, vertex processor 1705 performs the vertex processing stage of a 3D graphics pipeline. In at least one embodiment, one or more fragment processor(s) 1715A-1715N use geometric and

[0335] In at least one embodiment, graphics processor 1710 additionally includes one or more memory management units (MMUs) 1720A-1720B, one or more cache(s) 1725A-1725B, and one or more circuit interconnects 1730A-1730B. In at least one embodiment, one or more MMU(s) 1720A-1720B provide for virtual to physical address mapping for graphics processor 1710, including for vertex processor 1705 and / or fragment processor(s) 1715A-1715N, which can reference vertex or image / texture data stored in memory, in addition to vertex or image / texture data stored in one or more cache(s) 1725A-1725B. In at least one embodiment, one or more MMU(s) 1720A-1720B can be synchronized with one or more MMUs within Figure 16 one or more application processor(s) 1605, image processors 1615, and / or video processors 1620 such that each processor 1605-1620 can participate in a shared or unified virtual memory system. In at least one embodiment, one or more circuit interconnects 1730A-1730B enable graphics processor 1710 to interface with other IP cores within a SoC, via an internal bus, or via a direct connection.

[0336] In at least one embodiment, the graphics processor 1740 includes one or more shader cores 1755A-1755N (e.g., 1755A, 1755B, 1755C, 1755D, 1755E, 1755F to 1755N-1 and 1755N), such as Figure 17B As shown, it provides a unified shader core architecture, where a single core or type or core can execute all types of programmable shader code, including shader program code for implementing vertex shaders, fragment shaders, and / or compute shaders. In at least one embodiment, the number of shader cores can vary. In at least one embodiment, the graphics processor 1740 includes an inter-core task manager 1745, which acts as a thread dispatcher to assign execution threads to one or more shader cores 1755A-1755N and a tile unit 1758 to accelerate tile-based rendering operations, where scene rendering operations are subdivided in image space, for example, to take advantage of local spatial consistency within the scene or optimize the use of internal caches.

[0337] Inference and / or training logic 715 is used to perform inference and / or training operations associated with one or more embodiments. This document combines... Figure 7A and / or Figure 7B Details are provided regarding the inference and / or training logic 715. In at least one embodiment, the inference and / or training logic 715 may be integrated into an integrated circuit. Figure 17A and / or Figure 17B The above is used for inference or prediction operations based at least in part on weight parameters calculated using neural network training operations, neural network functions or architectures, or neural network use cases described herein.

[0338] In at least one embodiment, relative to Figure 17A and / or Figure 17B At least one component shown or described is used to achieve the combination Figure 1-6 The described techniques and / or functions. In at least one embodiment, the inference and / or training logic 715 includes and / or operates regarding... Figure 1 At least one aspect described (e.g., deep learning compiler 102, stream scheduler 110, memory allocator 112). In at least one embodiment, inference and / or training logic 715 uses a representation of a computer program to train at least one untrained or partially trained neural network, the representation of which may be as described regarding Figure 1-6 One or more of the operations and / or instructions described in the description are speculatively performed. In at least one embodiment, the inference and / or training logic uses a representation of a computer program to perform at least one inference operation, the representation of which may indicate, as per [the description of the inference and / or training logic]... Figure 1-6One or more of the operations and / or instructions described as being performed speculatively. In at least one embodiment, Figure 17A Graphics processor 1710 and / or Figure 17B The 1740 graphics processor was used to implement the combination Figure 1-6 The described technologies and / or functions.

[0339] Figure 18A-18B Additional exemplary graphics processor logic according to embodiments described herein is illustrated. In at least one embodiment, Figure 18A It shows that it can be included in Figure 16 The graphics core 1800 within the graphics processor 1610, and in at least one embodiment, may be as follows: Figure 17B The unified shader cores shown are 1755A-1755N. Figure 18B A highly parallel general-purpose graphics processing unit (“GPGPU”) 1830 suitable for deployment on a multi-chip module is shown in at least one embodiment.

[0340] In at least one embodiment, the graphics core 1800 includes a shared instruction cache 1802, texture units 1818, and cache / shared memory 1820, which are common to the execution resources within the graphics core 1800. In at least one embodiment, the graphics core 1800 may include multiple slices 1801A-1801N or partitions of each core, and the graphics processor may include multiple instances of the graphics core 1800. In at least one embodiment, slices 1801A-1801N may include supporting logic, including local instruction caches 1804A-1804N, thread schedulers 1806A-1806N, thread dispatchers 1808A-1808N, and a set of registers 1810A-1810N. In at least one embodiment, slices 1801A-1801N may include a set of additional functional units (AFU 1812A-1812N), floating-point units (FPU 1814A-1814N), integer arithmetic logic units (ALU 1816A-1816N), address calculation units (ACU 1813A-1813N), double-precision floating-point units (DPFPU 1815A-1815N), and matrix processing units (MPU 1817A-1817N).

[0341] In at least one embodiment, the FPU 1814A-1814N can perform single-precision (32-bit) and half-precision (16-bit) floating-point operations, while the DPFPU 1815A-1815N performs double-precision (64-bit) floating-point operations. In at least one embodiment, the ALU 1816A-1816N can perform variable-precision integer operations with 8-bit, 16-bit, and 32-bit precision, and can be configured for mixed-precision operations. In at least one embodiment, the MPU 1817A-1817N can also be configured for mixed-precision matrix operations, including half-precision floating-point operations and 8-bit integer operations. In at least one embodiment, the MPU 1817-1817N can perform various matrix operations to accelerate machine learning application frameworks, including enabling support for accelerated generalized matrix-to-matrix multiplication (GEMM). In at least one embodiment, the AFU 1812A-1812N can perform additional logical operations not supported by floating-point or integer units, including trigonometric operations (e.g., sine, cosine, etc.).

[0342] Inference and / or training logic 715 is used to perform inference and / or training operations associated with one or more embodiments. This is combined with... Figure 7A and / or Figure 7B Details are provided regarding the inference and / or training logic 715. In at least one embodiment, the inference and / or training logic 715 may be used in the graphics core 1800 to infer or predict operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures or neural network use cases described herein.

[0343] In at least one embodiment, relative to Figure 18A At least one component shown or described is used to achieve the combination Figure 1-6 The described techniques and / or functions. In at least one embodiment, the inference and / or training logic 715 includes and / or operates regarding... Figure 1 At least one aspect described (e.g., deep learning compiler 102, stream scheduler 110, memory allocator 112). In at least one embodiment, inference and / or training logic 715 uses a representation of a computer program to train at least one untrained or partially trained neural network, the representation of which may be as described regarding Figure 1-6 One or more of the operations and / or instructions described in the description are speculatively performed. In at least one embodiment, the inference and / or training logic uses a representation of a computer program to perform at least one inference operation, the representation of which may indicate, as per [the description of the inference and / or training logic]... Figure 1-6 One or more of the operations and / or instructions described as being performed speculatively. In at least one embodiment, Figure 18A The 1800 graphics core is used to achieve integration.Figure 1-6 The described techniques and / or functions.

[0344] Figure 18B A general processing unit (GPGPU) 1830 is shown in at least one embodiment, which can be configured to enable highly parallel computing operations to be performed by a group of graphics processing units. In at least one embodiment, GPGPU 1830 can be linked directly to other instances of GPGPU 1830 to create a multi-GPU cluster to improve speed of training for deep neural networks. In at least one embodiment, GPGPU 1830 includes a host interface 1832 to enable connection to a host processor. In at least one embodiment, host interface 1832 is a PCI Express interface. In at least one embodiment, host interface 1832 can be a proprietary communication interface or communication structure to a fabric. In at least one embodiment, GPGPU 1830 receives commands from a host processor to perform processing operations associated with those commands using a global scheduler 1834 to assign execution threads associated with those commands to a group of compute clusters 1836A-1836H. In at least one embodiment, compute clusters 1836A-1836H share a cache memory 1838. In at least one embodiment, cache memory 1838 can be used as a higher level cache for cache memories within compute clusters 1836A-1836H.

[0345] In at least one embodiment, GPGPU 1830 includes memory 1844A-1844B coupled with compute clusters 1836A-1836H via a group of memory controllers 1842A-1842B. In at least one embodiment, memory 1844A-1844B can include various types of memory devices including dynamic random access memory (DRAM) or graphics random access memory, such as synchronous graphics random access memory (SGRAM) including graphics double data rate (GDDR) memory.

[0346] In at least one embodiment, compute clusters 1836A-1836H each include a group of graphics cores, such as graphics core 1800 of FIG. 18, which can include multiple types of integer and floating point logic units that can perform computing operations across a range of precisions, including precisions suitable for machine learning computations. For example, in at least one embodiment, at least a subset of floating point units in each compute cluster 1836A-1836H can be configured to perform 16-bit or 32-bit floating point operations, while a different subset of floating point units can be configured to perform 64-bit floating point operations. Figure 18A

[0347] ​In at least one embodiment, multiple instances of GPGPU 1830 can be configured to function as a compute cluster. In at least one embodiment, communication for synchronization and data exchange for compute clusters 1836A-1836H varies between embodiments. In at least one embodiment, multiple instances of GPGPU 1830 communicate via host interface 1832. In at least one embodiment, GPGPU 1830 includes an I / O hub 1839 that couples the GPGPU 1830 with a GPU link 1840 that enables a direct connection to other instances of GPGPU 1830. In at least one embodiment, GPU link 1840 is coupled to a specialized GPU-to-GPU bridge that enables communication and synchronization between multiple instances of GPGP 1830. In at least one embodiment, GPU link 1840 is coupled with a high-speed interconnect to transmit and receive data to other GPGPUs or parallel processors. In at least one embodiment, multiple instances of GPGPU 1830 are located in separate data processing systems and communicate via a network device accessible via host interface 1832. In at least one embodiment, GPU link 1840 can be configured to enable connection to a host processor in addition to or as an alternative to host interface 1832.

[0348] In at least one embodiment, GPGPU 1830 can be configured to train neural networks. In at least one embodiment, GPGPU 1830 can be used within an inferencing platform. In at least one embodiment, where GPGPU 1830 is used for inferencing, GPGPU 1830 can include fewer compute clusters 1836A-1836H relative to when GPGPU 1830 is used to train neural networks. In at least one embodiment, memory technology associated with memory 1844A-1844B can differ between inferencing and training configurations, with higher bandwidth memory technology dedicated to training configurations. In at least one embodiment, an inferencing configuration of GPGPU 1830 can support inferencing specific instructions. For example, in at least one embodiment, an inferencing configuration can provide support for one or more 8-bit integer dot product instructions that can be used during inferencing operations for deployed neural networks.

[0349] Inference and / or training logic 715 are used to perform inferencing and / or training operations associated with one or more embodiments. In at least one embodiment, inference and / or training logic 715 include at least one of hardware logic elements. Figure 7A and / or Figure 7BDetails are provided regarding the inference and / or training logic 715. In at least one embodiment, the inference and / or training logic 715 may be used in the GPGPU 1830 for inferring or predicting operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architecture, or neural network use cases described herein.

[0350] In at least one embodiment, relative to Figure 18B At least one component shown or described is used to achieve the combination Figure 1-6 The described techniques and / or functions. In at least one embodiment, the inference and / or training logic 715 includes and / or operates regarding... Figure 1 At least one aspect described (e.g., deep learning compiler 102, stream scheduler 110, memory allocator 112). In at least one embodiment, inference and / or training logic 715 uses a representation of a computer program to train at least one untrained or partially trained neural network, the representation of which may be as described regarding Figure 1-6 One or more of the operations and / or instructions described in the description are speculatively performed. In at least one embodiment, the inference and / or training logic uses a representation of a computer program to perform at least one inference operation, the representation of which may indicate, as per [the description of the inference and / or training logic]... Figure 1-6 One or more of the operations and / or instructions described as being performed speculatively. In at least one embodiment, Figure 18B The GPGPU 1830 is used to implement the combination Figure 1-6 The described technologies and / or functions.

[0351] Figure 19A block diagram of a computer system 1900 is shown in accordance with at least one embodiment. In at least one embodiment, computer system 1900 includes a processing subsystem 1901 with one or more processors 1902 and a system memory 1904 communicating via an interconnection path 1905, which can include a memory hub 1905. In at least one embodiment, memory hub 1905 can be a separate component coupled with one or more processors 1902 via communication links 1906 to perform memory access operations; alternatively, memory hub 1905 can be integrated into one or more processors 1902. In at least one embodiment, memory hub 1905 couples with an I / O subsystem 1911 via one or more communication links 1907 to perform I / O operations. In at least one embodiment, I / O subsystem 1911 can include an I / O hub 1907 that enables communication between one or more processors 1902 and one or more I / O devices 1908. In at least one embodiment, one or more I / O devices 1908 can include, without limitation, audio devices, network interfaces, wireless transmitters / receivers (e.g., Bluetooth, 3G, Wi-Fi, etc.), graphics processors, and storage devices (e.g., optical, semiconductor, and so forth).

[0352] In at least one embodiment, processing subsystem 1901 includes one or more parallel processor(s) 1912 coupled to memory hub 1905 via a bus or other communication link 1913. In at least one embodiment, communication link 1913 can use any one of a number of standard communication links, such as, but not limited to, a PCI Express, or can be a vendor specific communications interface or communications structure. In at least one embodiment, one or more parallel processor(s) 1912 form a programmable processing sub-system that can include a number of processing cores, each of which can be a single-core processor (individual cores) or a plurality of cores (multi-core processors). In at least one embodiment, one or more parallel processor(s) 1912 form a stream processor or stream processing sub-system that can include one or more units for generating, initializing, and / or storing data for processing operations. In at least one embodiment, one or more parallel processor(s) 1912 form a vector processing unit or vector processing sub-system that can include a number of processor cores that are each capable of processing a vector of data elements. In at least one embodiment, one or more parallel processor(s) 1912 can form a plurality of programmable processing units, each of which can include a parallel processor 1912.

[0353] In at least one embodiment, system storage unit 1914 can connect to I / O hub 1907 to provide storage mechanisms to be used in conjunction with computer system 1900. In at least one embodiment, I / O switches 1916 can be used to provide an interface mechanism to enable connections between I / O hub 1907 and other components, such as network adapter 1918 and / or wireless network adapter 1919 that can be integrated into a platform, as well as various other devices that can be added via one or more expansion devices 1920. In at least one embodiment, network adapter 1918 can be an Ethernet adapter or another wired network adapter. In at least one embodiment, wireless network adapter 1919 can include one or more of Wi-Fi, Bluetooth, Near Field Communication (NFC), or other network devices including one or more radios.

[0354] In at least one embodiment, computer system 1900 can include other components not explicitly shown, including USB or other port connections, optical storage drives, video capture devices, etc., which can also be connected to I / O hub 1907. In at least one embodiment, interconnection of the various components of computer system 1900 can be accomplished using any suitable protocols, including PCI- based protocols (e.g., PCI-Express), or other bus or point-to-point communication interfaces and / or protocols, such as NV-Link high-speed interconnect, or interconnect protocols. Figure 19

[0355] In at least one embodiment, parallel processor 1912 includes circuitry such as, for example, video circuitry, constituting a graphics processing unit (GPU). In at least one embodiment, parallel processor 1912 includes circuitry optimized for general use such as, for example, one or more core units with one or more vector registers and math units for general parallel processing in the form of, for example, a single instruction multiple data (SIMD) architecture or a multiple instruction multiple data (MIMD) architecture. In at least one embodiment, components of computer system 1900 can be integrated with one or more other system elements on a single integrated circuit. For example, in at least one embodiment, parallel processor 1912, memory hub 1905, processor 1902, and I / O hub 1907 can be integrated together into a system on a chip (SoC) integrated circuit. In at least one embodiment, components of computer system 1900 can be integrated together into a single package to form a system in a package (SIP) configuration. In at least one embodiment, at least a portion of components of computer system 1900 can be integrated into a multi-chip module (MCM), which can be interconnected with other multi-chip modules to form a modular computing platform.

[0356] Inference and / or training logic 715 are used to perform inferencing and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 715 are provided below in conjunction with FIGS. 7 and 8. In at least one embodiment, inference and / or training logic 715 is used in conjunction with components of computer system 1900, components of computing device 700, components of computing device 800, and / or components of accelerator 1000, to perform inferencing and / or training operations associated with at least one embodiment. Figure 7A and / or​Figure 7B Details regarding inference and / or training logic 715 are provided. In at least one embodiment, inference and / or training logic 715 can be used in systems 1900 to infer or predict operations based, at least in part, on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein. Figure 19 Details regarding inference and / or training logic 715 are provided. In at least one embodiment, inference and / or training logic 715 can be used in systems 1900 to infer or predict operations based, at least in part, on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.

[0357] In at least one embodiment, at least one component shown or described as being implemented via hardware (e.g., via a processor) can be implemented at least partially using a computer program that Figure 19 In at least one embodiment, at least one component shown or described as being implemented via hardware (e.g., via a processor) can be implemented at least partially using a computer program that Figure 1-6 In at least one embodiment, at least one component shown or described as being implemented via hardware (e.g., via a processor) can be implemented at least partially using a computer program that Figure 1 In at least one embodiment, at least one component shown or described as being implemented via hardware (e.g., via a processor) can be implemented at least partially using a computer program that Figure 1-6 In at least one embodiment, at least one component shown or described as being implemented via hardware (e.g., via a processor) can be implemented at least partially using a computer program that Figure 1-6 In at least one embodiment, at least one component shown or described as being implemented via hardware (e.g., via a processor) can be implemented at least partially using a computer program that Figure 19 In at least one embodiment, at least one component shown or described as being implemented via hardware (e.g., via a processor) can be implemented at least partially using a computer program that Figure 1-6 In at least one embodiment, at least one component shown or described as being implemented via hardware (e.g., via a processor) can be implemented at least partially using a computer program that

[0358] In at least one embodiment, at least one component shown or described as being implemented via hardware (e.g., via a processor) can be implemented at least partially using a computer program that

[0359] Figure 20A A parallel processor 2000, according to at least one embodiment, is shown. In at least one embodiment, various components of parallel processor 2000 can be implemented using one or more integrated circuits, which can be programmable integrated circuits, application-specific integrated circuits (ASICs), or field-programmable gate arrays (FPGAs). Figure 19 In at least one embodiment, parallel processor 2000 is a variant of one or more parallel processors 1912 shown and described herein.

[0360] In at least one embodiment, parallel processor 2000 includes a parallel processing unit 2002. In at least one embodiment, parallel processing unit 2002 includes an I / O unit 2004 that enables communication with other devices, including other instances of parallel processing unit 2002. In at least one embodiment, I / O unit 2004 can be directly connected to other devices. In at least one embodiment, I / O unit 2004 connects with other devices via use of a hub or switch, for example, memory hub 2005. In at least one embodiment, connections between memory hub 2005 and I / O unit 2004 form a communication link 2013. In at least one embodiment, I / O unit 2004 connects with a host interface 2006 and a memory crossbar switch 2016, where host interface 2006 receives commands directed to processing operations and memory crossbar switch 2016 receives commands directed to memory operations.

[0361] In at least one embodiment, when host interface 2006 receives a command buffer via I / O unit 2004, host interface 2006 can direct work operations to execute those commands to front end 2008. In at least one embodiment, front end 2008 couples with a scheduler 2010, which is configured to assign commands or other work items to a processing cluster array 2012. In at least one embodiment, scheduler 2010 ensures that processing cluster array 2012 is correctly configured and in an active state before tasks are assigned to processing cluster array 2012. In at least one embodiment, scheduler 2010 is implemented by firmware logic executing on a microcontroller. In at least one embodiment, microcontroller- implemented scheduler 2010 is configurable to perform complex scheduling and work allocation operations for fine and coarse grain parallelism to enable quick preemption and context switching for threads executing on processing array 2012. In at least one embodiment, host software can prove a workload for scheduling on processing array 2012 through one of multiple graphics processing paths. In at least one embodiment, workload can then be automatically allocated by scheduler 2010 logic within microcontroller including scheduler 2010 on processing array 2012.

[0362] In at least one embodiment, processing cluster array 2012 can include up to “N” processing clusters (e.g., cluster 2014A, cluster 2014B, through cluster 2014N), where “N” represents a positive integer. In at least one embodiment, each cluster 2014A-2014N of processing cluster array 2012 can execute a large number of concurrent threads. In at least one embodiment, scheduler 2010 can allocate work to clusters 2014A-2014N of processing cluster array 2012 using various scheduling and / or work distribution algorithms, which can be determined at least in part by a workload assigned to processing cluster array 2012. In at least one embodiment, scheduling can be handled dynamically by scheduler 2010, or can be aided in part by compiler logic during compilation of program logic configured for execution by processing cluster array 2012. In at least one embodiment, different clusters 2014A-2014N of processing cluster array 2012 can be allocated for processing different types of programs or for performing different types of computations.

[0363] In at least one embodiment, processing cluster array 2012 can be configured to perform a variety of types of parallel processing operations. In at least one embodiment, processing cluster array 2012 is configured to perform general-purpose parallel compute operations. For example, in at least one embodiment, processing cluster array 2012 can include logic to perform processing tasks including filtering of video and / or audio data, performing modeling operations, including physics operations, and performing data transformations.

[0364] In at least one embodiment, processing cluster array 2012 is configured to perform parallel graphics processing operations. In at least one embodiment, processing cluster array 2012 can include additional logic to support the performance of such graphics processing operations including, but not limited to, texture sampling logic to perform texture operations, tessellation logic, and other vertex processing logic. In at least one embodiment, processing cluster array 2012 can be configured to execute graphics processing related shader programs, including, but not limited to, vertex shaders, tessellation shaders, geometry shaders, and pixel shaders, in at least one embodiment, parallel processing unit 2002 can transfer data from system memory for processing via I / O unit 2004. In at least one embodiment, data can be stored to on-chip memory (e.g., parallel processor memory 2022) for processing during

[0365] In at least one embodiment, when parallel processing unit 2002 is used to perform graphics processing, scheduler 2010 can be configured to divide incoming workloads into tasks of approximately equal size to better enable distribution of graphics processing operations across multiple clusters 2014A-2014N of processing cluster array 2012. In at least one embodiment, portions of processing cluster array 2012 can be configured to perform different types of processing. For example, in at least one embodiment, a first portion can be configured to

[0366] In at least one embodiment, processing cluster array 2012 can receive processing tasks to be executed from scheduler 2010, which receives commands defining the processing tasks from front end 2008. In at least one embodiment, processing tasks can comprise indices of data to be processed, e.g., surface (patch) data, primitive data, vertex data, and / or pixel data, as well as state parameters and commands (e.g., what programs to execute) that control how the data is to be processed. In at least one embodiment, scheduler 2010 can be configured to fetch the indices corresponding to a task, or can receive the indices from front end 2008. In at least one embodiment, front end 2008 can be configured to ensure that processing cluster array 2012 is configured in an effective state before a workload specified by an incoming command buffer (e.g., a batch- buffer, a push buffer, etc.) is initiated.

[0367] In at least one embodiment, each of one or more instances of parallel processing unit 2002 can be coupled to a parallel processor memory 2022. In at least one embodiment, parallel processor memory 2022 can be accessed by the memory crossbar 2016, which can receive memory requests from the processing cluster array 2012 and I / O units 2004. In at least one embodiment, memory crossbar 2016 can access parallel processor memory 2022 via a memory interface 2018. In at least one embodiment, memory interface 2018 can include a number of memory partitions (e.g., memory partition 2020A, memory partition 2020B, to memory partition 2020N), which can each be coupled to a portion of parallel processor memory 2022 (e.g., a memory unit). In at least one embodiment, a number of memory partitions 2020A-2020N is configured to be equal to a number of memory units, such that first memory partition 2020A has a corresponding first memory unit 2024A, second memory partition 2020B has a corresponding memory unit 2024B, and Nth memory partition 2020N has a corresponding Nth memory unit 2024N. In at least one embodiment, a number of memory partitions 2020A-2020N can not be equal to a number of memory units.

[0368] In at least one embodiment, memory units 2024A-2024N can include various types of memory devices including dynamic random access memory (DRAM) or graphics random access memory, such as a synchronous graphics random access memory (SGRAM), including a graphics double data rate (GDDR) memory. In at least one embodiment, memory units 2024A-2024N can also include 3D stacked memory including, but not limited to, high bandwidth memory (HBM). In at least one embodiment, rendering targets such as frame buffers or texture maps can be stored across memory units 2024A-2024N allowing partition units 2020A-2020N to write portions of each rendering target in parallel to effectively use available bandwidth of parallel processor memory 2022. In at least one embodiment, local instances of parallel processor memory 2022 can be excluded to facilitate a unified memory design that utilizes system memory in combination with local cache memory.

[0369] In at least one embodiment, any of clusters 2014A-2014N of processing cluster array 2012 can process data that is to be written into any of memory locations 2024A-2024N within parallel processor memory 2022. In at least one embodiment, memory crossbar 2016 can be configured to transmit outputs of each cluster 2014A-2014N to any partition unit 2020A-2020N or another cluster 2014A-2014N, which can perform additional processing operations on the outputs. In at least one embodiment, each cluster 2014A-2014N can communicate with memory interface 2018 through memory crossbar 2016 to read from or write to various external memory devices. In at least one embodiment, memory crossbar 2016 has a connection to memory interface 2018 to communicate with I / O unit 2004, and a local instance of connection to parallel processor memory 2022, to enable processing

[0370] In at least one embodiment, multiple instances of parallel processor 2002 can be provided on a single add-in card, or multiple add-in cards can be interconnected. In at least one embodiment, different instances of parallel processor 2002 can be configured to operate together as a single parallel processor 2002, even if the different instances have different numbers of processing cores, different amounts of local parallel processor memory, and / or other configuration differences.

[0371] Figure 20B is a block diagram of a partition unit 2020, in accordance with at least one embodiment. In at least one embodiment, partition unit 2020 is a Figure 20Ainstance of one of the partition units 2020A-2020N of FIG. 20. In at least one embodiment, partition unit 2020 includes an L2 cache 2021, a frame buffer interface 2025, and a ROP 2026 (raster operations unit). In at least one embodiment, L2 cache 2021 is a write-back cache that is configured to perform a load / store operation on a request by a processor 1902. In at least one embodiment, L2 cache 2021 outputs read misses and urgent writes to frame buffer interface 2025 for processing. In at least one embodiment, updates are also sent to frame buffer via frame buffer interface 2025 for processing. In at least one embodiment, frame buffer interface 2025 interacts with a memory unit of memory 2024A-2024N (e.g., within parallel processor memory 2022) to perform load and store memory operations. Figure 20A

[0372] In at least one embodiment, ROP 2026 is a processing unit that performs raster operations including, for example, fill, line, ellipse, triangle, and / or the like. In at least one embodiment, ROP 2026 is used for any purpose such as screenshot generation, preview, or the like.

[0373] In at least one embodiment, ROP 2026 includes, without limitation, compression logic to compress depth or color data to be written to memory and decompress depth or color data fetched from memory. In at least one embodiment, compression logic can be lossless compression logic that makes use of one or more of multiple compression algorithms. In at least one embodiment, a type of compression performed by ROP 2026 can vary based on a statistical property of the data to be compressed. For example, in at least one embodiment, delta color compression is performed on a per tile basis for depth and color data. Figure 19 In at least one embodiment, ROP 2026 is included within each processing cluster (e.g., clusters 2014A-2014N of FIG. 20A) instead of in the partition unit 2020. In at least one embodiment, read and write requests for pixel data are transmitted over memory crossbar 2016 instead of pixel fragment data. In at least one embodiment, processed graphics data can be displayed on a display device 1910, routed to a Figure 20A

[0374] Figure 20C FIG. 21 is a block diagram of a processing cluster 2014 within a parallel processing unit in accordance with at least one embodiment. In at least one embodiment, processing cluster is a Figure 20A ​​one of the processing clusters 2014A-2014N. In at least one embodiment, processing cluster 2014 can be configured to perform a number of threads in parallel, where a “thread” is an instance of a particular program executing on a particular set of input data. In at least one embodiment, single-instruction-multiple-data (SIMD) instruction issue techniques are used to support parallel execution of a large number of threads with no or negligible specification impact. In at least one embodiment, single-instruction-multiple-thread (SIMT) techniques are used to support parallel execution of a large number of generally synchronous threads, using a common instruction unit configured to issue instructions to a set of processing engines within each processing cluster.

[0375] In at least one embodiment, operation of processing cluster 2014 can be controlled via a pipeline manager 2032 that is assigned to processing task(s) by scheduler 2010. In at least one embodiment, pipeline manager 2032 receives instructions from scheduler 2010 and manages execution of those instructions by graphics multiprocessor 2034 and / or texture unit 2036. In at least one embodiment, graphics multiprocessor 2034 is an exemplary instance of a SIMT parallel processor. However, in at least one embodiment, various types of SIMT parallel processors of differing architectures can be included within processing cluster 2014. In at least one embodiment, one or more instances of graphics multiprocessor 2034 can be included within processing cluster 2014. In at least one embodiment, graphics multiprocessor 2034 can process data, and data crossbar 2040 can be used to distribute processed data to one of a number of possible destinations, including other shader units. In at least one embodiment, pipeline manager 2032 can facilitate distribution by specifying destinations for processed data as a function of the destination’s source in either a fixed function or state settable by shader program instructions. Figure 20A

[0376] In at least one embodiment, each graphics multiprocessor 2034 within processing cluster 2014 can include an identical set of functional execution logic (e.g., arithmetic logic units, load store units, etc.). In at least one embodiment, functional execution logic can be configured in a pipelined manner in which new instructions can be issued before previous instructions are complete. In at least one embodiment, functional execution logic supports a variety of operations including integer and floating point arithmetic, comparison operations, Boolean operations, shift operations, and a multitude of algebraic functions. In at least one embodiment, same functional-unit hardware can be leveraged to perform different operations using different settings of control bits in those instructions. Any combination of

[0377] ​In at least one embodiment, instructions delivered to processing cluster 2014 constitute a thread. In at least one embodiment, a set of threads executing across a set of parallel processing engines is a thread group. In at least one embodiment, a thread group executes a general program on different input data. In at least one embodiment, each thread within a thread group can be assigned to a different processing engine within graphics multiprocessor 2034. In at least one embodiment, a thread group can include fewer threads than are present in a plurality of processing engines. In at least one embodiment, when a thread group includes fewer threads than the number of processing engines present in graphics multiprocessor 2034, one or more of the processing engines can be idle during the execution of the threads in the thread group. In at least one embodiment, a thread group can also include more threads than are present in a plurality of processing engines. In at least one embodiment, when a thread group includes more threads than the number of processing engines present in graphics multiprocessor 2034, multiple threads in a thread group can be executed concurrently on multiple processing engines.

[0378] In at least one embodiment, graphics multiprocessor 2034 includes internal cache memory to perform load and store operations. In at least one embodiment, graphics multiprocessor 2034 can bypass internal cache and use cache memory within processing cluster 2014 (e.g., Ll cache 2048). In at least one embodiment, each graphics multiprocessor 2034 can also have access to L2 Cache within a partition unit (e.g., partition units 2020A-2020N) that is shared among multiple processing clusters 2014 and can be used to transfer data between threads. In at least one embodiment, graphics multiprocessor 2034 can also have access to off-chip global memory, which can include one or more of local parallel processor memory and / or system memory. In at least one embodiment, any memory external to parallel processor 2002 can be used as global memory. In at least one embodiment, processing cluster 2014 includes multiple instances of graphics multiprocessor 2034 that share shared cache memory that can be stored in Ll cache 2048. Figure 20A

[0379] In at least one embodiment, each processing cluster 2014 can include a memory management unit (MMU) 2045 configured to map virtual addresses to physical addresses as is known in the art. In at least one embodiment, one or more instances of MMU 2045 can reside in Figure 20A ​In at least one embodiment, MMU 2045 includes a set of page table entries (PTEs) used to map virtual addresses into physical addresses for task computation and optionally into cache line indices for tasks. In at least one embodiment, MMU 2045 can include an address translation lookaside buffer (TLB) or cache, which can reside in graphics multiprocessor 2034 or within Ll cache 2048 or processing cluster 2014. In at least one embodiment, processing physical addresses enables data access locality to be determined for efficient request interleaving between partition units.

[0380] In at least one embodiment, processing cluster 2014 can be configured such that each graphics multiprocessor 2034 is coupled to a texture unit 2036 for performing texture mapping operations, e.g., determining texture sample positions, reading texture data, and filtering texture data. In at least one embodiment, texture data is read from an internal texture Ll cache (not shown) or from an Ll cache within graphics multiprocessor 2034 as needed, and texture data is fetched from an L2 cache, an L3 cache, a local parallel processor memory, or system memory, as needed. In at least one embodiment, each graphics multiprocessor 2034 outputs processed tasks to data crossbar 2040 in order to provide processed task data to another processing cluster 2014 for further processing or to store processed task data in an L2 cache, an L3 cache, a local parallel processor memory, or system memory via memory crossbar 2016. In at least one embodiment, preROP 2042 (pre-raster operations unit) is configured to receive data from graphics multiprocessor 2034, direct data to ROP unit, which can be located within partition unit (e.g., partition unit 2020A-2020N as described herein) or within graphics multiprocessor 2034, in at least one embodiment. In at least one embodiment, preROP 2042 can perform optimizations to minimize or eliminate bandwidth usage, organize pixel color data, and perform address translations. Figure 20A

[0381] Inference and / or training logic 715 are used to perform inferencing and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 715 are provided herein in conjunction with FIGS. 7A and / or 7B. In at least one embodiment, inference and / or training logic 715 can be used in graphics processing cluster 2014 for inferencing or predicting operations based, at least in part, on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein. Figure 7A Figure 7B Details regarding inference and / or training logic 715 are provided herein. In at least one embodiment, inference and / or training logic 715 can be used in graphics processing cluster 2014 to perform inferencing or predicting operations based, at least in part, on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.

[0382] ​​In at least one embodiment, at least one component shown or described with respect to Figure 20A , 20B and / or 20C is used to implement techniques and / or functionality described in connection with Figures 1-6 . In at least one embodiment, inference and / or training logic 715 includes and / or runs at least one aspect described with respect to Figure 1 (e.g., deep learning compiler 102, stream scheduler 110, memory allocator 112). In at least one embodiment, inference and / or training logic 715 uses a representation of a computer program to train at least one untrained or partially trained neural network, the representation of the computer program indicating operations and / or instructions that can be speculatively performed as described with respect to one or more of Figures 1-6 . In at least one embodiment, inference and / or training logic uses a representation of a computer program to perform at least one inferencing operation, the representation of the computer program indicating operations and / or instructions that can be speculatively performed as described with respect to one or more of Figures 1-6 . In at least one embodiment, parallel processor 2000 of FIG. 20A is used to implement techniques and / or functionality described in connection with Figures 1-6 .

[0383] Figure 20D A graphics processing unit 2034 according to at least one embodiment is shown. In at least one embodiment, graphics processing unit 2034 is coupled with pipeline manager 2032 of processing cluster 2014. In at least one embodiment, graphics processing unit 2034 has a thread execution pipeline that includes, without limitation, an instruction cache 2052, an instruction unit 2054, an address mapping unit 2056, a register file 2058, one or more general-purpose GPU cores (GPGPU cores) 2062, and one or more load / store units 2066. In at least one embodiment, GPGPU cores 2062 and load / store units 2066 are coupled with shared memory 2070 and cache memory 2072 via a memory and cache interconnect 2068.

[0384] In at least one embodiment, instruction cache 2052 receives a stream of instructions 2050 to be executed by graphics processing engine 2030 from pipeline manager 2032. In at least one embodiment, instructions 2050 are cached in instruction cache 2052 and dispatched for execution by instruction unit 2054. In at least one embodiment, instruction unit 2054 can dispatch instructions to threads assigned to different execution units within GPGPU cores 2062 as a thread group (e.g., warp). In at least one embodiment, instructions can be accessed from an unified address space within a single program by each execution unit when needed with addresses translated on the fly or just before the access, depending on modes supported by graphics processing engine 2030. In at least one embodiment, address translation unit 2056 can be used to translate addresses from a unified address space into different memory addresses that can be accessed by load / store units 2066.

[0385] In at least one embodiment, register file 2058 provides a set of registers for functional units of graphics processing engine 2034. In at least one embodiment, register file 2058 provides temporary storage for operands of the data

[0386] In at least one embodiment, GPGPU cores 2062 can each include floating point

[0387] In at least one embodiment, GPGPU cores 2062 include SIMD logic capable of performing a single instruction on multiple sets of data. In at least one embodiment, GPGPU cores 2062 can physically execute SIMD 4, SIMD 8, and SIMD 16 instructions and logically execute a SIMD 1, SIMD 2, and SIMD 32 instructions. In at least one embodiment, SIMD instructions for GPGPU cores can be generated by a shader compiler during compilation of code written by a programmer. In at least one embodiment, a programmer writing code for the programmable processing unit 2000 can write SIMD code in a high level programming language, which is then compiled into multiple instruction packets that can include one or more SIMD instructions and / or one or more SIMD control instructions. In at least one embodiment, a single

[0388] In at least one embodiment, memory and cache interconnect 2068 is an interconnect network that connects each functional unit of graphics multiprocessor 2034 to register file 2058 and shared memory 2070. In at least one embodiment, memory and cache interconnect 2068 is a crossbar interconnect that allows load / store units 2066 to effect load and store operations between shared memory 2070 and register file 2058. In at least one embodiment, register file 2058 can operate at the same frequency as GPGPU cores 2062, so that there is very little latency in transferring data between GPGPU cores 2062 and register file 2058. In at least one embodiment, shared memory 2070 can be used to enable communication between threads executing on functional units within graphics multiprocessor 2034. In at least one embodiment, cache memory 2072 can be used to cache data stored in shared memory 2070, for example, to allow data to be accessed more quickly. In at least one embodiment, shared memory 2070 can also be used to store program metadata, for example, to store information about programs that are executed on graphics multiprocessor 2034. In at least one embodiment, in addition to automatically cached data stored in cache memory 2072, threads executing on GPGPU cores 2062 can store data in shared memory in a programmed manner.

[0389] In at least one embodiment, a parallel processor or GPGPU, as described herein, is communicatively coupled to a host / processor core to accelerate graphics operations, machine learning operations, pattern analysis operations, and various general-purpose GPU (GPGPU) functions. In at least one embodiment, the GPU may be communicatively coupled to the host processor / core via a bus or other interconnect (e.g., high-speed interconnects such as PCIe or NVLink). In at least one embodiment, the GPU may be integrated with the core on a package or chip and communicatively coupled to the core via an internal processor bus / interconnect (i.e., within the package or chip). In at least one embodiment, regardless of how the GPU is connected, the processor core may assign work to the GPU in the form of a sequence of commands / instructions contained in a job descriptor. In at least one embodiment, the GPU then uses dedicated circuitry / logic to efficiently process these commands / instructions.

[0390] The inference and / or training logic 715 is used to perform inference and / or training operations associated with one or more embodiments. The following is in conjunction with... Figure 7A and / or Figure 7B Details are provided regarding the inference and / or training logic 715. In at least one embodiment, the inference and / or training logic 715 may be used in a graphics multiprocessor 2034 to perform inference or prediction operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures or neural network use cases described herein.

[0391] In at least one embodiment, relative to Figure 20D At least one component shown or described is used to achieve the combination Figures 1-6 The described techniques and / or functions. In at least one embodiment, the inference and / or training logic 715 includes and / or operates regarding... Figure 1 At least one aspect described (e.g., deep learning compiler 102, stream scheduler 110, memory allocator 112). In at least one embodiment, inference and / or training logic 715 uses a representation of a computer program to train at least one untrained or partially trained neural network, the representation of which may be as described regarding Figures 1-6 One or more of the operations and / or instructions described in the description are speculatively performed. In at least one embodiment, the inference and / or training logic uses a representation of a computer program to perform at least one inference operation, the representation of which may indicate, as per [the description of the inference and / or training logic]... Figures 1-6 One or more of the operations and / or instructions described as being performed speculatively. In at least one embodiment, Figure 20D The 2034 graphics multiprocessor was used to implement the combination Figures 1-6 The aforementioned technologies and / or functions.

[0392] Figure 21 A multi-GPU computing system 2100 is shown in accordance with at least one embodiment. In at least one embodiment, multi-GPU computing system 2100 can include a processor 2102 coupled to a plurality of general purpose graphics processing units (GPGPUs) 2106A-D via a host interface switch 2104. In at least one embodiment, host interface switch 2104 is a PCI Express switch device that couples processor 2102 to a PCI Express bus over which processor 2102 can communicate with GPGPUs 2106A-D. In at least one embodiment, GPGPUs 2106A-D can be interconnected via a set of high-speed P2P GPU-to-GPU links 2116. In at least one embodiment, GPU-to-GPU links 2116 connect to each of GPGPUs 2106A-D via a dedicated GPU link. In at least one embodiment, P2P GPU links 2116 enable direct communication between each GPGPU 2106A-D without having to communicate through host interface bus 2104 to which processor 2102 is connected. In at least one embodiment, where GPU-to-GPU traffic is directed to P2P GPU links 2116, host interface bus 2104 remains available for system memory access or communication with other instances of multi-GPU computing system 2100, e.g., via one or more network devices. While in at least one embodiment GPGPUs 2106A-D are connected to processor 2102 via host interface switch 2104, in at least one embodiment processor 2102 includes direct support for P2P GPU links 2116 and can be directly connected to GPGPUs 2106A-D.

[0393] Inference and / or training logic 715 are used to perform inferencing and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 715 are provided below in conjunction with FIGS. 7A and / or 7B. In at least one embodiment, inference and / or training logic 715 can be used in multi-GPU computing system 2100 for inferencing or predicting operations based, at least in part, on weight parameters calculated using neural network training operations, neural network functions, and / or architectures, or neural network use cases described herein. Figure 7A and / or Figure 7B Details regarding inference and / or training logic 715 are provided. In at least one embodiment, inference and / or training logic 715 can be used in multi-GPU computing system 2100 for inferencing or predicting operations based, at least in part, on weight parameters calculated using neural network training operations, neural network functions, and / or architectures, or neural network use cases described herein.

[0394] In at least one embodiment, at least one component shown or described is used to implement techniques and / or functionality described in conjunction with Figure 21 At least one component shown or described is used to implement techniques and / or functionality described in conjunction with Figures 1-6 In at least one embodiment, inference and / or training logic 715 includes and / or is used in place of one or more components of one or more embodiments described herein. In at least one embodiment, inference and / or training logic 715 includes and / or is used in place of one or more components of one or more embodiments described herein. Figure 1The described at least one aspect (e.g., deep learning compiler 102, stream scheduler 110, memory allocator 112). In at least one embodiment, inference and / or training logic 715 uses a representation of a computer program to train at least one untrained or partially trained neural network, which representation of a computer program indicates operations and / or instructions that can be executed speculatively as described with respect to one or more of Figures 1-6 In at least one embodiment, inference and / or training logic uses a representation of a computer program to perform at least one inferencing operation, which representation of a computer program indicates operations and / or instructions that can be executed speculatively as described with respect to one or more of Figures 1-6 In at least one embodiment, Figure 21 GPU computing system 2100 of FIG. 21 is used to implement techniques and / or functionality described in conjunction with, for example, Figures 1-6

[0395] Figure 22 is a block diagram of graphics processor 2200 according to at least one embodiment. In at least one embodiment, graphics processor 2200 includes ring interconnect 2202, front-end pipeline 2204, media engine 2237, and graphics cores 2280A-2280N. In at least one embodiment, ring interconnect 2202 couples graphics processor 2200 to other processing units including other graphics processors or one or more general-purpose processor cores. In at least one embodiment, graphics processor 2200 is one of many processors integrated within a multi-core processing system.

[0396] In at least one embodiment, graphics processor 2200 receives batches of commands via ring interconnect 2202. In at least one embodiment, incoming commands are interpreted by a command streamer 2203 within pipeline front-end 2204. In at least one embodiment, graphics processor 2200 includes scalable execution logic to perform 3D geometry processing and media processing via the graphics cores 2280A-2280N. In at least one embodiment, for 3D geometry processing commands, command streamer 2203 supplies commands to geometry pipeline 2236. In at least one embodiment, for at least some media processing commands, command streamer 2203 supplies commands to a video front end 2234, which couples with a media engine 2237. In at least one embodiment, media engine 2237 includes a video quality engine (VQE) 2230 for video and image post-processing, and a multi-format encode / decode (MFX) 2233 engine to ​

[0397] In at least one embodiment, graphics processor 2200 includes a scalable thread execution resource including a graphics core 2280A-2280N (which can be modular and sometimes referred to as a core slice) featuring multiple sub-cores 2250A-2250N, 2260A-2260N (sometimes referred to as a core sub-slice). In at least one embodiment, graphics processor 2200 can have any number of graphics cores 2280A. In at least one embodiment, graphics processor 2200 includes graphics core 2280A with at least a first sub-core 2250A and a second sub-core 2260A. In at least one embodiment, graphics processor 2200 is a low power processor with a single sub-core (e.g., 2250A). In at least one embodiment, graphics processor 2200 includes multiple graphics cores 2280A-2280N each including a set of first sub-cores 2250A-2250N and a set of second sub-cores 2260A-2260N. In at least one embodiment, each sub-core in first sub-cores 2250A-2250N includes at least a first set of execution units 2252A-2252N and media / texture samplers 2254A-2254N. In at least one embodiment, each sub-core in second sub-cores 2260A-2260N includes at least a second set of execution units 2262A-2262N and samplers 2264A-2264N. In at least one embodiment, each sub-core 2250A-2250N, 2260A-2260N shares a set of shared resources 2270A-2270N. In at least one embodiment, shared resources include shared cache memory and pixel operation logic.

[0398] Inference and / or training logic 715 are used to perform inferencing and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 715 are provided below in conjunction with FIGS. 7A and / or 7B. In examples in which inference and / or training logic 715 are used for inferencing, inference and / or training logic 715 can be used to perform inferencing operations, for example, on neural network data. Figure 7A and / or Figure 7B Details regarding inference and / or training logic 715 are provided below in conjunction with FIGS. 7A and / or 7B. In examples in which inference and / or training logic 715 are used for inferencing, inference and / or training logic 715 can be used to perform inferencing operations, for example, on neural network data.

[0399] In at least one embodiment, at least one component shown or described as being implemented within inference and / or training logic 715 is implemented external to inference and / or training logic 715. In examples in which inference and / or training logic 715 are used for inferencing, inference and / or training logic 715 can be used to perform inferencing operations, for example, on neural network data. Figure 22 At least one component shown or described as being implemented within inference and / or training logic 715 is implemented external to inference and / or training logic 715. In examples in which inference and / or training logic 715 are used for inferencing, inference and / or training logic 715 can be used to perform inferencing operations, for example, on neural network data. Figures 1-6 In at least one embodiment, inference and / or training logic 715 include and / or is used in conjunction with components of graphics processor 2200, for example, to perform operations associated with Figure 1The described at least one aspect (e.g., deep learning compiler 102, stream scheduler 110, memory allocator 112). In at least one embodiment, inference and / or training logic 715 uses a representation of a computer program to train at least one untrained or partially trained neural network, which representation of a computer program indicates operations and / or instructions that can be executed speculatively as described with respect to one or more of Figures 1-6 In at least one embodiment, inference and / or training logic uses a representation of a computer program to perform at least one inference operation, which representation of a computer program indicates operations and / or instructions that can be executed speculatively as described with respect to one or more of Figures 1-6 In at least one embodiment, Figure 22 Graphics processor 2200 of FIG. 21 is used to implement techniques and / or functionality described in conjunction with, for example, Figures 1-6

[0400] Figure 23 is a block diagram illustrating microarchitecture of a processor 2300 that can include logic circuitry to execute instructions, according to at least one embodiment. In at least one embodiment, processor 2300 can execute instructions including x86 instructions, ARM instructions, specialized instructions for application specific integrated circuits (ASICs), and the like. In at least one embodiment, processor 2300 can include registers to store packed data, such as 64-bit wide MMX TM Registers enabled in microprocessors employing MMX technology by Intel Corporation of Santa Clara, California, as an example. In at least one embodiment, MMX registers available in integer and floating point form can operate with packed data elements that accompany single instruction multiple data (“SIMD”) and streaming SIMD extensions (“SSE”) instructions. In at least one embodiment, 128-bit wide XMM registers related to SSE2, SSE3, SSE4, AVX, or higher (generically referred to as “SSEx”) technology can hold such packed data operands. In at least one embodiment, processor 2300 can execute instructions to accelerate machine learning or deep learning algorithms, training, or inference.

[0401] ​In at least one embodiment, processor 2300 includes an in-order front-end (“front-end”) 2301 to fetch instructions to execute and prepare instructions to use in pipeline stages later in processor. In at least one embodiment, front-end 2301 can include several units. In at least one embodiment, instruction prefetcher 2326 fetches instructions from memory and provides instructions to instruction decoder 2328 which in turn decodes or interprets instructions. In at least one embodiment, instruction decoder 2328 decodes a received instruction into one or more operations called “micro-instructions” or “micro-operations” (also called “micro ops” or “uops”) that machine can execute. In at least one embodiment, instruction decoder 2328 parses instruction into operation code and corresponding data and control fields, which can be used by micro-architecture to perform operations according to at least one embodiment. In at least one embodiment, trace cache 2330 can assemble decoded uops into program ordered sequences or traces in uop queue 2334 for execution. In at least one embodiment, when trace cache 2330 encounters complex instruction, microcode ROM 2332 provides uops needed to complete operation.

[0402] In at least one embodiment, some instructions can be converted into single micro-op, while others can need several micro-ops to complete the instruction. In at least one embodiment, if more than four uops are needed to complete an instruction, then instruction decoder 2328 can access microcode ROM 2332 to execute the instruction. In at least one embodiment, instructions can be decoded into small number of uops for processing at instruction decoder 2328. In at least one embodiment, if multiple uops are needed to complete the instruction, then instruction can be stored in microcode ROM 2332. In at least one embodiment, trace cache 2330 references entry point programmable logic array (“PLA”) to determine correct microcode pointer for reading microcode sequence from microcode ROM 2332 to complete one or more instructions according to at least one embodiment. In at least one embodiment, after microcode ROM 2332 completes sequencing of micro-ops for an instruction, front-end 2301 of machine can resume fetching micro-ops from trace cache 2330.

[0403] In at least one embodiment, out-of-order execution engine (“out-of-order engine”) 2303 can prepare instructions for execution. In at least one embodiment, out-of-order execution logic has multiple buffers to smooth and reorder instruction flow to optimize performance as instructions are pipelined down and dispatched for execution. In at least one embodiment, out-of-order execution engine 2303 includes, without limitation, an allocator / register renamer 2340, a memory micro instruction queue 2342, an integer / float micro instruction queue 2344, a memory scheduler 2346, a fast scheduler 2302, a slow / general floating point scheduler (“slow / general FP scheduler”) 2304, and a simple floating point scheduler (“simple FP scheduler”) 2306. In at least one embodiment, fast scheduler 2302, slow / general floating point scheduler 2304, and simple floating point scheduler 2306 are also collectively referred to as “micro instruction schedulers 2302, 2304, 2306.” In at least one embodiment, allocator / register renamer 2340 allocates machine buffers and resources needed by each micro instruction to execute in sequence. In at least one embodiment, allocator / register renamer 2340 renames logical registers to entries in a register file. In at least one embodiment, allocator / register renamer 2340 also allocates an entry for each micro instruction in one of two micro instruction queues, memory micro instruction queue 2342 for memory operations and integer / float micro instruction queue 2344 for non-memory operations, in front of memory scheduler 2346 and micro instruction schedulers 2302, 2304, 2306. In at least one embodiment, micro instruction schedulers 2302, 2304, 2306 determine when micro instructions are ready to execute based on readiness of their dependent input register operand sources and availability of execution resource micro instructions needed to complete. In at least one embodiment, fast scheduler 2302 can schedule on every half of a main clock cycle, while slow / general floating point scheduler 2304 and simple floating point scheduler 2306 can schedule once per main processor clock cycle. In at least one embodiment, micro instruction schedulers 2302, 2304, 2306 arbitrate for a dispatch port to dispatch micro instructions for execution.

[0404] In at least one embodiment, execution block 2311 includes, without limitation, an integer register file / bypass network 2308, a floating point register file / bypass network (“FP register file / bypass network”) 2310, address generation units (“AGUs”) 2312 and 2314, fast arithmetic logic units (“fast ALUs”) 2316 and 2318, slow arithmetic logic unit (“slow ALU”) 2320, floating point ALU (“FP”) 2322, and floating point move unit (“FP move”) 2324. In at least one embodiment, integer register file / bypass network 2308 and floating point register file / bypass network 2310 are also referred to herein as “register files 2308, 2310.” In at least one embodiment, AGUs 2312 and 2314, fast ALUs 2316 and 2318, slow ALU 2320, floating point ALU 2322, and floating point move unit 2324 are also referred to herein as “execution units 2312, 2314, 2316, 2318, 2320, 2322, and 2324.” In at least one embodiment, execution block 2311 can include, without limitation, any number (including zero) and type of register files, bypass networks, address generation units, and execution units (in any combination).

[0405] In at least one embodiment, register networks 2308, 2310 can be arranged between micro-instruction schedulers 2302, 2304, 2306 and execution units 2312, 2314, 2316, 2318, 2320, 2322, and 2324. In at least one embodiment, integer register file / bypass network 2308 performs integer operations. In at least one embodiment, floating point register file / bypass network 2310 performs floating point operations. In at least one embodiment, each of register networks 2308, 2310 can include, without limitation, a bypass network that can bypass or forward a just-completed result that has not yet been written into a register file to a new dependee. In at least one embodiment, register networks 2308, 2310 can communicate data with each other. In at least one embodiment, integer register file / bypass network 2308 can include, without limitation, two separate register files, one for low order 32 bits data, a second for high order 32 bits data. In at least one embodiment, floating point register file / bypass network 2310 can include, without limitation, 128 bit-wi...

Claims

1. A processor, comprising: One or more circuits for executing a compiler, the compiler being used to speculatively execute graphics processing unit (GPU) code based at least in part on whether two or more central processing unit (CPU) code branch conditions are related to each other.

2. The processor of claim 1, wherein the instructions of the GPU code have been identified by the compiler, at least in part, as being to be executed in parallel speculatively based on the recognition of copy operations, and the one or more circuits are configured to execute the instructions, at least in part, based on receiving commands from another processor.

3. The processor of claim 1, wherein the instructions of the GPU code have been identified by the compiler, at least in part, as being presumably to be executed in parallel based on identifying copy operations between the parallel processing unit and the host computer system and marking safe operations after one or more identified copy operations.

4. The processor of claim 1, wherein the instructions of the GPU code include extended active periods of variables used by operations associated with instructions identified as to be speculatively executed in parallel.

5. The processor of claim 1, wherein the processor is part of a parallel processing unit, and the one or more circuits are configured to execute instructions of the GPU code after receiving a kernel boot command from a host computer system.

6. The processor of claim 1, wherein the instructions of the GPU code are part of a loop.

7. The processor of claim 1, wherein the instructions of the GPU code implement a portion of the inference operation using a recurrent neural network.

8. A system comprising: One or more processors are configured to execute a compiler, wherein the compiler is configured to speculatively execute graphics processing unit (GPU) code based at least in part on whether two or more CPU code branch conditions are related to each other; and One or more memories for storing one or more instructions of the GPU code.

9. The system of claim 8, wherein the instructions have been identified by the compiler, at least in part, as being presumably to be executed in parallel based on the recognition of copy operations from the parallel processing unit to the host computer system.

10. The system of claim 8, wherein the instructions have been identified by the compiler, at least in part, as being to be executed in parallel speculatively, based on the discovery of one or more conditional branches in the representation of a computer program using a neural network.

11. The system of claim 8, wherein the one or more processors are first one or more processors, and the system further comprises a second one or more processors for initiating the one or more instructions to be executed by the first one or more processors.

12. The system of claim 8, wherein the one or more processors are first one or more processors, the system further comprising a second one or more processors for initiating the one or more instructions to be executed by the first one or more processors, and the second one or more processors for stopping speculative initiation instructions in response to receiving a value of a condition prior to receiving the one or more instructions in a representation of a computer program via a copy operation.

13. The system of claim 8, wherein the instructions have been identified by the compiler, at least in part, as safe to perform based on flags, and are thus identified as to be executed speculatively in parallel.

14. The system of claim 8, wherein the instructions have been identified by the compiler, at least in part, as to be executed in parallel speculatively, based on searching for copy operations in the representation of the computer program and identifying operations following the copy operations that are speculatively safe to execute.

15. The system of claim 8, wherein the instruction is part of a loop that implements a portion of the inference operation using a neural network.

16. A method comprising: An execution compiler is used to speculatively execute one or more instructions of a graphics processing unit (GPU) code, based at least in part on whether two or more CPU code branch conditions are related to each other.

17. The method of claim 16, wherein the instructions have been identified by the compiler, at least in part, as operations that do not alter a random state, rewrite output, use signal instructions, or use wait instructions, as being speculatively to be executed in parallel.

18. The method of claim 16, wherein the instructions have been identified by the compiler, at least in part, as being to be executed in parallel speculatively, based on identifying conditional branches and selecting one path from a plurality of paths following the conditional branches.

19. The method of claim 16, wherein the instructions have been identified by the compiler, at least in part, as being presumably to be executed in parallel based on the recognition of copy operations.

20. The method of claim 16, wherein the instruction includes an extended active period of variables used in the speculatively performed operation.

21. The method of claim 16, wherein the instructions have been identified by the compiler, at least in part, as being to be executed in parallel speculatively based on the recognition of copy operations, and wherein the instructions utilize a neural network to implement a portion of the inference operation.

22. A machine-readable medium having a set of instructions stored thereon, which, if executed by one or more processors, causes the one or more processors to at least: Execute the compiler, where, The compiler is used to speculatively execute graphics processing unit (GPU) code based at least in part on whether two or more CPU code branch conditions are related to each other.

23. The machine-readable medium of claim 22, wherein if the set of instructions is executed by the one or more processors, the one or more processors further identify, at least in part, the GPU code as to be presumably executed in parallel based on recognizing copy operations between parallel processing units and a host computer system in the representation of the computer program.

24. The machine-readable medium of claim 22, wherein if the set of instructions is executed by the one or more processors, the one or more processors further enable the one or more processors to at least recognize a safe operation following a copy operation.

25. The machine-readable medium of claim 22, wherein if the set of instructions is executed by the one or more processors, the one or more processors are further made to mark at least speculatively perform an operation that is safe to perform.

26. The machine-readable medium of claim 22, wherein if the set of instructions is executed by the one or more processors, the one or more processors further cause the one or more processors to at least mark operations that are speculatively safe to perform and extend the active period of variables associated with the operations marked as speculatively safe to perform.

27. The machine-readable medium of claim 22, wherein if the set of instructions is executed by the one or more processors, the one or more processors further cause the one or more processors to at least: search in the representation of the computer program for a copy operation between the graphics processing unit and the host computer system; and identify an operation following the copy operation that is safe for speculative execution.

28. The machine-readable medium of claim 22, wherein if the set of instructions is executed by the one or more processors, the one or more processors further cause the one or more processors to at least extend the active period of variables associated with operations identified as to be speculatively executed in parallel.

29. The machine-readable medium of claim 22, wherein if the set of instructions is executed by the one or more processors, the one or more processors further cause the one or more processors to at least: find a conditional branch in the representation of the computer program based in part on identifying a copy operation in the representation of the computer program; select a path from a plurality of paths following the conditional branch; and identify instructions in the selected path that are safe for speculative execution.

30. A vehicle comprising: A computer vision system comprising one or more processors for identifying one or more trajectories of corresponding one or more objects by performing one or more inference operations, at least in part, based on causing one or more graphics processing units (GPUs) to use a representation of a computer program, the representation of the computer program comprising one or more instructions of GPU code, the GPU code being speculatively executed based at least in part on whether branch conditions of two or more central processing units (CPUs) are related to each other. as well as One or more of a propulsion system, a steering control system, and a vehicle operator notification system, for performing one or more actions based at least in part on one or more identified trajectories.

31. The vehicle of claim 30, wherein the one or more processors comprise one or more first processors of a host computer system and one or more second processors of the one or more GPUs, wherein the one or more second processors are configured to speculatively execute the one or more instructions of the GPU code based at least in part on receiving from the host computer system a command to initiate a kernel containing one or more instructions on the one or more GPUs.

32. The vehicle of claim 30, wherein one or more instructions of the GPU code have been identified by the compiler, at least in part, as being presumptively executed based on the recognition of copy operations.

33. The vehicle of claim 30, wherein one or more instructions of the GPU code have been identified by the compiler, at least in part, as being to be executed speculatively based on flag-safe operations.

34. The vehicle of claim 30, wherein one or more instructions of the GPU code include extended active periods of variables used by operations associated with instructions identified as to be speculatively executed in parallel.

35. The vehicle of claim 30, wherein one or more instructions of the GPU code implement a portion of the inference operation using a recurrent neural network.

Citation Information

Patent Citations

  • Prediction-based system and method for trajectory planning of autonomous vehicles

    CN111373458A