Optimizing compilation of program code

By identifying and omitting compilation passes that will not improve the program and utilizing the profile optimization mechanism, compiler optimization technology improves resource utilization efficiency and reduces compilation time, solving the resource and time deficiencies of existing compiler optimization technology and improving the performance of executable instructions.

CN120832148APending Publication Date: 2025-10-24NVIDIA CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510491961.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-04-18
Filing Date
2025-04-18
Publication Date
2025-10-24

AI Technical Summary

Technical Problem

Existing compiler optimization techniques have room for improvement in improving the performance of executable instructions, especially in terms of resource utilization efficiency and compilation time.

Method used

The compiler identifies compilation passes that will not improve the program and omits these compilation passes in subsequent compilations. Combined with a profile-based optimization mechanism, optimizations are selectively applied to reduce unnecessary resource consumption and compilation time.

Benefits of technology

Improves the compiler's resource utilization efficiency, reduces compilation time, and maintains or improves the performance of executable programs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120832148A_ABST
    Figure CN120832148A_ABST
Patent Text Reader

Abstract

The invention relates to optimizing compilation of program code. Apparatuses, systems, and techniques for selecting optimizations to be performed by a compiler. In at least one embodiment, a processor comprises one or more circuits, the one or more circuitry is to execute the compiler to select the one or more optimizations to the one or more first versions of the program based at least in part on a result of performing the one or more optimizations to the one or more second versions of the program.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] At least one embodiment is directed to performing compiler optimizations on program code.For example, at least one embodiment is directed to a processor or circuit that performs compiler optimizations. Background Art

[0002] Performing computational operations can use significant amounts of memory, time, or computing resources. Application programs are typically generated from source code, which is then compiled into executable instructions using a compiler. To improve the performance of the resulting executable instructions, the compiler may apply optimizations that remove redundant operations or combine or rearrange operations so that they can be executed more efficiently by the target processor. However, the techniques for applying these optimizations can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0003] FIG. 1 An example system for an optimizing compiler according to at least one embodiment is shown;

[0004] FIG. 2 An example system for optimizing a compiler based on modifications to intermediate code through optimization passes is shown in accordance with at least one embodiment;

[0005] FIG. 3 shows an example of profile-based optimization of a compilation pass in accordance with at least one embodiment;

[0006] FIG. 4 illustrates an example of profile-based optimization of compilation passes at the module level in accordance with at least one embodiment;

[0007] FIG. 5 is a flow chart of a technique for optimizing compiler optimization passes according to at least one embodiment;

[0008] FIG. 6A An example of a system including a driver and / or runtime including one or more libraries to provide one or more application programming interfaces (APIs) according to at least one embodiment is shown;

[0009] FIG. 6B is a block diagram illustrating an example of a processor and modules according to at least one embodiment;

[0010] FIG. 7 An example data center system is shown in accordance with at least one embodiment;

[0011] FIG. 8A An example of an autonomous vehicle according to at least one embodiment is shown;

[0012] FIG. 8B An example of a camera position and field of view of an autonomous vehicle is shown in accordance with at least one embodiment FIG. 8A An example of a camera position and field of view of an autonomous vehicle is shown in accordance with at least one embodiment

[0013] FIG. 8C A block diagram illustrating an example system architecture of an autonomous vehicle is shown in accordance with at least one embodiment FIG. 8A A block diagram illustrating an example system architecture of an autonomous vehicle is shown in accordance with at least one embodiment

[0014] FIG. 8D A diagram illustrating a system for communication between one or more cloud-based servers and an autonomous vehicle is shown in accordance with at least one embodiment FIG. 8A A diagram illustrating a system for communication between one or more cloud-based servers and an autonomous vehicle is shown in accordance with at least one embodiment

[0015] FIG. 9 A block diagram illustrating a computer system is shown in accordance with at least one embodiment

[0016] FIG. 10 A block diagram illustrating a computer system is shown in accordance with at least one embodiment

[0017] FIG. 11 A computer system is shown in accordance with at least one embodiment

[0018] FIG. 12 A computer system is shown in accordance with at least one embodiment

[0019] FIG. 13A A computer system is shown in accordance with at least one embodiment

[0020] FIG. 13B A computer system is shown in accordance with at least one embodiment

[0021] FIG. 13C A computer system is shown in accordance with at least one embodiment

[0022] FIG. 13D A computer system is shown in accordance with at least one embodiment

[0023] FIG. 13E A computer system is shown in accordance with at least one embodiment FIG. 13F A computer system is shown in accordance with at least one embodiment

[0024] FIG. 14 An exemplary integrated circuit and associated graphics processor are shown in accordance with at least one embodiment

[0025] FIG. 15A An exemplary integrated circuit and associated graphics processor are shown in accordance with at least one embodiment FIG. 15B An exemplary integrated circuit and associated graphics processor are shown in accordance with at least one embodiment

[0026] FIG. 16A An exemplary integrated circuit and associated graphics processor are shown in accordance with at least one embodiment FIG. 16BAn additional exemplary graphics processor logic is shown in accordance with at least one embodiment;

[0027] FIG. 17 A computer system is shown in accordance with at least one embodiment;

[0028] FIG. 18A A parallel processor is shown in accordance with at least one embodiment;

[0029] FIG. 18B A partition unit is shown in accordance with at least one embodiment;

[0030] FIG. 18C A processing cluster is shown in accordance with at least one embodiment;

[0031] FIG. 18D A graphics multiprocessor is shown in accordance with at least one embodiment;

[0032] FIG. 19 A multi-GPU system is shown in accordance with at least one embodiment;

[0033] FIG. 20 A graphics processor is shown in accordance with at least one embodiment;

[0034] FIG. 21 is a block diagram showing a processor micro-architecture for a processor in accordance with at least one embodiment;

[0035] FIG. 22 At least portions of a graphics processor are shown in accordance with one or more embodiments;

[0036] FIG. 23 At least portions of a graphics processor are shown in accordance with one or more embodiments;

[0037] FIG. 24 At least portions of a graphics processor are shown in accordance with one or more embodiments;

[0038] FIG. 25 is a block diagram of a graphics processing engine of a graphics processor in accordance with at least one embodiment;

[0039] FIG. 26 is a block diagram of at least portions of a graphics processor core in accordance with at least one embodiment;

[0040] FIG. 27A and FIG. 27B Thread execution logic including an array of processing elements of a graphics processor core is shown in accordance with at least one embodiment;

[0041] FIG. 28 A parallel processing unit ("PPU") is shown in accordance with at least one embodiment;

[0042] FIG. 29 A general processing cluster (“GPC”) is shown in accordance with at least one embodiment;

[0043] FIG. 30 A memory partition unit of a parallel processing unit (“PPU”) is shown in accordance with at least one embodiment;

[0044] FIG. 31 A streaming multiprocessor is shown in accordance with at least one embodiment;

[0045] FIG. 32 A network for communicating data within a 5G wireless communication network is shown in accordance with at least one embodiment;

[0046] FIG. 33 A network architecture for a 5G LTE wireless network is shown in accordance with at least one embodiment;

[0047] FIG. 34 A diagram showing some basic functions of a mobile telecommunications network / system operating according to LTE and 5G principles in accordance with at least one embodiment;

[0048] FIG. 35 A radio access network that can be part of a 5G network architecture is shown in accordance with at least one embodiment;

[0049] FIG. 36 An example diagram of a 5G mobile communications system in which multiple different types of devices are used is provided in accordance with at least one embodiment;

[0050] FIG. 37 An example high-level system is shown in accordance with at least one embodiment;

[0051] FIG. 38 An architecture of a network system is shown in accordance with at least one embodiment;

[0052] FIG. 39 Example components of a device are shown in accordance with at least one embodiment;

[0053] FIG. 40 Example interfaces of baseband circuitry are shown in accordance with at least one embodiment;

[0054] FIG. 41 Examples of uplink channels are shown in accordance with at least one embodiment;

[0055] FIG. 42 An architecture of a network system is shown in accordance with at least one embodiment;

[0056] FIG. 43 A control plane protocol stack is shown in accordance with at least one embodiment;

[0057] FIG. 44 A user plane protocol stack is shown in accordance with at least one embodiment;

[0058] FIG. 45 Components of a core network are shown in accordance with at least one embodiment;

[0059] FIG. 46 Components of a system that supports network function virtualization (NFV) are shown in accordance with at least one embodiment; and

[0060] FIG. 47 Components of a system that accesses a large language model are shown in accordance with at least one embodiment. DETAILED DESCRIPTION

[0061] In at least one embodiment, systems and methods implemented in accordance with the present disclosure are used to cause one or more circuits to execute a compiler to select one or more optimizations to one or more first versions of a program based at least in part on results of performing one or more optimizations to one or more second versions of the program.

[0062] In at least one embodiment, a compiler is a computer program that compiles source code to form an executable program, the source code being a representation of a given computer program in a programming language. In at least one embodiment, an executable program is a representation of a given computer program as processor instructions to be executed by a processor. In at least one embodiment, the compiler performs compilation of a program. In at least one embodiment, compilation of a program includes transforming source code of the program into an intermediate representation (IR), the intermediate representation being a data structure representing data objects and / or operations specified in the source code.

[0063] In at least one embodiment, during compilation of a program, the compiler performs one or more code optimizations (“optimizations”) on one or more IRs of the program. In at least one embodiment, each optimization transforms input IR into functionally equivalent output IR. In at least one embodiment, the compiler provides the output IR to a subsequent optimization, or, if no further optimizations are performed, to a target code generator that transforms the output IR into an executable program.

[0064] In at least one embodiment, code optimizations can improve an input IR by reducing resource usage of the input IR, such as CPU or memory usage of the input IR. In at least one embodiment, code optimizations transform an input IR into a functionally equivalent “output” IR having one or more improved characteristics. In at least one embodiment, a variety of different optimizations can be performed on an IR, and different optimizations improve an input IR to different degrees. In at least one embodiment, an optimization that improves an input IR generates an output IR having one or more differences from the input IR. In at least one embodiment, an optimization does not improve an input IR, e.g., because the input IR does not contain any statements to which the optimization applies, the input IR is too complex to optimize, or the optimization cannot improve the input IR for other reasons. In at least one embodiment, an optimization that does not improve an input IR generates an output IR that has no differences from the input IR.

[0065] In at least one embodiment, optimizations that do not change an IR of a program can be omitted from subsequent compilations of the program because the optimizations do not improve the program. In at least one embodiment, optimizations consume significant amounts of time and processor resources, so by omitting optimizations that do not change an IR, compilation time in subsequent compilations can be reduced. In at least one embodiment, a compiler identifies one or more optimizations that change an IR of a program during a compilation, and in subsequent compilations of the program, the one or more of the optimizations are performed without performing optimizations that do not change the IR of the program. In at least one embodiment, the subsequent compilations can be of a same version of the program or of a different version of the program. In at least one embodiment, a second version of the program can include similar statements or instructions as a first version, so optimizations that apply to the first version can also apply to the second version.

[0066] In at least one embodiment, a compiler performs configuration file based compile pass optimization to identify compile passes that do not improve a program being compiled and omit the identified compile passes in subsequent compilations of the program or other versions of the program. In at least one embodiment, the compiler generates a configuration file for the program and includes information in the configuration file identifying one or more optimizations that changed (and / or did not change) an IR of the program during compilation. In at least one embodiment, to perform a subsequent compilation of the program, the compiler retrieves information identifying one or more optimizations that changed (and / or did not change) an IR of the program and performs the one or more optimizations on the IR of the program without performing optimizations that did not change the IR of the program. In at least one embodiment, the compiler identifies one or more optimizations that changed an IR of a module (or other portion of the program) and, at a subsequent compilation of the program, performs the one or more optimizations on the IR but does not perform other optimizations on the module (or other portion of the program) because the IR was not changed by the other optimizations in a previous compilation for that module (or other portion of the program).

[0067] FIG. 1An example system 100 for optimizing a compiler 102 is shown, in accordance with at least one embodiment. In at least one embodiment, system 100 is used to execute compiler 102 to select one or more optimizations to one or more first versions of a program based, at least in part, on results of performing the one or more optimizations to one or more second versions of the program. In at least one embodiment, one or more first versions of a program include source code 104, IR (initial IR 118, input IR 120, output IR 124), and / or one or more executable programs 150. In at least one embodiment, one or more second versions of a program include source code 104, IR (initial IR 118, input IR 120, output IR 124), and / or one or more executable programs 150. In at least one embodiment, the program refers to a computer program that specifies operations of a program to be performed by one or more processors 160 of a computer system. In at least one embodiment, the program is also referred to as a computer program and is a set of instructions that, if executed, cause one or more processors 160 to perform one or more computing operations. In at least one embodiment, the computer program has two or more representations that specify the operations of the program and the representations include source code 104, input IR 120, output IR 124, and / or executable program 150. In at least one embodiment, source code 104, input IR 120, and executable program 150 are computer programs, for example. In at least one embodiment, source code 104 is a representation of the computer program in a programming language.

[0068] In at least one embodiment, system 100 includes one or more computing devices or systems (e.g., one or more servers). In at least one embodiment, processor 160 accesses memory 162, for example, to store data in and / or receive data from the memory 162. In at least one embodiment, the memory can be one or more non-transitory processor-readable media. In at least one embodiment, processor 160 and / or memory 162 are components of a computing device, such as a server.

[0069] In at least one embodiment, compiler 102 is a set of software instructions that, if executed, cause one or more processors 160 to generate one or more executable programs 150 based at least in part on source code 104. In at least one embodiment, compiler 102 is a computer program. In at least one embodiment, compiler 102 receives source code 104 to compile from one computing language to another computing language to generate executable program 150. In at least one embodiment, compiler 102 receives source code 104 to compile from one computing language to another computing language to generate output. In at least one embodiment, this output is received as input by a linker (not shown). In at least one embodiment, the linker creates executable program 150. In at least one embodiment, source code 104 is one or more instructions and / or other commands that will be compiled or otherwise assembled into executable program 150. In at least one embodiment, source code 104 is received by processor 160 to be read and used to generate said executable program 150 specific to processor 160. In at least one embodiment, source code 104 includes instructions in any programming language and / or instruction set, such as an instruction set of processor 160. In at least one embodiment, one or more instructions include any processor instructions or programming language statements.

[0070] In at least one embodiment, compiler 102 is a set of instructions which, if executed, cause one or more processors 160 to generate one or more outputs, such as object code to be input to a linker (not shown), an output intermediate representation (output IR) 124 of code to be additionally compiled, such as by a just-in-time compiler, and / or an executable program 150 to be executed by one or more processors 160. In at least one embodiment, compiler 102 generates an output by converting one or more inputs in one format, such as source code 104, into one or more outputs in another format, such as executable code. In at least one embodiment, compiler 102 is, by way of example, one or more of a traditional compiler (e.g., C, C++, or Pascal), an interpreter (e.g., LISP, SNOBOL, or Java 2.0), a cross compiler, an incremental compiler, a translator (e.g., COBOL to C++), a just-in-time (JIT) compiler (e.g., Java, Microsoft.NET), a single-pass compiler, a multi-pass compiler, an ahead-of-time (AOT) compiler (e.g.,.NET ngen), or a binary compiler, or any other compiler described further herein. In at least one embodiment, compiler 102 is, by way of example, of a programming language or a variant thereof that it receives as source code 104 or outputs as executable program 150 is Python, JavaScript, Java, C#, C, C++, GO, R, Swift, PHP, Dart, Kotlin, MATLAB, Perl, Ruby, Rust, or Scala, or any other programming language. In at least one embodiment, computer program of compiler 102 includes one or more of a parser 106, an intermediate code generator 108, and a configuration file-based optimizer (“optimizer”) 110.

[0071] In at least one embodiment, parser 106 is a set of instructions which, if executed, cause one or more processors to parse source code 104. In at least one embodiment, parser 106 parses source code 104 to determine whether source code 104 is in a correct format, such as the syntax of a programming language. In at least one embodiment, source code 104 is parsed by constructing a data structure (also referred to as a parse tree or syntax tree) built from a pre-defined grammar (e.g., language rules) of a programming language, by way of example. In at least one embodiment, the parse tree is a hierarchical structure that represents source code 104 in a tree structure that corresponds to the grammar of a programming language. In at least one embodiment, the parse tree is an abstract syntax tree, which is a hierarchical structure that represents the source code 104 in a tree structure that corresponds to a simplified grammar (e.g., language rules) of a programming language.

[0072] In at least one embodiment, parser 106 generates output for use by a semantic analyzer (not shown) and / or intermediate code generator 108. In at least one embodiment, a semantic analyzer is a set of instructions which, if executed, cause one or more processors 160 to verify the semantic correctness of a declaration or statement of a computer program. In at least one embodiment, a semantic analyzer, if executed, causes one or more processors 160 to perform type checking to verify that each operator output by parser 106 contains matching operands.

[0073] In at least one embodiment, output of parser 106 and / or a semantic analyzer is to be used as input for intermediate code generator 108. In at least one embodiment, intermediate code generator 108 is software instructions which, if executed, cause one or more processors 160 to convert a set of tokens representing source code 104 to IR, such as initial IR 118. In at least one embodiment, the tokens are generated by parser 106 and / or a semantic analyzer. In at least one embodiment, IR is data used to represent individual data objects and / or operations indicated in source code 104. In at least one embodiment, IR includes intermediate code representing individual data objects and / or operations indicated in source code 104. In at least one embodiment, initial IR 118 is an abstract syntax tree, where nodes represent operations such as addition, multiplication, expressions, variables used in the expressions, statements that perform actions using the expressions, and other entities of a computer program.

[0074] In at least one embodiment, intermediate code generator 108 converts tokenized source code 104 into intermediate code. In at least one embodiment, intermediate code is high-level IR, code similar to source language. In at least one embodiment, intermediate code is low-level IR (e.g., code similar to target machine language). In at least one embodiment, intermediate code is either high-level IR (e.g., code similar to source language) or low-level IR (e.g., code similar to target machine language of processor 160). In at least one embodiment, intermediate code generator 108 generates language-independent code, such as an architecture-neutral output. In at least one embodiment, output from intermediate code generator 108 is input to optimizer 110.

[0075] In at least one embodiment, optimizer 110 is a set of instructions that, if executed, cause one or more processors to apply one or more optimizations to initial IR 118 generated from source code 104. In at least one embodiment, optimizer 110 receives initial IR 118 output by intermediate code generator 108. In at least one embodiment, optimizer 110 causes one or more processors to perform one or more optimizations on one or more IRs 120, including the initial IR 118. In at least one embodiment, for each optimization, optimizer 110 causes one or more processors to apply one or more transformations to improve initial IR 118. In at least one embodiment, each optimization transforms input IR 120 to a functionally equivalent output IR 124. In at least one embodiment, compiler 102 includes a target code generator 136. In at least one embodiment, the target code generator 136 is program code instructions that, if executed, cause one or more processors to generate code specific to a computing architecture. In at least one embodiment, optimizer 110 provides each output IR 124 to a subsequent optimization, or, if no further optimizations are to be performed, to a target code generator 136 that converts the output IR to an executable program.

[0076] In at least one embodiment, improving intermediate code or IR includes reducing resource usage, such as CPU or memory resources, to produce machine code that executes faster. In at least one embodiment, code optimization is a process of transforming one or more code segments into another functional equivalent to improve one or more characteristics. In at least one embodiment, optimizer 110 includes built-in knowledge of one or more processor-specific functions, such as intrinsic functions. In at least one embodiment, optimizer 110 causes one or more processors to optimize intermediate code, such as optimizations specific to a target processor architecture. In at least one embodiment, optimizer 110 causes one or more processors to apply and / or insert one or more optimized functions, such as intrinsic functions, from an optimized function library generated for a processor into input IR 120. In at least one embodiment, a compiler intrinsic function is a processor-specific function. In at least one embodiment, a processor function is a set of instructions that, if executed, cause a processor to perform one or more computation operations optimized specifically for that processor.

[0077] In at least one embodiment, compiler 102 is capable of performing one or more optimizations. In at least one embodiment, optimizations that compiler 102 is capable of performing include strength reduction, which replaces an operation with a more efficient expression that performs the same task; common subexpression elimination, which identifies mathematical expressions that produce the same result and computes the result of the expression once instead of multiple times; and / or constant folding, which precomputes constant values during compilation instead of at runtime. In at least one embodiment, optimizations are performed by optimizer 110 of compiler 102.

[0078] In at least one embodiment, optimizations that compiler 102 is capable of performing include one or more loop optimizations, one or more dataflow optimizations, one or more code generator optimizations, or one or more interprocedural optimizations, or other optimizations. In at least one embodiment, the loop optimizations include, for example, loop unrolling, which replaces a loop statement with multiple repeated instances of a loop’s body. In at least one embodiment, the dataflow optimizations include, for example, common subexpression elimination. In at least one embodiment, code generator optimizations include, for example, register allocation, which determines which variables to store in processor registers (which can be accessed faster than memory) and which variables to store in memory, to get the benefit of register storage for variables that are frequently accessed.

[0079] In at least one embodiment, other optimizations include inlining expansion of procedure calls, which involves replacing a call to a procedure with program code that will be executed by the procedure, thereby avoiding the time and memory cost of the procedure call operation. In at least one embodiment, inter-procedural optimizations include optimizations performed, for example, on an entire program, without being limited to a particular procedure (e.g., a unit of program code). In at least one embodiment, inter-procedural optimizations can be performed, for example, on inlining expansion of procedure calls. In at least one embodiment, optimizer 110 and / or compiler 102 performs one or more of the optimizations on input IR 120, and the optimizations transform input IR 120 into output IR 124.

[0080] In at least one embodiment, optimizer 110 performs a sequence of one or more optimization passes 122 to transform initial IR 118 into output IR 124. In at least one embodiment, each optimization pass 122 performs a respective optimization. In at least one embodiment, each optimization pass 122 receives input IR 120, performs an optimization on the input IR 120 to generate output IR 124, and outputs the output IR 124. In at least one embodiment, the output IR 124 from a current optimization pass 122 is used as input IR 120 for a subsequent optimization pass 122. In at least one embodiment, the input IR 120 for the subsequent optimization pass 122 includes the output IR 124 from the current optimization pass 122. In at least one embodiment, the input IR 120 for the subsequent optimization pass 122 is the output IR 124 from the current optimization pass 122. In at least one embodiment, a last optimization pass 122 in a sequence of optimization passes 122 generates a last output IR 124, which is provided by optimizer 110 as input to target code generator 136.

[0081] In at least one embodiment, compiler 102 can receive an input optimization configuration file 112 as input. In at least one embodiment, the input optimization configuration file 112 specifies one or more optimization pass names 114 that identify optimization passes and / or optimizations to be performed by optimizer 110 of compiler 102. In at least one embodiment, for example, the optimization pass names 114 are Pass A, Pass B, and so on, up to Pass N. In at least one embodiment, the optimization pass names 114 are descriptive pass names and / or optimization names, such as strength reduction, subexpression elimination, and constant folding. In at least one embodiment, optimizer 110 performs each optimization specified by the optimization pass names 114. In at least one embodiment, optimizer 110 performs optimizations in an order specified in the optimization pass names 114. In at least one embodiment, if the input optimization configuration file 112 and / or the optimization pass names 114 are not specified, optimizer 110 performs a default set of optimization passes, such as strength reduction, subexpression elimination, and constant folding, and / or other optimization passes. In at least one embodiment, input optimization configuration file 112 can be specified in a file provided to compiler 102. In at least one embodiment, the file can contain a list of optimization pass names 114 to perform, or a list of optimization pass names 114 that compiler 102 will not perform. In at least one embodiment, input optimization configuration file 112 can be omitted, and compiler 102 can perform using a default list of optimization passes 122. In at least one embodiment, the default list can include optimization passes 122 that can improve performance of frequently used program code structures, such as loops, function calls, mathematical expressions, and the like. In at least one embodiment, the default list can be provided to compiler 102 as a configuration argument or parameter by a user. In at least one embodiment, the input optimization configuration file 112 can be an output optimization configuration file 140 from a previous run of compiler 102. In at least one embodiment, input optimization configuration file 112 can be represented as structured configuration data (e.g., YAML (Yet Another Markup Language), and / or the like) and / or, for example, command line parameters.

[0082] In at least one embodiment, as an example, in a first optimization pass A 122A, optimizer 110 receives input IR 120A. In at least one embodiment, the input IR 120A includes the initial IR 118 from the intermediate code generator 108. In at least one embodiment, the input IR 120A is the initial IR 118. In at least one embodiment, in the first optimization pass A 122A, optimizer 110 performs an optimization on the input IR 120A. In at least one embodiment, for example, the optimization is strength reduction. In at least one embodiment, the optimization generates output IR 124A as a result of performing the optimization on the input IR 120A.

[0083] In at least one embodiment, in a second optimization pass B 122B, the optimizer receives input IR 120B. In at least one embodiment, the input IR 120B is or includes the output IR 124A generated by optimization pass A 122A. In at least one embodiment, in the second optimization pass B 122B, optimizer 110 performs an optimization on the input IR 120B. In at least one embodiment, for example, the optimization is common subexpression elimination. In at least one embodiment, the optimization generates output IR 124B as a result of performing the optimization on the input IR 120B.

[0084] In at least one embodiment, optimizer 110 can perform one or more additional passes before performing a last optimization pass N 122N. In at least one embodiment, in the optimization pass N 122N, the optimizer receives input IR 120N. In at least one embodiment, the input IR 120N is or includes the output of a previous compilation pass 122 performed before the optimization pass N 122N. In at least one embodiment, for example, the input IR 120N includes the output IR 124B from the optimization pass B 122B. In at least one embodiment, in optimization pass N 122N, optimizer 110 performs an optimization on the input IR 120N. In at least one embodiment, for example, the optimization is constant folding. In at least one embodiment, the optimization generates output IR 124N as a result of performing the optimization on the input IR 120N. In at least one embodiment, target code generator 136 can generate executable program 150 based on the output IR 124N.

[0085] In at least one embodiment, the computer program has one or more versions. In at least one embodiment, each version of the computer program has corresponding source code 104, so that a version of the computer program corresponds to a version of source code 104. In at least one embodiment, a version of the computer program is identified by a version identifier, e.g., a first version is 1.0, a second version is 2.0, and so on. In at least one embodiment, source code 104 for different versions of the computer program differ by at least one statement and / or data value. In at least one embodiment, executable program 150 for different versions of the computer program differ by at least one instruction and / or data value. In at least one embodiment, because compiler 102 generates output IR 124 from source code 104, different versions of the program have at least one different output IR 124.

[0086] In at least one embodiment, an optimization (e.g., an optimization performed by optimization pass 122) changes input IR 120 and generates output IR 124 that is different from the input IR 120. In at least one embodiment, an optimization that changes input IR 120 improves the input IR 120 and generates output IR 124 that is improved compared to the input IR 120. In at least one embodiment, an optimization that improves input IR 120 generates output IR 124 that is improved compared to the input IR 120. In at least one embodiment, an optimization that does not change input IR 120 does not improve the input IR 120.

[0087] In at least one embodiment, an optimization that changes input IR 120 for a first version of a program changes input IR 120 for a second version of the program. In at least one embodiment, an optimization that changes input IR 120 for a first version changes input IR 120 for a second version. In at least one embodiment, an optimization that changes input IR 120 for a first version of a program does not change input IR 120 for a second version of the program. In at least one embodiment, an optimization that does not change input IR 120 for a first version of a program does not change input IR 120 for a second version of the program. In at least one embodiment, an optimization that improves input IR 120 for a first version of a program also improves input IR 120 for a second version of the program. In at least one embodiment, an optimization that changes input IR 120 for a first version of a program does not improve input IR 120 for a second version of the program.

[0088] In at least one embodiment, optimizations that do not change the IR of the program can be omitted from subsequent compilations of the program because the optimizations will not improve the program. In at least one embodiment, performing optimizations consumes significant amounts of time and processor resources, so by omitting optimizations that do not change the IR, compilation time in subsequent compilations can be reduced. In at least one embodiment, optimizations that do not change the IR of a first version of a program can be omitted from subsequent compilations of a second version of the program because the second version of the program can include similar statements or instructions as the first version and optimizations that are applicable to the first version can also be applicable to the second version.

[0089] In at least one embodiment, compiler 102 includes optimization profile generator 126. In at least one embodiment, optimization profile generator 126 is an instruction that, if executed, causes one or more processors 160 to generate an output optimization profile 140 that includes one or more optimization pass names 142. In at least one embodiment, the optimization pass names 142 identify one or more optimization passes 122 that changed the IR of a program and / or optimization passes 122 that did not change the IR of a program. In at least one embodiment, the program is represented by source code 104 received by compiler 102. In at least one embodiment, optimization profile generator 126 identifies one or more optimizations that changed the IR of a program and / or one or more optimizations that did not change the IR of a program during compilation of the program. In at least one embodiment, the optimization profile generator 126 identifies a set of passes that changed the IR 130 of the program and / or a set of passes that did not change the IR 132 of the program and includes names or other identifiers of the passes that changed the IR 130 and / or the passes that did not change the IR 132 in the output optimization profile 140.

[0090] In at least one embodiment, the optimizer 110 generates one or more optimization pass change indicators 116 that indicate which optimizations the optimizer 110 performed on a program. In at least one embodiment, for example, the optimization pass change indicators 116 indicate which optimizations the optimizer 110 performed during compilation of the source code 104 by the compiler 102. In at least one embodiment, the optimization pass change indicators 116 include an optimization pass change indicator 116 for each optimization pass 122 performed by the optimizer 110 on a program. In at least one embodiment, each optimization pass change indicator 116 indicates that a corresponding optimization pass 122 was performed. In at least one embodiment, if a compilation pass was performed by the optimizer 110 that is not included in the optimization pass change indicators 116, then the compilation pass was not performed during compilation of the source code 104. For example, if the optimizer 110 performed optimization pass A 122A, optimization pass B 122B, and optimization pass N 122N, then the optimization pass change indicators 116 include a pass A indicator 116A, a pass B indicator 116B, and a pass N indicator 116B. In at least one embodiment, for example, if the optimizer 110 performed optimization passes 122 named pass A and pass N but not pass B, then the optimization pass change indicators 116 that indicate which optimization passes were performed include the pass A indicator 116A and the pass B indicator 116B, but not the pass B indicator 116B. In at least one embodiment, the optimizer 110 does not generate optimization pass change indicators 116. In at least one embodiment, the optimizer 110 generates other information that indicates which optimizations the optimizer 110 performed on a program. In at least one embodiment, the optimizer 110 provides information indicating which optimizations the optimizer 110 performed on a program and / or which optimizations the optimizer 110 did not perform on the program using any suitable data format.

[0091] In at least one embodiment, the optimizer 110 generates one or more optimization pass change indicators 116 that indicate which optimizations the optimizer 110 did not perform on a program. In at least one embodiment, for example, the optimization pass change indicators 116 indicate which optimizations the optimizer 110 did not perform during compilation of the source code 104 by the compiler 102. In at least one embodiment, the optimization pass change indicators 116 include an optimization pass change indicator 116 for each optimization pass 122 that the optimizer 110 did not perform on a program. In at least one embodiment, each optimization pass change indicator 116 indicates that a corresponding optimization pass 122 was not performed. In at least one embodiment, if a compilation pass by the optimizer 110 is not included in the optimization pass change indicators 116, then the compilation pass was performed during compilation of the source code 104. In at least one embodiment, for example, if the optimizer 110 performed optimization pass A 122A, optimization pass B 122B, and optimization pass N 122N, then the optimizer 110 does not generate optimization pass change indicators 116. In at least one embodiment, for example, if the optimizer 110 performed optimization passes 122 named pass A and pass N, but did not perform pass B, then the optimization pass change indicators 116 that indicate which optimization passes were not performed include a pass B indicator 116B, but do not include a pass A indicator 116A or a pass N indicator 116N.

[0092] In at least one embodiment, optimization profile generator 126 uses one or more optimization pass change indicators 116 received from optimizer 110 to identify passes that changed IR 130 and / or passes that did not change IR 132. In at least one embodiment, optimization profile generator 126 generates a list of passes that changed IR 130, including one or more optimization passes 122 identified by respective optimization pass change indicators 116 that indicate which optimizations were performed by optimizer 110 on a program. In at least one embodiment, optimization profile generator 126 generates a list of passes that did not change IR 132 as a difference between a list of optimization passes that optimizer 110 was capable of performing and the optimization pass change indicators 116 that indicate which optimization passes were performed by optimizer 110. In at least one embodiment, if no passes changed IR 130, then the set of passes that changed IR 130 is empty. In at least one embodiment, if all passes changed IR 130, then the set of passes that did not change IR 132 is empty. In at least one embodiment, for example, if optimization pass A 122A changed input IR 120A and optimization pass N 122N changed input IR 120N but did not change input IR 120B, then the set of passes that changed IR 130 includes a name or other identifier of pass A and pass N, and the set of passes that did not change IR 132 includes a name or other identifier of pass B. In at least one embodiment, for example, the passes that changed IR 130 can include names such as “loop unrolling” and “constant folding,” and the passes that did not change IR 132 can include a name such as “common subexpression elimination.” In at least one embodiment, optimization profile generator 126 uses information other than optimization result indicators 116, such as other information received from optimizer 110, to identify passes that changed IR 130 and / or passes that did not change IR 132. In at least one embodiment, optimization profile generator 126 uses information in any suitable data format to determine which optimizations were performed by optimizer 110 on a program and / or which optimizations were not performed by optimizer 110 on the program.

[0093] In at least one embodiment, optimization profile generator 126 generates a list of passes that did not change IR 132, including one or more optimization passes 122 identified by respective optimization pass change indicators 116 that indicate which optimizations were not performed by optimizer 110 on a program. In at least one embodiment, optimization profile generator 126 generates a list of passes that changed IR 130 as a difference between a list of optimization passes that optimizer 110 was capable of performing and the optimization pass change indicators 116 that indicate which optimization passes were not performed by optimizer 110 on a program.

[0094] In at least one embodiment, the optimization profile generator 126 generates an output optimization profile 140 that includes a list of optimization pass names 142. In at least one embodiment, the optimization pass names 142 include names of optimizations or optimization passes that changed the IR of a program associated with the source code 104. In at least one embodiment, for example, the optimization pass names 142 can include pass A and pass N to indicate that optimization pass A 122A and optimization pass N 122N changed the IR of the program.

[0095] In at least one embodiment, the optimization pass names 142 include names of optimizations or optimization passes that did not change the IR of a program associated with the source code 104. In at least one embodiment, for example, the optimization pass names 142 can include pass B to indicate that optimization pass B 122B did not change the IR of the program.

[0096] In at least one embodiment, in a subsequent compilation of the program, the optimizer 110 selects optimizations to perform on the program or another version of the program based on the output optimization profile 140. In at least one embodiment, the output optimization profile 140 is provided to the compiler 102 as an input optimization profile 112 for use in a subsequent compilation of the source code 104 or another version of the source code 104.

[0097] In at least one embodiment, to perform a subsequent compilation of the program, the compiler 102 retrieves a list of optimization pass names 114 from the input optimization profile 112 that identifies one or more optimization passes. In at least one embodiment, the optimization pass names 114 can identify optimization passes to perform. In at least one embodiment, the optimization pass names 114 can identify optimization passes not to perform.

[0098] In at least one embodiment, the list of optimization pass names 114 specifies optimization passes that the compiler 102 is to perform. In at least one embodiment, the compiler 102 performs one or more optimization passes 122 identified in the list of optimization pass names 114 to perform. In at least one embodiment, the compiler 102 does not perform one or more optimization passes 122 not identified in the list of optimization pass names 114 to perform.

[0099] In at least one embodiment, the list of optimization pass names 114 specifies optimization passes that the compiler 102 is not to perform. In at least one embodiment, the compiler 102 does not perform one or more optimization passes 122 identified in the list of optimization pass names 114 not to perform. In at least one embodiment, the compiler 102 performs one or more optimization passes not identified in the list of optimization pass names 114 not to perform.

[0100] In at least one embodiment, during compilation of source code 104 of a program, an optimization profile generator 126 generates one or more passes of IR 130 that change modules (and / or functions or other portions) of the program, and further identifies a set of changed modules (and / or functions or other portions) changed by each of the optimization passes. In at least one embodiment, the optimization profile generator 126 receives the one or more sets of optimization passes that change IR and the set of changed modules (and / or functions or other portions) from an optimizer 110. In at least one embodiment, the optimizer 110 provides optimization pass change indicators 116 to the optimization profile generator 126 that indicate optimization passes that change IR. In at least one embodiment, the optimization pass change indicators 116 include an optimization pass change indicator 116 for each optimization pass 122 that changes IR 120 of a module (and / or function or other portion) of the program. In at least one embodiment, each optimization pass change indicator 116 indicates an optimization pass 122, and further indicates a set of changed modules (and / or functions or other portions) changed by the optimization pass 122. In at least one embodiment, for example, if an optimization pass named pass A 122A changes module Ml of the program, an optimization pass named pass B 122B does not change any modules of the program, and an optimization pass named pass N 122N changes modules Ml and M2 of the program, the optimization pass change indicators 116 include a pass A indicator 116A that indicates pass A and module Ml, and a pass N indicator 116N that indicates pass N and modules Ml and M2, but does not include a pass B indicator 116B.

[0101] In at least one embodiment, the optimization profile generator 126 uses the identified sets of optimization passes that change IR 130 of modules (and / or functions or other portions) of the program to generate an output optimization profile 140 that specifies optimizations to perform. In at least one embodiment, the output optimization profile 140 that specifies optimizations to perform specifies one or more optimizations to perform on one or more modules (and / or functions or other portions) of the program. In at least one embodiment, the output optimization profile 140 that specifies optimizations to perform specifies the one or more optimization passes that change IR 120, and further specifies each module (and / or function or other portion) of the input IR 120 of the program changed by the one or more optimization passes. In at least one embodiment, for example, the output optimization profile 140 that specifies optimizations to perform includes a data structure such as a table that associates each optimization with each module (and / or function or other portion) of the program on which to perform the optimization.

[0102] In at least one embodiment, during compilation of source code 104 of a program, optimization profile generator 126 generates one or more optimization passes that do not change IR 130 of modules (and / or functions or other portions) of the program, and further identifies a set of unchanged modules (and / or functions or other portions) that are not changed by each of the optimization passes. In at least one embodiment, optimizer 110 provides optimization pass change indicators 116 to the optimization profile generator 126 that indicate optimization passes that do not change IR. In at least one embodiment, the optimization pass change indicators 116 include an optimization pass change indicator 116 for each optimization pass 122 that does not change IR 120 of a module (and / or function or other portion) of the program. In at least one embodiment, each optimization pass change indicator 116 indicates an optimization pass 122, and further indicates a set of unchanged modules (and / or functions or other portions) that are not changed by the optimization pass 122. In at least one embodiment, for example, if an optimization pass named pass A 122A does not change module M2 of the program, an optimization pass named pass B 122B does not change any modules of the program, and an optimization pass named pass N 122N changes all modules of the program, the optimization pass change indicators 116 include a pass A indicator 116A that indicates pass A and module M2, and a pass B indicator 116B that indicates all modules of the program (e.g., modules Ml and M2), but does not include a pass N indicator 116N.

[0103] In at least one embodiment, the optimization profile generator 126 uses the identified set of optimization passes that do not change IR of modules (and / or functions or other portions) of the program and the set of unchanged modules (and / or functions or other portions) to generate an output optimization profile 140 that specifies optimizations not to perform. In at least one embodiment, the output optimization profile 140 that specifies optimizations not to perform specifies one or more optimizations not to perform on one or more modules (and / or functions or other portions) of the program. In at least one embodiment, for example, the output optimization profile 140 that specifies optimizations not to perform specifies the one or more optimization passes that do not change IR 120 of the program, and further specifies each module (and / or function or other portion) of the input IR 120 of the program that is not changed by the one or more optimization passes. In at least one embodiment, for example, the output optimization profile 140 that specifies optimizations not to perform includes a data structure such as a table that associates each optimization with each module (and / or function or other portion) of the program for which the optimization is not to be performed.

[0104] In at least one embodiment, to perform a subsequent compilation of the program or another version of the program, the output optimization profile 140 specifying optimizations to be performed is provided to the compiler 102 as an input optimization profile 112 specifying optimizations to be performed. In at least one embodiment, the compiler 102 selects one or more optimizations to be performed based on one or more modules (and / or functions or other portions) of the program specified in the input optimization profile 112 specifying optimizations to be performed. In at least one embodiment, the compiler 102 performs the selected one or more optimizations to be performed on the modules (and / or functions or other portions) of the program.

[0105] In at least one embodiment, to perform a subsequent compilation of the program or another version of the program, the output optimization profile 140 specifying optimizations not to be performed is provided to the compiler 102 as an input optimization profile 112 specifying optimizations not to be performed. In at least one embodiment, the compiler 102 selects one or more optimizations to be performed on the program based on one or more modules (and / or functions or other portions) specified in the input optimization profile 112 specifying optimizations not to be performed. In at least one embodiment, for example, the compiler 102 selects one or more optimizations to be performed on the modules (and / or functions or other portions) of the program by generating a set of optimizations to be performed on the modules (and / or functions or other portions) of the program. In at least one embodiment, the set of optimizations includes one or more optimizations that the compiler 102 is capable of performing but excludes one or more optimizations specified by the input optimization profile 112 as not to be performed on one or more modules (and / or functions or other portions) of the program. In at least one embodiment, the compiler 102 performs the set of optimizations to be performed on the modules (and / or functions or other portions) of the program. In at least one embodiment, the compiler 102 does not perform optimizations specified in the input optimization profile 112 specifying optimizations not to be performed on one or more modules (and / or functions or other portions) of the program.

[0106] In at least one embodiment, target code generator 136 is a set of software instructions which, if executed, cause one or more processors 160 to generate code specific to a computing architecture. In at least one embodiment, target code generator 136 causes one or more processors 160 to generate target code specific to a processor architecture. In at least one embodiment, target code generator 136 causes one or more processors 160 to generate target code in an intermediate format for further compilation by a JIT compiler. In at least one embodiment, target code generator 136 causes one or more processors 160 to generate target code for interpretation by an interpreter, as further described herein. In at least one embodiment, target code generator 136 can cause one or more processors 160 to perform instruction selection, register allocation, and / or instruction ordering. In at least one embodiment, target code generator 136 causes one or more processors 160 to generate executable program 150.

[0107] In at least one embodiment, output of compiler 102, such as output of target code generator 136, is input to a linker (not shown). In at least one embodiment, a linker is a set of instructions which, if executed, cause one or more processors 160 to convert data output from compiler 102 into executable program 150. In at least one embodiment, a linker receives data in a relocatable format and converts it to an absolute format specific to processor 160 or processor architecture. In at least one embodiment, a linker causes one or more processors 160 to combine one or more object files into executable program 150. In at least one embodiment, executable program 150 is a set of instructions to be executed by one or more processors 160.

[0108] FIG. 2 An example system 200 is shown that optimizes compiler 102 based on changes made to input IR 120 by optimization pass 122, in accordance with at least one embodiment. In at least one embodiment, system 100 is used to execute compiler 102 to select one or more optimizations to one or more first versions of a program based at least in part on results of performing one or more optimizations to one or more second versions of the program. In at least one embodiment, one or more first versions of a program include source code 104, IR (initial IR 118, input IR 120, output IR 124), and / or one or more executable programs 150. In at least one embodiment, one or more second versions of a program include source code 104, IR (initial IR 118, input IR 120, output IR 124), and / or one or more executable programs 150. In at least one embodiment, compiler optimization system 200 is similar to compiler optimization system 100, described above.FIG. 1 Compiler optimization system 100, but compiler optimization system 200 includes an optimization profile generator 226. In at least one embodiment, optimization profile generator 226 is an instruction that, if executed, causes one or more processors 160 to compare a current IR 234 (which is output IR 124 from optimization pass 122) to a previous IR 232 (which is input IR 120 for the optimization pass 122) and generate an output optimization profile 240 based on a result of the comparison. In at least one embodiment, optimization profile generator 226 generates an output optimization profile 140 that includes one or more optimization pass names 142 that identify one or more optimization passes 122 that changed the IR of the program and / or one or more optimization passes 122 that did not change the IR of the program. In at least one embodiment, optimization profile generator 226 identifies passes that changed the IR 130 based on changes made to input IR 120 by one or more optimization passes 122. In at least one embodiment, optimization profile generator 226 compares output IR 124 from optimization pass 122 to input IR 120 for the optimization pass 122. In at least one embodiment, if optimization profile generator 226 identifies one or more differences 236 between the output IR 124 and the input IR 120, optimization profile generator 226 includes a name or other identifier for the optimization pass 122 in a list of passes that changed the IR 130. In at least one embodiment, optimization profile generator 226 generates an output optimization profile 240 based on passes that changed the IR 130. In at least one embodiment, optimization profile generator 226 includes a name or identifier for a pass that changed the IR 130 in the output optimization profile 240.

[0109] In at least one embodiment, compiler 102 includes optimization profile generator 226. In at least one embodiment, optimization profile generator 226 is a set of instructions which, if executed, cause one or more processors to generate an optimization profile 130 based on a plurality of IRs 120 generated from source code 104. In at least one embodiment, optimization profile generator 226 receives, from optimizer 210, an input IR 120 for each optimization pass 122. In at least one embodiment, optimization profile generator 226 compares a previous IR 232 to a current IR 234 for each optimization pass 122. In at least one embodiment, previous IR 232 is an input IR 120 received from optimizer 210, while current IR 234 is an output IR 124 received from optimizer 210 for optimization pass 122. In at least one embodiment, in response to receiving current IR 234 and previous IR 232 for optimization pass 122, optimization profile generator 226 compares previous IR 232 to current IR 234. In at least one embodiment, if optimization profile generator 226 determines that there is one or more differences between previous IR 232 and current IR 234, optimization profile generator 226 includes a name or other identifier of optimization pass 122 in a list of passes that changed IR 130. In at least one embodiment, if there are no differences between previous IR 232 and current IR 234, optimization profile generator 226 does not include a name or other identifier of optimization pass 122 in the list of passes that changed IR 130. In at least one embodiment, if optimization profile generator 226 determines that there are no differences between previous IR 232 and current IR 234, optimization profile generator 226 includes a name or other identifier of optimization pass 122 in a list of passes that did not change IR 132.

[0110] In at least one embodiment, IR is represented as a data structure such as a tree. In at least one embodiment, to determine whether one or more differences exist between a previous IR 232 and a current IR 234, optimization profile generator 226 compares each element (e.g., each tree node) of previous IR 232 to each corresponding element (e.g., corresponding tree node) of current IR 234. In at least one embodiment, if at least one element of previous IR 232 is different from a corresponding element of current IR 234, optimization profile generator 226 determines that one or more differences exist between previous IR 232 and current IR 234. In at least one embodiment, an element of previous IR 232 is different from a corresponding element of current IR 234 if, for example, the element specifies a different statement, variable name, data type, or other element of IR than the corresponding element of current IR 234. In at least one embodiment, IR is represented as text or other sequence of symbols, and optimization profile generator 226 determines whether any differences exist between previous IR 232 and current IR 234 by comparing text or other symbols of previous IR 232 to current IR 234.

[0111] In at least one embodiment, optimization profile generator 226 generates a hash value based on each output IR 124 and compares the hash value to a hash value of input IR 120. In at least one embodiment, optimization profile generator 226 uses a hash value of an output IR 124 in an optimization pass 122 as a hash value of an input IR 120 in a subsequent optimization pass 122. In at least one embodiment, the hash value can be computed based on a sequence of bytes representing the IR. In at least one embodiment, the hash value can be computed based on values of elements (e.g., tree nodes) of the data structure representation of the IR. In at least one embodiment, if a hash value of an output IR 124 matches (e.g., is equal to) a hash value of an input IR 120, the output IR 124 is the same IR as the input IR 120. In at least one embodiment, if the hash values are not equal, the output IR 124 is not the same IR as the input IR 120. In at least one embodiment, a hash value can be a cyclic redundancy code (CRC), a cryptographic hash (e.g., a message digest), or other hash value computed using a suitable hash algorithm.

[0112] In at least one embodiment, optimization profile generator 226 generates output optimization profile 240 based on a list of iterations that changed IR 130 and / or a list of iterations that did not change IR 132. In at least one embodiment, output optimization profile 240 includes one or more optimization iteration names 242. In at least one embodiment, optimization iteration names 242 are a list of iterations that changed IR of a program being compiled, and optimization profile generator 226 adds each optimization iteration name or identifier in the list of iterations that changed IR 130 to optimization iteration names 242 of output optimization profile 240. In at least one embodiment, optimization iteration names 242 of output optimization profile 240 are a list of iterations that did not change IR 132 of a program being compiled, and optimization profile generator 226 adds each optimization iteration name or identifier in the list of iterations that did not change IR 132 to optimization iteration names 242 of output optimization profile 240.

[0113] In at least one embodiment, compiler optimization system 200 provides output optimization profile 240 as input optimization profile 112 to a subsequent run of compiler 102 on source code 104 or another version of source code 104. In at least one embodiment, the subsequent run of compiler 102 performs optimization iterations 122 specified by input optimization profile 112, as described herein with respect to FIG. 1

[0114] In at least one embodiment, in an example compilation of source code 104, optimization profile generator 226 compares output 124 of each optimization iteration 122 to input of the optimization iteration 122. In at least one embodiment, for example, input optimization profile 112 includes a list of optimization iteration names 114 that specify iteration A, iteration B, and iteration N. In at least one embodiment, for example, iteration A is “strength reduction,” iteration B is “subexpression elimination,” and iteration N is “constant folding.” In at least one embodiment, compiler 102 receives input optimization profile 112 with optimization iteration names 114 “strength reduction,” “subexpression elimination,” and “constant folding.” In at least one embodiment, for each iteration name specified by optimization iteration names 114, compiler 102 performs optimization iteration 122 using the optimization specified by the iteration name.

[0115] ​In at least one embodiment, for pass A having a pass name of “strength reduction,” the optimizer 210 performs an optimization pass A 122A, which is a strength reduction optimization. In at least one embodiment, the optimizer 210 provides the initial IR 118 as an input IR 120A to the optimization pass A 122A. In at least one embodiment, in the optimization pass A 122A, the optimizer 210 performs a strength reduction optimization on the input IR 120A to generate an output IR 124A. In at least one embodiment, for example, the strength reduction optimization improves the input IR 120A and generates an output IR 124A that is different from the input IR 120A.

[0116] In at least one embodiment, the optimizer 210 provides the input IR 120A and the output IR 124A as a previous IR 232 and a current IR 234, respectively, to the optimization profile generator 226. In at least one embodiment, the optimization profile generator 226 compares the previous IR 232 to the current IR 234 and determines that the previous IR 232 is different from the current IR 234. In at least one embodiment, because the previous IR 232 is different from the current IR 234, the optimization profile generator 226 includes the name or identifier of pass A, e.g., “strength reduction,” in the optimization profile 240 as one of the optimization passes of the changed IR 130.

[0117] In at least one embodiment, for pass B having a pass name of “subexpression elimination,” the optimizer 210 performs an optimization pass B 122B, which is a subexpression elimination optimization. In at least one embodiment, the optimizer 210 provides the output IR 124A from the optimization pass A 122A as an input IR 120B to the optimization pass B 122B. In at least one embodiment, in the optimization pass B 122B, the optimizer 210 performs a strength reduction optimization on the input IR 120B to generate an output IR 124B. In at least one embodiment, for example, the subexpression elimination optimization does not improve the input IR 120B and generates an output IR 124B that is not different from the input IR 120B.

[0118] In at least one embodiment, the optimizer 210 provides the input IR 120B and the output IR 124B to the optimization profile generator 226 as a previous IR 232 and a current IR 234, respectively. In at least one embodiment, the optimization profile generator 226 compares the previous IR 232 to the current IR 234 and determines that the previous IR 232 is not different from the current IR 234. In at least one embodiment, because the previous IR 232 is not different from the current IR 234, the optimization profile generator 226 does not include a name or identifier of the pass B in the pass that changed the IR 130. In at least one embodiment, the optimization profile generator 226 includes a name or identifier of the pass B, e.g., “subexpression elimination,” in the pass that did not change the IR 132. In at least one embodiment, if the optimization pass name 242 of the output optimization profile 240 is to include optimization pass names of the unchanged IR, the 226 includes a name or identifier of the pass B, e.g., “subexpression elimination,” in the optimization pass name 242 of the unchanged IR.

[0119] In at least one embodiment, for the pass N with a pass name of “constant folding,” the optimizer 210 performs the optimization pass N 122N, which is a constant folding optimization. In at least one embodiment, the optimizer 210 provides the output IR 124B from the optimization pass B 122B as the input IR 120N to the optimization pass B 122B. In at least one embodiment, in the optimization pass N 122N, the optimizer 210 performs a constant folding optimization on the input IR 120N to generate an output IR 124N. In at least one embodiment, for example, the constant folding optimization improves the input IR 120B and generates an output IR 124N that is different from the input IR 120B.

[0120] In at least one embodiment, the optimizer 210 provides the input IR 120B and the output IR 124B to the optimization profile generator 226 as a previous IR 232 and a current IR 234, respectively. In at least one embodiment, the optimization profile generator 226 compares the previous IR 232 to the current IR 234 and determines that the previous IR 232 is different from the current IR 234. In at least one embodiment, because the previous IR 232 is different from the current IR 234, the optimization profile generator 226 includes a name or identifier of the pass N, e.g., "constant folding," in the output optimization profile 240 as one of the changed-IR optimization pass names 242. In at least one embodiment, after performing the optimization pass N 122N, the optimization profile generator 226 generates the output optimization profile 240 including "strength reduction" and "constant folding" as optimization pass names 242. In at least one embodiment, the target code generator 136 of the compiler 102 generates the executable program 150 based on the output IR 124N of the last optimization pass 122N performed by the optimizer 210.

[0121] In at least one embodiment, the compiler optimization system 200 receives a request to perform another compilation of the source code 104 after generating the output optimization profile 240 for the source code 104 or a previous version of the source code 104. In at least one embodiment, for the subsequent compilation, the compiler optimization system 200 provides the output optimization profile 240 to the compiler 102 as an input optimization profile 112 for the subsequent compilation of the source code 104 or the previous version of the source code 104. Thus, in at least one embodiment, the input optimization profile 112 has optimization pass names 114 of "strength reduction" and "constant folding."

[0122] In at least one embodiment, to perform the subsequent compilation, compiler 102 receives the input optimization profile 112 with optimization pass names 114 “strength reduction” and “constant folding.” In at least one embodiment, for each pass name specified by optimization pass names 114, compiler 102 performs an optimization pass 122 using the optimization specified by the pass name. In at least one embodiment, compiler 102 performs optimization pass A 122A on input IR 120A that includes a strength reduction optimization. In at least one embodiment, input IR 120A is initial IR 118 generated from source code 104. In at least one embodiment, the strength reduction optimization generates output IR 124A that is different from the input IR 120A, so optimization profile generator 226 includes optimization name “strength reduction” in the list of passes that changed IR 130 and / or the list of optimization pass names 242 that changed IR.

[0123] In at least one embodiment, in the subsequent compilation, compiler 102 performs optimization pass B 122B on input IR 120B that includes a constant folding optimization. In at least one embodiment, the input IR 120B includes output IR 124A from optimization pass A 122A. In at least one embodiment, the constant folding optimization generates output IR 124B that is different from the input IR 120A, so optimization profile generator 226 includes optimization name “constant folding” in the list of passes that changed IR 130 and / or the list of optimization pass names 242 that changed IR.

[0124] In at least one embodiment, in the subsequent compilation, compiler 102 does not perform any further optimization passes after optimization pass B 122B. Thus, in at least one embodiment, compiler 102 does not perform optimization passes 122 in the subsequent compilation. In at least one embodiment, object code generator 136 generates executable program 150 based on IR generated by the last optimization pass performed, in this example optimization pass B 122B. In at least one embodiment, compiler optimization system 200 can provide the list of passes that changed IR 130 (including “strength reduction” and “constant folding”) as input to another subsequent compilation of source code 104 (or another version of source code 104) to cause optimizer 210 to perform these optimizations.

[0125] FIG. 3An example 300 of profile-based optimization of a compilation pass is shown, in accordance with at least one embodiment. In at least one embodiment, example 300 shows a profile-based compiler run 302A in which compiler 102 compiles source code VI 304A of a first version (VI) of a program. In at least one embodiment, compiler 102 performs optimizations specified by an input optimization profile 312A. In at least one embodiment, input optimization profile 312A specifies optimization passes named pass A, pass B, pass C, and pass D, so compiler 102 performs optimization passes pass A, pass B, pass C, and pass D during compilation of source code 304A.

[0126] In at least one embodiment, in profile-based compiler run 302A, optimization profile generator 126 of compiler 102 determines which of the optimization passes changed IR of the program. In at least one embodiment, optimization profile generator 126 generates a list of passes that changed IR 330A, which includes pass A and pass C. In at least one embodiment, optimization profile generator 126 generates a list of passes that did not change IR 332A, which includes pass B and pass D.

[0127] In at least one embodiment, in profile-based compiler run 302A, compiler 102 generates an output optimization profile 340A that includes a list of optimization passes that changed IR during compilation of source code VI 304A. In at least one embodiment, output optimization profile 340A is the same list of passes as the list of passes that changed IR 330A. In at least one embodiment, output optimization profile 340A specifies pass A and pass C because pass A and pass C changed IR of the program. In at least one embodiment, compiler 102 generates executable program VI 350A as a result of compiling source code VI 304A. In at least one embodiment, executable program VI 350A has been optimized by pass A and pass C.

[0128] In at least one embodiment, example 300 also shows a profile-based compiler run 302B in which compiler 102 compiles source code V2 304B for a second version (V2) of the program. In at least one embodiment, in profile-based compiler run 302B, compiler 102 receives output optimization profile 340A as input, so output optimization profile 340A is also shown as input optimization profile 312B for profile-based compiler run 302B. In at least one embodiment, in profile-based compiler run 302B, compiler 102 performs the optimizations specified by input optimization profile 312B. In at least one embodiment, input optimization profile 312B specifies optimization passes named pass A and pass C, so compiler 102 performs optimization passes pass A and pass C during compilation of source code 304B.

[0129] In at least one embodiment, in profile-based compiler run 302B, optimization profile generator 126 of compiler 102 determines which of the optimization passes change the IR of the program. In at least one embodiment, optimization profile generator 126 generates a list of passes that change IR 330B, which includes pass A and pass C. In at least one embodiment, since both pass A and pass C change the IR, and no other passes are performed, the list of passes that do not change IR 332B is empty or is not generated by optimization profile generator 126.

[0130] In at least one embodiment, in profile-based compiler run 302B, compiler 102 generates output optimization profile 340B, which includes a list of optimization passes performed by compiler 102 during compilation of source code 304B, e.g., from the list of passes that change IR 330B. In at least one embodiment, since pass A and pass C change the IR of the program, output optimization profile 340B specifies pass A and pass C. In at least one embodiment, compiler 102 generates executable program V2 350B as a result of compiling source code 304B. In at least one embodiment, executable program V2 350B has been optimized by pass A and pass C, and optimization passes B and D were not performed during compilation of source code 304B.

[0131] FIG. 4An example 400 of configuration file-based optimization of a compilation pass at the module level is shown, in accordance with at least one embodiment. In at least one embodiment, example 400 shows a configuration file-based compiler run 402A in which compiler 102 compiles source code VI 404A of a first version (VI) of a program. In at least one embodiment, source code 404A includes module M1 406A and module M2 408A. In at least one embodiment, each of the modules 406A, 408A includes one or more functions and / or other portions of source code 404A.

[0132] In at least one embodiment, compiler 102 performs optimizations specified by input optimization configuration file 412A. In at least one embodiment, input optimization configuration file 412A associates optimization passes named Pass A, Pass B, Pass C, and Pass D with modules M1 and M2 of source code 404A. In at least one embodiment, for example, input optimization configuration file 412A is a table in which each row represents an optimization pass 122, each column represents a module of source code 404A, and each entry in the rows and columns of the table is a “yes” or “no” (or true or false) value indicating whether the optimization pass corresponding to the row is to be performed on the module corresponding to the column. In at least one embodiment, all entries in the table representing input optimization configuration file 412A are “yes” to indicate that all four optimization passes A, B, C, and D are to be performed on both modules M1 and M2. In at least one embodiment, in configuration file-based compiler run 402A, compiler 102 receives input optimization configuration file 412A and performs optimization passes on the modules of the program as specified by input optimization configuration file 412A. In at least one embodiment, compiler 102 performs optimization passes A, B, C, and D on modules M1 and M2 during compilation of source code 404A.

[0133] In at least one embodiment, in profile-based compiler run 402A, optimization profile generator 126 of compiler 102 determines which of the optimization passes changed the IR of the modules of the IR. In at least one embodiment, optimization profile generator 126 generates a list of the passes that changed IR 430A and an indication of which module(s) each pass changed. In at least one embodiment, optimization profile generator 126 generates an association between the optimization passes that changed the IR and the modules of the IR. In at least one embodiment, the passes that changed IR modules 430A include an association between pass A and module Ml, and an association between pass C and modules Ml and M2, to indicate that pass A changed module Ml and pass C changed modules Ml and M2. In at least one embodiment, the passes that changed IR modules 430A is a table where an entry at row pass A and column Ml has a value of “yes,” and entries at row pass C and columns Ml and M2 each have a value of “yes.” In at least one embodiment, entries with a value of “no” indicate that the optimization pass corresponding to the row of the entry did not change the module corresponding to the column of the entry.

[0134] In at least one embodiment, in profile-based compiler run 402A, compiler 102 generates output optimization profile 440A that includes a list of the optimization passes that changed the IR and an indication of which module(s) each pass changed during compilation of source code 404A. In at least one embodiment, output optimization profile 440A includes a table where entries indicate which optimization passes were performed for which modules, as described herein with reference to the table of passes that changed IR modules 430A. In at least one embodiment, the table of output optimization profile 440A is the same as the table of passes that changed IR modules 430A, and includes an association between pass A and module Ml, and an association between pass C and modules Ml and M2, to indicate that pass A changed module Ml and pass C changed modules Ml and M2. In at least one embodiment, compiler 102 generates executable program Vl 450A as a result of compiling source code Vl 404A. In at least one embodiment, in executable program Vl 450A, module Ml has been optimized by pass A and modules Ml and M2 have been optimized by pass C.

[0135] In at least one embodiment, example 400 also shows a profile-based compiler run 402B in which compiler 102 compiles source code V2 404B for a second version (V2) of the program. In at least one embodiment, in profile-based compiler run 402B, compiler 102 receives output optimization profile 440A as input, so output optimization profile 440A is also shown as input optimization profile 412B for profile-based compiler run 402B. In at least one embodiment, in profile-based compiler run 402B, compiler 102 performs optimizations specified by input optimization profile 412B. In at least one embodiment, input optimization profile 412B includes an association between optimization pass A and module Ml, and an association between pass C and modules Ml and M2, to indicate that pass A is to be performed on module Ml, and that pass C is to be performed on modules Ml and M2.

[0136] In at least one embodiment, in profile-based compiler run 402B, optimization profile generator 126 of compiler 102 determines which of the optimization passes change IR for the program. In at least one embodiment, optimization profile generator 126 generates a list of passes that change IR 430B, which includes pass A and pass C, and an association between optimization pass A and module Ml, an association between pass C and module Ml, and an association between pass C and module M2.

[0137] In at least one embodiment, in profile-based compiler run 402B, compiler 102 generates output optimization profile 440B, which includes a list of optimization passes and associated modules on which compiler 102 is to perform the passes during compilation of source code 404B. In at least one embodiment, output optimization profile 440B is the same as the passes that change IR module 430B. In at least one embodiment, compiler 102 generates executable program V2 450B as a result of compiling source code 404B. In at least one embodiment, in executable program V2 450B, module Ml has been optimized by pass A, and modules Ml and M2 have been optimized by pass C.

[0138] FIG. 5 FIG. 500 is a flow diagram of techniques for profile-based compiler pass optimization, in accordance with at least one embodiment. In at least one embodiment, techniques 500 are performed by at least one circuit, at least one system, at least one processor, at least one graphics processing unit, at least one parallel processor, and / or at least some other processor or component as described and / or illustrated herein. In at least one embodiment, at least one aspect of techniques 500 is performed by a compiler, such as compiler 102, as described and / or illustrated herein. FIG. 1by computer system 100 (e.g., processor 160). In at least one embodiment, technique 500 is performed at least in part by using (e.g., by computer system 100) FIG. 1 by one or more processors of computer system 100 (e.g., processor 160) and / or any other suitable processor such as those described or shown herein. In at least one embodiment, technique 500 is performed by one or more processors executing a set of instructions (e.g., from a non-transitory machine-readable medium) that are equivalent to FIG. 1 by compiler 102 of computer system 100 or using one or more processors to perform FIG. 1 by compiler 102 of computer system 100. In at least one embodiment, executing a set of instructions includes executing a set of instructions (e.g., using one or more processors). In at least one embodiment, technique 500 is performed by FIG. 1 processor 160 of computer system 100.

[0139] In at least one embodiment, compiler includes instructions that, if executed, generate one or more computer programs. In at least one embodiment, a computer program is a thread, a group of threads, a cooperative thread array (CTA), a kernel, a task, a node in a graph, or any other organization of software instructions further described herein, and includes software instructions that, if executed, cause one or more processors 160 to perform a computational operation. In at least one embodiment, during generation of one or more computer programs, instructions of a compiler, if executed, cause one or more processors and / or systems to generate an input intermediate representation (IR) based on source code of a program. In at least one embodiment, instructions of a compiler, if executed, cause one or more processors to perform an optimization pass to generate an optimized IR based on the input IR, as described in connection with FIG. 2 In at least one embodiment, one or more processors executing a compiler include the one or more processors executing instructions of the compiler. In at least one embodiment, one or more circuits executing a compiler include the one or more circuits executing instructions of the compiler. In at least one embodiment, one or more processors are used to execute a compiler to generate an input intermediate representation (IR) based on source code of a program. In at least one embodiment, one or more circuits are used to execute a compiler to perform an optimization pass to generate an optimized IR based on the input IR.

[0140] In at least one embodiment, at block 502, processor 160 receives an input optimization profile 112 that indicates one or more optimization passes 122. In at least one embodiment, as described in connection with FIG. 1 In at least one embodiment, input optimization profile 112 can indicate one or more optimization pass names 114 that identify optimization passes 122 to be performed by compiler 102.

[0141] In at least one embodiment, at block 504, the processor 160 generates an input IR 120A based on the source code 104 of the program. In at least one embodiment, the intermediate code generator 108 converts a set of tokens representing the source code 104 into an initial IR 118, as described with respect to FIG. 1 In at least one embodiment, the initial IR 118 is used as the input IR 120A, as described with respect to FIG. 1 As stated.

[0142] In at least one embodiment, at block 506, the processor 160 performs an optimization pass 122 to generate an optimized output IR 124 based on the input IR 120. In at least one embodiment, the processor 160 performs block 506 for each of the optimization pass names 114 such that a corresponding optimization pass 122 is performed for each respective optimization pass name in the optimization pass names 114 using a corresponding optimization specified by or otherwise associated with the respective optimization pass name. In at least one embodiment, in each respective optimization pass 122, the processor 160 performs the corresponding optimization to transform the respective input IR 120 into the respective output IR 124.

[0143] In at least one embodiment, after each respective call of block 506, the processor 160 executes block 508 to determine whether the respective output IR 124 generated by the respective optimization pass 122 is different from the respective input IR 120 of the respective optimization pass 122. In at least one embodiment, block 506 compares the output IR 124 to the input IR 120, as described with respect to FIG. FIG. 2 A comparison of the IR data structure of the optimization profile generator 226 is described.

[0144] In at least one embodiment, if the processor 160 determines at block 508 that the optimized output IR 124 is different from the input IR 120, the processor 160 executes block 510. In at least one embodiment, at block 510, the processor 160 updates the output optimization profile 140 based on the results of the corresponding optimization pass 122. In at least one embodiment, to update the output optimization profile 140, the processor 160 adds the name or other identifier of the corresponding optimization pass 122 to the output optimization profile 140. In at least one embodiment, after executing block 510, the processor 160 executes block 512.

[0145] In at least one embodiment, if processor 160 determines at block 508 that the respective optimized output IR 124 is not different from the respective input IR 120, processor 160 performs block 512. In at least one embodiment, at block 512, processor 160 determines whether another optimization pass 122 is to be performed. In at least one embodiment, optimization passes 122 will be performed for each of optimization pass names 114 specified in input optimization configuration file 112. In at least one embodiment, if processor 160 determines at block 512 that another optimization pass 122 is to be performed, processor 160 performs block 516.

[0146] In at least one embodiment, at block 516, processor 160 provides the respective output IR 124 to a next optimization pass 122 as an input IR 120. In at least one embodiment, the next optimization pass 122 is to be performed by processor 160 at block 506. In at least one embodiment, after performing block 516, processor 160 performs block 506 to perform the next optimization pass 122.

[0147] In at least one embodiment, if processor 160 determines at block 512 that another optimization pass 122 is not to be performed, processor 160 performs block 514. In at least one embodiment, at block 514, processor 160 generates executable program 150 based on optimized output IR 124. In at least one embodiment, at block 514, processor 160 generates program code, such as object code, as described with reference to target code generator 136. FIG. 1

[0148] In at least one embodiment, for example, if optimization pass names 114 include names “strength reduction,” “sub-expression elimination,” and “constant folding,” processor 160 performs a first invocation of block 506 to perform a strength reduction optimization, performs a second invocation of block 506 to perform a sub-expression elimination optimization, and performs a third invocation of block 506 to perform a constant folding optimization.

[0149] ​In at least one embodiment, the first invocation of block 506 causes block 506 to perform optimization pass A 122A, which performs strength reduction optimization on input IR 120A and generates output IR 124A. In at least one embodiment, input IR 120A is or includes the initial IR 118. In at least one embodiment, after the first invocation of block 506, processor 160 performs a first invocation of block 508. In at least one embodiment, for example, the first invocation of block 508 determines that the input IR 120A is different from the output IR 124A, and causes processor 160 to perform a first invocation of block 510. In at least one embodiment, the first invocation of block 510 updates output optimization profile 240 to include a name of the optimization pass A 122A, namely “strength reduction”. In at least one embodiment, after the first invocation of block 510, processor 160 performs a first invocation of block 512, which determines that another optimization pass (subexpression elimination) is to be performed. In at least one embodiment, after performing block 512, processor 160 causes a first invocation of block 516 to be performed. In at least one embodiment, the first invocation of block 516 provides the output IR 124A to a next optimization pass as input IR 120B. In at least one embodiment, the next optimization pass is optimization pass B 122B. In at least one embodiment, after performing block 516, processor 160 performs a second invocation of block 506.

[0150] In at least one embodiment, the second invocation of block 506 causes block 506 to perform optimization pass B 122B, which performs subexpression elimination optimization on input IR 120B and generates output IR 124B. In at least one embodiment, input IR 120B is or includes the output IR 124A. In at least one embodiment, after the second invocation of block 506, processor 160 performs a second invocation of block 508. In at least one embodiment, for example, the second invocation of block 508 determines that the output IR 124B is not different from the input IR 120B, and causes processor 160 to perform a second invocation of block 512. In at least one embodiment, the second invocation of block 512 determines that another optimization pass (constant folding) is to be performed. In at least one embodiment, after performing block 512, processor 160 causes a second invocation of block 516 to be performed. In at least one embodiment, the second invocation of block 516 provides the output IR 124B to a next optimization pass as input IR 120N. In at least one embodiment, the next optimization pass is optimization pass N 122N. In at least one embodiment, after performing block 516, processor 160 performs a third invocation of block 506.

[0151] In at least one embodiment, the third invocation of block 506 causes block 506 to perform optimization pass N 122N, which performs a constant folding optimization on input IR 120N and generates output IR 124N. In at least one embodiment, input IR 120N is or includes output IR 124B. In at least one embodiment, after the third invocation of block 506, processor 160 performs a third invocation of block 508. In at least one embodiment, the third invocation of block 508 determines that output IR 124N is different from input IR 120N and causes processor 160 to perform a second invocation of block 510. In at least one embodiment, the second invocation of block 510 updates output optimization profile 240 to include a name of optimization pass N 122N, i.e., “constant folding”. In at least one embodiment, after the second invocation of block 510, processor 160 performs a third invocation of block 512, which determines that no further optimization passes are to be performed. In at least one embodiment, after performing block 512, processor 160 performs block 514. In at least one embodiment, at block 514, processor 160 generates an executable program based on output IR 124N. In at least one embodiment, output optimization profile 240 can be used as optimization pass name 114 in a subsequent run of compiler 102 to compile source code 104 or another version of source code 104.

[0152] FIG. 6A An example of a system 600 is shown, including one or more drivers and / or one or more runtimes (shown as reference number 604) that include one or more libraries 606 to provide one or more application programming interfaces (“APIs”) 610, in accordance with at least one embodiment. In at least one embodiment, system 600 includes drivers 604 and / or runtimes 604 that include one or more libraries 606 to provide to APIs 610. In at least one embodiment, APIs 610 are sets of software instructions that, if executed, cause one or more processors (e.g., processors 202) to perform one or more operations. FIG. 6BThe one or more APIs 610 perform one or more computing operations. In at least one embodiment, one or more of the APIs 610 are distributed or otherwise provided as part of one or more of the libraries 606, one or more of the runtime 604, one or more of the drivers 604, and / or one or more components of any other software and / or executable code set further described herein. In at least one embodiment, one or more of the APIs 610 perform one or more computing operations in response to an invocation by one or more of the software programs 602.

[0153] In at least one embodiment, one or more of the software programs 602 are software modules and / or include one or more software modules. In at least one embodiment, software modules are collections of software objects and / or data that perform one or more related FIG. 6B In at least one embodiment, one or more of the software programs 602 are software modules and / or include one or more software modules. In at least one embodiment, software modules are collections of software objects and / or data that perform one or more related

[0154] In at least one embodiment, one or more of the APIs 610 are one or more hardware interfaces to one or more circuits to perform one or more computing operations. In at least one embodiment, one or more of the APIs 610 described herein are implemented as one or more circuits to perform one or more of the techniques described in conjunction with FIG. 1 to FIG. 5 In at least one embodiment, one or more of the software programs 602 include instructions that, if executed, cause one or more hardware devices and / or circuits to perform one or more of the techniques described in conjunction with FIG. 1 to FIG. 5 In at least one embodiment, one or more of the software programs 602 include instructions that, if executed, cause one or more hardware devices and / or circuits to perform one or more of the techniques described in conjunction with FIG. 1 In at least one embodiment, one or more of the software programs 602 include instructions that, if executed, cause one or more hardware devices and / or circuits to perform one or more of the techniques described in conjunction with

[0155] In at least one embodiment, the software program 602 (such as a user-implemented software program) utilizes one or more of the APIs 610 to perform various computational operations, such as memory reservations, matrix multiplications, arithmetic operations, and / or any computational operations performed by a PPU (such as a GPU), as further described herein. In at least one embodiment, the functions 612 include a set of callable functions provided by one or more of the APIs 610, which are referred to herein as APIs, API functions, software functions, and / or functions, which individually perform one or more computational operations, such as computational operations related to parallel computing. In at least one embodiment, one or more of the APIs 610 provide functions 612 to perform selection of program code optimizations, and / or to perform other operations described herein (e.g., with FIG. 1 to FIG. 5 related).

[0156] In at least one embodiment, one or more of the software programs 602 interact or otherwise communicate with one or more of the APIs 610 to utilize one or more processors (e.g., FIG. 6B 622), such as one or more PPUs (such as a GPU), to perform one or more computing operations. In at least one embodiment, the one or more computing operations using the one or more PPUs include at least one or more groups of computing operations that are accelerated at least in part by the one or more PPUs. In at least one embodiment, one or more of the software programs 602 interact with one or more of the APIs 610 to select program code optimizations based on the results of performing the optimizations on the software program 602, and / or perform other operations described herein (e.g., with FIG. 1 to FIG. 5 related).

[0157] In at least one embodiment, an interface is software instructions that, if executed, enable access to one or more of the functions 612 provided by one or more of the APIs 610. In at least one embodiment, when a software developer compiles one or more of the software programs 602 in conjunction with one or more of the libraries 606 that include or otherwise provide access to one or more of the APIs 610, the one or more of the software programs 602 use a native interface. In at least one embodiment, one or more of the software programs 602 are statically compiled in conjunction with one or more precompiled libraries of the libraries 606 and / or include uncompiled source code that executes instructions of one or more of the APIs 610. In at least one embodiment, one or more of the software programs 602 are dynamically compiled and the dynamically compiled software programs are linked to one or more precompiled libraries of the libraries 606 including one or more of the APIs 610 using a linker.

[0158] In at least one embodiment, one or more of the software programs 602 use a remote interface when a software developer executes a software program that utilizes or otherwise communicates with at least one of the libraries 606 including one or more of the APIs 610 over a network or other remote communication medium. In at least one embodiment, one or more of the libraries 606 including one or more of the APIs 610 are to be executed by a remote computing service, such as a computing resource service provider. In at least one embodiment, one or more of the libraries 606 including one or more particular APIs (of the APIs 610) are to be executed by any other computing host that provides the particular APIs to one or more of the software programs 602.

[0159] In at least one embodiment, a processor executing or using one or more particular software programs of the software programs 602 (e.g., FIG. 6BThe one or more PPUs (such as GPUs) or any other accelerators or processors further described herein. In at least one embodiment, one or more of software programs 602 request that one or more neural networks perform signal processing using one or more of functions 612 provided by one or more of APIs 610. In at least one embodiment, memory (such as memory 162 used by processor 160 of computing device of system 100) implements memory 614.

[0160] In at least one embodiment, one or more of APIs 610 are APIs for facilitating parallel computing. In at least one embodiment, one or more of APIs 610 are any other APIs further described herein. In at least one embodiment, one or more of APIs 610 are provided by one or more of drivers 604 and / or one or more of runtimes 604. In at least one embodiment, one or more of APIs 610 are provided by a CUDA user mode driver. In at least one embodiment, one or more of APIs 610 are provided by a CUDA runtime. In at least one embodiment, one or more of drivers 604 are data values and software instructions that, if executed, perform and / or otherwise facilitate operations of one or more of functions 612 of one or more of APIs 610 during loading and execution of one or more portions of at least one of software programs 602. In at least one embodiment, one or more of runtimes 604 are data values and / or software instructions that, if executed, perform or otherwise facilitate operations of one or more of functions 612 of one or more of APIs 610 during execution of at least one of software programs 602. In at least one embodiment, one or more particular ones of software programs 602 utilize one or more of APIs 610 implemented and / or otherwise provided by one or more of drivers 604 and / or one or more of runtimes 604 to perform combined arithmetic operations during execution by one or more PPUs (such as GPUs) of the particular software programs.

[0161] In at least one embodiment, one or more of software programs 602 utilize one or more of APIs 610 provided by one or more of drivers 604 and / or one or more of runtimes 604 to perform combined arithmetic operations of one or more PPUs, such as a GPU. In at least one embodiment, one or more of APIs 610 provide combined arithmetic operations by one or more of drivers 604 and / or one or more of runtimes 604, as described above. In at least one embodiment, one or more of software programs 602 utilize one or more of APIs 610 provided by one or more of drivers 604 and / or one or more of runtimes 604 to select program code optimizations based on results of performing optimizations of software programs 602 and / or to perform one or more compilers that select program code optimizations based on results of performing optimizations of software programs 602.

[0162] In at least one embodiment, to improve availability and / or improve performance of one or more particular software programs of software programs 602, one or more portions of particular software programs are to be accelerated by one or more PPUs, such as a GPU. In at least one embodiment, one or more of functions 612 receive one or more input parameters that indicate one or more inputs of one or more neural networks and / or other data to be utilized by neural networks, such as one or more hyperparameters of neural networks. In at least one embodiment, input parameters include one or more inputs and / or other data. In at least one embodiment, input parameters include one or more pointers to one or more memory locations that store inputs and / or other data.

[0163] In at least one embodiment, system 600 includes at least one processor (e.g., processor 622, shown in FIG. 6) that includes one or more circuits to execute one or more software programs to combine two or more of APIs 610 into a single API. In at least one embodiment, system 600 includes at least one processor (e.g., processor 622, shown in FIG. 6) that uses one or more of APIs 610 to select program code optimizations based on results of performing optimizations of software programs 602 and / or to otherwise perform operations described herein. In at least one embodiment, system 600 includes at least one processor (e.g., processor 622, shown in FIG. 6) that uses one or more of APIs 610 to perform combined arithmetic operations of one or more PPUs, such as a GPU. FIG. 6B FIG. 6B In at least one embodiment, system 600 includes at least one processor (e.g., processor 622, shown in FIG. 6) that includes one or more circuits to execute one or more software programs to combine two or more of APIs 610 into a single API. In at least one embodiment, system 600 includes at least one processor (e.g., processor 622, shown in FIG. 6) that uses one or more of APIs 610 to select program code optimizations based on results of performing optimizations of software programs 602 and / or to otherwise perform operations described herein. In at least one embodiment, system 600 includes at least one processor (e.g., processor 622, shown in FIG. 6) that uses one or more of APIs 610 to perform combined arithmetic operations of one or more PPUs, such as a GPU. FIG. 6B ​) that uses one or more of the APIs 610 to perform operations related to FIG. 1 To one or more of the operations shown and / or described in FIG. 8 , such as FIG. 1 To one or more processes or portions thereof shown in FIG8. In at least one embodiment, the system 600 includes at least one processor (e.g., FIG. 6B ) to perform one or more of the functions 612, such as in conjunction with FIG. 1 to FIG. 5 In at least one embodiment, one or more of the APIs 610 will be combined with FIG. 1 to FIG. 47 The hardware described is implemented.

[0164] FIG. 6B is a block diagram 620 illustrating an example processor 622 and the modules 624 according to at least one embodiment. FIG. 6B In at least one embodiment, the processor 622 may be implemented by the processor 160. In at least one embodiment, the processor 622 may perform one or more processes, such as the processes described herein with respect to selection of program code optimization, and / or may otherwise perform the operations described herein. In at least one embodiment, the processor 622 may perform one or more processes, such as in conjunction with FIG. 4 to FIG. 5 The process described.

[0165] In at least one embodiment, the processor 622 includes one or more processors, such as a processor in combination with FIG. 9 to FIG. 47 In at least one embodiment, the processor 622 can be any suitable processing unit and / or combination of processing units, such as one or more CPUs, GPUs, DPUs, GPGPUs, PPUs, and / or variants thereof. The processor 622 includes the module 624, which may include a model compiler module 626 and a deep learning compiler module 628. The modules 624 can be distributed among multiple processors that communicate via a bus, a network, by writing to a shared memory, and / or any suitable communication process, such as the communication process described herein. In at least one embodiment, the module 624 may include processor-executable instructions that implement the selection of program code optimization based on the results of performing the optimization on the software program 602.

[0166] As used in any implementation described herein, unless otherwise clear from context, a module refers to any combination of software logic, firmware logic, hardware logic, and / or circuitry configured to provide the functionality described herein. Software can be embodied as a software package, code and / or instructions set or instruction set. As used in any implementation described herein, "hardware" can include, for example, hardwired circuitry, programmable circuitry, state machine circuitry, fixed function circuitry, execution unit circuitry, and / or firmware that stores instructions executed by programmable circuitry. The modules may, collectively or individually, be embodied as circuitry that forms part of a larger system, for example, an integrated circuit (IC), system on-chip (SoC), and the like. The modules are a part of the processing unit and / or a combination of the processing unit and any suitable processing unit(s) (such as one or more CPUs, GPUs, GPGPUs, DPUs, PPUs, and / or variations thereof) that execute one or more processes.

[0167] In at least one embodiment, as used in any implementation described herein, unless otherwise clear from context, the terms such as "module" and nominalized verbs (e.g., image manager, image analyzer, analysis engine, controller, and / or other terms) each refer to any combination of software logic, firmware logic, hardware logic, and / or circuitry configured to provide the functionality described herein. In at least one embodiment, software can be embodied as a software package, code and / or instruction set or instructions, and "hardware" (as used in any implementation described herein) can include, for example, hardwired circuitry, programmable circuitry, state machine circuitry, fixed function circuitry, execution unit circuitry, and / or firmware that stores instructions executed by programmable circuitry. In at least one embodiment, modules may, collectively or individually, be embodied as circuitry that forms part of a larger system, for example, an integrated circuit (IC), system on-chip (SoC), and the like.

[0168] Data center

[0169] FIG. 7 An example data center 700 that can use at least one embodiment is shown. In at least one embodiment, data center 700 includes a data center infrastructure layer 710, a framework layer 720, a software layer 730, and an application layer 740.

[0170] In at least one embodiment, as FIG. 7As shown, the data center infrastructure layer 710 can include a resource orchestrator 712, grouped computing resources 714, and node computing resources (“node C.R.s”) 716(1)-716(N), where “N” represents any integer, positive integer. In at least one embodiment, node C.R.s 716(1)-716(N) can include, but are not limited to, any number of central processing units (“CPUs” or “processors”), including accelerators, field programmable gate arrays (FPGAs), graphics processors, etc., memory devices (e.g., dynamic read-only memory), storage devices (e.g., solid state or disk drives), network input / output (“NW I / O”) devices, network switches, virtual machines (“VMs”), power modules, and cooling modules, etc. In at least one embodiment, one or more node C.R.s of node C.R.s 716(1)-716(N) can be a server having one or more of above-described computing resources.

[0171] In at least one embodiment, grouped computing resources 714 can include individual groups of node C.R.s housed within one or more racks (not shown), or housed within a number of racks (also not shown) within various geographic locations of a data center. In at least one embodiment, individual groups of node C.R.s within grouped computing resources 714 can include groups of computing, network, memory, or storage resources that can be configured or allocated to support one or more workloads. In at least one embodiment, several node C.R.s including CPUs or processors can be grouped within one or more racks to provide computing resources to support one or more workloads. In at least one embodiment, one or more racks can also include any number of power modules, cooling modules, and network switches, in any combination.

[0172] In at least one embodiment, resource orchestrator 712 can configure or otherwise control one or more node C.R.s 716(1)-716(N) and / or grouped computing resources 714. In at least one embodiment, resource orchestrator 712 can include a software design infrastructure (“SDI”) management entity for data center 700. In at least one embodiment, resource orchestrator can include hardware, software, or some combination thereof.

[0173] In at least one embodiment, as FIG. 7As shown, the framework layer 720 includes a job scheduler 732, a configuration manager 734, a resource manager 736, and a distributed file system 738. In at least one embodiment, the framework layer 720 can include a framework that supports the software 732 of the software layer 730 and / or one or more applications 742 of the application layer 740. In at least one embodiment, the software 732 or the applications 742 can include web-based service software or applications, respectively, such as services or applications provided by Amazon Web Services, Google Cloud, and Microsoft Azure. In at least one embodiment, the framework layer 720 can be, but is not limited to, a type of free and open-source software web application framework such as the Apache Spark TM (“Spark”) that can utilize the distributed file system 738 for large-scale data processing (e.g., “big data”). In at least one embodiment, the job scheduler 732 can include a Spark driver to facilitate scheduling workloads supported by various layers of the data center 700. In at least one embodiment, the configuration manager 734 can be capable of configuring different layers, such as the software layer 730 and the framework layer 720 including Spark and the distributed file system 738 for supporting large-scale data processing. In at least one embodiment, the resource manager 736 can be capable of managing clustered or grouped computing resources mapped to or allocated for supporting the distributed file system 738 and the job scheduler 732. In at least one embodiment, the clustered or grouped computing resources can include the grouped computing resources 714 on the data center infrastructure layer 710. In at least one embodiment, the resource manager 736 can coordinate with the resource orchestrator 712 to manage these mapped or allocated computing resources.

[0174] In at least one embodiment, the software 732 included in the software layer 730 can include software used by at least a portion of the node C.R.s 716(1)-716(N), the grouped computing resources 714, and / or the distributed file system 738 of the framework layer 720. In at least one embodiment, one or more types of software can include, but are not limited to, Internet web page search software, email virus scanning software, database software, and streaming video content software.

[0175] In at least one embodiment, one or more applications 742 included in application layer 740 can include one or more types of applications used by at least portions of node C.R.s 716(1)-716(N), grouped computing resources 714, and / or distributed file system 738 of framework layer 720. In at least one embodiment, one or more types of applications can include, but are not limited to, any number and type of genomics applications, cognitive computing and machine learning applications including training or inferencing software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), or other machine learning applications used in conjunction with one or more embodiments.

[0176] In at least one embodiment, any of configuration manager 734, resource manager 736, and resource orchestrator 712 can implement any number and type of self-modification actions based on any amount and type of data acquired in any technically feasible fashion. In at least one embodiment, self-modification actions can relieve data center operators of data center 700 from making possibly poor configuration decisions and can avoid underutilization and / or poorly performing portions of a data center.

[0177] In at least one embodiment, data center 700 can include tools, services, software, or other resources to train one or more machine learning models or use one or more machine learning models to predict or infer information in accordance with one or more embodiments described herein. For example, in at least one embodiment, a machine learning model can be trained by computing weight parameters according to a neural network architecture using software and computing resources described above with respect to data center 700. In at least one embodiment, using weight parameters computed by one or more training techniques described herein, a trained machine learning model corresponding to one or more neural networks can be used to infer or predict information using resources described above with respect to data center 700.

[0178] In at least one embodiment, data center 700 can use CPUs, application specific integrated circuits (ASICs), GPUs, FPGAs, or other hardware to perform training and / or inference using resources described above. Moreover, one or more software and / or hardware resources described above can be configured as a service to allow users to train or perform information inference, such as image recognition, speech recognition, or other artificial intelligence services.

[0179] In at least one embodiment, FIG. 7 One or more systems shown in FIG. 6 are used to utilize various algorithms, formulas, and processes (e.g., in conjunction with one or more embodiments described herein) to train and / or infer information. FIG. 1 to FIG. 2those described herein) to perform selection of program code optimizations, and / or otherwise perform the operations described herein. In at least one embodiment, FIG. 7 One or more systems shown in the foregoing are used to implement one or more systems and / or processes (e.g., in conjunction with FIG. 1 to FIG. 6B those described herein), such as performing selection of program code optimizations in one or more compilers, and / or otherwise performing the operations described herein.

[0180] FIG. 8A An example of an autonomous vehicle 800 according to at least one embodiment is shown. In at least one embodiment, autonomous vehicle 800 (alternatively referred to herein as "vehicle 800") can be, but is not limited to, a passenger vehicle, such as a car, truck, bus, and / or another type of vehicle that can accommodate one or more passengers. In at least one embodiment, vehicle 800 can be a semi-tractor-trailer for hauling cargo. In at least one embodiment, vehicle 800 can be an aircraft, a robotic vehicle, or another type of vehicle.

[0181] An autonomous vehicle may be described according to the levels of automation defined by the National Highway Traffic Safety Administration ("NHTSA") and the Society of Automotive Engineers ("SAE") under the U.S. Department of Transportation, "Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles" (e.g., Standard No. J3016-201806, issued on June 15, 2018, Standard No. J3016-201609, issued on September 30, 2016, and previous and future versions of this standard). In one or more embodiments, the vehicle 800 may be capable of functioning according to one or more of the levels 1 to 5 of the autonomous driving levels. For example, in at least one embodiment, the vehicle 800 may be capable of conditional automation (Level 3), high automation (Level 4), and / or full automation (Level 5), depending on the embodiment.

[0182] In at least one embodiment, the vehicle 800 may include, but is not limited to, components such as a chassis, a body, wheels (e.g., 2, 4, 6, 8, 18, etc.), tires, axles, and other components of the vehicle. In at least one embodiment, the vehicle 800 may include, but is not limited to, a propulsion system 850, such as an internal combustion engine, a hybrid power plant, an all-electric engine, and / or another type of propulsion system. In at least one embodiment, the propulsion system 850 may be connected to a drive train of the vehicle 800, which may include, but is not limited to, a transmission, to enable propulsion of the vehicle 800. In at least one embodiment, the propulsion system 850 may be controlled in response to receiving a signal from a throttle / accelerator 852.

[0183] In at least one embodiment, when the propulsion system 850 is operating (e.g., when the vehicle is traveling), a steering system 854 (which may include, but is not limited to, a steering wheel) is used to steer the vehicle 800 (e.g., along a desired path or route). In at least one embodiment, the steering system 854 may receive signals from a steering actuator 856. In at least one embodiment, a steering wheel may be optional for fully automated (Level 5) functionality. In at least one embodiment, a brake sensor system 846 may be used to operate the vehicle brakes in response to signals received from a brake actuator 848 and / or brake sensors.

[0184] In at least one embodiment, the controller 836 may include, but is not limited to, one or more system-on-chips ("SoCs") ( FIG. 8Aand / or a graphics processing unit (“GPU”) to provide signals (e.g., representative of commands) to one or more components and / or systems of vehicle 800. For example, in at least one embodiment, controller(s) 836 can send signals to operate vehicle brakes by brake actuator(s) 848, to operate steering system 854 by one or more steering actuators 856, to operate propulsion system 850 by one or more throttle / accelerator 852. In at least one embodiment, controller(s) 836 can include one or more on-board (e.g., integrated) computing devices (e.g., supercomputers) that process sensor signals and output operational commands (e.g., signals representative of commands) to enable autonomous driving and / or assist a human driver in driving vehicle 800. In at least one embodiment, controller(s) 836 can include a first controller 836 for autonomous driving functionality, a second controller 836 for functional safety functionality, a third controller 836 for artificial intelligence functionality (e.g., computer vision), a fourth controller 836 for infotainment functionality, a fifth controller 836 for redundancy in emergency situations, and / or other controllers. In at least one embodiment, a single controller 836 can handle two or more of above-described functionalities, two or more controllers 836 can handle a single functionality, and / or any combination thereof.

[0185] In at least one embodiment, controller(s) 836 provide signals for controlling one or more components and / or systems of vehicle 800 in response to sensor data received from one or more sensors (e.g., sensor inputs). In at least one embodiment, sensor data can be received from sensor types such as, but not limited to, one or more global navigation satellite system (“GNSS”) sensors 858 (e.g., one or more global positioning system sensors), one or more RADAR sensors 860, one or more ultrasonic sensors 862, one or more LIDAR sensors 864, one or more inertial measurement unit (IMU) sensors 866 (e.g., one or more accelerometers, one or more gyroscopes, one or more magnetometers, one or more magnetorquers, etc.), one or more microphones 896, one or more stereo cameras 868, one or more wide-view cameras 870 (e.g., fisheye cameras), one or more infrared cameras 872, one or more surround cameras 874 (e.g., 360 degree cameras), long-range cameras (not shown in FIG. 8), mid-range cameras (not shown in FIG. 8), and / or other sensors. FIG. 8A In at least one embodiment, controller(s) 836 provide signals for controlling one or more components and / or systems of vehicle 800 in response to sensor data received from one or more sensors (e.g., sensor inputs). In at least one embodiment, sensor data can be received from sensor types such as, but not limited to, one or more global navigation satellite system (“GNSS”) sensors 858 (e.g., one or more global positioning system sensors), one or more RADAR sensors 860, one or more ultrasonic sensors 862, one or more LIDAR sensors 864, one or more inertial measurement unit (IMU) sensors 866 (e.g., one or more accelerometers, one or more gyroscopes, one or more magnetometers, one or more magnetorquers, etc.), one or more microphones 896, one or more stereo cameras 868, one or more wide-view cameras 870 (e.g., fisheye cameras), one or more infrared cameras 872, one or more surround cameras 874 (e.g., 360 degree cameras), long-range cameras (not shown in FIG. 8), mid-range cameras (not shown in FIG. 8), and / or other sensors. FIG. 8A), one or more speed sensors 844 (e.g., for measuring the speed of the vehicle 800), one or more vibration sensors 842, one or more steering sensors 840, one or more brake sensors (e.g., as part of a brake sensor system 846), and / or other sensor types are received.

[0186] In at least one embodiment, one or more controllers 836 may receive input (e.g., represented by input data) from a dashboard 832 of the vehicle 800 and provide output (e.g., represented by output data, display data, etc.) via a human machine interface ("HMI") display 834, an audible annunciator, a speaker, and / or other components of the vehicle 800. In at least one embodiment, the output may include information such as vehicle speed, velocity, time, map data (e.g., high definition map ( FIG. 8A ), location data (e.g., the location of the vehicle 800, such as on a map), directions, the locations of other vehicles (e.g., occupancy barriers), information about objects and the states of objects sensed by the one or more controllers 836, etc. For example, in at least one embodiment, the HMI display 834 can display information about the presence of one or more objects (e.g., road signs, warning signs, traffic light changes, etc.) and / or information about the driving maneuvers the vehicle has made, is making, or will make (e.g., changing lanes now, taking exit 34B in two miles, etc.).

[0187] In at least one embodiment, the vehicle 800 further includes a network interface 824 that can communicate over one or more networks using one or more wireless antennas 826 and / or one or more modems. For example, in at least one embodiment, the network interface 824 may be capable of communicating over Long Term Evolution ("LTE"), Wideband Code Division Multiple Access ("WCDMA"), Universal Mobile Telecommunications System ("UMTS"), Global System for Mobile Communications ("GSM"), IMT-CDMA Multi-Carrier ("CDMA2000") networks, etc. In at least one embodiment, the one or more wireless antennas 826 can also use one or more local area networks (e.g., Bluetooth, Bluetooth Low Energy (LE), Z-Wave, ZigBee, etc.) and / or one or more low power wide area networks (hereinafter referred to as "LPWAN") (e.g., protocols such as LoRaWAN, SigFox, etc.) to enable communication between objects in the environment (e.g., vehicles, mobile devices).

[0188] In at least one embodiment, FIG. 8A One or more systems shown in the foregoing are used to utilize various algorithms, formulas, and processes (e.g., in conjunction with FIG. 1 to FIG. 2those described herein) to perform selection of program code optimizations, and / or otherwise perform the operations described herein. In at least one embodiment, FIG. 8A One or more systems shown in the foregoing are used to implement one or more systems and / or processes (e.g., in conjunction with FIG. 1 to FIG. 6B those described herein), such as performing selection of program code optimizations in one or more compilers, and / or otherwise performing the operations described herein.

[0189] FIG. 8B According to at least one embodiment, FIG. 8A An example of camera positions and fields of view for an autonomous vehicle 800 is shown. In at least one embodiment, the cameras and respective fields of view are an example embodiment and are not intended to be limiting. For example, in at least one embodiment, additional and / or alternative cameras may be included and / or the cameras may be located in different locations on the vehicle 800.

[0190] In at least one embodiment, the camera type used for the camera may include, but is not limited to, a digital camera that may be suitable for use with components and / or systems of the vehicle 800. In at least one embodiment, one or more cameras may operate at Automotive Safety Integrity Level ("ASIL") B and / or other ASILs. In at least one embodiment, the camera type may have any image capture rate, such as 60 frames per second (fps), 120fps, 240fps, etc., depending on the embodiment. In at least one embodiment, the camera may be capable of using a rolling shutter, a global shutter, another type of shutter, or a combination thereof. In at least one embodiment, the color filter array may include a red-clear-clear ("RCCC") filter array, a red-clear-clear-blue ("RCCB") filter array, a red-blue-green-clear ("RBGC") filter array, a Foveon X3 filter array, a Bayer sensor ("RGGB") filter array, a monochrome sensor filter array, and / or other types of filter arrays. In at least one embodiment, a clear pixel camera, such as one having an RCCC, RCCB, and / or RBGC color filter array, may be used in an effort to increase photosensitivity.

[0191] In at least one embodiment, one or more cameras can be used to perform advanced driver assistance system ("ADAS") functions (e.g., as part of a redundant or fail-safe design). For example, in at least one embodiment, a multi-function mono camera can be installed to provide functions including lane departure warning, traffic sign assistance, and intelligent headlight control. In at least one embodiment, one or more cameras (e.g., all cameras) can simultaneously record and provide image data (e.g., video).

[0192] In at least one embodiment, one or more cameras can be mounted in a mounting assembly, such as a custom designed (three-dimensional ("3D") printed) assembly, so as to cut out stray light and reflections from within the car (e.g., reflections from the dashboard reflecting in the windshield mirror) that may interfere with the camera's ability to capture image data. With respect to the rearview mirror mounting assembly, in at least one embodiment, the rearview mirror assembly can be 3D printed custom so that the camera mounting plate matches the shape of the rearview mirror. In at least one embodiment, one or more cameras can be integrated into the rearview mirror. In at least one embodiment, for side-view cameras, one or more cameras can also be integrated into the four pillars at each corner of the car.

[0193] In at least one embodiment, a camera (e.g., a forward-facing camera) having a field of view that includes a portion of the environment in front of the vehicle 800 can be used for surround vision, as well as to help identify the forward path and obstacles with the assistance of one or more controllers 836 and / or control SoCs, thereby providing information that is critical for generating an occupancy grid and / or determining a preferred vehicle path. In at least one embodiment, the forward-facing camera can be used to perform many of the same ADAS functions as LIDAR, including but not limited to emergency braking, pedestrian detection, and collision avoidance. In at least one embodiment, the forward-facing camera can also be used for ADAS functions and systems, including but not limited to lane departure warning ("LDW"), automatic cruise control ("ACC"), and / or other functions (e.g., traffic sign recognition).

[0194] In at least one embodiment, various cameras can be used in a forward-facing configuration, including, for example, a monocular camera platform including a CMOS ("Complementary Metal Oxide Semiconductor") color imager. In at least one embodiment, a wide-angle camera 870 can be used to sense objects entering from the periphery (e.g., pedestrians, people crossing the road, or bicycles). Although FIG. 8B Only one wide-angle camera 870 is shown in FIG. 1 , however, in other embodiments, any number (including zero) of wide-angle cameras can be present on the vehicle 800. In at least one embodiment, any number of remote cameras 898 (e.g., a remote stereo camera pair) can be used for depth-based object detection, particularly for objects for which a neural network has not yet been trained. In at least one embodiment, the remote cameras 898 can also be used for object detection and classification, as well as basic object tracking.

[0195] In at least one embodiment, any number of stereo cameras 868 may also be included in the forward configuration. In at least one embodiment, one or more stereo cameras 868 may include an integrated control unit including a scalable processing unit that may provide programmable logic ("FPGA") and a multi-core microprocessor with a controller area network ("CAN") or Ethernet interface integrated on a single chip. In at least one embodiment, such a unit may be used to generate a 3D map of the environment of the vehicle 800, including distance estimates for all points in the image. In at least one embodiment, one or more stereo cameras 868 may include, but are not limited to, a compact stereo vision sensor, which may include, but are not limited to, two camera lenses (one on each side) and an image processing chip that may measure the distance from the vehicle 800 to the target object and use the generated information (e.g., metadata) to activate autonomous emergency braking and lane departure warning functions. In at least one embodiment, other types of stereo cameras 868 may be used in addition to those described herein.

[0196] In at least one embodiment, a camera having a field of view of a portion of the environment including the sides of the vehicle 800 (e.g., a side-view camera) can be used for surround viewing to provide information for creating and updating occupancy grids and generating side collision warnings. For example, in at least one embodiment, the surround camera 874 (e.g., FIG. 8B Four surround cameras 874 (shown) can be positioned on the vehicle 800. In at least one embodiment, the one or more surround cameras 874 can include, but are not limited to, any number and combination of wide-angle cameras 870, one or more fish-eye lenses, one or more 360-degree cameras, and / or the like. For example, in at least one embodiment, four fish-eye lens cameras can be located on the front, rear, and sides of the vehicle 800. In at least one embodiment, the vehicle 800 can use three surround cameras 874 (e.g., left, right, and rear), and can utilize one or more other cameras (e.g., a forward-facing camera) as a fourth surround-view camera.

[0197] In at least one embodiment, a camera having a field of view that includes a portion of the environment behind the vehicle 800 (e.g., a rearview camera) can be used for parking assistance, surround view, rear collision warning, and creating and updating occupancy grids. In at least one embodiment, a variety of cameras can be used, including but not limited to cameras that are also suitable as one or more forward-facing cameras (e.g., long-range camera 898 and / or one or more mid-range cameras 876, one or more stereo cameras 868, one or more infrared cameras 872, etc.), as described herein.

[0198] In at least one embodiment, FIG. 8BOne or more systems shown in the foregoing are used to utilize various algorithms, formulas, and processes (e.g., in conjunction with FIG. 1 to FIG. 2 those described herein) to perform selection of program code optimizations, and / or otherwise perform the operations described herein. In at least one embodiment, FIG. 8B One or more systems shown in the foregoing are used to implement one or more systems and / or processes (e.g., in conjunction with FIG. 1 to FIG. 6B those described herein), such as performing selection of program code optimizations in one or more compilers, and / or otherwise performing the operations described herein.

[0199] FIG. 8C According to at least one embodiment, FIG. 8A A block diagram of an example system architecture for an autonomous vehicle 800 is provided. In at least one embodiment, FIG. 8C Each of one or more components, one or more features, and one or more systems of vehicle 800 is shown as being connected via bus 802. In at least one embodiment, bus 802 may include, but is not limited to, a CAN data interface (alternatively referred to herein as a "CAN bus"). In at least one embodiment, CAN may be a network internal to vehicle 800 that helps control various features and functions of vehicle 800, such as brake actuation, acceleration, braking, steering, wipers, etc. In one embodiment, bus 802 may be configured to have dozens or even hundreds of nodes, each with its own unique identifier (e.g., a CAN ID). In at least one embodiment, bus 802 may be read to find steering wheel angle, ground speed, engine revolutions per minute ("RPM"), button position, and / or other vehicle status indicators. In at least one embodiment, bus 802 may be an ASIL B compliant CAN bus.

[0200] In at least one embodiment, FlexRay and / or Ethernet can be used in addition to or instead of CAN. In at least one embodiment, there can be any number of buses 802, which can include, but are not limited to, zero or more CAN buses, zero or more FlexRay buses, zero or more Ethernet buses, and / or zero or more other types of buses that use other protocols. In at least one embodiment, two or more buses 802 can be used to perform different functions, and / or can be used for redundancy. For example, a first bus 802 can be used for collision avoidance functions, and a second bus 802 can be used for actuation control. In at least one embodiment, each bus 802 can be in communication with any component of vehicle 800, and two or more of buses 802 can be in communication with the same component. In at least one embodiment, each of any number of system on a chip (“SoC”) 804, each of one or more controllers 836, and / or each computer within a vehicle can have access to the same input data (e.g., input from sensors of vehicle 800), and can be connected to a common bus, such as a CAN bus.

[0201] In at least one embodiment, vehicle 800 can include one or more controllers 836, such as those described herein with respect to FIG. 8A In at least one embodiment, controllers 836 can be used for a variety of functions. In at least one embodiment, controllers 836 can be coupled to any of various other components and systems of vehicle 800, and can be used to control vehicle 800, artificial intelligence of vehicle 800, infotainment of vehicle 800, and / or other functions.

[0202] In at least one embodiment, vehicle 800 can include any number of SoCs 804. Each of SoCs 804 can include, without limitation, central processing units (“one or more CPUs”) 806, graphics processing units (“one or more GPUs”) 808, one or more processors 810, one or more caches 812, one or more accelerators 814, one or more data stores 816, and / or other non- shown components and features. In at least one embodiment, one or more SoCs 804 can be used to control vehicle 800 in a variety of platforms and systems. For example, in at least one embodiment, one or more SoCs 804 can be combined with a high definition (“HD”) map 822 in a system (e.g., a system of vehicle 800) that can obtain map refreshes and / or updates from one or more servers (not shown in FIG. 8) via a network interface 824. FIG. 8C

[0203] ​In at least one embodiment, CPU(s) 806 can include a CPU cluster or CPU complex (alternatively referred to herein as a “CCPLEX”). In at least one embodiment, CPU(s) 806 can include multiple cores and / or level two (“L2”) caches. For example, in at least one embodiment, CPU(s) 806 can include eight cores in a multi-processor configuration coupled to one another. In at least one embodiment, CPU(s) 806 can include four dual-core clusters with each cluster having a dedicated L2 cache (e.g., 2 MB L2 cache). In at least one embodiment, CPU(s) 806 (e.g., CCPLEX) can be configured to support simultaneous cluster operation such that any combination of clusters of CPU(s) 806 can be active at any given time.

[0204] In at least one embodiment, CPU(s) 806 can implement power management functionality including, without limitation, one or more of the following features: individual hardware modules can be automatically clock-gated at idle to conserve dynamic power; each core clock can be gated when that core is not actively executing instructions due to execution of a wait for interrupt (“WFI”) / wait for event (“WFE”) instruction; each core can be independently powered; each core cluster can be independently clock-gated when all cores are clock-gated or power-gated; and / or each core cluster can be independently power-gated when all cores are power-gated. In at least one embodiment, CPU(s) 806 can further implement enhanced algorithms for managing power states with allowed power states and expected wake-up times specified and hardware / microcode determining optimal power states for core, cluster, and CCPLEX inputs. In at least one embodiment, processing cores can support a simplified power state input sequence in software with work offloaded to microcode. In at least one embodiment, processing cores are referred to as compute units or arithmetic logic units.

[0205] In at least one embodiment, GPU(s) 808 can include an integrated GPU (also referred to herein as an “iGPU”). In at least one embodiment, GPU(s) 808 can be programmable and can be efficient for parallel workloads. In at least one embodiment, GPU(s) 808 can use an enhanced tensor instruction set. In one embodiment, GPU(s) 808 can include one or more streaming microprocessors, where each streaming microprocessor can include a level one (“LI”) cache (e.g., an LI cache with at least 96 KB of storage capacity), and two or more streaming microprocessors can share an L2 cache (e.g., an L2 cache with 512 KB storage capacity). In at least one embodiment, GPU(s) 808 can include at least eight streaming microprocessors. In at least one embodiment, GPU(s) 808 can use a compute application programming interface (“API”). In at least one embodiment, GPU(s) 808 can use one or more parallel computing platforms and / or programming models (e.g., NVIDIA’s CUDA).

[0206] In at least one embodiment, GPU(s) 808 can be power-optimized to achieve best performance in automotive and embedded use cases. For example, in one embodiment, GPU(s) 808 can be fabricated on a finned field effect transistor (“FinFET”). In at least one embodiment, each streaming microprocessor can contain a plurality of mixed-precision processing cores divided into a plurality of blocks. For example, and without limitation, 64 PF32 cores and 32 PF64 cores can be divided into four processing blocks. In at least one embodiment, each processing block can be allocated 16 FP32 cores, 8 FP64 cores, 16 INT32 cores, two mixed-precision NVIDIA Tensor Cores for deep learning matrix arithmetic, a level zero (“L0”) instruction cache, a warp scheduler, a dispatch unit, and / or a 64 KB register file. In at least one embodiment, a streaming microprocessor can include independent parallel integer and floating point data paths to provide efficient execution of workloads that mix compute and address operations. In at least one embodiment, a streaming microprocessor can include independent thread scheduling capabilities to enable finer-grain synchronization and cooperation between parallel threads. In at least one embodiment, a streaming microprocessor can include a combined LI data cache and shared memory unit to enable improved performance while simplifying programming.

[0207] In at least one embodiment, one or more GPU(s) 808 can include high bandwidth memory (“HBM”) and / or 16 GB HBM2 memory subsystems to provide, in some examples, a peak memory bandwidth of about 900 GB / sec. In at least one embodiment, in addition to, or instead of, HBM memory, synchronous graphics random access memory (“SGRAM”) can be used, for example, graphics double data rate type five synchronous random access memory (“GDDR5”).

[0208] In at least one embodiment, one or more GPU(s) 808 can include unified memory technology. In at least one embodiment, address translation services (“ATS”) support can be used to allow one or more GPU(s) 808 to directly access one or more CPU(s) 806 page tables. In at least one embodiment, when one or more GPU(s) 808 memory management unit (“MMU”) experiences a miss, an address translation request can be sent to one or more CPU(s) 806. In response, one or more CPU(s) 806 can look up a virtual-to-physical mapping for an address in their page tables and transmit the translation back to one or more GPU(s) 808, in at least one embodiment. In at least one embodiment, unified memory technology can allow a single unified virtual address space to be used for memory of both one or more CPU(s) 806 and one or more GPU(s) 808, simplifying programming for one or more GPU(s) 808 and porting applications to one or more GPU(s) 808.

[0209] In at least one embodiment, one or more GPU(s) 808 can include any number of access counters that can track how frequently one or more GPU(s) 808 is accessing memory of other processors. In at least one embodiment, one or more access counters can help ensure that memory pages are moved into physical memory of a processor that most frequently accesses the page, improving efficiency of memory ranges shared between processors.

[0210] In at least one embodiment, one or more SoC(s) 804 can include any number of caches 812, including those described herein. For example, in at least one embodiment, one or more cache(s) 812 can include a level three (“L3”) cache that can be available to one or more CPU(s) 806 and one or more GPU(s) 808 (e.g., connected to CPU(s) 806 and GPU(s) 808). In at least one embodiment, one or more cache(s) 812 can include a write-back cache that can track state of lines, for example, by using a cache coherence protocol (e.g., MEI, MESI, MSI, etc.). In at least one embodiment, although smaller cache sizes can be used, L3 cache can include 4MB or more, depending on embodiment.

[0211] In at least one embodiment, one or more SoC(s) 804 can include one or more accelerator(s) 814 (e.g., hardware accelerators, software accelerators, or a combination thereof). In at least one embodiment, one or more SoC(s) 804 can include a hardware acceleration cluster that can include optimized hardware accelerators and / or large on-chip memory. In at least one embodiment, large on-chip memory (e.g., 4MB of SRAM) can enable hardware acceleration cluster to accelerate neural networks and other computations. In at least one embodiment, hardware acceleration cluster can be used to supplement and offload some tasks of one or more GPU(s) 808 (e.g., freeing up more cycles of one or more GPU(s) 808 to perform other tasks). In at least one embodiment, one or more accelerator(s) 814 can be used for target workloads (e.g., perception, convolutional neural networks (“CNNs”), recurrent neural networks (“RNNs”), etc.) that are stable enough to withstand the speedup test. In at least one embodiment, CNNs can include region-based or region with convolutional neural networks (“RCNNs”) and fast RCNNs (e.g., as used for object detection) or other types of CNNs.

[0212] In at least one embodiment, one or more accelerators 814 (e.g., hardware acceleration clusters) can include one or more deep learning accelerators (“DLAs”). One or more DLAs can include, without limitation, one or more Tensor Processing Units (“TPUs”) that can be configured to provide an additional 100 trillion operations per second for deep learning applications and inferencing. In at least one embodiment, a TPU can be an accelerator configured and optimized for performing image processing functions (e.g., for CNNs, RCNNs, etc.). One or more DLAs can be further optimized for a particular set of neural network types and floating point operations, and inferencing. In at least one embodiment, design of one or more DLAs can provide higher performance per mm than a typical general purpose GPU, and often significantly outperform CPUs. In at least one embodiment, one or more TPUs can perform several functions including support for INT8, INT16, and FP16 data types for features and weights, single instance convolution functionality, and post-processor functionality, for example. In at least one embodiment, one or more DLAs can quickly and efficiently execute neural networks, especially CNNs, on processed or unprocessed data for any of a variety of functions including, for example and without limitation: CNNs for object recognition and detection using data from camera sensors; CNNs for distance estimation using data from camera sensors; CNNs for emergency vehicle detection, and identification and detection using data from microphones 896; CNNs for facial recognition and vehicle owner identification using data from camera sensors; and / or CNNs for safety and / or safety related events.

[0213] In at least one embodiment, a DLA can perform any of functions of one or more GPU(s) 808, and by using an inferencing accelerator, for example, a designer can target one or more DLAs or one or more GPU(s) 808 for any function. For example, in at least one embodiment, a designer can concentrate processing and floating point operations for CNNs on one or more DLAs, and leave other functions to one or more GPU(s) 808 and / or other one or more accelerators 814.

[0214] In at least one embodiment, one or more accelerators 814 (e.g., hardware acceleration clusters) can include programmable vision accelerators (“PVAs”), which can alternatively be referred to herein as computer vision accelerators. In at least one embodiment, one or more PVAs can be designed and configured to accelerate computer vision algorithms for advanced driver assistance systems (“ADAS”) 838, autonomous driving, augmented reality (“AR”) applications, and / or virtual reality (“VR”) applications. In at least one embodiment, one or more PVAs can strike a balance between performance and flexibility. For example, in at least one embodiment, each of one or more PVAs can include, without limitation, any number of reduced instruction set computer (“RISC”) cores, direct memory access (“DMA”), and / or any number of vector processors.

[0215] In at least one embodiment, RISC cores can interact with image sensors (e.g., image sensors of any camera described herein), image signal processors, etc. In at least one embodiment, each RISC core can include any number of memories. In at least one embodiment, RISC cores can use any of a number of protocols, depending on embodiment. In at least one embodiment, RISC cores can execute a real-time operating system (“RTOS”). In at least one embodiment, RISC cores can be implemented using one or more integrated circuit devices, application specific integrated circuits (“ASICs”), and / or memory devices. For example, in at least one embodiment, RISC cores can include instruction caches and / or tightly coupled RAM.

[0216] In at least one embodiment, DMA can enable components of a PVA to access system memory independently of one or more CPUs 806. In at least one embodiment, DMA can support any number of features for providing optimizations to a PVA, including, without limitation, support for multi-dimensional addressing and / or circular addressing. In at least one embodiment, DMA can support up to six or more dimensions of addressing, which can include, without limitation, block width, block height, block depth, horizontal block stride, vertical block stride, and / or depth stride.

[0217] In at least one embodiment, vector processors can be programmable processors that can be designed to efficiently and flexibly execute programming for computer vision algorithms and provide signal processing capabilities. In at least one embodiment, a PVA can include a PVA core and two vector processing subsystem partitions. In at least one embodiment, a PVA core can include a processor subsystem, DMA engines (e.g., two DMA engines), and / or other peripherals. In at least one embodiment, a vector processing subsystem can function as a primary processing engine for a PVA and can include a vector processing unit (“VPU”), an instruction cache, and / or a vector memory (e.g., “VMEM”). In at least one embodiment, a VPU core can include a digital signal processor, such as a single instruction multiple data (“SIMD”), very long instruction word (“VLIW”) digital signal processor. In at least one embodiment, a combination of SIMD and VLIW can improve throughput and speed.

[0218] In at least one embodiment, each vector processor can include an instruction cache and can be coupled to a dedicated memory. As a result, in at least one embodiment, each vector processor can be configured to execute independently of other vector processors. In at least one embodiment, vector processors included in a particular PVA can be configured to employ data parallelism. For example, in at least one embodiment, multiple vector processors included in a single PVA can execute the same computer vision algorithm, except on different regions of an image. In at least one embodiment, vector processors included in a particular PVA can execute different computer vision algorithms on the same image simultaneously, or even different algorithms on sequential images or portions of images. In at least one embodiment, any number of PVAs can be included in a hardware acceleration cluster, and any number of vector processors can be included in each PVA, among other things. In at least one embodiment, a PVA can include additional error correcting code (“ECC”) memory to enhance overall system security.

[0219] In at least one embodiment, one or more accelerators 814 (e.g., hardware acceleration clusters) can include an on-chip computer vision network and static random access memory (“SRAM”) for providing high bandwidth, low latency SRAM for one or more accelerators 814. In at least one embodiment, on-chip memory can include at least 4 MB of SRAM that includes, for example and without limitation, eight field-programmable memory blocks that are accessible by both PVA and DLA. In at least one embodiment, each pair of memory blocks can include an advanced peripheral bus (“APB”) interface, configuration circuitry, a controller, and a multiplexer. In at least one embodiment, any type of memory can be used. In at least one embodiment, PVA and DLA can access memory via a backbone that provides PVA and DLA with high-speed access to memory. In at least one embodiment, a backbone can include an on-chip computer vision network that interconnects PVA and DLA to memory (e.g., using APB).

[0220] In at least one embodiment, an on-chip computer vision network can include an interface that determines that both PVA and DLA provide ready and valid signals before transmitting any control signals / addresses / data. In at least one embodiment, an interface can provide separate phases and separate channels for transmitting control signals / addresses / data, as well as burst-type communication for continuous data transfer. In at least one embodiment, although other standards and protocols can be used, an interface can comply with International Organization for Standardization (“ISO”) 26262 or International Electrotechnical Commission (“IEC”) 61508 standards.

[0221] In at least one embodiment, one or more SoC 804 can include a real-time line-of-sight tracking hardware accelerator. In at least one embodiment, a real-time line-of-sight tracking hardware accelerator can be used to quickly and efficiently determine locations and ranges of objects (e.g., within a world model) to generate real-time visualizations simulations for RADAR signal interpretation, for sound propagation synthesis and / or analysis, for simulations of SONAR systems, for general wave propagation simulations, for comparison with LIDAR data for positioning and / or other functions, and / or for other uses.

[0222] In at least one embodiment, one or more accelerators 814 (e.g., hardware acceleration clusters) have broad use for autonomous driving. In at least one embodiment, a PVA can be a programmable vision accelerator that is used for key processing stages in ADAS and autonomous cars. In at least one embodiment, capabilities of a PVA at low power and low latency are well matched to algorithmic domains that require predictable processing. In other words, PVAs excel at semi-dense or dense regular computations, even on small data sets that can require predictable runtimes with low latency and low power. In at least one embodiment, autonomous vehicles, such as in vehicle 800, PVAs can be designed to run classic computer vision algorithms as they can be efficient at object detection and integer math operations.

[0223] For example, in accordance with at least one embodiment of technology, a PVA is used to perform computer stereo vision. In at least one embodiment, a semi-global matching based algorithm can be used in some examples, although this is not meant to be limiting. In at least one embodiment, applications for level 3-5 autonomous driving use dynamic estimation / stereo matching in run (e.g., structure from motion, pedestrian recognition, lane detection, etc.). In at least one embodiment, a PVA can perform computer stereo vision functions on inputs from two monocular cameras.

[0224] In at least one embodiment, a PVA can be used to perform dense optical flow. For example, in at least one embodiment, a PVA can process raw RADAR data (e.g., using a 4D fast Fourier transform) to provide processed RADAR data. In at least one embodiment, a PVA is used for time-of-flight depth processing, e.g., by processing raw time-of-flight data to provide processed time-of-flight data.

[0225] In at least one embodiment, DLA can be used to run any type of network to enhance control and driving safety, including, for example and without limitation, a neural network that outputs a confidence level for each object detection. In at least one embodiment, confidence level can be represented or interpreted as a probability, or as providing a relative “weight” of each detection relative to other detections. In at least one embodiment, confidence level enables system to make further decisions as to which detections should be considered as true positive detections and not false positive detections. In at least one embodiment, system can set a threshold for confidence level, and only consider detections that exceed threshold as true positive detections. In embodiments using automatic emergency braking (“AEB”) systems, false positive detections would result in vehicle automatically performing emergency braking, which is obviously undesirable. In at least one embodiment, highly confident detections can be considered as triggers for AEB. In at least one embodiment, DLA can run a neural network for regression of confidence values. In at least one embodiment, neural network can take as its input at least some subset of parameters, such as bounding box size, ground plane estimate obtained (e.g., from another subsystem), output of one or more IMU sensors 866 related to object’s vehicle 800 direction, distance, 3D position estimate obtained from neural network and / or other sensors (e.g., one or more LIDAR sensors 864 or one or more RADAR sensors 860), etc.

[0226] In at least one embodiment, one or more SoC(s) 804 (e.g., hardware acceleration cluster(s)) can include one or more data storage devices 816 (e.g., memory). In at least one embodiment, one or more data storage 816 can be on-chip memory of one or more SoC(s) 804, which can store neural networks to be executed on one or more GPU(s) 808 and / or DLA. In at least one embodiment, one or more data storage 816 can have capacity large enough to store multiple instances of a neural network for redundancy and safety. In at least one embodiment, one or more data storage 812 can include L2 or L3 cache.

[0227] In at least one embodiment, one or more SoC(s) 804 can include any number of processor(s) 810 (e.g., embedded processors). In at least one embodiment, one or more processor(s) 810 can include a boot and power management processor that can be a dedicated processor and subsystem to handle boot power and management functions and related security enforcement. In at least one embodiment, a boot and power management processor can be part of a one or more SoC(s) 804 boot sequence and can provide run-time power management services. In at least one embodiment, a boot power and management processor can provide clock and voltage programming, assist system low power state transitions, one or more SoC(s) 804 thermal and temperature sensor management, and / or one or more SoC(s) 804 power state management. In at least one embodiment, each temperature sensor can be implemented as a ring oscillator whose output frequency is proportional to temperature, and one or more SoC(s) 804 can use ring oscillators to detect temperature of one or more CPU(s) 806, one or more GPU(s) 808, and / or one or more accelerator(s) 814. In at least one embodiment, if a temperature is determined to exceed a threshold, a boot and power management processor can enter a temperature fault routine and put one or more SoC(s) 804 into a lower power state and / or put vehicle 800 into a safe park pattern for the driver (e.g., cause vehicle 800 to safely park).

[0228] In at least one embodiment, one or more processor(s) 810 can further include a set of embedded processors that can function as an audio processing engine. In at least one embodiment, an audio processing engine can be an audio subsystem that is capable of providing full hardware support for multi-channel audio to hardware through a number of interfaces as well as a broad and flexible range of audio I / O interfaces. In at least one embodiment, an audio processing engine is a dedicated processor core with a digital signal processor with dedicated RAM.

[0229] In at least one embodiment, one or more processor(s) 810 can further include an always-on processor engine. In at least one embodiment, an always-on processor engine can provide necessary hardware features to support low-power sensor management and wake-up use cases. In at least one embodiment, a processor on an always-on processor engine can include, but is not limited to, a processor core, tightly coupled RAM, support peripherals (e.g., timers and interrupt controllers), various I / O controller peripherals, and routing logic.

[0230] In at least one embodiment, one or more processors 810 can further include a safety cluster engine including, without limitation, a dedicated processor subsystem for handling safety management for automotive applications. In at least one embodiment, safety cluster engine can include, without limitation, two or more processor cores, tightly coupled RAM, supporting peripherals (e.g., timers, interrupt controllers, etc.), and / or routing logic. In a safety mode, in at least one embodiment, two or more cores can operate in a lockstep mode and can function as a single core with comparison logic to detect any differences between their operations. In at least one embodiment, one or more processors 810 can further include a real-time camera engine that can include, without limitation, a dedicated processor subsystem for handling real-time camera management. In at least one embodiment, one or more processors 810 can further include a high dynamic range signal processor that can include, without limitation, an image signal processor that is a hardware engine that is part of a camera processing pipeline.

[0231] In at least one embodiment, one or more processors 810 can include a video image compositor that can be a processing block (e.g., implemented on a microprocessor) that implements video post-processing functions needed by a video playback application to produce a final video to produce a final image for a player window. In at least one embodiment, video image compositor can perform lens distortion correction on one or more wide-view cameras 870, one or more surround cameras 874, and / or one or more in-cabin monitoring camera sensors. In at least one embodiment, preferably, in-cabin monitoring camera sensors are monitored by a neural network running on another instance of SoC 804 that is configured to identify cabin events and respond accordingly. In at least one embodiment, in-cabin systems can perform, without limitation, lip reading to activate cellular service and place a phone call, dictate an email, change a destination of a vehicle, activate or change an infotainment system and settings of a vehicle, or provide voice-activated web surfing. In at least one embodiment, certain functionality is available to a driver when a vehicle is operating in an autonomous mode, otherwise it is disabled.

[0232] In at least one embodiment, video image compositor can include enhanced temporal noise reduction for simultaneous spatial and temporal noise reduction. For example, in at least one embodiment, where motion occurs in a video, noise reduction appropriately weights spatial information, reducing a weight of information provided by adjacent frames. In at least one embodiment, where an image or portion of an image does not include motion, temporal noise reduction performed by video image compositor can use information from a previous image to reduce noise in a current image.

[0233] In at least one embodiment, video image compositor can also be configured to perform stereo correction on input stereoscopic lens frames. In at least one embodiment, when using an operating system desktop, video image compositor can also be used for user interface composition and one or more GPUs 808 are not required to continuously render new surfaces. In at least one embodiment, when one or more GPUs 808 are powered and active for 3D rendering, video image compositor can be used to offload one or more GPUs 808 to improve performance and responsiveness.

[0234] In at least one embodiment, one or more SoCs in SoC 804 can further include a mobile industry processor interface (“MIPI”) camera serial interface for receiving video and input from cameras, a high-speed interface, and / or a video input block that can be used for camera and related pixel input functionality. In at least one embodiment, one or more SoCs 804 can further include an input / output controller that can be controlled by software and can be used to receive I / O signals that are not committed to a particular role.

[0235] In at least one embodiment, one or more SoCs in SoC 804 can further include a wide range of peripheral interfaces to enable communication with peripherals, audio encoders / decoders (“codecs”), power management, and / or other devices. One or more SoCs 804 can be used to process data from cameras (e.g., over a gigabit multimedia serial link and an Ethernet connection), sensors (e.g., one or more LIDAR sensors 864, one or more RADAR sensors 860, etc., which can be connected over an Ethernet bus), data from bus 802 (e.g., speed of vehicle 800, steering wheel position, etc.), data from one or more GNSS sensors 858 (e.g., connected over an Ethernet bus or a CAN bus), etc. In at least one embodiment, one or more of SoCs 804 can further include a dedicated high-performance mass storage controller that can include their own DMA engine and can be used to free one or more CPUs 806 from regular data management tasks.

[0236] In at least one embodiment, SoC(s) 804 can be an end-to-end platform with a flexible architecture that spans automation levels 3-5, providing a comprehensive functional safety architecture that leverages and efficiently uses computer vision and ADAS technology for diversity and redundancy, which provides a platform that can provide a flexible, reliable driving software stack, as well as deep learning tools. In at least one embodiment, SoC(s) 804 can be faster, more reliable, and even more energy and spatial efficient than conventional systems. For example, in at least one embodiment, accelerator(s) 814, when combined with CPU(s) 806, GPU(s) 808, and data storage(s) 816, can provide a fast, efficient platform for level 3-5 autonomous vehicles.

[0237] In at least one embodiment, computer vision algorithms can be executed on CPUs, which can be configured using high-level programming languages (e.g., C programming language) to perform a variety of processing algorithms on a variety of visual data. However, in at least one embodiment, CPUs typically cannot meet performance requirements of many computer vision applications, such as performance requirements related to execution time and power consumption. In at least one embodiment, many CPUs cannot execute complex object detection algorithms in real-time, which are used in on-board ADAS applications and actual level 3-5 autonomous vehicles.

[0238] Embodiments described herein allow for simultaneous and / or sequential execution of multiple neural networks, and allow for results to be combined together to enable level 3-5 autonomous driving functionality. For example, in at least one embodiment, CNNs executed on DLAs or discrete GPUs (e.g., GPU(s) 820) can include text and word recognition, allowing a supercomputer to read and understand traffic signs, including signs that a neural network has not been specifically trained for. In at least one embodiment, DLAs can also include neural networks capable of recognizing, interpreting, and providing semantic understanding of symbols, and passing this semantic understanding to a path planning module running on a CPU Complex.

[0239] In at least one embodiment, multiple neural networks can be run simultaneously for a level 3, 4, or 5 drive. For example, in at least one embodiment, a warning sign consisting of a “Caution: flashing lights indicate icy conditions” sign together with a flashing light connected to an electrical light can be interpreted independently or collectively by multiple neural networks. In at least one embodiment, the sign itself can be recognized as a traffic sign by a first deployed neural network (e.g., a neural network that has been trained), the text “flashing lights indicate icy conditions” can be interpreted by a second deployed neural network that informs vehicle’s path planning software (preferably executing on a CPU Complex) that icy conditions exist when flashing lights are detected. In at least one embodiment, the flashing lights can be recognized by a third deployed neural network operating over multiple frames, informing vehicle’s path planning software of the existence (or non-existence) of flashing lights. In at least one embodiment, all three neural networks can be run simultaneously, e.g., within a DLA and / or on one or more GPU(s) 808.

[0240] In at least one embodiment, a CNN for facial recognition and vehicle owner identification can use data from a camera sensor to identify presence of an authorized driver and / or owner of vehicle 800. In at least one embodiment, when an owner approaches a driver door and opens a light, a normally open sensor processor engine can be used to unlock the vehicle, and, in a safe mode, when the owner leaves the vehicle, can be used to disable the vehicle. In this way, one or more SoC(s) 804 provide a safeguard against theft and / or carjacking.

[0241] In at least one embodiment, a CNN for emergency vehicle detection and identification can use data from microphones 896 to detect and identify emergency vehicle sirens. In at least one embodiment, one or more SoCs 804 use a CNN to classify ambient and urban sounds, as well as to classify visual data. In at least one embodiment, a CNN running on a DLA is trained to identify relative proximity of an emergency vehicle (e.g., by using Doppler effect). In at least one embodiment, a CNN can also be trained to identify emergency vehicles for regions in which a vehicle is operating, as identified by one or more GNSS sensors 858. In at least one embodiment, when operating in Europe, a CNN will seek to detect European sirens, while in the United States, a CNN will seek to identify only North American sirens. In at least one embodiment, once an emergency vehicle is detected, a control program can be used to execute emergency vehicle safety routines, slow vehicle down, pull vehicle to side of road, stop, and / or idle vehicle until emergency vehicle passes, with assistance from one or more ultrasonic sensors 862.

[0242] In at least one embodiment, vehicle 800 can include one or more CPUs 818 (e.g., one or more discrete CPUs or one or more dCPUs) that can be coupled to one or more SoCs 804 via a high-speed interconnect (e.g., PCIe). In at least one embodiment, one or more CPUs 818 can include an X86 processor, such as one or more CPUs 818 can be used to perform any of a variety of functions, such as including potentially arbitrating inconsistent results between ADAS sensors and one or more SoCs 804, and / or one or more supervisory controllers 836 state and health and / or an information system on a chip (“Info SoC”) 830.

[0243] In at least one embodiment, vehicle 800 can include one or more GPUs 820 (e.g., one or more discrete GPUs or one or more dGPUs) that can be coupled to one or more SoCs 804 via a high-speed interconnect (e.g., NVIDIA’s NVLINK). In at least one embodiment, one or more GPUs 820 can provide additional artificial intelligence functionality, such as by executing redundant and / or different neural networks, and can be used to train and / or update neural networks based at least in part on input (e.g., sensor data) from sensors of vehicle 800.

[0244] In at least one embodiment, vehicle 800 can further include network interface 824, which can include, without limitation, one or more wireless antennas 826 (e.g., one or more wireless antennas 826 for different communication protocols, such as cellular antennas, Bluetooth antennas, etc.). In at least one embodiment, network interface 824 can be used to enable wireless connectivity with other vehicles and / or computing devices (e.g., client devices of passengers) through the Internet with a cloud (e.g., employing servers and / or other network devices). In at least one embodiment, to communicate with other vehicles, a direct link can be established between vehicle 800 and other vehicles and / or an indirect link can be established (e.g., through a network and the Internet). In at least one embodiment, a direct link can be provided using a vehicle-to-vehicle communication link. In at least one embodiment, a vehicle-to-vehicle communication link can provide vehicle 800 with information about vehicles in a vicinity of vehicle 800 (e.g., vehicles in front of, to the side of, and / or behind vehicle 800). In at least one embodiment, this aforementioned functionality can be part of a cooperative adaptive cruise control functionality of vehicle 800.

[0245] In at least one embodiment, network interface 824 can include a SoC that provides modulation and demodulation functionality and enables one or more controllers 836 to communicate over wireless networks. In at least one embodiment, network interface 824 can include a radio frequency front end for up-conversion from baseband to radio frequency and down-conversion from radio frequency to baseband. In at least one embodiment, frequency conversion can be performed in any technically feasible way. For example, frequency conversion can be performed through well-known processes and / or using a superheterodyne process. In at least one embodiment, radio frequency front end functionality can be provided by a separate chip. In at least one embodiment, a network interface can include wireless functionality to communicate over LTE, WCDMA, UMTS, GSM, CDMA2000, Bluetooth, Bluetooth LE, Wi-Fi, Z-Wave, ZigBee, LoRaWAN, and / or other wireless protocol.

[0246] In at least one embodiment, vehicle 800 can further include one or more data stores 828, which can include, without limitation, off-chip (e.g., of SoC(s) 804) storage. In at least one embodiment, one or more data stores 828 can include, without limitation, one or more storage elements including RAM, SRAM, dynamic random access memory (“DRAM”), video random access memory (“VRAM”), flash memory, hard disks, and / or other components and / or devices that can store at least one bit of data.

[0247] In at least one embodiment, vehicle 800 can further include one or more GNSS sensors 858 (e.g., GPS and / or assisted GPS sensors) to assist in mapping, perception, occupancy grid generation, and / or path planning functions. In at least one embodiment, any number of GNSS sensors 858 can be used, including, for example and without limitation, a GPS using a USB connector with an Ethernet connection to a serial interface (e.g., RS-232) bridge.

[0248] In at least one embodiment, vehicle 800 can further include one or more RADAR sensors 860. One or more RADAR sensors 860 can be used by vehicle 800 for long-range vehicle detection, even in darkness and / or adverse weather conditions. In at least one embodiment, a RADAR functional safety level can be ASIL B. One or more RADAR sensors 860 can use CAN bus and / or bus 802 (e.g., to transfer data generated by one or more RADAR sensors 860) for control and access to object tracking data, and in certain examples can access an Ethernet channel for access to raw data. In at least one embodiment, a wide variety of RADAR sensor types can be used. For example and without limitation, one or more of RADAR sensors 860 can be suitable for front, rear, and side RADAR use. In at least one embodiment, one or more RADAR sensors 860 are pulse Doppler RADAR sensors.

[0249] In at least one embodiment, one or more RADAR sensors 860 can include different configurations, such as long-range with narrow field of view, short-range with wide field of view, short-range side coverage, etc. In at least one embodiment, long-range RADAR can be used for adaptive cruise control functionality. In at least one embodiment, a long-range RADAR system can provide a wide field of view achieved through two or more independent scans (e.g., over 250 m range). In at least one embodiment, one or more RADAR sensors 860 can help distinguish between static and moving objects, and can be used by ADAS system 838 for emergency brake assist and forward collision warning. In at least one embodiment, one or more sensors 860 included in a long-range RADAR system can include, without limitation, a single-site multi-mode RADAR with multiple (e.g., six or more) fixed RADAR antennas, as well as high-speed CAN and FlexRay interfaces. In at least one embodiment, with six antennas, a central four antennas can create a focused beam pattern designed to record the environment around vehicle 800 at higher speeds with minimal traffic interference from adjacent lanes. In at least one embodiment, other two antennas can expand the field of view, which can allow for quick detection of vehicles entering or leaving vehicle 800’s lane.

[0250] In at least one embodiment, as an example, a mid-range RADAR system can include a range of up to 160 m (front) or 80 m (rear), for example, and a field of view of up to 42 degrees (front) or 150 degrees (rear). In at least one embodiment, a short-range RADAR system can include, without limitation, any number of RADAR sensors 860 designed to be mounted at both ends of a rear bumper. When mounted at both ends of a rear bumper, in at least one embodiment, a RADAR sensor system can produce two beams that constantly monitor a vehicle’s rear and near blind spots. In at least one embodiment, a short-range RADAR system can be used in ADAS system 838 for blind spot detection and / or lane change assist.

[0251] In at least one embodiment, vehicle 800 can further include one or more ultrasonic sensors 862. In at least one embodiment, one or more ultrasonic sensors 862, which can be positioned on front, back, and / or sides of vehicle 800, can be used for parking assist and / or to create and update an occupancy grid. In at least one embodiment, a wide variety of ultrasonic sensors 862 can be used, and different ultrasonic sensors 862 can be used for different detection ranges (e.g., 2.5 m, 4 m). In at least one embodiment, ultrasonic sensors 862 can operate at a functional safety level of ASIL B.

[0252] In at least one embodiment, vehicle 800 can include one or more LIDAR sensors 864. One or more LIDAR sensors 864 can be used for object and pedestrian detection, emergency braking, collision avoidance, and / or other functions. In at least one embodiment, one or more LIDAR sensors 864 can be functional safety level ASIL B. In at least one embodiment, vehicle 800 can include multiple (e.g., two, four, six, etc.) LIDAR sensors 864 that can use Ethernet (e.g., provide data to a Gigabit Ethernet switch).

[0253] In at least one embodiment, one or more LIDAR sensors 864 can be capable of providing a list of objects and their distances for a 360 degree field of view. In at least one embodiment, one or more LIDAR sensors 864 can be commercially available, for example, with an advertised range of about 100 m, with an accuracy of 2 cm - 3 cm, and supporting a 100 Mbps Ethernet connection. In at least one embodiment, one or more non-protruding LIDAR sensors 864 can be used. In such embodiments, one or more LIDAR sensors 864 can be implemented as small devices embedded into front, back, sides, and / or corners of vehicle 800. In at least one embodiment, one or more LIDAR sensors 864, in such embodiments, can provide a horizontal field of view of up to 120 degrees and a vertical field of view of 35 degrees, with a range of 200 m, even for low reflectivity objects. In at least one embodiment, forward-facing one or more LIDAR sensors 864 can be configured for a horizontal field of view between 45 degrees and 135 degrees.

[0254] In at least one embodiment, LIDAR technology such as 3D Flash LIDAR can also be used. 3D Flash LIDAR uses a laser flash as a transmission source to illuminate about 200 m around vehicle 800. In at least one embodiment, a flash LIDAR unit includes, without limitation, a receiver that records laser pulse travel time and reflected light on each pixel, which in turn corresponds to a range from vehicle 800 to an object. In at least one embodiment, flash LIDAR can allow for generation of highly accurate and distortion-free images of surrounding environment with each laser flash. In at least one embodiment, four flash LIDAR sensors can be deployed, one on each side of vehicle 800. In at least one embodiment, a 3D flash LIDAR system includes, without limitation, a solid-state 3D staring array LIDAR camera with no moving parts other than a fan (e.g., a non-scanning LIDAR device). In at least one embodiment, a flash LIDAR device can use 5 nanosecond Class I (eye-safe) laser pulses per frame, and can capture reflected laser light in the form of a 3D ranging point cloud and co-registered intensity data.

[0255] In at least one embodiment, vehicle 800 can also include one or more IMU sensors 866. In at least one embodiment, one or more IMU sensors 866 can be located at a center of a rear axle of vehicle 800. In at least one embodiment, one or more IMU sensors 866 can include, without limitation, one or more accelerometers, one or more magnetometers, one or more gyroscopes, one or more magnetic compasses, and / or other sensor types. In at least one embodiment, one or more IMU sensors 866 can include, without limitation, an accelerometer and a gyroscope, such as in a six-axis application. In at least one embodiment, one or more IMU sensors 866 can include, without limitation, an accelerometer, a gyroscope, and a magnetometer, such as in a nine-axis application.

[0256] In at least one embodiment, one or more IMU sensors 866 can be implemented as a miniature, high-performance GPS-aided inertial navigation system (“GPS / INS”) that combines micro-electro-mechanical systems (“MEMS”) inertial sensors, high-sensitivity GPS receiver, and advanced Kalman filtering algorithms to provide estimates of position, velocity, and attitude; in at least one embodiment, one or more IMU sensors 866 can enable vehicle 800 to estimate heading without requiring input from a magnetic sensor by directly observing and correlating changes in velocity from GPS to one or more IMU sensors 866. In at least one embodiment, one or more IMU sensors 866 and one or more GNSS sensors 858 can be combined in a single integrated unit.

[0257] In at least one embodiment, vehicle 800 can include one or more microphones 896 placed within and / or around vehicle 800. In at least one embodiment, additionally, one or more microphones 896 can be used for emergency vehicle detection and identification.

[0258] In at least one embodiment, vehicle 800 can further include any number of camera types including one or more stereo cameras 868, one or more wide-view cameras 870, one or more infrared cameras 872, one or more surround cameras 874, one or more long-range cameras 898, one or more mid-range cameras 876, and / or other camera types. In at least one embodiment, cameras can be used to capture image data around an entire periphery of vehicle 800. In at least one embodiment, a type of camera used depends on vehicle 800. In at least one embodiment, any combination of camera types can be used to provide necessary coverage around vehicle 800. In at least one embodiment, a number of cameras deployed can vary from embodiment to embodiment. For example, in at least one embodiment, vehicle 800 can include six cameras, seven cameras, ten cameras, twelve cameras, or other number of cameras. In at least one embodiment, cameras can support Gigabit Multimedia Serial Link (“GMSL”) and / or Gigabit Ethernet, by way of example and without limitation. In at least one embodiment, each camera can be described in greater detail previously with reference to FIG. 8A and FIG. 8B Each camera can be described in greater detail.

[0259] In at least one embodiment, vehicle 800 can further include one or more vibration sensors 842. In at least one embodiment, one or more vibration sensors 842 can measure vibrations of components of vehicle 800 (e.g., axles). For example, in at least one embodiment, changes in vibration can be indicative of changes in a road surface. In at least one embodiment, when two or more vibration sensors 842 are used, differences between vibrations can be used to determine friction or slippage of a road surface (e.g., when there is a difference in vibration between a power driven axle and a free spinning axle).

[0260] In at least one embodiment, vehicle 800 can include an ADAS system 838. ADAS system 838 can include, without limitation, a SoC. In at least one embodiment, ADAS system 838 can include, without limitation, any number of adaptive / autonomous / automatic cruise control (“ACC”) systems, cooperative adaptive cruise control (“CACC”) systems, forward collision warning (“FCW”) systems, automatic emergency braking (“AEB”) systems, lane departure warning (“LDW”) systems, lane keep assist (“LKA”) systems, blind spot warning (“BSW”) systems, rear cross traffic warning (“RCTW”) systems, collision warning (“CW”) systems, lane centering (“LC”) systems, and / or other systems, features, and / or functionality, and combinations thereof.

[0261] In at least one embodiment, ACC system can use one or more RADAR sensors 860, one or more LIDAR sensors 864, and / or any number of cameras. In at least one embodiment, ACC system can include a longitudinal ACC system and / or a lateral ACC system. In at least one embodiment, a longitudinal ACC system monitors and controls a distance to a vehicle immediately ahead of vehicle 800 and automatically adjusts speed of vehicle 800 to maintain a safe distance from the vehicle ahead. In at least one embodiment, a lateral ACC system performs distance keeping and suggests lane changes for vehicle 800 when needed. In at least one embodiment, lateral ACC is relevant to other ADAS applications, such as LC and CW.

[0262] In at least one embodiment, a CACC system uses information from other vehicles, which can be received from other vehicles via a wireless link or indirectly via a network connection (e.g., via the Internet) via network interface 824 and / or one or more wireless antennas 826. In at least one embodiment, a direct link can be provided by a vehicle-to-vehicle (“V2V”) communication link, while an indirect link can be provided by an infrastructure-to-vehicle (“I2V”) communication link. In general, a V2V communication concept provides information about vehicles immediately ahead (e.g., vehicles immediately ahead of and in same lane as vehicle 800), while an I2V communication concept provides information about traffic further ahead. In at least one embodiment, a CACC system can include one or both of I2V and V2V information sources. In at least one embodiment, a CACC system can be more reliable with information about vehicles ahead of vehicle 800, and has potential to improve smoothness of traffic flow and reduce road congestion.

[0263] In at least one embodiment, a FCW system is designed to warn a driver of a hazard so that the driver can take corrective action. In at least one embodiment, a FCW system uses a forward-facing camera and / or one or more RADAR sensors 860, coupled to a dedicated processor, DSP, FPGA, and / or ASIC, electrically coupled to driver feedback such as a display, speaker, and / or vibration component. In at least one embodiment, a FCW system can provide a warning, such as in the form of a sound, visual warning, vibration, and / or quick brake pulse.

[0264] In at least one embodiment, an AEB system detects an impending forward collision with another vehicle or other object and can automatically apply brakes if a driver does not take corrective action within a specified time or distance parameter. In at least one embodiment, AEB systems can use one or more forward facing cameras and / or one or more RADAR sensors 860 coupled to dedicated processors, DSPs, FPGAs, and / or ASICs. In at least one embodiment, when an AEB system detects a hazard, the AEB system typically first warns a driver to take corrective action to avoid a collision, and if that driver does not take corrective action, the AEB system can automatically apply brakes in an attempt to prevent or at least mitigate the effects of a predicted collision. In at least one embodiment, AEB systems can include technologies such as dynamic brake support and / or pre-crash braking.

[0265] In at least one embodiment, a LDW system provides visual, audible, and / or tactile warnings, such as steering wheel or seat vibrations, to warn the driver when vehicle 800 crosses lane markers. In at least one embodiment, a LDW system is not active when the driver indicates an intentional lane departure, such as by activating turn signals. In at least one embodiment, a LDW system can use a front facing camera coupled to dedicated processors, DSPs, FPGAs, and / or ASICs that is electrically coupled to driver feedback such as displays, speakers, and / or vibrating components. In at least one embodiment, an LKA system is a variation of a LDW system. If vehicle 800 begins to drift out of a lane, an LKA system provides steering input or braking to correct vehicle 800.

[0266] In at least one embodiment, a BSW system detects and warns vehicle drivers of vehicles in a car’s blind spot. In at least one embodiment, a BSW system can provide visual, audible, and / or tactile alerts to indicate that merging or changing lanes is unsafe. In at least one embodiment, a BSW system can provide additional warnings when a driver is using turn signals. In at least one embodiment, a BSW system can use one or more rear facing cameras and / or one or more RADAR sensors 860 coupled to dedicated processors, DSPs, FPGAs, and / or ASICs that are electrically coupled to driver feedback such as displays, speakers, and / or vibrating components.

[0267] In at least one embodiment, when an object is detected outside of a rear camera range while vehicle 800 is backing up, RCTW system can provide a visual, audible, and / or tactile notification. In at least one embodiment, RCTW system includes an AEB system to ensure application of vehicle brakes to avoid a collision. In at least one embodiment, RCTW system can use one or more rear-facing RADAR sensors 860 coupled to a dedicated processor, DSP, FPGA, and / or ASIC that is electrically coupled to driver feedback such as a display, speaker, and / or vibrating component.

[0268] In at least one embodiment, conventional ADAS systems can be prone to false positives, which can annoy and distract drivers, but are typically not catastrophic because conventional ADAS systems alert the driver and allow that driver to decide whether a safety condition is truly present and take appropriate action. In at least one embodiment, in case of a result conflict, vehicle 800 itself decides whether to follow the results of the primary or secondary computer (e.g., first controller 836 or second controller 836). For example, in at least one embodiment, ADAS system 838 can be a backup and / or secondary computer for providing perception information to a backup computer plausibility module. In at least one embodiment, a backup computer plausibility monitor can run redundant varieties of software on hardware components to detect faults in perception and dynamic driving tasks. In at least one embodiment, output from ADAS system 838 can be provided to a supervisory MCU. In at least one embodiment, if outputs from primary and secondary computers conflict, supervisory MCU decides how to reconcile the conflict to ensure safe operation.

[0269] In at least one embodiment, a primary computer can be configured to provide a confidence score to a supervisory MCU to indicate a confidence of the primary computer in a selected result. In at least one embodiment, if the confidence score exceeds a threshold, the supervisory MCU can follow the primary computer’s indication regardless of whether the secondary computer provides conflicting or inconsistent results. In at least one embodiment, in case the confidence score does not satisfy the threshold, and in case the primary and secondary computers indicate different results (e.g., conflict), the supervisory MCU can arbitrate between the computers to determine an appropriate result.

[0270] In at least one embodiment, a supervisory MCU can be configured to run a neural network trained and configured to determine conditions under which an auxiliary computer provides false alarms based at least in part on outputs from a host computer and an auxiliary computer. In at least one embodiment, a neural network in a supervisory MCU can learn when to trust outputs of an auxiliary computer, and when not to. For example, in at least one embodiment, when the auxiliary computer is a RADAR-based FCW system, a neural network in a supervisory MCU can learn when the FCW system identifies metal objects that are not actually dangerous, such as drain grates or manhole covers that would trigger an alert. In at least one embodiment, when the auxiliary computer is a camera-based LDW system, a neural network in a supervisory MCU can learn to override LDW when a bicyclist or pedestrian is present and it is actually safest to lane depart. In at least one embodiment, a supervisory MCU can include at least one of a DLA or GPU suitable for running a neural network with associated memory. In at least one embodiment, a supervisory MCU can include and / or be included as a component of one or more SoCs 804.

[0271] In at least one embodiment, ADAS system 838 can include an auxiliary computer that performs ADAS functions using traditional computer vision rules. In at least one embodiment, this auxiliary computer can use classic computer vision rules (if-then), and presence of a neural network in a supervisory MCU can improve reliability, safety, and performance. For example, in at least one embodiment, diverse implementation and intentional non-identity make the overall system more fault-tolerant, especially to faults caused by software (or software-hardware interface) functionality. For example, in at least one embodiment, if there is a software bug or error in software running on a host computer, and non-identical software code running on an auxiliary computer provides the same overall result, then a supervisory MCU can be more confident that the overall result is correct, and the bug in software or hardware on that host computer did not cause a significant error.

[0272] In at least one embodiment, outputs of ADAS system 838 can be input into a perception module of a host computer and / or a dynamic driving task module of a host computer. For example, in at least one embodiment, if ADAS system 838 indicates a forward collision warning due to an object directly in front, then a perception block can use this information in identifying the object. In at least one embodiment, as described herein, an auxiliary computer can have its own neural network trained such that risk of false positives is reduced.

[0273] In at least one embodiment, vehicle 800 can further include infotainment SoC 830 (e.g., an in-vehicle infotainment system (IVI)). Although illustrated and described as a SoC, in at least one embodiment, infotainment system 830 can not be a SoC, and can include, without limitation, two or more discrete components. In at least one embodiment, infotainment SoC 830 can include, without limitation, a combination of hardware and software that can be used to provide audio (e.g., music, personal digital assistant, navigation instructions, news, radio, etc.), video (e.g., television, movies, streaming media, etc.), telephony (e.g., hands-free calling), network connectivity (e.g., LTE, WiFi, etc.), and / or information services (e.g., navigation systems, rear park assist, radio data system, vehicle related information such as fuel level, total range, brake fuel level, oil level, doors open / close, air filter information, etc.) to vehicle 800. For example, infotainment SoC 830 can include a radio, disc player, navigation system, video player, USB and Bluetooth connectivity, car, car entertainment system, WiFi, steering wheel audio controls, hands-free voice controls, heads-up display (“HUD”), HMI display 834, telematics equipment, control panel (e.g., for controlling and / or interacting with various components, features, and / or systems), and / or other components. In at least one embodiment, infotainment SoC 830 can be further used to provide information (e.g., visual and / or audible) to a user of vehicle 800, such as information from ADAS system 838, autonomous driving information (such as planned vehicle maneuvers), trajectory, surrounding environment information (e.g., intersection information, vehicle information, road information, etc.), and / or other information.

[0274] In at least one embodiment, infotainment SoC 830 can include any number and type of GPU functionality. In at least one embodiment, infotainment SoC 830 can communicate with other devices, systems, and / or components of vehicle 800 over bus 802 (e.g., CAN bus, Ethernet, etc.). In at least one embodiment, infotainment SoC 830 can be coupled to a monitoring MCU such that a GPU of the infotainment system can perform some autonomous driving functions in the event of a failure of primary controller 836 (e.g., a primary and / or backup computer of vehicle 800). In at least one embodiment, infotainment SoC 830 can cause vehicle 800 to enter a driver-to-safe-stop mode, as described herein.

[0275] In at least one embodiment, vehicle 800 can further include an instrument cluster 832 (e.g., a digital instrument cluster, an electronic instrument cluster, a digital instrument cluster panel, etc.). In at least one embodiment, instrument cluster 832 can include, without limitation, a controller and / or supercomputer (e.g., a discrete controller or supercomputer). In at least one embodiment, instrument cluster 832 can include, without limitation, any number and combination of gauges such as a speedometer, fuel level, oil pressure, tachometer, odometer, turn indicator, shift position indicator, one or more seatbelt warning lights, one or more parking brake warning lights, one or more engine malfunction lights, auxiliary restraint system (e.g., airbag) information, lighting controls, safety system controls, navigation information, etc. In some examples, information can be displayed and / or shared between infotainment SoC 830 and instrument cluster 832. In at least one embodiment, instrument cluster 832 can be included as part of infotainment SoC 830, and vice versa.

[0276] In at least one embodiment, one or more systems shown in FIG. 8C are used to perform selection of program code optimizations using various algorithms, formulas, and processes (such as those described in connection with FIG. 1 to FIG. 2 In at least one embodiment, one or more systems shown in FIG. 8C are used to implement one or more systems and / or processes (such as those described in connection with FIG. 1 to FIG. 6B In at least one embodiment, one or more systems shown in

[0277] FIG. 8D are used to implement one or more systems and / or processes (such as those described in connection with FIG. 8Aa diagram of a system 877 that enables communication between autonomous vehicles 800. In at least one embodiment, system 877 can include, without limitation, one or more servers 878, one or more networks 890, and any number and type of vehicles, including vehicles 800. One or more servers 878 can include, without limitation, a plurality of GPUs 884(A)-884(H) (collectively referred to herein as GPUs 884), PCIe switches 882(A)-882(D) (collectively referred to herein as PCIe switches 882), and / or CPUs 880(A)-880(B) (collectively referred to herein as CPUs 880). GPUs 884, CPUs 880, and PCIe switches 882 can be interconnected with high-speed connection lines such as, but not limited to, NVLink interfaces 888 developed by NVIDIA and / or PCIe connections 886. GPUs 884 are connected by NVLink and / or NVSwitch SoC connections, and GPUs 884 and PCIe switches 882 are connected by PCIe interconnects. In at least one embodiment, although eight GPUs 884, two CPUs 880, and four PCIe switches 882 are shown, this is not intended to be limiting. In at least one embodiment, each of one or more servers 878 can include, without limitation, any number of GPUs 884, CPUs 880, and / or PCIe switches 882 in any combination. For example, in at least one embodiment, one or more servers 878 can each include eight, sixteen, thirty-two, and / or more GPUs 884.

[0278] In at least one embodiment, one or more servers 878 can receive, over one or more networks 890 and from vehicles, image data representative of images showing unexpected or changing road conditions, such as road work that has recently begun. In at least one embodiment, one or more servers 878 can transmit, over one or more networks 890 and to vehicles, neural networks 892, updated neural networks 892, and / or map information 894 including, without limitation, information about traffic and road conditions. In at least one embodiment, updates to map information 894 can include, without limitation, updates to HD map 822, such as information about construction sites, potholes, detours, flooding, and / or other obstacles. In at least one embodiment, neural networks 892, updated neural networks 892, and / or map information 894 can be the result of new training and / or experience represented in data received from any number of vehicles in an environment, and / or based at least on training performed at a data center (e.g., using one or more servers 878 and / or other servers).

[0279] In at least one embodiment, one or more servers 878 can be used to train machine learning models (e.g., neural networks) based at least in part on training data. In at least one embodiment, training data can be generated by vehicles, and / or can be generated in simulations (e.g., using a game engine). In at least one embodiment, any amount of training data is labeled (e.g., where an associated neural network benefits from supervised learning) and / or undergoes other pre-processing. In at least one embodiment, no training data is labeled and / or pre-processed (e.g., where an associated neural network does not require supervised learning). In at least one embodiment, once a machine learning model is trained, a machine learning model can be used by vehicles (e.g., transferred to vehicles by one or more networks 890, and / or a machine learning model can be used by one or more servers 878 to monitor vehicles remotely.

[0280] In at least one embodiment, one or more servers 878 can receive data from vehicles and apply the data to up-to-date, real-time neural networks for real-time intelligent inference. In at least one embodiment, one or more servers 878 can include deep-learning supercomputers and / or specialized Al computers powered by one or more GPUs 884, such as DGX and DGX Station machines developed by NVIDIA. However, in at least one embodiment, one or more servers 878 can include deep learning infrastructure of a data center powered using CPUs.

[0281] In at least one embodiment, deep learning infrastructure of one or more servers 878 can be capable of fast, real-time inference, and can use this capability to assess and validate health of processors, software, and / or related hardware in vehicles 800. For example, in at least one embodiment, deep learning infrastructure can receive periodic updates from vehicles 800, such as sequences of images and / or objects located by vehicles 800 in the sequence of images (e.g., through computer vision and / or other machine learning object classification techniques). In at least one embodiment, deep learning infrastructure can run its own neural networks to identify objects and compare them to objects identified by vehicles 800, and, if results do not match and deep learning infrastructure concludes that Al in vehicles 800 is malfunctioning, one or more servers 878 can send a signal to vehicles 800 instructing a failsafe computer of vehicles 800 to take control, notify passengers, and complete a safe parking operation.

[0282] In at least one embodiment, one or more servers 878 may include one or more GPUs 884 and one or more programmable inference accelerators (e.g., NVIDIA's TensorRT 3). In at least one embodiment, the combination of GPU-powered servers and inference acceleration can enable real-time responses. In at least one embodiment, for example, in situations where performance is less critical, servers powered by CPUs, FPGAs, and other processors can be used for inference.

[0283] Computer system

[0284] FIG. 9 9 is a block diagram illustrating an exemplary computer system according to at least one embodiment, which may be a system of interconnected devices and components, a system on a chip (SOC), or some combination thereof formed with a processor that may include execution units to execute instructions. In at least one embodiment, in accordance with the present disclosure, such as the embodiments described herein, computer system 900 may include, but is not limited to, components such as processor 902, whose execution units include logic to execute algorithms for processing data. In at least one embodiment, computer system 900 may include a processor such as the Intel® processor 902 available from Intel Corporation of Santa Clara, California. Processor family, Xeon TM 、 XScale TM and / or StrongARM TM , Core TM or Nervana TM microprocessor, although other systems (including PCs with other microprocessors, engineering workstations, set-top boxes, etc.) may also be used. In at least one embodiment, computer system 900 may execute a version of the WINDOWS operating system available from Microsoft Corporation of Redmond, Wash., although other operating systems (e.g., UNIX and Linux), embedded software, and / or graphical user interfaces may also be used.

[0285] Embodiments can be used in other devices such as handheld devices and embedded applications. Some examples of handheld devices include cellular phones, Internet Protocol devices, digital cameras, personal digital assistants ("PDAs"), and handheld PCs. In at least one embodiment, embedded applications can include a microcontroller, a digital signal processor ("DSP"), a system on a chip, a network computer ("NetPC"), a set-top box, a network hub, a wide area network ("WAN") switch, or any other system that can perform one or more instructions in accordance with at least one embodiment.

[0286] In at least one embodiment, computer system 900 can include, but is not limited to, processor 902, which can include, but is not limited to, one or more execution units 908 to perform, e.g., machine learning model training and / or inferencing, in accordance with techniques described herein. In at least one embodiment, system 900 is a single processor desktop or server system, but in another embodiment, system 900 can be a multiprocessor system. In at least one embodiment, processor 902 can include, but is not limited to, a complex instruction set computer ("CISC") microprocessor, a reduced instruction set computing ("RISC") microprocessor, a very long instruction word ("VLIW") microprocessor, a processor implementing a combination of instruction sets, or any other processor device, such as a digital signal processor. In at least one embodiment, processor 902 can be coupled to a processor bus 910 that can transmit data signals between processor 902 and other components in computer system 900.

[0287] In at least one embodiment, processor 902 can include, but is not limited to, level 1 ("Ll") internal cache memory ("cache") 904. In at least one embodiment, processor 902 can have a single -level internal cache or multi-level internal cache. In at least one embodiment, cache memory can reside in the processor 902's external. Other embodiments can include a combination of internal and external cache memory depending on the specific implementation and requirements. In at least one embodiment, register file 906 can store different types of data such as integer, floating point, status, and instruction pointer registers in various registers within processor 902.

[0288] In at least one embodiment, execution unit 908 includes, without limitation, logic to perform integer and floating point operations, including bit- wide operations. In at least one embodiment, processor 902 can also include microcode (“ucode”) read-only memory (“ROM”), which stores microcode for certain macroinstructions. In at least one embodiment, execution unit 908 can also include logic to handle a packed data instruction set 909. In at least one embodiment, by including packed data instruction set 909 in a general-purpose processor, many multimedia applications can be accelerated by a general purpose processor 902 including packed data instruction set 909. In one or more embodiments, by using a full width of a processor’s data bus for one or more of operations on packed data, many multimedia applications can be executed more efficiently, which can not require transferring smaller units of data to a processor for one or more operations on one data element at a time.

[0289] In at least one embodiment, execution unit 908 can also be used in microcontrollers, embedded processors, graphics devices, DSPs, and other types of logic circuits. In at least one embodiment, computer system 900 can include, without limitation, memory 920. In at least one embodiment, memory 920 can be implemented as a Dynamic Random Access Memory (“DRAM”) device, a Static Random Access Memory (“SRAM”) device, a flash memory device, or other memory device. In at least one embodiment, memory 920 can store instruction(s) 919 and / or data 921 represented by data signals that can be executed by processor 902.

[0290] In at least one embodiment, a system logic chip can be coupled to processor bus 910 and memory 920. In at least one embodiment, system logic chip can include, without limitation, a memory controller hub (“MCH”) 916 and processor 902 can communicate with MCH 916 via processor bus 910. In at least one embodiment, MCH 916 can provide a high bandwidth memory path 918 to memory 920 for instruction and data storage and for storage of graphics commands, data, and textures. In at least one embodiment, MCH 916 can direct data signals between processor 902, memory 920, and other components in computer system 900, and can

[0291] In at least one embodiment, computer system 900 can use system I / O 922, which is a proprietary hub interface bus to couple MCH 916 to I / O controller hub (“ICH”) 930. In at least one embodiment, ICH 930 can provide a direct connection to some I / O devices and can indirect connect other devices via an

[0292] In at least one embodiment, FIG. 9 A system including interconnected hardware devices or “chips” is shown, while in other embodiments, FIG. 9 A system on a chip (SoC) can be shown. In at least one embodiment, FIG. 9The devices shown in can be interconnected with a proprietary interconnect, a standardized interconnect (e.g., PCIe), or some combination thereof. In at least one embodiment, one or more components of system 900 are interconnected using a Compute Express Link (CXL) interconnect.

[0293] In at least one embodiment, FIG. 9 One or more systems shown in the foregoing are used to utilize various algorithms, formulas, and processes (e.g., in conjunction with FIG. 1 to FIG. 2 those described herein) to perform selection of program code optimizations, and / or otherwise perform the operations described herein. In at least one embodiment, FIG. 9 One or more systems shown in the foregoing are used to implement one or more systems and / or processes (e.g., in conjunction with FIG. 1 to FIG. 6B those described herein), such as performing selection of program code optimizations in one or more compilers, and / or otherwise performing the operations described herein.

[0294] FIG. 10 1 is a block diagram illustrating an electronic device 1000 for utilizing a processor 1010 according to at least one embodiment. In at least one embodiment, the electronic device 1000 may be, for example, but not limited to, a notebook computer, a tower server, a rack server, a blade server, a laptop computer, a desktop computer, a tablet computer, a mobile device, a phone, an embedded computer, or any other suitable electronic device.

[0295] In at least one embodiment, system 1000 may include, but is not limited to, a processor 1010 communicatively coupled to any suitable number or kind of components, peripherals, modules, or devices. In at least one embodiment, processor 1010 is coupled using a bus or interface, such as an I2C bus, a system management bus ("SMBus"), a low pin count (LPC) bus, a serial peripheral interface ("SPI"), a high-definition audio ("HDA") bus, a serial advanced technology attachment ("SATA") bus, a universal serial bus ("USB") (versions 1, 2, 3, etc.), or a universal asynchronous receiver / transmitter ("UART") bus. In at least one embodiment, FIG. 10 shows a system comprising interconnected hardware devices or "chips", while in other embodiments, FIG. 10 An exemplary system on a chip (SoC) may be shown. In at least one embodiment, FIG. 10 The devices shown in can be interconnected with proprietary interconnects, standardized interconnects (e.g., PCIe), or some combination thereof. In at least one embodiment, FIG. 10 One or more components of the system are interconnected using Compute Express Link (CXL) interconnect lines.

[0296] In at least one embodiment, FIG. 10The components can include a display 1024, a touch screen 1025, a touch pad 1030, a near field communication unit ("NFC") 1045, a sensor hub 1040, a thermal sensor 1039, an express chipset ("EC") 1035, a trusted platform module ("TPM") 1038, a BIOS / firmware / flash memory ("BIOS, FW Flash") 1022, a DSP 1060, a drive "SSD or HDD" 1020 (e.g., solid state disk ("SSD") or hard disk drive ("HDD")), a wireless local area network unit ("WLAN") 1050, a Bluetooth unit 1052, a wireless wide area network unit ("WWAN") 1056, a global positioning system (GPS) 1055, a camera ("USB 3.0 camera") 1054 (e.g., USB 3.0 camera), or a low power double data rate ("LPDDR") memory unit ("LPDDR3") 1015 implemented in, for example, LPDDR3 standard. These components can each be implemented in any suitable manner.

[0297] In at least one embodiment, other components can be communicatively coupled to processor 1010 by components as described above. In at least one embodiment, an accelerometer 1041, an ambient light sensor ("ALS") 1042, a compass 1043, and a gyroscope 1044 can be communicatively coupled to sensor hub 1040. In at least one embodiment, thermal sensor 1039, fan 1037, keyboard 1036, and touch pad 1030 can be communicatively coupled to EC 1035. In at least one embodiment, speaker 1063, earpiece 1064, and microphone ("mic") 1065 can be communicatively coupled to an audio unit ("audio codec and class D amplifier") 1064, which in turn can be communicatively coupled to DSP 1060. In at least one embodiment, audio unit 1064 can include, for example and without limitation, an audio coder / decoder ("codec") and a class D amplifier. In at least one embodiment, a SIM card ("SIM") 1057 can be communicatively coupled to WWAN unit 1056. In at least one embodiment, components such as WLAN unit 1050 and Bluetooth unit 1052, as well as WWAN unit 1056, can be implemented as a next generation form factor (NGFF).

[0298] In at least one embodiment, FIG. 10 One or more systems shown in FIG. 11 are used to perform selection of program code optimization with various algorithms, formulas, and processes (e.g., those described in connection with FIG. 1 to FIG. 2 In at least one embodiment, FIG. 10 One or more systems shown in FIG. 11 are used to implement one or more systems and / or processes (e.g., those described in connection with FIG. 1 to FIG. 6Bthose described herein), such as performing selection of program code optimizations in one or more compilers, and / or otherwise performing the operations described herein.

[0299] FIG. 11 A computer system 1100 is shown in accordance with at least one embodiment. In at least one embodiment, the computer system 1100 is configured to implement the various processes and methods described throughout this disclosure.

[0300] In at least one embodiment, computer system 1100 includes, but is not limited to, at least one central processing unit ("CPU") 1102 connected to a communication bus 1110 implemented using any suitable protocol, such as PCI ("Peripheral Component Interconnect"), Peripheral Component Interconnect Express ("PCI-Express"), AGP ("Accelerated Graphics Port"), HyperTransport, or any other bus or point-to-point communication protocol. In at least one embodiment, computer system 1100 includes, but is not limited to, main memory 1104 and control logic (e.g., implemented as hardware, software, or a combination thereof), and data may be stored in main memory 1104 in the form of random access memory ("RAM"). In at least one embodiment, a network interface subsystem ("network interface") 1122 provides an interface to other computing devices and networks for receiving data from computer system 1100 and transmitting data to other systems.

[0301] In at least one embodiment, computer system 1100 includes, but is not limited to, input device 1108, parallel processing system 1112, and display device 1106, which can be implemented using conventional cathode ray tubes ("CRTs"), liquid crystal displays ("LCDs"), light emitting diodes ("LEDs"), plasma displays, or other suitable display technologies. In at least one embodiment, user input is received from input device 1108 (such as a keyboard, mouse, touchpad, microphone, etc.). In at least one embodiment, each of the aforementioned modules can be located on a single semiconductor platform to form a processing system.

[0302] In at least one embodiment, FIG. 11 One or more systems shown in the foregoing are used to utilize various algorithms, formulas, and processes (e.g., in conjunction with FIG. 1 to FIG. 2 those described herein) to perform selection of program code optimizations, and / or otherwise perform the operations described herein. In at least one embodiment, FIG. 11 One or more systems shown in the foregoing are used to implement one or more systems and / or processes (e.g., in conjunction with FIG. 1 to FIG. 6B those described herein), such as performing selection of program code optimizations in one or more compilers, and / or otherwise performing the operations described herein.

[0303] FIG. 12 A computer system 1200 is shown in accordance with at least one embodiment. In at least one embodiment, computer system 1200 includes, but is not limited to, a computer 1210 and a USB stick 1220. In at least one embodiment, computer 1210 may include, but is not limited to, any number and type of processors (not shown) and memory (not shown). In at least one embodiment, computer 1210 includes, but is not limited to, a server, a cloud instance, a laptop computer, and a desktop computer.

[0304] In at least one embodiment, the USB stick 1220 includes, but is not limited to, a processing unit 1230, a USB interface 1240, and USB interface logic 1250. In at least one embodiment, the processing unit 1230 can be any instruction execution system, apparatus, or device capable of executing instructions. In at least one embodiment, the processing unit 1230 can include, but is not limited to, any number and type of processing cores (not shown). In at least one embodiment, the processing core 1230 includes an application-specific integrated circuit ("ASIC") that is optimized to perform any number and type of operations associated with machine learning. For example, in at least one embodiment, the processing core 1230 is a tensor processing unit ("TPC") that is optimized to perform machine learning inference operations. In at least one embodiment, the processing core 1230 is a vision processing unit ("VPU") that is optimized to perform machine vision and machine learning inference operations.

[0305] In at least one embodiment, USB interface 1240 can be any type of USB connector or USB receptacle. For example, in at least one embodiment, USB interface 1240 is a USB 3.0 Type-C receptacle for data and power. In at least one embodiment, USB interface 1240 is a USB 3.0 Type-A connector. In at least one embodiment, USB interface logic 1250 can include any number and type of logic that enables processing unit 1230 to connect to a device (e.g., computer 1210) via USB connector 1240.

[0306] In at least one embodiment, FIG. 12 One or more systems shown in the foregoing are used to utilize various algorithms, formulas, and processes (e.g., in conjunction with FIG. 1 to FIG. 2 those described herein) to perform selection of program code optimizations, and / or otherwise perform the operations described herein. In at least one embodiment, FIG. 12 One or more systems shown in the foregoing are used to implement one or more systems and / or processes (e.g., in conjunction with FIG. 1 to FIG. 6BThose described, such as performing a selection of program code optimizations in one or more compilers, and / or otherwise performing the operations described herein.

[0307] FIG. 13A An exemplary architecture is shown in which a plurality of GPUs 1310-1313 are communicatively coupled to a plurality of multi-core processors 1305-1306 over high-speed links 1340-1343 (e.g., buses, point-to-point interconnects, etc.). In one embodiment, the high-speed links 1340-1343 support a communication throughput of 4GB / s, 30GB / s, 80GB / s or higher. Various interconnect protocols can be used including, but not limited to, PCIe 4.0 or 5.0 and NVLink 2.0.

[0308] Further, in one embodiment, two or more of the GPUs 1310-1313 are interconnected over high-speed links 1329-1330, which can be implemented using the same or different protocol / links as those used for high-speed links 1340-1343. Similarly, two or more of the multi-core processors 1305-1306 can be connected over a high-speed link 1328, which can be an SMP bus running at 20GB / s, 30GB / s, 120GB / s or higher. Alternatively, the same protocol / link used by high-speed links 1340-1343 (e.g., over common interconnect fabric) can be used for the high-speed link 1328 connecting FIG. 13A all communication between the various system components shown in FIG. 13.

[0309] In one embodiment, each multi-core processor 1305-1306 is communicatively coupled to processor memory 1301-1302 via memory interconnects 1326-1327, respectively, and each GPU 1310-1313 is communicatively coupled to GPU memory 1320-1323 over GPU memory interconnects 1350-1353, respectively. The memory interconnects 1326-1327 and 1350-1353 can utilize the same or different memory access technologies. By way of example and without limitation, the processor memory 1301-1302 and GPU memory 1320-1323 can be volatile memory such as dynamic random access memory (DRAM) (including stacked DRAM), graphics DDR SDRAM (GDDR) (e.g., GDDR5, GDDR6), or high bandwidth memory (HBM), and / or can be non-volatile memory such as 3D XPoint or Nano-Ram. In one embodiment, certain portions of the processor memory 1301-1302 can be volatile memory while another portion can be non-volatile memory (e.g., using a two-level memory (2LM) hierarchy).

[0310] As described herein, although the various processors 1305-1306 and GPUs 1310-1313 can be physically coupled to specific memories 1301-1302, 1320-1323, respectively, a unified memory architecture can be implemented in which the same virtual system address space (also referred to as an "effective address" space) is distributed across the various physical memories. For example, processor memories 1301-1302 can each contain 64 GB of system memory address space, and GPU memories 1320-1323 can each contain 32 GB of system memory address space (resulting in a total of 256 GB of addressable memory size in this example).

[0311] FIG. 13B Additional details for an interconnect between multi-core processor 1307 and graphics acceleration module 1346 are shown according to one example embodiment. Graphics acceleration module 1346 can include one or more GPU chips integrated on a line card that is coupled via a high-speed link 1340 to processor 1307. Alternatively, graphics acceleration module 1346 can be integrated on the same package or chip as processor 1307.

[0312] In at least one embodiment, processor 1307 is shown including multiple cores 1360A-1360D, each with a translation lookaside buffer 1361A-1361D and one or more caches 1362A-1362D. In at least one embodiment, cores 1360A-1360D can include various other components not shown, for executing instructions and processing data. In at least one embodiment, caches 1362A-1362D can include level one (LI) and level two (L2) caches. In addition, one or more shared caches 1356 can be included in caches 1362A-1362D and shared by groups of cores 1360A-1360D. For example, one embodiment of processor 1307 includes 24 cores, each with its own LI cache, twelve shared L2 caches, and twelve shared L3 caches. In this embodiment, two adjacent cores share one or more L2 and L3 caches. Processor 1307 and graphics acceleration module 1346 are connected with system memory 1314, which can include processor memories 1301-1302. FIG. 13A

[0313] ​Consistency for data and instructions stored in the respective caches 1362A-1362D, 1356, and system memory 1314 is maintained via inter-core communication over the coherence bus 1364. For example, each cache can have cache coherency logic / circuitry associated therewith to communicate over the coherence bus 1364 in response to detecting a read or write to a particular cache line. In one implementation, a cache snoop protocol is implemented over the coherence bus 1364 to snoop cache accesses.

[0314] In one embodiment, the agent circuit 1325 communicatively couples the graphics acceleration module 1346 to the coherence bus 1364, allowing the graphics acceleration module 1346 to participate in the cache coherence protocol as a peer to the cores 1360A-1360D. The interface 1335 provides connectivity from the agent circuit 1325 over a high-speed link 1340 (e.g., a PCIe bus, NVLink, etc.) to a switch 1348. The switch 1348 is coupled to a bus 1350 that operates as a hub or switch for the cores 1360A-1360D. The interface 1337 connects the graphics acceleration module 1346 to the link 1340.

[0315] In one implementation, the accelerator integration circuit 1336 provides the graphics acceleration module 1346 with cache management, memory access, context management, and interrupt management services. The graphics processing engines 1331, 1332, N can each comprise a separate graphics processing unit (GPU). Alternatively, the graphics processing engines 1331, 1332, N can selectively comprise different types of graphics processing engines within a GPU, such as graphics execution units, media processing engines (e.g., video encoders / decoders), samplers, and blit engines. In at least one embodiment, the graphics acceleration module 1346 can be a GPU with a plurality of graphics processing engines 1331-1332, N, or the graphics processing engines 1331-1332, N can be individual GPUs integrated on a common package, line card, or chip.

[0316] In one embodiment, the accelerator integrated circuit 1336 includes a memory management unit (MMU) 1339 for performing various memory management functions, such as virtual-to-physical memory translation (also known as effective-to-real memory translation), and a memory access protocol for accessing system memory 1314. The MMU 1339 may also include a translation lookaside buffer ("TLB") (not shown) for caching virtual / effective-to-physical / real address translations. In one implementation, cache 1338 may store commands and data for efficient access by the graphics processing engines 1331-1332, N. In at least one embodiment, data stored in cache 1338 and graphics memory 1333-1334, M is kept consistent with the core caches 1362A-1362D, 1356 and system memory 1314. As previously described, this task may be accomplished via proxy circuitry 1325 acting on behalf of cache 1338 and graphics memory 1333-1334, M (e.g., sending updates related to modifications / accesses of cache lines on processor caches 1362A-1362D, 1356 to cache 1338 and receiving updates from cache 1338).

[0317] A set of registers 1345 stores context data for threads executed by graphics processing engines 1331, 1332, N, and context management circuitry 1348 manages thread contexts. For example, context management circuitry 1348 can perform save and restore operations to save and restore the context of each thread during a context switch (e.g., where a first thread is saved and a second thread is stored so that the second thread can be executed by the graphics processing engine). For example, context management circuitry 1348 can store current register values ​​to a designated area in memory (e.g., identified by a context pointer) upon a context switch. The register values ​​can then be restored upon returning to the context. In one embodiment, interrupt management circuitry 1347 receives and processes interrupts received from system devices.

[0318] In one implementation, the MMU 1339 translates virtual / effective addresses from the graphics processing engines 1331 into real / physical addresses in system memory 1314. One embodiment of the accelerator integration circuit 1336 supports multiple (e.g., 4, 8, 16) graphics accelerator modules 1346 and / or other accelerator devices. The graphics accelerator modules 1346 can be dedicated to a single application executing on the processor 1307 or can be shared between multiple applications. In one embodiment, a virtualized graphics execution environment is presented in which resources of the graphics processing engines 1331-1332, N are shared with multiple applications or virtual machines (VMs). In at least one embodiment, resources can be subdivided into“slices” that are allocated to different VMs and / or applications based on processing requirements and priorities associated with the VMs and / or applications.

[0319] In at least one embodiment, the accelerator integration circuit 1336 performs as a bridge to the system for a system of graphics acceleration modules 1346 and provides address translation and system memory cache services. In addition, the accelerator integration circuit 1336 can provide virtualization facilities for a host processor to manage virtualization of graphics processing engines 1331-1332, N, interrupts, and memory management.

[0320] Because the hardware resources of the graphics processing engines 1331-1332, N are explicitly mapped to the real address space seen by the host processor 1307, any host processor can directly address these resources using effective address values. In at least one embodiment, one function of the accelerator integration circuit 1336 is to physically separate the graphics processing engines 1331-1332, N so that they appear to be independent units to the system.

[0321] In at least one embodiment, one or more graphics memories 1333-1334, M are coupled to each of the graphics processing engines 1331-1332, N, respectively. The graphics memories 1333-1334, M store instructions and data for processing by each of the graphics processing engines 1331-1332, N. The graphics memories 1333-1334, M can be volatile memory, such as DRAM (including stacked DRAM), GDDR memory (e.g., GDDR5, GDDR6), or HBM, and / or can be non-volatile memory, such as 3D XPoint or Nano-Ram.

[0322] In one embodiment, to reduce the data traffic on link 1340, a biasing technique can be used to ensure that data stored in graphics memory 1333-1334, N is most frequently used by graphics processing engines 1331-1332, N and least frequently used by cores 1360A-1360D. Similarly, in at least one embodiment, the biasing mechanism attempts to keep data needed by the cores (and preferably not graphics processing engines 1331-1332, N) in the caches 1362A-1362D, 1356 and system memory 1314 of the cores.

[0323] FIG. 13C Another exemplary embodiment is shown in which accelerator integration circuit 1336 is integrated within processor 1307. In this embodiment, graphics processing engines 1331-1332, N communicate directly over high-speed link 1340 to accelerator integration circuit 1336 via interface 1337 and interface 1335 (which can also utilize any form of bus or interface protocol). Accelerator integration circuit 1336 can perform same operations as those described with regard to FIG. 13B But can have higher throughput due to its close proximity to coherence bus 1364 and caches 1362A-1362D, 1356. One embodiment supports different programming models, including a dedicated process programming model (no graphics acceleration module virtualization) and a shared programming model (with virtualization), which can include programming models controlled by accelerator integration circuit 1336 and programming models controlled by graphics acceleration module 1346.

[0324] In at least one embodiment, graphics processing engines 1331-1332, N are dedicated to a single application or process under a single operating system. In at least one embodiment, a single application can funnel other application requests to graphics processing engines 1331-1332, N, providing virtualization within a VM / partition.

[0325] In at least one embodiment, graphics processing engines 1331-1332, N can be shared by multiple VMs / application partitions. In at least one embodiment, a shared model can use a hypervisor to virtualize the graphics processing engines 1331-1332, N to allow access by each operating system. For a single partition system without a hypervisor, the operating system owns the graphics processing engines 1331-1332, N. In at least one embodiment, the operating system can virtualize the graphics processing engines 1331-1332, N to provide access to each process or application.

[0326] In at least one embodiment, graphics acceleration module 1346 or individual graphics processing engines 1331-1332, N use a process handle to select a process element. In one embodiment, process elements are stored in system memory 1314 and can be addressed using effective to real address translation techniques described herein. In at least one embodiment, a process handle can be an implementation specific value provided to a host process when registering its context with a graphics processing engine 1331-1332, N (i.e., calling system software to add a process element to a process element linked list). In at least one embodiment, lower 16 bits of a process handle can be an offset into a process element linked list for a process.

[0327] FIG. 13D An exemplary accelerator integration slice 1390 is shown. As used herein, a “slice” comprises a specified portion of processing resources of accelerator integration circuit 1336. An application is an effective address space 1382 in system memory 1314 that stores process elements 1383. In one embodiment, process elements 1383 are stored in response to GPU invocations 1381 from applications 1380 executing on processor 1307. Process elements 1383 contain process state for respective applications 1380. A work descriptor (WD) 1384 contained in process element 1383 can be a single job requested by an application or can contain a pointer to a queue of jobs. In at least one embodiment, WD 1384 is a pointer to a job request queue in an address space 1382 of an application.

[0328] Graphics acceleration module 1346 and / or individual graphics processing engines 1331-1332, N can be shared by all or a subset of processes in a system. In at least one embodiment, a process subset can include a hypervisor.

[0329] In at least one embodiment, a dedicated process programming model is implementation specific. In this model, a single process owns a graphics acceleration module 1346 or individual graphics processing engines 1331. As graphics acceleration module 1346 is owned by a single process, a hypervisor initializes accelerator integration circuit for an owned partition, and an operating system initializes accelerator integration circuit 1336 for an owned process when graphics acceleration module 1346 is assigned.

[0330] In operation, a WD fetch unit 1391 in accelerator integration slice 1390 fetches the next WD 1384, which includes an indication of work to be completed by one or more graphics processing engines of graphics acceleration module 1346. Data from WD 1384 can be stored in registers 1345 and used by MMU 1339, interrupt management circuit 1347, and / or context management circuit 1348, as shown. For example, one embodiment of MMU 1339 includes segment / page walk circuitry to access segment / page tables 1386 within OS virtual address space 1385. Interrupt management circuit 1347 can handle interrupt events 1392 received from graphics acceleration module 1346. When performing graphics operations, effective addresses 1393 generated by graphics processing engines 1331-1332, N are translated to real addresses by MMU 1339.

[0331] In one embodiment, a same set of registers 1345 is replicated for each graphics processing engine 1331-1332, N and / or graphics acceleration module 1346, and the registers 1345 can be initialized by a hypervisor or operating system. Each of these replicated registers can be included in accelerator integration slice 1390. Exemplary registers that can be initialized by a hypervisor are shown in Table 1.

[0332]

[0333]

[0334] Exemplary registers that can be initialized by an operating system are shown in Table 2.

[0335]

[0336] In one embodiment, each WD 1384 is specific to a particular graphics acceleration module 1346 and / or graphics processing engine 1331-1332, N. It contains all information needed for the graphics processing engines 1331-1332, N to complete the work, or it can be a pointer to a memory location where an application has set up a command queue of work to be completed.

[0337] FIG. 13E Additional details are shown for one exemplary embodiment of a shared model. This embodiment includes a hypervisor real address space 1398 in which is stored a list of process elements 1399. The hypervisor real address space 1398 is accessible via hypervisor 1396, which virtualizes the graphics acceleration module engines for operating system 1395.

[0338] In at least one embodiment, a shared programming model allows all processes or a subset of processes from all partitions or a subset of partitions in a system to use graphics acceleration module 1346. There are two programming models in which graphics acceleration module 1346 is shared by multiple processes and partitions: time-sliced sharing and graphics-directed sharing.

[0339] In this model, system hypervisor 1396 owns graphics acceleration module 1346 and makes its functionality available to all operating systems 1395. For graphics acceleration module 1346 to support virtualization by system hypervisor 1396, graphics acceleration module 1346 can adhere to the following: (1) application job requests must be autonomous (i.e., state need not be maintained between jobs), or graphics acceleration module 1346 must provide a context save and restore mechanism, (2) graphics acceleration module 1346 guarantees that an application’s job request will complete in a specified amount of time, including any translation faults, or graphics acceleration module 1346 provides the ability to preempt job processing, (3) in directed shared programming models, fairness between graphics acceleration module 1346 processes must be ensured.

[0340] In at least one embodiment, application 1380 is required to use a graphics acceleration module 1346 type, a work descriptor (WD), an authority mask register (AMR) value, and a context save / restore area pointer (CSRP) for an operating system 1395 system call. In at least one embodiment, the graphics acceleration module 1346 type describes a target acceleration function for the system call. In at least one embodiment, the graphics acceleration module 1346 type can be a system-specific value. In at least one embodiment, the WD is formatted specifically for graphics acceleration module 1346 and can take the form of a graphics acceleration module 1346 command, a valid address pointer to a user-defined structure, a valid address pointer to a command queue, or any other data structure describing work to be done by graphics acceleration module 1346. In one embodiment, the AMR value is the AMR state for the current process. In at least one embodiment, the value passed to the operating system is similar to how an application program sets the AMR. If the accelerator integration circuit 1336 and graphics acceleration module 1346 implementation does not support a user authority mask override register (UAMOR), the operating system can apply the current UAMOR value to the AMR value before passing the AMR in a hypervisor call. The hypervisor 1396 can selectively apply the current authority mask override register (AMOR) value before placing the AMR in the process element 1383. In at least one embodiment, the CSRP is one of registers 1345 that contains a valid address of a region in the application’s address space 1382 for graphics acceleration module 1346 to save and restore context state. This pointer is optional if there is no need to save state between jobs or when a job is preempted. In at least one embodiment, the context save / restore region can be a fixed system memory.

[0341] Upon receiving the system call, operating system 1395 can verify that application 1380 is registered and has been granted authority to use graphics acceleration module 1346. Operating system 1395 then invokes hypervisor 1396 using the information shown in Table 3.

[0342]

[0343] Upon receiving the hypervisor call, hypervisor 1396 verifies that operating system 1395 is registered and has been granted authority to use graphics acceleration module 1346. Hypervisor 1396 then places the process element 1383 in a process element linked list of the corresponding graphics acceleration module 1346 type. The process element can include the information shown in Table 4.

[0344]

[0345] ​

[0346] In at least one embodiment, the hypervisor initializes the plurality of accelerator integrated slice 1390 registers 1345 .

[0347] like FIG. 13F As shown, in at least one embodiment, a unified memory is used that is addressable via a common virtual memory address space for accessing physical processor memories 1301-1302 and GPU memories 1320-1323. In this implementation, operations executed on GPUs 1310-1313 utilize the same virtual / effective memory address space to access processor memories 1301-1302, and vice versa, thereby simplifying programmability. In at least one embodiment, a first portion of the virtual / effective address space is allocated to processor memory 1301, a second portion is allocated to second processor memory 1302, a third portion is allocated to GPU memory 1320, and so on. In at least one embodiment, the entire virtual / effective memory space (sometimes referred to as the effective address space) is thus distributed across each of processor memories 1301-1302 and GPU memories 1320-1323, thereby allowing any processor or GPU to access that memory using a virtual address mapped to any physical memory.

[0348] In one embodiment, bias / coherency management circuitry 1394A-1394E within one or more MMUs 1339A-1339E ensures cache coherency between the caches of one or more host processors (e.g., 1305) and GPUs 1310-1313 and implements biasing techniques that indicate the physical memory where certain types of data should be stored. FIG. 13F Multiple instances of bias / coherence management circuits 1394A- 1394E are shown in , but bias / coherence circuits may be implemented within an MMU of one or more host processors 1305 and / or within an accelerator integrated circuit 1336 .

[0349] One embodiment allows GPU-attached memory 1320-1323 to be mapped as part of system memory and accessed using shared virtual memory (SVM) techniques, but without suffering the performance penalties associated with full system cache coherency. In at least one embodiment, the ability to access GPU-attached memory 1320-1323 as system memory without the heavy cache coherency overhead provides a favorable operating environment for GPU offload. This arrangement allows software of host processor 1305 to set operands and access computation results without the overhead of traditional I / O DMA data copies. Such traditional copies include driver calls, interrupts, and memory-mapped I / O (MMIO) accesses, which are all less efficient than simple memory accesses. In at least one embodiment, the ability to access GPU-attached memory 1320-1323 without cache coherency overhead can be critical to the execution time of offloaded computations. For example, in cases with a large amount of streaming write memory traffic, cache coherency overhead can significantly reduce the effective write bandwidth seen by GPU 1310-1313. In at least one embodiment, the efficiency of operand setup, the efficiency of result access, and the efficiency of GPU computation can all play a role in determining the effectiveness of GPU offload.

[0350] In at least one embodiment, the selection of GPU bias and host processor bias is driven by a bias tracker data structure. For example, a bias table can be used, which can be a page-granularity structure (e.g., controlled at the granularity of a memory page) that includes a memory page 1 or 2 bits per GPU attachment. In at least one embodiment, with or without a bias cache in GPU 1310-1313 (e.g., to cache frequently / recently used entries of the bias table), the bias table can be implemented in the stolen memory range of one or more GPU-attached memory 1320-1323. Alternatively, the entire bias table can be maintained within the GPU.

[0351] In at least one embodiment, prior to actually accessing GPU memory, the bias table entry associated with each access to GPU-attached memory 1320-1323 is accessed, resulting in the following operations. First, local requests from GPUs 1310-1313 that find their pages in GPU bias are forwarded directly to the corresponding GPU memory 1320-1323. Local requests from GPUs that find their pages in host bias are forwarded to processor 1305 (e.g., over a high-speed link as described above). In one embodiment, requests from processor 1305 that find requested pages in host processor bias complete the request similarly to a normal memory read. Alternatively, requests that point to GPU-biased pages can be forwarded to GPUs 1310-1313. In at least one embodiment, if the page is not currently in use by the GPU, the GPU can then migrate the page to host processor bias. In at least one embodiment, the bias state of a page can be changed by software-based mechanisms, hardware-assisted software-based mechanisms, or in limited cases purely hardware-based mechanisms.

[0352] One mechanism for changing bias state employs an API call (e.g., OpenCL) that in turn invokes a device driver of a GPU, which in turn sends a message (or causes a command descriptor to be enqueued) to the GPU, directing the GPU to change the bias state, and in certain migrations to perform a cache flush operation in the host. In at least one embodiment, the cache flush operation is used for migrations from host processor 1305 bias to GPU bias, but not for the reverse.

[0353] In one embodiment, cache coherency is maintained by temporarily rendering GPU-biased pages that cannot be cached by host processor 1305. To access these pages, processor 1305 can request access from GPU 1310, which can or can not grant access immediately. Thus, to reduce communication between processor 1305 and GPU 1310, it is beneficial to ensure that GPU-biased pages are pages that are needed by the GPU but not by host processor 1305, and vice versa.

[0354] FIG. 14 Exemplary integrated circuits and associated graphics processors in accordance with various embodiments described herein are shown, which can be fabricated using one or more IP cores. In addition to the illustrations, other logic and circuitry can be included in at least one embodiment, including additional graphics processors / cores, peripheral interface controllers or general-purpose processor cores.

[0355] FIG. 14is a block diagram illustrating an exemplary system on a chip integrated circuit 1400 that can be fabricated using one or more IP cores, according to at least one embodiment. In at least one embodiment, integrated circuit 1400 includes one or more application processor(s) 1405 (e.g., CPUs), at least one graphics processor 1410, and can additionally include an image processor 1415 and / or a video processor 1420, any of which can be a modular IP core. In at least one embodiment, integrated circuit 1400 includes peripheral or bus logic including a USB controller 1425, a UART controller 1430, an SPI / SDIO controller 1435, and an I2S / I2C controller 1440. In at least one embodiment, integrated circuit 1400 can include a display device 1445 coupled to one or more of a high-definition multimedia interface (HDMI) controller 1450 and a mobile industry processor interface (MIPI) display interface 1455. In at least one embodiment, storage can be provided by a flash memory subsystem 1460 including flash memory and a flash memory controller. In at least one embodiment, memory interface can be provided via a memory controller 1465 for access to SDRAM or SRAM memory devices. In at least one embodiment, some integrated circuits also include an embedded security engine 1470.

[0356] In at least one embodiment, FIG. 14 One or more systems shown in FIG. 11 are used to perform selection of program code optimizations with various algorithms, formulas, and processes such as those described in conjunction with FIG. 1 to FIG. 2 In at least one embodiment, FIG. 14 One or more systems shown in FIG. 11 are used to implement one or more systems and / or processes such as those described in conjunction with FIG. 1 to FIG. 6B In at least one embodiment,

[0357] FIG. 15A-FIG. 15B Exemplary integrated circuits and associated graphics processors according to various embodiments described herein are shown, which can be fabricated using one or more IP cores. In addition to what is illustrated, other logic and circuitry can be included in the at least one embodiment, including additional graphics processors / cores, peripheral interface controllers, or general purpose processor cores.

[0358] FIG. 15A-FIG. 15B is a block diagram illustrating an exemplary graphics processor used within a SoC according to embodiments described herein. FIG. 15A An exemplary graphics processor 1510 of a system on a chip integrated circuit, which can be fabricated using one or more IP cores, according to at least one embodiment is shown.FIG. 15B Another exemplary graphics processor 1540 of a system on a chip integrated circuit, in accordance with at least one embodiment, is shown which can be fabricated using one or more IP cores. In at least one embodiment, graphics processor 1510 is a low power graphics processor core. In at least one embodiment, graphics processor 1540 is a higher performance graphics processor core. In at least one embodiment, each graphics processor 1510, 1540 can be a variant of graphics processor 1410. FIG. 15A Graphics processor 1510 of FIG. 14A is a low power graphics processor core. In at least one embodiment, graphics processor 1540 of FIG. 14B is a higher performance graphics processor core. In at least one embodiment, each graphics processor 1510, 1540 can be a variant of graphics processor 1410. FIG. 15B Graphics processor 1510 of FIG. 14A is a low power graphics processor core. In at least one embodiment, graphics processor 1540 of FIG. 14B is a higher performance graphics processor core. In at least one embodiment, each graphics processor 1510, 1540 can be a variant of graphics processor 1410. FIG. 14 Variants of graphics processor 1410 of FIG. 14A.

[0359] In at least one embodiment, graphics processor 1510 includes a vertex processor 1505 and one or more fragment processor(s) 1515A-1515N (e.g., 1515A, 1515B, 1515C, 1515D, through 1515N-1, and 1515N). In at least one embodiment, graphics processor 1510 can execute different shader programs via separate logical

[0360] In at least one embodiment, graphics processor 1510 additionally includes one or more memory management units (MMUs) 1520A-1520B, one or more caches 1525A-1525B, and one or more circuit interconnects 1530A-1530B. In at least one embodiment, one or more MMUs 1520A-1520B provide a mapping of virtual to physical addresses for graphics processor 1510, including for vertex processor 1505 and / or fragment processors 1515A-1515N, which may reference vertex or image / texture data stored in memory, in addition to vertex or image / texture data stored in one or more caches 1525A-1525B. In at least one embodiment, one or more MMUs 1520A-1520B may synchronize with other MMUs within the system, including with other MMUs. FIG. 14 One or more MMUs associated with one or more application processors 1405, graphics processor 1415, and / or video processor 1420 enable each processor 1405-1420 to participate in a shared or unified virtual memory system. In at least one embodiment, one or more circuit interconnects 1530A-1530B enable graphics processor 1510 to connect to other IP cores within the SoC via an internal bus of the SoC or via a direct connection.

[0361] In at least one embodiment, graphics processor 1540 includes FIG. 15A One or more MMUs 1520A-1520B, caches 1525A-1525B, and circuit interconnects 1530A-1530B of the graphics processor 1510. In at least one embodiment, the graphics processor 1540 includes one or more shader cores 1555A-1555N (e.g., 1555A, 1555B, 1555C, 1555D, 1555E, 1555F through 1555N-1 and 1555N) that provide a unified shader core architecture in which a single core or type or core can execute all types of programmable shader code, including shader program code for implementing vertex shaders, fragment shaders, and / or compute shaders. In at least one embodiment, the number of shader cores can vary. In at least one embodiment, the graphics processor 1540 includes an inter-core task manager 1545 that acts as a thread dispatcher to dispatch execution threads to one or more shader cores 1555A-1555N and a tiling unit 1558 to accelerate tile-based rendering operations in which rendering operations of a scene are subdivided in image space, for example, to exploit local spatial coherence within a scene or to optimize the use of internal caches.

[0362] In at least one embodiment, FIG. 15A-15BOne or more systems shown in the foregoing are used to utilize various algorithms, formulas, and processes (e.g., in conjunction with FIG. 1 to FIG. 2 those described herein) to perform selection of program code optimizations, and / or otherwise perform the operations described herein. In at least one embodiment, FIG. 15A-15B One or more systems shown in the foregoing are used to implement one or more systems and / or processes (e.g., in conjunction with FIG. 1 to FIG. 6B those described herein), such as performing selection of program code optimizations in one or more compilers, and / or otherwise performing the operations described herein.

[0363] FIG. 16A-FIG. 16B Additional exemplary graphics processor logic according to embodiments described herein is shown. In at least one embodiment, FIG. 16A Shows that can be included in FIG. 14 Graphics core 1600 within graphics processor 1410 of FIG. 1 and, in at least one embodiment, may be such as FIG. 15B Unified shader cores 1555A-1555N are shown. FIG. 16B A highly parallel, general-purpose graphics processing unit 1630 suitable for deployment on a multi-chip module in at least one embodiment is shown.

[0364] In at least one embodiment, graphics core 1600 includes a shared instruction cache 1602, texture units 1618, and cache / shared memory 1620, which are common to execution resources within graphics core 1600. In at least one embodiment, graphics core 1600 may include multiple slices 1601A-1601N, or partitions of each core, and the graphics processor may include multiple instances of graphics core 1600. Slices 1601A-1601N may include support logic including local instruction caches 1604A-1604N, thread schedulers 1606A-1606N, thread dispatchers 1608A-1608N, and a set of registers 1610A-1610N. In at least one embodiment, slices 1601A-1601N may include a set of additional function units (AFUs 1612A-1612N), floating point units (FPUs 1614A-1614N), integer arithmetic logic units (ALUs 1616A-1616N), address calculation units (ACUs 1613A-1613N), double-precision floating point units (DPFPUs 1615A-1615N), and matrix processing units (MPUs 1617A-1617N).

[0365] In at least one embodiment, FPUs 1614A-1614N can perform single-precision (32-bit) and half-precision (16-bit) floating point operations, while DPFPUs 1615A-1615N perform double-precision (64-bit) floating point operations. In at least one embodiment, ALUs 1616A-1616N can perform variable precision integer operations including 8-bit, 16-bit, and 32-bit integer operations at integer ALU 1616A-1616N, which can be configured to perform operations including addition, subtraction, multiplication, and shift; and can be configured to perform a mix of operations including integer and floating point operations as well. In at least one embodiment, MPUs 1617A-1617N can also be configured for mixed precision matrix operations including half-precision floating point operations and 8-bit integer operations. In at least one embodiment, MPUs 1617-1617N can perform various matrix operations to accelerate machine learning application frameworks, including enabling support for accelerated General Matrix to Matrix Multiplication (GEMM). In at least one embodiment, AFUs 1612A-1612N can perform additional logical operations not supported by floating point or integer units, including trigonometric operations (e.g., Sine, Cosine, etc.).

[0366] In at least one embodiment, one or more systems shown in FIG. 16A are used to perform selection of program code optimizations using various algorithms, formulas, and processes (such as those described in reference to FIG. 1 to FIG. 2 In at least one embodiment, one or more systems shown in FIG. 16A are used to implement one or more systems and / or processes (such as those described in reference to FIG. 1 to FIG. 6B In at least one embodiment, one or more systems shown in

[0367] FIG. 16BA general purpose processing unit (GPGPU) 1630 is shown in at least one embodiment, which can be configured to enable highly parallel computing operations to be performed by a group of graphics processing units. In at least one embodiment, GPGPU 1630 can be directly linked to other instances of GPGPU 1630 to create a multi-GPU cluster to increase the training speed for deep neural networks. In at least one embodiment, GPGPU 1630 includes a host interface 1632 to enable connection to a host processor. In at least one embodiment, host interface 1632 is a PCI Express interface. In at least one embodiment, host interface 1632 can be a vendor-specific communication interface or communication structure. In at least one embodiment, GPGPU 1630 receives commands from the host processor and uses a global scheduler 1634 to assign execution threads associated with those commands to a group of compute clusters 1636A-1636H. In at least one embodiment, compute clusters 1636A-1636H share cache memory 1638. In at least one embodiment, cache memory 1638 may be used as a higher level cache for cache memory within compute clusters 1636A-1636H.

[0368] In at least one embodiment, GPGPU 1630 includes memory 1644A-1644B coupled to compute clusters 1636A-1636H via a set of memory controllers 1642A-1642B. In at least one embodiment, memory 1644A-1644B may include various types of memory devices, including dynamic random access memory (DRAM) or graphics random access memory, such as synchronous graphics random access memory (SGRAM), including graphics double data rate (GDDR) memory.

[0369] In at least one embodiment, computing clusters 1636A-1636H each include a set of graphics cores, e.g. FIG. 16A The graphics core 1600 may include multiple types of integer and floating-point logic units that can perform computational operations at various precision ranges, including precision suitable for machine learning computations. For example, in at least one embodiment, at least a subset of the floating-point units in each of the compute clusters 1636A-1636H may be configured to perform 16-bit or 32-bit floating-point operations, while a different subset of the floating-point units may be configured to perform 64-bit floating-point operations.

[0370] In at least one embodiment, multiple instances of GPGPU 1630 can be configured to function as a compute cluster. In at least one embodiment, the communications used by compute clusters 1636A-1636H for synchronization and data exchange vary between embodiments. In at least one embodiment, multiple instances of GPGPU 1630 communicate via host interface 1632. In at least one embodiment, GPGPU 1630 includes an I / O hub 1639 that couples GPGPU 1630 to GPU link 1640, enabling direct connections to other instances of GPGPU 1630. In at least one embodiment, GPU link 1640 is coupled to a dedicated GPU-to-GPU bridge that enables communication and synchronization between multiple instances of GPGPU 1630. In at least one embodiment, GPU link 1640 is coupled to a high-speed interconnect to send and receive data to other GPGPUs or parallel processors. In at least one embodiment, multiple instances of GPGPU 1630 reside in separate data processing systems and communicate via a network device accessible through host interface 1632. In at least one embodiment, GPU link 1640 may be configured to enable connection to a host processor in addition to or as an alternative to host interface 1632 .

[0371] In at least one embodiment, GPGPU 1630 can be configured to train neural networks. In at least one embodiment, GPGPU 1630 can be used within an inference platform. In at least one embodiment, where GPGPU 1630 is used for inference, the GPGPU can include fewer compute clusters 1636A-1636H than when the GPGPU is used to train a neural network. In at least one embodiment, the memory technology associated with memory 1644A-1644B can differ between the inference and training configurations, with higher-bandwidth memory technology being dedicated to the training configuration. In at least one embodiment, the inference configuration of GPGPU 1630 can support inference-specific instructions. For example, in at least one embodiment, the inference configuration can provide support for one or more 8-bit integer dot product instructions, which can be used during inference operations of a deployed neural network.

[0372] In at least one embodiment, FIG. 16B One or more systems shown in the foregoing are used to utilize various algorithms, formulas, and processes (e.g., in conjunction with FIG. 1 to FIG. 2 those described herein) to perform selection of program code optimizations, and / or otherwise perform the operations described herein. In at least one embodiment, FIG. 16B One or more systems shown in the foregoing are used to implement one or more systems and / or processes (e.g., in conjunction with FIG. 1 to FIG. 6BThose described, such as performing a selection of program code optimizations in one or more compilers, and / or otherwise performing the operations described herein.

[0373] FIG. 17 A block diagram of a computer system 1700 is shown, in accordance with at least one embodiment. In at least one embodiment, computer system 1700 includes a processing subsystem 1701, with one or more processor(s) 1702, and a system memory 1704, communicating via an interconnection path 1705, which can include a memory hub 1705. In at least one embodiment, the memory hub 1705 can be a separate component, or it can be integrated into one or more of the processor(s) 1702. In at least one embodiment, memory hub 1705 couples with processor(s) 1702 via communication links 1706 to perform lookups of

[0374] In at least one embodiment, processing subsystem 1701 includes one or more parallel processor(s) 1712, coupled to memory hub 1705 via a bus or other communication link 1713. In at least one embodiment, communication link 1713 can be any of a number of standards-based communication links, such as, but not limited to, a PCI Express, or it can be a vendor specific communications interface or communications structure. In at least one embodiment, one or more parallel processor(s) 1712 form a programmable processing sub system, in which case, one or more parallel processor(s) 1712 can execute programs as defined by program instructions contained in program

[0375] In at least one embodiment, system storage unit 1714 can connect to I / O hub 1707 to provide storage mechanisms for computing system 1700. In at least one embodiment, I / O switches 1716 can be used to provide an interface mechanism to enable connections between I / O hub 1707 and other components, such as network adapter 1718 and / or wireless network adapter 1717 that can be integrated into a platform, as well as various other devices that can be added via one or more add-in devices 1720. In at least one embodiment, network adapter 1718 can be an Ethernet adapter or another wired network adapter. In at least one embodiment, wireless network adapter 1719 can include one or more of Wi-Fi, Bluetooth, Near Field Communication (NFC), or other network devices that include one or more radio(s).

[0376] In at least one embodiment, computer system 1700 can include other components not explicitly shown, including USB or other port connections, optical storage drives, video capture devices, etc., which can also be connected to I / O hub 1707. In at least one embodiment, interconnection of the various components of computing system 1700 can be implemented using any suitable protocols, including PCI- based protocols (e.g., PCI Express), or other bus or point-to-point communication interfaces and / or protocols. FIG. 17

[0377] In at least one embodiment, parallel processor(s) 1712 include circuitry optimized for graphics and video processing, including for example video output circuitry, and are configured for use in a gaming console, a personal computer, or other system. In at least one embodiment, parallel processor(s) 1712 incorporate circuitry optimized for general use computational processing, which can comprise parallel processor(s) 1712 configured for use in a server or a personal computer. In at least one embodiment, system memory 1706 stores data and software for use by central processor 1702 and parallel processor(s) 1712. In at least one embodiment, system memory 1706 stores an operating system 1720, application programs 1722, and any data 1724 used by the operating system 1720 and application programs 1722. In at least one embodiment, system memory 1706 stores a programming module 1726 for use by parallel processor(s) 1712. In at least one embodiment, system memory 1706 stores a programming module 1728 for use by parallel processor(s) 1712.

[0378] In at least one embodiment, FIG. 17 ​One or more systems shown in FIG. 11 are used to implement one or more systems and / or processes (e.g., those described in connection with FIG. 1 to FIG. 2 The described operations can be performed using a variety of algorithms, formulas, and processes, such as those described in connection with FIG. 17 One or more systems shown in FIG. 11 are used to implement one or more systems and / or processes (e.g., those described in connection with FIG. 1 to FIG. 6B The described operations can be performed using a variety of algorithms, formulas, and processes, such as those described in connection with

[0379] Processor

[0380] FIG. 18A A parallel processor 1800 according to at least one embodiment is shown. In at least one embodiment, various components of parallel processor 1800 can be implemented using one or more integrated circuits, which can be programmable processors, application specific integrated circuits (ASICs), or field programmable gate arrays (FPGAs). In at least one embodiment, parallel processor 1800 is a graphics processor that can be used in graphics processing unit (GPU) 1714. FIG. 17 Variations of one or more parallel processors 1712 shown in FIG. 17.

[0381] In at least one embodiment, parallel processor 1800 includes a parallel processing unit 1802. In at least one embodiment, parallel processing unit 1802 includes an I / O unit 1804 that enables communication with other devices, including other instances of parallel processing unit 1802. In at least one embodiment, I / O unit 1804 can be directly connected to other devices. In at least one embodiment, I / O unit 1804 connects with other devices using a hub or switch interface, such as memory hub 1805. In at least one embodiment, connections between memory hub 1805 and I / O unit 1804 form a communication link. In at least one embodiment, I / O unit 1804 connects with a host interface 1806 and a memory crossbar 1816, where host interface 1806 receives commands directed to processing operations and memory crossbar 1816 receives commands directed to memory operations.

[0382] In at least one embodiment, when host interface 1806 receives a command buffer via I / O unit 1804, host interface 1806 can direct a work operation to execute those commands to front end 1808. In at least one embodiment, front end 1808 is coupled with scheduler 1810, which is configured to assign commands or other work items to processing cluster array 1812. In at least one embodiment, scheduler 1810 ensures that processing cluster array 1812 is properly configured and in an active state before assigning tasks to processing cluster array 1812. In at least one embodiment, scheduler 1810 is implemented by firmware logic executing on a microcontroller. In at least one embodiment, microcontroller- implemented scheduler 1810 is configurable to perform complex scheduling and work distribution operations with coarse and fine grain granularity, enabling fast preemption and context switching for threads executing on processing array 1812. In at least one embodiment, host software can prove a workload for scheduling on processing array 1812 through one of a number of graphics processing doorbells. In at least one embodiment, workload can then be automatically distributed on processing array 1812 by scheduler 1810 logic within microcontroller that includes scheduler 1810.

[0383] In at least one embodiment, processing cluster array 1812 can include up to “N” processing clusters (e.g., cluster 1814A, cluster 1814B, through cluster 1814N). In at least one embodiment, each cluster 1814A-1814N of processing cluster array 1812 can execute a large number of concurrent threads. In at least one embodiment, scheduler 1810 can use various scheduling and / or work distribution algorithms to assign work to clusters 1814A-1814N of processing cluster array 1812, which can vary depending on workload produced by each type of program or computation. In at least one embodiment, scheduling can be handled by scheduler 1810 dynamically, or can be assisted in part by compiler logic during compilation of program logic configured for execution by processing cluster array 1812. In at least one embodiment, different clusters 1814A-1814N of processing cluster array 1812 can be allocated for processing different types of programs or for performing different types of computations.

[0384] In at least one embodiment, processing cluster array 1812 can be configured to perform various types of parallel processing operations. In at least one embodiment, processing cluster array 1812 is configured to perform general-purpose parallel compute operations. For example, in at least one embodiment, processing cluster array 1812 can include logic to perform processing tasks including filtering of video and / or audio data, performing modeling operations, including physics operations, and performing data transformations.

[0385] In at least one embodiment, processing cluster array 1812 is configured to perform parallel graph processing operations. In at least one embodiment, processing cluster array 1812 can include additional logic to support performance of such graph processing operations, including but not limited to texture mapping logic to perform texture operations, and tessellation logic, and other vertex processing logic. In at least one embodiment, processing cluster array 1812 can be configured to execute shader programs related to graphics processing, such as but not limited to vertex shaders, tessellation shaders, geometry shaders, and pixel (or fragment) shaders. In at least one embodiment, parallel processor 1802 can transfer data to be processed from system memory via I / O unit 1804. In at least one embodiment, data

[0386] In at least one embodiment, when parallel processor 1802 is used to perform graphics processing, scheduler 1810 can be configured to divide the processing workload into approximately equal sized tasks to better enable distribution of the graphics processing operations across multiple clusters 1814A-1814N of processing cluster array 1812. In at least one embodiment, portions of processing cluster array 1812 can be configured to perform different types of processing. For example, in at least one embodiment, a first portion can be configured to perform vertex shading and topology generation, a second portion can be configured to perform tessellation and geometry shading, and a third portion can be configured to perform pixel shading or other screen space operations to produce a rendered image for display. In at least one embodiment, intermediate data produced by one or more of clusters 1814A-1814N can be stored in buffers to allow transmission of the intermediate data between clusters 1814A-1814N for further processing.

[0387] In at least one embodiment, processing cluster array 1812 can receive processing tasks to be executed via scheduler 1810, which receives commands defining the processing tasks from front end 1808. In at least one embodiment, a processing task can include an index of data to be processed, such as surface (patch) data, primitive data, vertex data, and / or pixel data, as well as state parameters and commands defining how the data is to be processed (e.g., what program is to be executed). In at least one embodiment, scheduler 1810 can be configured to fetch the index corresponding to a task, or can receive the index from front end 1808. In at least one embodiment, front end 1808 can be configured to ensure that processing cluster array 1812 is configured in an effective state before launching a workload specified by an incoming command buffer (e.g., a batch-buffer, a push buffer, etc.).

[0388] In at least one embodiment, each of one or more instances of parallel processing unit 1802 can be coupled to a parallel processor memory 1822. In at least one embodiment, parallel processor memory 1822 can be accessed by the memory crossbar 1816, which can receive memory requests from the processing cluster array 1812 and the I / O unit 1804. In at least one embodiment, memory crossbar 1816 can access parallel processor memory 1822 via a memory interface 1818. In at least one embodiment, memory interface 1818 can include a number of memory ports, for example, a first memory port 1820A, a second memory port 1820B up to an nth memory port 1820N. In at least one embodiment, each memory port can be coupled to a respective memory unit in parallel processor memory 1822, such as first memory unit 1824A, second memory unit 1824B up to Nth memory unit 1824N, which can provide storage for parallel processing units 1802. In at least one embodiment, memory crossbar 1816 can be configured to route an operation modified data received from processing cluster array 1812 or I / O unit 1804, to a memory port that can route the operation to a respective memory unit in parallel processor memory 1822.

[0389] In at least one embodiment, memory units 1824A-1824N can include various types of memory devices including dynamic random access memory (DRAM) or graphics random access memory, such as synchronous graphics random access memory (SGRAM). In at least one embodiment, memory units 1824A-1824N can also include 3D stacked memory including but not limited to high bandwidth memory (HBM). In at least one embodiment, rendering targets such as frame buffers or texture maps can be stored across memory units 1824A-1824N allowing partition units 1820A-1820N to write portions of each rendering target in parallel to effectively use available bandwidth of parallel processor memory 1822. In at least one embodiment, local instances of parallel processor memory 1822 can be excluded in favor of a unified memory design that utilizes system memory in combination with local cache memory.

[0390] In at least one embodiment, any of clusters 1814A-1814N of processing cluster array 1812 can process data to be written into any of memory locations 1824A-1824N within parallel processor memory 1822. In at least one embodiment, memory crossbar 1816 can be configured to transmit outputs of each cluster 1814A-1814N to any partition unit 1820A-1820N or another cluster 1814A-1814N, which can perform additional processing operations on the outputs. In at least one embodiment, each cluster 1814A-1814N can communicate with memory interface 1818 through memory crossbar 1816 to read from or write to various external memory devices. In at least one embodiment, memory crossbar 1816 has a connection to memory interface 1818 to communicate with I / O unit 1804, and a local instance of connection to parallel processor memory 1822, to enable processing clusters 1814A-1814N within different processing clusters 1814A-1814N to communicate with system memory or other memories not local to the parallel processing units 1802. In at least one embodiment, memory crossbar 1816 can use virtual channels to separate traffic streams between clusters 1814A-1814N and partition units 1820A-1820N.

[0391] In at least one embodiment, multiple instances of parallel processing unit 1802 can be provided on a single add-in card, or multiple add-in cards can be interconnected. In at least one embodiment, different instances of parallel processing unit 1802 can be configured to operate together as a single parallel processing unit 1802, even if the different instances have different numbers of processing cores, different amounts of local parallel processor memory, and / or other configuration differences. In at least one embodiment, some instances of parallel processing unit 1802 can be configured to operate as processing clusters of a multi -cluster parallel processor. In at least one embodiment, instances of parallel processing unit 1802 can be configured to operate in a locked state, in which case external devices cannot access the internal state of the instance. In at least one embodiment, instances can operate in a unlocked state, in which an external device can access the internal state.

[0392] FIG. 18B is a block diagram of a partition unit 1820, in accordance with at least one embodiment. In at least one embodiment, partition unit 1820 is a FIG. 18Aone of the partition units 1820A-1820N. In at least one embodiment, partition unit 1820 includes an L2 cache 1821, a frame buffer interface 1825, and a ROP 1826 (Raster Operations Unit). L2 cache 1821 is a read / write cache that is configured to perform load and store operations received from memory crossbar 1816 and ROP 1826. In at least one embodiment, L2 cache 1821 outputs read misses and urgent write-backs to frame buffer interface 1825 for processing. In at least one embodiment, updates can also be sent to a frame buffer via frame buffer interface 1825 for processing. In at least one embodiment, frame buffer interface 1825 interacts with one of memory units 1824A-1824N (e.g., within parallel processor memory 1822) in parallel processor memory. FIG. 18A

[0393] In at least one embodiment, ROP 1826 is a processing unit that performs raster operations including, for example, fill, line, ellipse, triangle, and / or the like. In at least one embodiment, ROP 1826 is configured to execute shaders consumed by parallel processor 1800. In at least one embodiment, ROP 1826 includes support for variable length storage mode objects, multiple vertex shader instruction caches, multiple fragment shader instruction caches, and / or the like. In at least one embodiment, ROP 1826 is configured to support programing techniques, such as those described in U.S. Patent 7,988,773, entitled "Programmable Shader Storage Mode," which issued on July 26, 2011, the disclosure of which is incorporated herein by reference in its entirety.

[0394] In at least one embodiment, ROP 1826 is included within each processing cluster 1814A-1814N (e.g., of parallel processor 1800) instead of in the partition unit 1820. In at least one embodiment, read and write requests for pixel data are transmitted over memory crossbar 1816 instead of pixel fragment data transfers. In at least one embodiment, processed graphics data can be displayed on display device(s) 1710, routed to a graphics processor for further processing, or routed to one of the processing entities within parallel processor 1800 for further processing. FIG. 18A FIG. 17 FIG. 18A

[0395] FIG. 18C is a block diagram of a processing cluster 1814 within a parallel processing unit according to at least one embodiment. In at least one embodiment, processing cluster 1814 is a FIG. 18A ​​​​one of the processing clusters 1814A-1814N. In at least one embodiment, processing cluster 1814 can be configured to perform a number of threads in parallel, where the term “thread” refers to an instance of a particular program executing on a particular set of input data. In at least one embodiment, Single-Instruction, Multiple-Data (SIMD) instruction issue techniques are used to support parallel execution of a large number of threads with no or negligible program overhead when switching from one thread to another. In at least one embodiment, Single-Instruction, Multiple Thread (SIMT) techniques are used to support parallel execution of a large number of generally synchronous threads, using a common instruction unit configured to issue instructions to a set of processing engines within each processing cluster.

[0396] In at least one embodiment, operation of processing cluster 1814 can be controlled via a pipeline manager 1832 that is assigned to processing task(s) by scheduler 1810. In at least one embodiment, pipeline manager 1832 receives instructions from scheduler 1810, and manages execution of those instructions by graphics multiprocessor 1834 and / or texture unit 1836. In at least one embodiment, graphics multiprocessor 1834 is an exemplary instance of a SIMT parallel processor. However, in at least one embodiment, various types of SIMT parallel processors of differing architectures can be included within processing cluster 1814. In at least one embodiment, one or more instances of graphics multiprocessor 1834 can be included within processing cluster 1814. In at least one embodiment, graphics multiprocessor 1834 can process data, and a data crossbar 1840 can be used to distribute processed data to one of a number of possible destinations, including other shader units. In at least one embodiment, pipeline manager 1832 can facilitate distribution by specifying destinations for processed data as a function of its source within pipeline. FIG. 18A

[0397] In at least one embodiment, each graphics multiprocessor 1834 within processing cluster 1814 can include an identical set of functional execution logic (e.g., arithmetic logic, load store units, etc.). In at least one embodiment, functional execution logic can be configured in a pipelined manner in which new instructions can be issued before previous instructions are complete. In at least one embodiment, functional execution logic supports a variety of operations including integer and floating point arithmetic, comparison operations, Boolean operations, shift operations, and a multitude of algebraic functions. In at least one embodiment, same functional-unit hardware can be leveraged to perform different operations using different control signals. Any combination of

[0398] ​In at least one embodiment, instructions delivered to processing cluster 1814 constitute a thread. In at least one embodiment, a set of threads executing across a set of parallel processing engines constitutes a warp. In at least one embodiment, a thread group is a group of threads executing the same program, although each thread within a thread group can be at different instruction points within that program. In at least one embodiment, a thread group is associated with a unique slice of data in global memory.

[0399] In at least one embodiment, graphics multiprocessor 1834 includes internal cache memory, to perform load and store operations. In at least one embodiment, graphics multiprocessor 1834 can bypass internal cache and use cache memory within processing cluster 1814 (e.g., Ll cache 1848). In at least one embodiment, each graphics multiprocessor 1834 can also have access to L2 Cache within a partition unit (e.g., partition unit 1820A-1820N) that is shared among all processing clusters 1814 and can be used to transfer data between threads. In at least one embodiment, graphics multiprocessor 1834 can also have access to off-chip global memory, which can include one or more of local parallel processor memory and / or system memory. In at least one embodiment, any memory external to parallel processor 1802 can be used as global memory. In at least one embodiment, processing cluster 1814 includes multiple instances of graphics multiprocessor 1834 that share common instructions and data, which can be stored in Ll cache 1848. FIG. 18A

[0400] In at least one embodiment, each processing cluster 1814 can include a memory management unit (MMU) 1845 to map virtual addresses into physical addresses, as is known to those skilled in the art. In at least one embodiment, one or more instances of MMU 1845 can reside within graphics multiprocessor 1834. FIG. 18A ​In at least one embodiment, MMU 1845 includes a set of page table entries (PTEs) for mapping virtual addresses into physical addresses for task computation, as well as optionally into cache line indices for tasks that can be cached. In at least one embodiment, MMU 1845 can include address translation lookaside buffer (TLB) or can reside within graphics multiprocessor 1834 or Ll cache or processing cluster 1814. In at least one embodiment, processing physical addresses enables local access to data in on-chip memory. In at least one embodiment, cache line indices enable caching of tasks in Ll cache memory for faster access.

[0401] In at least one embodiment, processing cluster 1814 can be configured such that each graphics multiprocessor 1834 is coupled to a texture unit 1836 for performing texture mapping operations, e.g., determining texture sample positions, reading texture data, and filtering texture data. In at least one embodiment, texture data is read from an internal texture Ll cache (not shown) or from an Ll cache within graphics multiprocessor 1834 as needed, and texture data is fetched from an L2 cache, local parallel processor memory, or system memory, as needed. In at least one embodiment, each graphics multiprocessor 1834 outputs processed tasks to data crossbar 1840 in order to provide processed task data to another processing cluster 1814 for further processing or to store processed task data in an L2 cache, local parallel processor memory, or system memory via memory crossbar 1816. In at least one embodiment, PreROP 1842 (pre-raster operations unit) is configured to receive data from graphics multiprocessor 1834, direct data to ROP unit, which can be located within partition unit (e.g., partition units 1820A-1820N as described herein) or within graphics multiprocessor 1834, in at least one embodiment. In at least one embodiment, PreROP 1842 unit can perform optimizations for color blending, organize pixel color data, and perform address translation. FIG. 18A

[0402] In at least one embodiment, one or more systems shown in FIG. 18A-18C are used to perform selection of program code optimizations with various algorithms, formulas, and processes, such as those described in conjunction with FIG. 1 to FIG. 2 In at least one embodiment, one or more systems shown in FIG. 18A-18C are used to implement one or more systems and / or processes, such as those described in conjunction with FIG. 1 to FIG. 6B for example, performing selection of program code optimizations in one or more compilers, and / or otherwise performing operations described herein.​

[0403] FIG. 18D A graphics processing unit 1834 according to at least one embodiment is shown. In at least one embodiment, graphics processing unit 1834 is coupled with a pipeline manager 1832 of processing cluster 1814. In at least one embodiment, graphics processing unit 1834 has a thread execution pipeline that includes, without limitation, an instruction cache 1852, an instruction unit 1854, an address mapping unit 1856, a register file 1858, one or more general-purpose GPU (GPGPU) cores 1862, and one or more load / store units 1866. GPGPU cores 1862 and load / store units 1866 are coupled with a shared memory 1872 and shared cache memory 1870 via a memory and cache interconnect 1868.

[0404] In at least one embodiment, instruction cache 1852 receives a stream of instructions 1833 to execute from pipeline manager 1832. In at least one embodiment, instructions are cached in instruction cache 1852 and dispatched for execution by instruction unit 1854. In one embodiment, instruction unit 1854 can dispatch instructions to be executed as groups of threads (e.g., warps), with each thread in the group assigned to a different execution unit within GPGPU cores 1862. In at least one embodiment, instructions can access any of multiple different address spaces (local, shared, or global) via addresses specified in Unified Addressing Space (UAS). In at least one embodiment, address mapping unit 1856 can be used to translate an address in UAS to an address that can be accessed by load / store units 1866.

[0405] In at least one embodiment, register file 1858 provides a set of registers for functional units of graphics processing unit 1834. In at least one embodiment, register file 1858 provides temporary storage for operands of the data paths connected to functional units (e.g., GPGPU cores 1862, load / store units 1866) of graphics processing unit 1834. In at least one embodiment, register file 1858 is partitioned between different thread groups being executed by graphics processing unit 1834, with each thread group being allocated a dedicated portion of register file 1858. In at least one embodiment, register file 1858 is partitioned between different warps being executed by graphics processing unit 1834.

[0406] In at least one embodiment, GPGPU cores 1862 can each include floating point units (FPUs) and / or integer arithmetic logic units (ALUs) that are capable of performing instructions for processing graphics, simulations, and / or other applications. GPGPU cores 1862 can be similar to the GPGPU cores 1652 described herein. In at least one embodiment, first portion of GPGPU cores 1862 includes single precision FPUs and integer ALUs, while second portion of GPGPU cores 1862 includes double precision FPUs. In at least one embodiment, FPUs can implement IEEE 754-2008 standard for floating point arithmetic or enable variable precision floating point arithmetic. In at least one embodiment, GPGPU cores 1862 can additionally include one or more fixed function or special purpose logic units to perform specific functions or

[0407] In at least one embodiment, GPGPU cores 1862 include SIMD logic capable of performing a single instruction on multiple sets of data. In at least one embodiment, GPGPU cores 1862 can physically execute SIMD4, SIMD8, and SIMD16 instructions and logically execute SIMD1, SIMD2, and SIMD32 instructions. In at least one embodiment, the SIMD instructions for GPGPU cores 1862 can be generated at compile time by a shader compiler or automatically generated upon execution of a programmable application with a graphics processing unit capable of a SIMD group size not supported by the application. In at least one embodiment, multiple threads with the same instruction pointer are available to a thread group and feed an instruction queue for a SIMD slice or group of threads. In at least one embodiment, threads in a thread group can be split into a number of thread lanes and multiple threads with different instruction pointers can be assigned to a thread lane. In at least one embodiment, a thread lane is coupled with a corresponding result register in an LOQ and can be mapped to a specific SIMD execution unit for an instruction.

[0408] In at least one embodiment, memory and cache interconnect 1868 is an interconnect network that connects each functional unit of graphics multiprocessor 1834 to register file 1858 and shared memory 1870. In at least one embodiment, memory and cache interconnect 1868 is a crossbar interconnect that allows load / store units 1866 to implement load and store operations between shared memory 1870 and register file 1858. In at least one embodiment, register file 1858 can operate at same frequency as GPGPU cores 1862, such that latency for data transfers between GPGPU cores 1862 and register file 1858 is very low. In at least one embodiment, shared memory 1870 can be used to enable communication between threads executing on functional units within graphics multiprocessor 1834. In at least one embodiment, cache memory 1872 can be used as, e.g., a data cache to cache texture data communicated between functional units and texture unit 1836. In at least one embodiment, shared memory 1870 can also be used as a program managed cache. In at least one embodiment, in addition to auto-cached data stored in cache memory 1872, threads executing on GPGPU cores 1862 can also store data in shared memory in a programmed manner.

[0409] In at least one embodiment, parallel processor or GPGPU as described herein is communicatively coupled to host / processor cores to accelerate graphics operations, machine learning operations, pattern analysis operations, and various general purpose GPU (GPGPU) functions. In at least one embodiment, GPU can be communicatively coupled to host processor / cores by a bus or other interconnect (e.g., a high speed

[0410] In at least one embodiment, FIG. 18D One or more systems shown in FIG. 11 are used to perform selection of program code optimizations with various algorithms, formulas, and processes (e.g., those described in conjunction with FIG. 1 to FIG. 2 In at least one embodiment, FIG. 18D One or more systems shown in FIG. 11 are used to implement one or more systems and / or processes (e.g., those de...

Claims

1. A processor comprising: one or more circuits to execute a compiler to select one or more optimizations to one or more first versions of a program based at least in part on results of performing the one or more optimizations on one or more second versions of the program.

2. The processor of claim 1, wherein, the compiler to perform the one or more optimizations on the one or more first versions of the program.

3. The processor of claim 1, wherein, the one or more optimizations to change at least one representation of at least one of the one or more second versions of the program.

4. The processor of claim 1, wherein, performing the one or more optimizations on the one or more second versions of the program includes performing the one or more optimizations using one or more optimization passes that change one or more intermediate representations of the program.

5. The processor of claim 1, wherein, the compiler to select the one or more optimizations from a plurality of second optimizations performed on the one or more second versions of the program, wherein one or more of the second optimizations that do not change at least one representation of at least one of the second versions of the program are not included in the one or more optimizations.

6. The processor of claim 1, wherein, the results of performing the one or more optimizations on the one or more second versions of the program include an optimization profile that specifies the one or more optimizations, and the compiler to select the one or more optimizations from the optimization profile.

7. The processor of claim 1, wherein, the one or more circuits to further execute the compiler to perform a plurality of second optimizations on the one or more second versions of the program, wherein the one or more optimizations are selected from the plurality of second optimizations based on one or more changes made to one or more intermediate representations of the program by the plurality of second optimizations.

8. The processor of claim 1, wherein, at least one of the one or more second versions of the program is the same as at least one of the one or more first versions of the program.

9. A system comprising: one or more processors to execute a compiler to select one or more optimizations to one or more first versions of a program based at least in part on results of performing the one or more optimizations on one or more second versions of the program.

10. The system of claim 9, wherein, the compiler to perform the one or more optimizations on the one or more first versions of the program.

11. The system of claim 9, wherein, the one or more optimizations to change at least one representation of at least one of the one or more second versions of the program.

12. The system of claim 9, wherein, performing the one or more optimizations on the one or more second versions of the program includes performing the one or more optimizations using one or more optimization passes that change one or more intermediate representations of the program.

13. The system of claim 9, wherein, The compiler selects the one or more optimizations from a plurality of second optimizations performed on the one or more second versions of the program, wherein one or more of the second optimizations that do not change at least one representation of at least one of the second versions of the program are not included in the one or more optimizations.

14. The system of claim 9, wherein, The results of performing the one or more optimizations on the one or more second versions of the program include an optimization profile that specifies the one or more optimizations, and the compiler selects the one or more optimizations from the optimization profile.

15. The system of claim 9, wherein, The one or more processors are further to execute the compiler to perform a plurality of second optimizations on the one or more second versions of the program, wherein the one or more optimizations are selected from the plurality of second optimizations based on one or more changes made to one or more intermediate representations of the program by the plurality of second optimizations.

16. The system of claim 9, wherein, At least one of the second versions of the program is the same as at least one of the first versions of the program.

17. A method comprising: A compiler is executed to select one or more optimizations for one or more first versions of a program based at least in part on results of performing the one or more optimizations on one or more second versions of the program.

18. The method of claim 17, wherein, The compiler performs the one or more optimizations on the one or more first versions of the program.

19. The method of claim 17, wherein, The one or more optimizations change at least one representation of at least one of the one or more second versions of the program.

20. The method of claim 17, wherein, Performing the one or more optimizations on the one or more second versions of the program includes performing the one or more optimizations using one or more optimization passes, wherein the one or more optimization passes change one or more intermediate representations of the program.

21. The method of claim 17, wherein, The compiler selects the one or more optimizations from a plurality of second optimizations performed on the one or more second versions of the program, wherein one or more of the second optimizations that do not change at least one representation of at least one of the second versions of the program are not included in the one or more optimizations.

22. The method of claim 17, wherein, The results of performing the one or more optimizations on the one or more second versions of the program include an optimization profile that specifies the one or more optimizations, and the compiler selects the one or more optimizations from the optimization profile.

23. The method of claim 17, further comprising: A plurality of second optimizations are performed on the one or more second versions of the program, wherein the one or more optimizations are selected from the plurality of second optimizations based on one or more changes made to one or more intermediate representations of the program by the plurality of second optimizations.

24. The method of claim 17, wherein, A second version of the program is the same as a first version of the program.