Application programming interface for configuring processor

By introducing processor setting profiles in the data center and dynamically adjusting processor performance preferences and priorities based on job characteristics, the problem of low computing resource utilization in existing technologies is solved, and more efficient computing resource management and job execution are achieved.

CN120803625APending Publication Date: 2025-10-17NVIDIA CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510449938.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-04-10
Filing Date
2025-04-10
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

In the prior art, job schedulers in data centers fail to effectively utilize computing resources when scheduling jobs, resulting in waste of computing resources and low efficiency.

Method used

By introducing processor setting profiles, the performance preferences and priorities of processors are dynamically adjusted based on the characteristic information of jobs to optimize the allocation and utilization of computing resources.

Benefits of technology

It improves the computing resource utilization of the data center, optimizes job execution time, and achieves more efficient computing resource management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120803625A_ABST
    Figure CN120803625A_ABST
Patent Text Reader

Abstract

The invention relates to an application programming interface for configuring a processor. Apparatus, systems, and techniques that execute an application programming interface (API) to identify processor settings to be used in executing one or more software workloads. For example, one or more processors including one or more circuits execute an API to identify processor settings to be used to configure a processor assigned to execute a software workload based at least in part on one or more characteristics of the software workload.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross Reference to Related Applications

[0002] The disclosure of the following applications is incorporated by reference herein in its entirety: co-pending U.S. Patent Application No. 18 / 632,247 , entitled "APPLICATION PROGRAMMING INTERFACE TO INDICATE A COMPUTING RESOURCE"; co-pending U.S. Patent Application No. 18 / 632,255 , entitled "APPLICATION PROGRAMMING INTERFACE TO INDICATE A PRIORITY"; U.S. Patent Application No. 18 / 632,260 , entitled "APPLICATION PROGRAMMING INTERFACE TO IDENTIFY PROCESSOR SETTINGS", filed concurrently herewith; co-pending U.S. Patent Application No. 18 / 632,267 , entitled "APPLICATION PROGRAMMING INTERFACE TO CONFIGURE A PROCESSOR USING COMPUTING RESOURCE INPUTS", filed concurrently herewith; U.S. Patent Application No. 18 / 632,270 , entitled "APPLICATION PROGRAMMING INTERFACE TO CONFIGURE A PROCESSOR USING PRIORITY", filed concurrently herewith; co-pending U.S. Patent Application No. 18 / 632,274 , entitled "APPLICATION PROGRAMMING INTERFACE TO IDENTIFY SETTINGS TO CONFIGURE A PROCESSOR", and U.S. Patent Application No. 18 / 632,279 , entitled "APPLICATION PROGRAMMING INTERFACE TO PERFORM INSTRUCTIONS USING PROCESSOR SETTINGS". TECHNICAL FIELD

[0003] At least one embodiment relates to processing resources for identifying processor settings. At least one embodiment relates to a processor or computing system for identifying processor settings based at least in part on characteristics of a job. BACKGROUND

[0004] A data center can include software for scheduling jobs performed by processors in the data center. For example, a job scheduler can schedule jobs to be launched according to a priority of each job, but this does not necessarily always result in efficient utilization of computing resources. As part of the job scheduling process, the amount of computing resources and time used to execute a job can be improved. BRIEF DESCRIPTION OF DRAWINGS

[0005] Figure 1 A block diagram of a system for configuring a processor using processor settings based at least in part on one or more job characteristics is shown in accordance with at least one embodiment;

[0006] Figure 2 A block diagram of a system for configuring a processor using processor settings based at least in part on one or more job characteristics is shown in accordance with at least one embodiment;

[0007] Figure 3 A block diagram of a system including a data structure for identifying processor settings for a processor to use to execute a job based at least in part on one or more job characteristics is shown in accordance with at least one embodiment;

[0008] Figure 4 A block diagram of a system including a data structure for identifying processor settings for a processor to use to execute a job based at least in part on one or more job characteristics is shown in accordance with at least one embodiment;

[0009] Figure 5 A block diagram of a system including a job scheduler and a processor settings configuration file adjusted by a bias value is shown in accordance with at least one embodiment;

[0010] Figure 6 A block diagram of a process for configuring a processor based at least in part on one or more job characteristics is shown in accordance with at least one embodiment;

[0011] Figure 7 A block diagram of a system including a driver and / or runtime for identifying processor settings based at least in part on one or more job characteristics is shown in accordance with at least one embodiment;

[0012] Figure 8An invocation flow diagram of a system for identifying processor settings based at least in part on one or more job characteristics, in accordance with at least one embodiment, is shown;

[0013] Figure 8A A diagram of a schedule job API invocation, in accordance with at least one embodiment, is shown;

[0014] Figure 8B A diagram of a get all available configuration file API invocation, in accordance with at least one embodiment, is shown;

[0015] Figure 8C A diagram of a schedule set specific configuration file on processor API invocation is shown.

[0016] Figure 9 A block diagram of a system for identifying power policies for a plurality of nodes, in accordance with at least one embodiment, is shown;

[0017] Figure 10 A block diagram of a system for identifying processor setting configuration files estimated to meet a per watt target performance value, in accordance with at least one embodiment, is shown;

[0018] Figure 11 A block diagram of a system for identifying processor setting configuration files estimated to meet a per watt target performance value, in accordance with at least one embodiment, is shown;

[0019] Figure 12 A block diagram of a system for providing power information to a scheduler, in accordance with at least one embodiment, is shown;

[0020] Figure 13 A block diagram of a system for identifying power policies to apply to one or more nodes shared by a plurality of users, in accordance with at least one embodiment, is shown;

[0021] Figure 14 A block diagram of a system for identifying power policies to apply to a node based on whether the node can be partitioned, in accordance with at least one embodiment, is shown;

[0022] Figure 15 A block diagram of a system for identifying power policies to apply to a plurality of nodes, in accordance with at least one embodiment, is shown;

[0023] Figure 16 A block diagram of a system for identifying power policies to apply to one or more nodes based at least in part on a number of graphics processing units (GPUs) required to perform a job, in accordance with at least one embodiment, is shown;

[0024] Figure 17A block diagram illustrating a system for identifying a power policy to apply to one or more nodes based at least in part on a number of GPUs required to perform a job is shown, in accordance with at least one embodiment;

[0025] Figure 18 A block diagram illustrating a system for identifying a processor settings profile using user input in a single exclusive access scenario is shown, in accordance with at least one embodiment;

[0026] Figure 19 A block diagram illustrating a system for identifying a processor settings profile based on a system profiling a job is shown, in accordance with at least one embodiment;

[0027] Figure 20 A block diagram illustrating a system for indicating which jobs to connect to a bus is shown, in accordance with at least one embodiment;

[0028] Figure 21 A block diagram illustrating a system for indicating how a job is to be performed with respect to a rate of change of current over time is shown, in accordance with at least one embodiment;

[0029] Figure 22 A block diagram illustrating a system for identifying a power profile and GPU allocation is shown, in accordance with at least one embodiment;

[0030] Figure 23 A block diagram illustrating a system for allocating GPUs based at least in part on power telemetry of a plurality of GPUs is shown, in accordance with at least one embodiment;

[0031] Figure 24 A block diagram illustrating a system for identifying a processor settings profile and allocating GPUs is shown, in accordance with at least one embodiment;

[0032] Figure 25 A block diagram illustrating a system for identifying a processor settings profile and allocating GPUs using GPU telemetry data is shown, in accordance with at least one embodiment;

[0033] Figure 26 A block diagram illustrating a system for identifying a processor settings profile based at least in part on GPUs shared between jobs is shown, in accordance with at least one embodiment;

[0034] Figure 27 A block diagram illustrating a system for identifying a processor settings profile based at least in part on partitioned nodes is shown, in accordance with at least one embodiment;

[0035] Figure 28 A block diagram illustrating a system for identifying a processor settings profile based at least in part on partitioned nodes is shown, in accordance with at least one embodiment;

[0036] Figure 29 A block diagram illustrating a system for identifying a processor settings configuration file based, at least in part, on job characteristics, in accordance with at least one embodiment is shown;

[0037] Figure 30 A block diagram illustrating a system for identifying a processor settings configuration file based on system profiling of a job, in accordance with at least one embodiment is shown;

[0038] Figure 31 A block diagram illustrating a system for causing a scheduler to receive power information for a plurality of GPUs, in accordance with at least one embodiment is shown;

[0039] Figure 32 An exemplary data center, in accordance with at least one embodiment is shown;

[0040] Figure 33 A processing system, in accordance with at least one embodiment is shown;

[0041] Figure 34 A computer system, in accordance with at least one embodiment is shown;

[0042] Figure 35 A system, in accordance with at least one embodiment is shown;

[0043] Figure 36 An exemplary integrated circuit, in accordance with at least one embodiment is shown;

[0044] Figure 37 A computing system, in accordance with at least one embodiment is shown;

[0045] Figure 38 An APU, in accordance with at least one embodiment is shown;

[0046] Figure 39 A CPU, in accordance with at least one embodiment is shown;

[0047] Figure 40 An exemplary accelerator integration slice, in accordance with at least one embodiment is shown;

[0048] Figure 41A and Figure 41B An exemplary graphics processor, in accordance with at least one embodiment is shown;

[0049] Figure 42A A graphics core, in accordance with at least one embodiment is shown;

[0050] Figure 42B A GPGPU, in accordance with at least one embodiment is shown;

[0051] Figure 43A A parallel processor, in accordance with at least one embodiment is shown;

[0052] Figure 43B A processing cluster is shown in accordance with at least one embodiment;

[0053] Figure 43C A graphics multiprocessor is shown in accordance with at least one embodiment;

[0054] Figure 44 A graphics processor is shown in accordance with at least one embodiment;

[0055] Figure 45 A processor is shown in accordance with at least one embodiment;

[0056] Figure 46 A processor is shown in accordance with at least one embodiment;

[0057] Figure 47 A graphics processor core is shown in accordance with at least one embodiment;

[0058] Figure 48 A PPU is shown in accordance with at least one embodiment;

[0059] Figure 49 A GPC is shown in accordance with at least one embodiment;

[0060] Figure 50 A streaming multiprocessor is shown in accordance with at least one embodiment;

[0061] Figure 51 A software stack of a programming platform is shown in accordance with at least one embodiment;

[0062] Figure 52 A CUDA implementation of the software stack of Figure 51 is shown in accordance with at least one embodiment;

[0063] Figure 53 A ROCm implementation of the software stack of Figure 51 is shown in accordance with at least one embodiment;

[0064] Figure 54 An OpenCL implementation of the software stack of Figure 51 is shown in accordance with at least one embodiment;

[0065] Figure 55 Software supported by a programming platform is shown in accordance with at least one embodiment;

[0066] Figure 56 Compiled code executed on a programming platform of Figures 51-54 is shown in accordance with at least one embodiment;

[0067] Figure 57 A programming platform is shown in accordance with at least one embodiment;Figures 51-54 More detailed compiled code executed on the programming platform;

[0068] Figure 58 Transforming source code before compiling it according to at least one embodiment is shown;

[0069] Figure 59A A system configured to compile and execute CUDA source code using different types of processing units is shown in accordance with at least one embodiment;

[0070] Figure 59B A method configured to compile and execute a program using a CPU and a CUDA-enabled GPU according to at least one embodiment is shown. Figure 59A CUDA source code system;

[0071] Figure 59C A method configured to compile and execute using a CPU and a non-CUDA enabled GPU according to at least one embodiment is shown. Figure 59A CUDA source code system;

[0072] Figure 60 According to at least one embodiment, Figure 59C An example kernel converted by the CUDA to HIP conversion tool;

[0073] Figure 61 More details are shown according to at least one embodiment. Figure 59C a non-CUDA-enabled GPU; and

[0074] Figure 62 shows how threads of an exemplary CUDA grid are mapped to Figure 61 Different computing units;

[0075] Figure 63 shows how to migrate existing CUDA code to data parallel C++ code according to at least one embodiment; and

[0076] Figure 64 Components of a system for accessing large language models in accordance with at least one embodiment are shown. DETAILED DESCRIPTION

[0077] In the following description, numerous specific details are set forth to provide a more thorough understanding of at least one embodiment. However, it will be understood by those skilled in the art that the inventive concept can be implemented without one or more of these specific details, and that any two or more aspects of any one or more embodiments described herein can be combined.

[0078] In at least one embodiment, a processor performs operations of a workload scheduler (e.g., a job scheduler) of a data center that allows a user to provide information about a software workload to the workload scheduler. In at least one embodiment, information provided by a user to the workload scheduler includes an indication of a processor performance preference, a type of software workload to be scheduled, a priority of a software workload to be scheduled, or some combination thereof. In at least one embodiment, a processor performs operations of a workload scheduler to manage when and how a software workload is to be executed by a data center or other processors of any facility including computing devices (e.g., computers, servers, processors) and networking devices (e.g., routers, switches). In at least one embodiment, a processor performs operations of a workload scheduler to manage when and how a software workload is to be executed by other processors is referred to as scheduling. In at least one embodiment, a processor performs operations of a workload scheduler to instruct processor settings for other processors to use when executing a software workload as part of a workload scheduling process. In at least one embodiment, a processor performs operations of a workload scheduler to instruct processor settings for other processors to use when executing a particular software workload such that a processor management application of a data center is able to set those processor settings on other processors before executing that particular software workload. In at least one embodiment, a processor performs operations of a workload scheduler to cause a processor management application to check whether processor settings for executing a software workload have been set before executing that software workload. In at least one embodiment, a processor performs operations of a workload scheduler to cause a processor management application to adjust processor settings of other processors executing a software workload using processor performance metrics observed during execution of that software workload.

[0079] In at least one embodiment, a processor performs operations of a processor management application to receive or otherwise obtain information about a software workload from a workload scheduler in order to set processor settings before other processors execute that software workload. In at least one embodiment, a processor performs operations of a processor management application to use information about a software workload to identify a combination of processor settings from a data structure (e.g., a data table, a lookup table) that will cause other processors of a data center to execute that software workload in accordance with a user’s performance preference, in accordance with a priority of that software workload, within processor performance constraints, within data center constraints, or some combination thereof, as further described herein. In at least one embodiment, a combination of processor settings is referred to as a processor settings profile or a processor profile.

[0080] In at least one embodiment, a processor executes an application programming interface (API) function to cause one or more processors to be configured to run at one or more clock frequencies based at least in part on one or more inputs to the API. In at least one embodiment, an API function is referred to as an API. In at least one embodiment, a processor executes an API to indicate one or more computing resources to be used by one or more instructions based at least in part on one or more inputs to the API. In at least one embodiment, a processor executes an API to indicate a priority for executing one or more instructions based at least in part on one or more inputs to the API. In at least one embodiment, a processor executes an API to identify one or more settings to be used to configure one or more processors to run at one or more processor clock frequencies based at least in part on one or more clock frequency inputs to the API. In at least one embodiment, a processor executes an API to identify one or more settings to be used to configure one or more processors to run at one or more processor clock frequencies based at least in part on one or more priority inputs to the API. In at least one embodiment, a processor executes an API to identify one or more settings to be used to configure one or more processors to run at one or more processor clock frequencies based at least in part on one or more processors to be used. In at least one embodiment, a processor executes an API to cause one or more instructions to be executed based at least in part on one or more processor settings inputs to the API.

[0081] Figure 1 A block diagram of a system 100 including one or more processors including one or more circuits to receive or otherwise obtain information from a user regarding a software workload and identify a processor setting configuration file for executing the software workload is shown. In at least one embodiment, one or more aspects of one or more embodiments described herein incorporate at least those aspects described in connection with Figure 1 one or more aspects of one or more embodiments described herein include at least those aspects described in connection with Figures 2-31 In at least one embodiment, one or more processors perform one or more operations of system 100. In at least one embodiment, one or more processors performing one or more operations of system 100 are any one or combination of processors described herein, including processor 108, Figure 2 processors of processor complex 208, Figure 3 processors of processor complex 308, Figure 4 processors of processor complex 408, processors of processor complex 508, Figure 38APU 3800, in conjunction with Figure 41A CPU 4100, in conjunction with Figure 43A graphics processor 4310, in conjunction with Figure 50 parallel processing unit (“PPU”) 5000, or Figure 49 one or more SMs 5114. In at least one embodiment, processor 108 performs operations used by system 100, such as operations of processor configuration file module 104. In at least one embodiment, processor 108 performs one or more operations described in conjunction with Figure 2 one or more operations described in conjunction with Figure 3 one or more operations described in conjunction with Figure 4 one or more operations described in conjunction with Figure 5 one or more operations described in conjunction with Figure 6 one or more operations described in conjunction with Figure 7 one or more operations described in conjunction with Figure 8 one or more operations described in conjunction with Figures 9-31 one or more operations described in conjunction with

[0082] In at least one embodiment, as used in any implementation described herein, the terms “system,” “device,” “component,” and “module,” as well as any variations thereof, are intended to refer to any combination of software, firmware, hardware, and / or circuitry configured to provide the functionality described herein, unless otherwise indicated. In at least one embodiment, software, firmware, hardware, and / or circuitry configured to provide the functionality described herein comprises one or more components implemented collectively or individually as part of a larger system, such as an integrated circuit (IC), system on a chip (SoC), and / or the like. In at least one embodiment, any of the circuitry of one or more modules can be represented as a register transfer level (RTL) representation and / or another representation of another abstraction level, which can be authorized and / or used for a tape-out, which is a final stage of an IC design before it is sent to a fabrication facility for fabrication of ICs.

[0083] In at least one embodiment, system 100 is any computing system including one or more data centers or other facilities housing computing and networking equipment. In at least one embodiment, system 100 is used to perform high performance computing tasks, neural network training, neural network inference, or some combination thereof. In at least one embodiment, system 100 includes an edge computing system, an accelerated computing system, a cloud computing system, a hybrid cloud computing system, or some combination thereof. In at least one embodiment, system 100 is a computing system including multiple distributed components connected by a network, such as the Internet network. In at least one embodiment, system 100 is used in fields such as healthcare, genomics, engineering, aerospace, urban planning, graphics processing, finance, data storage and management, online commerce, meteorology, physical modeling, or some combination thereof. In at least one embodiment, system 100 is used to perform artificial intelligence (AI) tasks such as image classification, image segmentation, autonomous driving, manufacturing defect identification, or some combination thereof. In at least one embodiment, a neural network is a type of AI.

[0084] In at least one embodiment, the system 100 includes a user interface 102 through which a user provides input that provides information about one or more software workloads. In at least one embodiment, the software workloads are referred to as jobs, a term used further herein. In at least one embodiment, the user interface 102 is a user interface for a job scheduler 110. In at least one embodiment, at least a portion of the job scheduler 110 is implemented on a computing device operating the user interface 102. In at least one embodiment, the user interface 102 is a user interface for a processor management application. In at least one embodiment, the processor management application is any combination of hardware, firmware, or software (e.g., a data center processor management module 120 described further herein). In at least one embodiment, at least a portion of the data center processor management module 120 is implemented on a computing device operating the user interface 102.

[0085] In at least one embodiment, user interface 102 is communicatively connected to network 104. In at least one embodiment, network 104 can be one or more of any type of network, such as a hosted network (e.g., an enterprise network), a cloud network, the Internet, a local private network, or some combination thereof. In an embodiment, network 104 is a local network. In at least one embodiment, network 104 is communicatively connected to one or more components of data center 106.

[0086] In at least one embodiment, the system 100 includes a data center 106. In at least one embodiment, the data center 106 is one or more data centers. In at least one embodiment, the data center 106 is a combination of at least Figure 34 At least a portion of data center 3400 is described. In at least one embodiment, a data center is any facility that houses computers and network equipment. In at least one embodiment, the data center includes processors that execute operations in parallel to process massive data sets across multiple dimensions. In at least one embodiment, the data center performs one or more AI tasks. In at least one embodiment, a user remotely accesses at least a portion of the computing resources of data center 106 via network 104 to schedule and execute jobs.

[0087] In at least one embodiment, system 100 includes a processor 108, which is any one or combination of processors described herein, including processor group 508, Figure 38 APU 3800, combined with Figure 41A Description of the CPU 4100, combined with Figure 43A The graphics processor 4310 described, and the combination Figure 50In at least one embodiment, any processor described herein, including processor 108, includes one or more circuits. In at least one embodiment, processor 108 is one or more processors implemented in a computing system designed to perform AI tasks, such as image classification, autonomous driving, or some combination thereof. In at least one embodiment, processor 108 is one or more processors implemented in an edge computing device, a workstation, a server, or some combination thereof, such as DGX TM In at least one embodiment, the processor 108 is one or more Epics TM Embedded processor and / or one or more A100 TM GPU.In at least one embodiment, processor 108 is one or more different types of processors implemented as part of a heterogeneous computing device.

[0088] In at least one embodiment, the processor 108 is a group of processors. In at least one embodiment, two or more processors 108 are installed in different locations, such as two different data centers connected by network communication. In at least one embodiment, the processor 108 is one or more graphics processing units (GPUs) in a group of GPUs. In at least one embodiment, a group of GPUs is referred to as a GPU cluster. In at least one embodiment, the processor 108 is one or more portions of one or more GPUs, wherein each portion includes a portion of GPU memory and a portion of GPU computing hardware that are configured to operate as an independent, separate, and complete GPU. In at least one embodiment, a portion of the GPU computing hardware is a portion of a streaming multiprocessor (SM) of a GPU, such as Figure 51 SM 5114. In at least one embodiment, processor 108 is part of one or more GPUs, referred to as partitions. In at least one embodiment, processor 108 is a GPU partition system (e.g., A portion of one or more GPUs in a Multi-Instance GPU (MIG) configuration.

[0089] In at least one embodiment, system 100 includes a job scheduler 110. In at least one embodiment, job scheduler 110 is implemented on processor 108. In at least one embodiment, processor 108 performs one or more operations of job scheduler 110. In at least one embodiment, any description of a scheduler or module performing an operation refers to a processor performing that scheduler or module to perform that operation. In at least one embodiment, job scheduler is referred to as a software workload scheduler, a scheduling software application, or a scheduler. In at least one embodiment, job scheduler 110 is a job scheduler 3432 in Figure 34 In at least one embodiment, job scheduler 110 is any combination of hardware, firmware, or software for managing when and how one or more processors perform jobs. In at least one embodiment, a job is any software workload, software instruction, or set of software instructions identifiable as a unit of work to be performed by one or more processors. In at least one embodiment, a software instruction is referred to as an instruction. In at least one embodiment, job scheduler 110 is at least part of a compute management system, such as Grid Engine, Scheduler, Spectrum LSF, or some combination thereof. In at least one embodiment, job scheduler 110 is at least part of a distributed resource management (DRM) system. In at least one embodiment, a job is any computing workload defined by a user or an application. In at least one embodiment, a job is referred to as a set of one or more tasks, processes, or operations. In at least one embodiment, a job is a kernel, which is a set of instructions executed in parallel by one or more processors. In at least one embodiment, a job is a container, which is a set of instructions that can be executed by one or more processors in different computing environments using different hardware, firmware, software, or some combination thereof. In at least one embodiment, different types of jobs include jobs related to physical modeling, image classification, cloud-based document management, web hosting, or some combination thereof.

[0090] In at least one embodiment, system 100 includes a job scheduler database 112, which is one or more data storage devices that store information about jobs, such as job IDs, processor performance preferences, job types, job priorities, specific processors assigned to perform specific jobs, or some combination thereof, as further described herein, including in connection with Figure 2 In at least one embodiment, job scheduler database 112 is implemented as part of job scheduler 110. In at least one embodiment, job scheduler database 112 is updated using information received from user interface 102.

[0091] In at least one embodiment, system 100 includes job information 114. In at least one embodiment, information about jobs stored on job scheduler database 112 is job information 114. In at least one embodiment, job information 114 includes one or more indications of information about jobs. In at least one embodiment, information about jobs is referred to as one or more characteristics of the job. In at least one embodiment, job information 114 includes an indication of a particular job to be scheduled, e.g., a job ID. In at least one embodiment, job information 114 includes an indication of a particular job that has been scheduled. In at least one embodiment, job information 114 includes an indication of a job type, e.g., compute-bound or memory-bound. In at least one embodiment, a compute-bound job is referred to as a compute-intensive, tensor-core-intensive, math-bound, or arithmetic-intensive job. In at least one embodiment, a memory-bound job is referred to as a memory-intensive job. In at least one embodiment, a memory transfer rate and / or a memory amount available on a processor is an operational specification of the processor that impacts whether a particular job type should be executed by that processor. In at least one embodiment, a job type describes a number and type of mathematical operations to be performed as part of the job, a number and type of data formats to be used during execution of the job, or some combination thereof. In at least one embodiment, a job type describes a number and type of memory transfers required to execute a job. In at least one embodiment, job information 114 includes one or more indications of a job priority, at least in combination with Figures 10-14 Further described herein. In at least one embodiment, job information 114 includes one or more indications of a particular processor assigned by job scheduler 110 to execute a particular job (e.g., a GPU handle). In at least one embodiment, a processor assigned by a job scheduler to execute a job is referred to as a job scheduler-allocated processor to execute a job.

[0092] In at least one embodiment, system 100 includes a predetermined job 116. In at least one embodiment, predetermined job 116 is any combination of hardware, firmware, or software implemented as part of job scheduler 110. In at least one embodiment, predetermined job 116 includes a storage device that stores an indication of a job that has been scheduled to be executed according to one or more factors, such as latency, triggers, available computing resources, or some combination thereof. In at least one embodiment, job scheduler 110 schedules jobs on a first-in, first-out (FIFO) basis. In at least one embodiment, predetermined job 116 is a job queue. In at least one embodiment, predetermined job 116 includes an information indication regarding a job described herein. In at least one embodiment, predetermined job 116 includes an indication of a processor configuration file for executing a particular job described further herein.

[0093] In at least one embodiment, system 100 includes a data center processor management module 120. In at least one embodiment, data center processor management module 120 is hardware, firmware, software, or some combination thereof, for setting processor setting values for processors of a data center, such as processors 108. In at least one embodiment, a processor setting is a value used by data center processor management module 120 to configure one or more processors to run at or within a range of processor setting values. In at least one embodiment, a range of processor setting values is calculated based on a percentage of a processor setting value. In at least one embodiment, data center processor management module 120 configures one or more processors to run at or within a range of processor setting values by managing or modifying how a processor inputs or executes instructions; by causing a device (e.g., a microcontroller, a voltage regulator module, a switch) to control power consumption, fan speed; by causing a particular circuit or circuit portion of a processor to be used; by physically modifying some aspect of a processor (e.g., modifying a logic component); by using techniques known to those of ordinary skill in the art; or some combination thereof.

[0094] In at least one embodiment, processor settings values are referred to as processor settings. In at least one embodiment, data center processor management module 120 is referred to as a compute resource manager, resource manager (RM), or processor management application. In at least one embodiment, data center processor management module 120 includes any combination of hardware, firmware, or software that manages communication between components of a data center, such as communication between a job scheduler and a processor, using a communication protocol. In at least one embodiment, one or more portions of data center processor management module 120 that manage communication between components of a data center are implemented as separate modules. In at least one embodiment, one or more portions of data center processor management module 120 are implemented on a computing network, computing facility, node, or some combination thereof that is separate from another computing network, computing facility, node, or some combination thereof on which another portion of data center processor management module 120 is implemented. In at least one embodiment, a portion of a module that is implemented separately from another portion of that module or other modules is referred to as being out-of-band, remote, or distributed.

[0095] In at least one embodiment, data center processor management module 120 includes In at least one embodiment, data center GPU manager (DCGM) includes one or more API functions of this system. In at least one embodiment, at least a portion of data center processor management module 120 manages processor settings and / or processor configuration at a low level. In at least one embodiment, low level management of a processor refers to management that includes commands and / or instructions sent to and usable by a processor driver. In at least one embodiment, a portion of data center processor management module 120 that performs low level management of a processor is implemented as a separate module. In at least one embodiment, at least a portion of data center processor management module 120 includes an interface (e.g., a user interface) and an API library that a user or application (e.g., a job scheduler) can use with a portion of data center processor management module 120 that performs low level management of a processor. In at least one embodiment, any one or more portions of data center processor management module 120 that perform low level management of a processor, including an interface that performs low level management of a processor, an API function that performs low level management of a processor, or some combination thereof, are referred to as a resource manager system management interface (RMSMI). In at least one embodiment, a RMSMI is included in multiple embodiments described herein, including in connection with Figures 10-14 In at least one embodiment, a portion of a RMSMI is One or more portions of a System Management Interface (SMI) system, including one or more API functions of the system. In at least one embodiment, a portion of an RMSMI is located on one or more portions of a processor management library (e.g., a ROCm SMI library or management library (NVML)).

[0096] In at least one embodiment, at least a portion of data center processor management module 120 is a baseboard management controller (BMC) that at least partially monitors and controls processors of a computing system. In at least one embodiment, at least a portion of data center processor management module 120 is an interface (e.g., a user interface) and API library that a user or application (e.g., a job scheduler) can use with a portion of data center processor management module 120 that performs baseboard management. Any one or more portions of data center processor management module 120 that perform baseboard management, including an interface that performs baseboard management, including an API function that performs baseboard management, or some combination thereof, is referred to as a resource manager baseboard management interface (RMBMCI). In at least one embodiment, a RMBMCI is included in multiple embodiments described herein, including at least embodiments described in connection with Figure 2 In at least one embodiment, a portion of a RMBMCI is one or more portions of a baseboard management controller (BMC), including one or more API functions of the system, or similar functions.

[0097] In at least one embodiment, at least a portion of data center processor management module 120 is one or more processor drivers, such as a GPU driver. In at least one embodiment, a processor driver is a driver such as driver 226 of Figure 7 and driver 704 of Figure 14 In at least one embodiment, a processor executes a processor driver to configure the processor and / or another processor according to a selected and / or identified processor setting otherwise described herein. In at least one embodiment, a processor driver of data center processor management module 120 is referred to as a resource manager driver (RM driver). In at least one embodiment, a RMSMI is included in multiple embodiments described herein, including at least embodiments described in connection with Figure 3 In at least one embodiment, a portion of a RMBMCI is one or more portions of a baseboard management controller (BMC), including one or more API functions of the system, or similar functions.

[0098] ​In at least one embodiment, data center processor management module 120 is implemented on a processor of one or more processors 108 that is different from a processor of one or more processors 108 on which job scheduler 110 is implemented. In at least one embodiment, data center processor management module 120 is implemented on a different computing device (e.g., server) than a computing device on which job scheduler 110 is implemented. In at least one embodiment, data center processor management module 120 is implemented in a different data center than a data center in which job scheduler 110 is implemented.

[0099] In at least one embodiment, system 100 includes job priority processor configuration file API module 122. In at least one embodiment, job priority processor configuration file API module 122 is one or more API functions for receiving information about a job to be scheduled, as described further herein. In at least one embodiment, job priority processor configuration file API module 122 is one or more API functions for identifying a processor configuration file to be used to execute a job based on information about the job, including processor performance preferences, job type, job priority, operating specifications of a particular processor, or some combination thereof, as described further herein. In at least one embodiment, job priority processor configuration file API module 122 is one or more API functions for ensuring that a processor configuration file identified by other API functions is set on a processor prior to execution of a particular job, as described further herein.

[0100] In at least one embodiment, API of job priority processor configuration file API module 122 identifies a processor configuration file based at least in part on a job priority and constraints provided by a user or an application. In at least one embodiment, constraints include any parameter, metric, measurement, specification, value, or some combination thereof to be followed and / or satisfied when executing a job on a processor group. In at least one embodiment, a constraint is a value of a performance metric that a portion of a processor, an entire processor, or a data center is not to exceed during execution of a job, such as a performance metric like Fmax, maxTGP, Vmax, or some combination thereof. In at least one embodiment, a constraint is a minimum value of a performance metric that a portion of a processor, an entire processor, or a data center is to meet or exceed during execution of a job, such as a performance metric like a minimum clock frequency or a minimum power consumption. In at least one embodiment, a constraint includes a hardware constraint, such as one or more types of processors to be used to execute a job.

[0101] In at least one embodiment, job priority is an indication of how urgent a job should be executed by one or more processors. In at least one embodiment, job scheduler computes job priority of a job based on factors such as job type, amount of time a job has been waiting in a queue, user’s history of using computing resources, available computing resources, or some combination of these factors. In at least one embodiment, job scheduler computes job priority based on estimated time required to complete execution of a job. In at least one embodiment, job scheduler computes job priority based on estimated amount of power required to complete execution of a job. In at least one embodiment, job scheduler computes job priority based on estimated power consumption required to complete execution of a job. In at least one embodiment, a processor executes operations of job scheduler 110 to assign a background job whose job priority is lower than job priority of a more urgent job. In at least one embodiment, a lower priority background job is a job that performs nightly updates of a database, while a higher priority job is a job that performs AI-assisted medical segmentation in medical images to help detect cancer cells. In at least one embodiment, job priority is any value suitable to indicate job priority, such as a numerical value, a string of letters and numbers, or a word such as low, medium, or high.

[0102] In at least one embodiment, a processor in processor 108 executes one or more API functions of job priority processor configuration file API module 122 to cause other processors in processor 108 to execute a job according to information about the job, including processor performance preferences, job type, job priority, and type of processor assigned to execute the job. In at least one embodiment, processor settings of a processor configuration file are any parameters that can be modified to affect processor performance, such as maximum operating frequency (Fmax or Fmax cap), maximum total graphics power (maximum TGP), clock frequency ratio between two devices connected to a crossbar (Xbar ratio), memory clock frequency (MCLK), maximum voltage allowed to be consumed (Vmax), fan speed, or some combination thereof. In at least one embodiment, Xbar ratio is referred to as Cbar ratio. In at least one embodiment, Xbar ratio is a ratio between a graphics processing cluster clock and a crossbar clock. In at least one embodiment, Xbar ratio is a ratio between two crossbar clocks.

[0103] In at least one embodiment, one or more API functions of job priority processor configuration file API module 122 are executed by a processor of processors 108 to cause an indication of a processor configuration file to be used when executing a job to be sent to job scheduler 110. In at least one embodiment, the indication of the processor configuration file is stored in scheduled job 116 to correspond to a particular job to be executed by a processor 108 of data center 106.

[0104] In at least one embodiment, system 100 includes processor configuration file database 124. In at least one embodiment, processor configuration file database 124 is any combination of hardware, firmware, or software for storing processor configuration files or indications of such processor configuration files. In at least one embodiment, processor configuration file database 124 is one or more data structures (e.g., tables, trees) that associate processor configuration files with information about jobs that the processor configuration files are to be used for, which is described further herein. In at least one embodiment, processor configuration file database 124 includes a lookup table that includes processor configuration files, functions, and biases, such as configuration file and bias interaction lookup table 306 (lookup table 306) of Figure 2

[0105] Figure 2 A block diagram of system 200 is shown in at least one embodiment, which includes a job scheduler of a data center to schedule jobs for processors to execute according to processor setup configuration files based at least in part on job information. In at least one embodiment, one or more aspects of one or more embodiments described in connection with Figure 1 include at least in connection with aspects described in connection with Figures 3-31 and Figure 1 In at least one embodiment, one or more processors perform one or more operations of system 200. In at least one embodiment, the one or more processors that perform one or more operations of system 200 are any one or combination of processors described herein, including processors 108 of Figure 3 processor group 208 of Figure 4 processors 308 of Figure 38 processors 408, processor group 508 of Figure 41A APU 3800 of Figure 43A CPU 4100 described in connection with Figure 50 graphics processor 4310 described in connection with Figure 49 PPU 5000 described in connection with Figure 3 ​one or more SMs 5114. In at least one embodiment, one or more processors of processor complex 208 perform one or more operations described in conjunction with, for example, generating a new processor configuration file 330. In at least one embodiment, one or more processors of processor complex 208 perform one or more operations described in conjunction with, for example, generating a new processor configuration file 330. Figure 4 one or more operations described in conjunction with, for example, selecting a processor configuration file 427 using job priority. In at least one embodiment, one or more processors of processor complex 208 perform one or more operations described in conjunction with, for example, selecting a processor configuration file 427 using job priority. Figure 5 one or more operations described in conjunction with, for example, selecting a processor configuration file 427 using job priority. In at least one embodiment, one or more processors of processor complex 208 perform one or more operations described in conjunction with, for example, selecting a processor configuration file 427 using job priority. Figure 6 one or more operations described in conjunction with, for example, job scheduler 510 receiving job priority from user processor complex 508. In at least one embodiment, one or more processors of processor complex 208 perform one or more operations described in conjunction with, for example, job scheduler 510 receiving job priority from user processor complex 508. Figure 7 one or more operations described in conjunction with, for example, accessing processor configuration files stored in a data structure using operation 604. In at least one embodiment, one or more processors of processor complex 208 perform one or more operations described in conjunction with, for example, accessing processor configuration files stored in a data structure using operation 604. Figure 8 one or more operations of API 710 described in conjunction with, for example, API 710. In at least one embodiment, one or more processors of processor complex 208 perform one or more operations of API 710 described in conjunction with, for example, API 710. Figures 9-31 one or more operations described in conjunction with, for example, identifying best processor settings from a database. In at least one embodiment, one or more processors of processor complex 208 perform one or more operations described in conjunction with, for example, identifying best processor settings from a database. Figure 1 one or more operations described in conjunction with, for example, identifying best processor settings from a database. In at least one embodiment, one or more processors of processor complex 208 perform one or more operations described in conjunction with, for example, identifying best processor settings from a database.

[0106] In at least one embodiment, system 200 includes data center 206. In at least one embodiment, data center 206 is data center 106 described in conjunction with, for example, Figure 32 In at least one embodiment, data center 206 is one or more data centers. In at least one embodiment, data center 206 is at least a portion of data center 3200 described in conjunction with, for example, Figure 1 In at least one embodiment, any component or module of data center 206 is implemented on any other component or module of data center 206. In at least one embodiment, any component or module of data center 206 is communicatively connected with any other component or module of data center 206. In at least one embodiment, any component or module of data center 206 is part of a distributed computing system in which any two components and / or modules are each implemented on different computing systems (e.g., data centers, servers) connected over a network.

[0107] In at least one embodiment, the system 200 includes a job scheduler 210. In at least one embodiment, the job scheduler 210 is Figure 1 In at least one embodiment, the job scheduler 210 receives a request from a user via a user interface (e.g., Figure 1 In at least one embodiment, the job scheduler 210 receives a command to schedule a job, wherein the command includes an indication of a particular job (e.g., a job ID), a processor performance preference, a job type, a job priority, or some combination thereof. In at least one embodiment, the job scheduler 210 receives a command to submit a job for execution by a processor, wherein the command includes as input an indication of a particular job. In at least one embodiment, the command to submit a job for execution is a command of the job scheduler, e.g., Slurm job scheduler. In at least one embodiment, the command entered through the user interface is represented in pseudocode as srun –max_perf –job priority0, where srun is a command to submit a job for execution, max_perf is an indication of a processor performance preference, and jobpriority0 is an indication that the job priority is 0, where such indications are further described herein. In at least one embodiment, the command to submit a job to the job scheduler 210 includes an indication of the job, such as a job ID. In at least one embodiment, the command to submit a job to the job scheduler 210 includes an indication of the job type. In at least one embodiment, the command to submit a job to the job scheduler 210 includes one or more indications of constraints to be applied to the execution of the job, such as a constraint or limit on the amount of power used to complete the job.

[0108] In at least one embodiment, the commands of the job scheduler are called API functions. In at least one embodiment, the API of the job scheduler 210 is stored in the job priority handler profile API module 222. In at least one embodiment, the job priority handler profile API module 222 is Figure 52The job priority handler configuration file API module 122. In at least one embodiment, at least a portion of the job priority handler configuration file API module 222 is implemented as part of the job scheduler 210. In at least one embodiment, at least a portion of the job priority handler configuration file API module 222 is implemented as part of the data center processor management module 220. In at least one embodiment, when a user or application program enters a command, this is referred to as invoking an API of the job priority handler configuration file API module 222. In at least one embodiment, when the job scheduler receives a command, this refers to a user or application program entering a command line containing text that causes the processor to execute one or more API functions. In at least one embodiment, an API function is referred to as an API.

[0109] In at least one embodiment, a processor performance preference is a processor target metric. In at least one embodiment, a processor target metric is a processor metric that a processor attempts to achieve or maintain during execution of a job. In at least one embodiment, a processor performance metric includes one or more clock frequencies at which one or more processors are to run. In at least one embodiment, a processor metric measured by a job test run includes any metric used to measure a performance characteristic of a set of processors executing a job. In at least one embodiment, a processor metric includes any type of throughput metric that measures a number of operations performed in a given time period. In at least one embodiment, a processor metric is a type of measurement related to power consumed and / or temperature reached by one or more processors. In at least one embodiment, a processor metric is referred to as a performance metric.

[0110] In at least one embodiment, a processor performance preference is a user preference of a user of how a processor should perform a job. In at least one embodiment, a processor performance preference is a pre-set combination of two or more processor metrics stored in a database. In at least one embodiment, a processor performance preference is a processor configuration file. In at least one embodiment, a processor performance preference is referred to as maximum performance, or max perf in pseudocode, where such a preference is associated with setting one or more processor settings so as to estimate that a job will be performed in a given time (e.g., the shortest possible time) and / or by consuming a particular amount of power (e.g., maximum power). In at least one embodiment, a processor performance preference is referred to as energy efficiency, or energy effi ciency in pseudocode, where such a preference is associated with setting one or more processor settings so as to estimate that a job will be performed in a given time with the least power consumption. In at least one embodiment, a processor performance preference is referred to as tensor cores, or tensor core in pseudocode, where such a preference is associated with setting one or more processor settings to maximize tensor core performance according to certain metrics (e.g., floating point operations per second (FLOPS)). In at least one embodiment, a tensor core is a part of a GPU that is specifically designed to perform mathematical operations using tensors, and at least in conjunction with Figure 52 In at least one embodiment, a tensor core is a processing core of a GPU, such as processing core 5210 in FIG. 5B. In at least one embodiment, a tensor core is a processing core of a GPU, such as processing core 5210 in FIG. 5B. Figure 1 In at least one embodiment, a tensor core is a processing core of a GPU, such as processing core 5210 in FIG. 5B. In at least one embodiment, a tensor core is a processing core of a GPU, such as processing core 5210 in FIG. 5B. TM In at least one embodiment, a tensor core is a processing core of a GPU, such as processing core 5210 in FIG. 5B. In at least one embodiment, a tensor core is a processing core of a GPU, such as processing core 5210 in FIG. 5B.

[0111] In at least one embodiment, the job scheduler database 212 stores information about job priorities. In at least one embodiment, priorities are referred to as priority levels. In at least one embodiment, two jobs indexed as job 0 and job 2 have a default priority of 0. In at least one embodiment, job priorities are represented by integers, where a lower number (e.g., -1023) represents the lowest possible priority and a higher integer (e.g., 1024) represents the highest possible priority. In at least one embodiment, the default job priority is represented by the integer 0. In at least one embodiment, the job priority represents the urgency with which a job is to be executed. In at least one embodiment, the factors included in calculating the job priority based on urgency are the computing resources required for the job, the amount of time the job has been waiting in the job queue for execution, when the job must be completed, or some combination of these factors.

[0112] In at least one embodiment, in response to receiving information about a job to be scheduled through a command or otherwise obtaining information about the job, the job scheduler 210 assigns one or more processors to execute the job. In at least one embodiment, the job scheduler 210 instructs the one or more processors to execute the job by generating and storing one or more identifiers (e.g., GPU handles) of the one or more processors in a job scheduler database such that the one or more identifiers are associated with the job.

[0113] In at least one embodiment, in response to receiving a command to schedule a particular job based on information about the job, the job scheduler 210 stores the information in the job scheduler database 212. In at least one embodiment, the job scheduler database 212 is Figure 2 In at least one embodiment, the job scheduler database 212 is located in Figure 1 is depicted as having five jobs stored in the queue, at positions 0 through 4. In at least one embodiment, the job scheduler database 212 stores an indication of a processor performance preference corresponding to a particular job. In at least one embodiment, the processor performance preference includes maximum performance (max perf), energy efficiency (energy efficiency), tensor cores (tensor cores), compute (compute), or some combination thereof.

[0114] In at least one embodiment, in response to receiving a command to schedule a particular job based on information about the job, the job scheduler 210 enters a command or otherwise calls an API of the data center processor management module 220 to identify one or more processor profiles based on the information about the job. In at least one embodiment, the data center processor management module 220 is Figure 3the data center processor management module 120. In at least one embodiment, the job scheduler 210 sends information about a job from the job scheduler database 212 to the data center processor management module 220. In at least one embodiment, the data center processor management module 220 receives or otherwise obtains information about a job from the job scheduler database 212. In at least one embodiment, one or more APIs of the data center management module 220 receive or otherwise obtain as input an indication of processor performance preferences, a job type, a job priority, a processor type to execute the job, or some combination thereof from the job scheduler database 212. In at least one embodiment, the job scheduler database 212 is implemented as part of the job scheduler 210.

[0115] In at least one embodiment, a processor performs operations of the data center processor telemetry module 223 to transmit processor performance metrics of one or more processors of the processor group 208 to the data center processor management module 220. In at least one embodiment, the data center processor management module 220 receives or otherwise obtains processor performance metrics from the data center processor telemetry module 223. In at least one embodiment, the processor performance metrics are metrics observed while the one or more processors are executing a job. In at least one embodiment, the processor performance metrics include an indication of tensor core activity (Tensor_Active), a percentage of SMs that are in use (SM_utilization), a fraction of cycles that use FP64 cores (FP64_utilization), a number of vector instructions executed per cycle (#of vector instructions executed per cycle), a number of instances of operations that had to wait to execute due to memory limitations (#Mem stalls per cycle), a percentage of data transfers that were provided by L2 cache instead of DRAM (L2_Hit_rate), or some combination thereof.

[0116] In at least one embodiment, data center processor management module 220 uses processor performance metrics of processors executing a job to identify a job type of a job being executed. In at least one embodiment, processor performance metrics are used, at least in part, to identify that a job being executed requires processor utilization of tensor cores, and thus, identifies the job as a tensor core intensive job. In at least one embodiment, when data center processor management module 220 identifies a job type of a job being executed by a processor, data center processor management module 220 uses this identification to modify processor settings to cause these processors to more optimally execute the job according to one or more metrics (e.g., FLOPS). In at least one embodiment, data center processor management module 220 uses job type identification based on data center processor telemetry metrics to generate new performance profiles, as described herein in connection with at least FIG. 7. Figure 7 Further described.

[0117] In at least one embodiment, at least a portion of processor driver / firmware 226 is implemented on a processor. In at least one embodiment, at least a portion of processor driver / firmware 226 is implemented on data center processor management module 220. In at least one embodiment, driver / firmware 226 includes at least a portion of driver 704 of Figure 8 In at least one embodiment, at least a portion of processor driver / firmware 226 is driver / runtime 804 of Figure 1 In at least one embodiment, process driver / firmware 226 performs one or more operations of data center processor management module 220, such as identifying processor profiles based on information described herein regarding jobs. In at least one embodiment, at least a portion of job priority processor profile API module 222 is implemented as part of processor driver / firmware 226. In at least one embodiment, processor driver / firmware 226 includes one or more APIs described herein.

[0118] In at least one embodiment, processor profile database 224 is accessible by data center processor management module 220, job priority processor profile API module 222, processor driver / firmware 226, or some combination thereof. In at least one embodiment, processor profile database 224 is a database of processor profiles for different types of jobs, such as tensor core intensive jobs, general purpose computing on graphics processing units (GPGPU) jobs, and so on. Figure 3In at least one embodiment, any module or component of the data center 206 accesses the processor profile database 224 to identify, at least in part, a processor profile for a processor to use when executing a particular job. In at least one embodiment, one or more data structures of the processor profile database 224 are accessible by the data center processor management module 220, the job priority processor profile API module 222, the processor driver / firmware 226, or some combination thereof.

[0119] In at least one embodiment, the processor profile database 224 includes one or more data structures for storing information about a job such that one or more processor profiles are associated with the information. In at least one embodiment, the data structure is a lookup table 228. In at least one embodiment, the lookup table 228 is any one or more data structures, such as a hash table, an index, a graph, or some combination thereof, that associate information about a job with a processor profile. In at least one embodiment, the lookup table 228 stores an indication of information about a job. In at least one embodiment, the lookup table 228 associates an indication or combination of indications of information about a job with a processor profile. In at least one embodiment, the lookup table 228 is a table that includes Figure 40 In at least one embodiment, the data center processor management module 220, the job priority processor profile API module 222, the processor driver / firmware 226, or some combination thereof can access the lookup table 228 to identify one or more processor profiles for a processor to use when executing a particular job.

[0120] In at least one embodiment, the data center 206 includes a processor group 208. In at least one embodiment, the processor 108 includes the processor group 208. In at least one embodiment, the processor group 208 includes processors 208a-208n. In at least one embodiment, one or more of the processors 208a-208n are portions of a processor, such as partitions of a streaming multiprocessor and memory of a GPU, each configured to function as an independent GPU, as further described herein. In at least one embodiment, the processor group 208 is a cluster of computing resources, such as a combination of Figure 44 Describes the thread block cluster, combined Figure 44 Multi-GPU cluster described, Figure 45 Computing clusters 4436A-4436H, Figure 50 Cluster 4514A-4514N cluster, Figure 51 GPC 5018 general processing cluster ("GPC"), Figure 3Data Processing Cluster (DPC) 5106, or some combination thereof.

[0121] Figure 3 A block diagram of a system 300 is shown that includes a lookup table for identifying one or more processor profiles based on information about a job as part of a job scheduling process. Figures 1-2 One or more aspects of one or more embodiments described herein may be combined with at least Figures 4-31 and Figure 1 In at least one embodiment, one or more processors perform one or more operations of system 300. In at least one embodiment, the one or more processors that perform one or more operations of system 300 are any one of the processors or combinations of processors described herein, including Figure 2 processor 108, Figure 4 Processor group 208, processor 308, Figure 38 Processor 408, processor group 508, Figure 41A APU 3800, combined with Figure 43A Description of the CPU 4100, combined with Figure 50 The described graphics processor 4310, combined with Figure 49 PPU 5000 as described, or Figure 1 In at least one embodiment, the processor 308 executes one or more SMs 5114. Figure 2 One or more operations of the system 100, such as the operation of the job scheduler 110. In at least one embodiment, the processor 308 executes Figure 4 One or more operations of the system 200, such as an operation for identifying a processor profile using the lookup table 228. In at least one embodiment, the processor 308 performs operations in conjunction with Figure 5 One or more operations described, such as selecting a selected processor profile 427 using job priority. In at least one embodiment, the processor 308 performs a combination of Figure 6 One or more operations described, such as the operation of the job scheduler 510 for receiving job priorities from the user processor group 508. In at least one embodiment, the processor 308 performs the operation in conjunction with Figure 7 One or more operations described, such as the operation of accessing a processor configuration file stored in a data structure using operation 604. In at least one embodiment, processor 308 performs Figure 8In at least one embodiment, the processor 308 performs one or more operations in conjunction with the API 710. Figures 9-31 In at least one embodiment, processor 308 performs one or more operations described herein, such as identifying a processor setting from a database. Figure 1 One or more operations described.

[0122] In at least one embodiment, system 300 includes a data center, such as Figure 1 In at least one embodiment, the system 300 includes a lookup table 328. In at least one embodiment, the lookup table 328 is Figure 2 At least a portion of the lookup table 228 of FIG. In at least one embodiment, the lookup table 328 is visually depicted using rows and columns. In at least one embodiment, the lookup table includes processor profiles or indications thereof, such as maximum performance, energy efficiency, or tensor core density. In at least one embodiment, each processor profile includes values ​​and / or formulas for setting and / or controlling processor settings.

[0123] In at least one embodiment, lookup table 328 includes any information regarding any one or more processor settings that may affect the processor's execution of a job. In at least one embodiment, lookup table 328 includes values ​​related to maximum processor clock frequency (Fmax cap), maximum total graphics power (max TGP), crossbar ratio (Xbar ratio), maximum memory clock (max Mclk), maximum operating voltage (Vmax), or some combination thereof. In at least one embodiment, lookup table 328 includes an algorithm for calculating fan speed, referred to as a fan control algorithm. In at least one embodiment, the input to the fan control algorithm is a bias value, as further described herein. In at least one embodiment, lookup table 328 includes performance adjustment coefficients, which are used in an algorithm used to set processor settings. In at least one embodiment, the performance adjustment coefficients are used as coefficients in the fan control algorithm. In at least one embodiment, lookup table 328 includes constraints on at least a portion of the neural network weights used to execute the job. In at least one embodiment, the neural network weights are deep learning (DL) weights. In at least one embodiment, the DL weight constraint list (also referred to as DL weights) indicates minimum and / or maximum values ​​for weights to be used during mathematical operations. In at least one embodiment, weight constraints are important because a processor may perform mathematical operations more slowly or more quickly depending on the range of values ​​and / or data formats used during those operations.

[0124] In at least one embodiment, lookup table 328 includes a bias. In at least one embodiment, a bias is a value that indicates a percentage increase or decrease in one or more processor settings. In at least one embodiment, a bias is a value that is inserted into a fan control algorithm to adjust processor fan speed. In at least one embodiment, a bias is generated by a data center processor management module (e.g., data center processor management module 220 of Figure 3 FIG. 1) to cause one or more processor settings of a default processor profile to increase or decrease. In at least one embodiment, a default processor profile includes one or more default processor setting values. In at least one embodiment, in conjunction with Figure 3 and other figures, a data center processor management module is referred to as a resource manager (RM). In at least one embodiment, when a data center processor must execute multiple jobs with the same priority, a resource manager uses a bias to more finely define job priority. In at least one embodiment, a resource manager uses a bias to more finely define job priority of two different jobs. In at least one embodiment, using a bias allows for better modification of processor settings to optimize execution of multiple jobs according to various factors, such as power usage efficiency, amount of time required to complete a job, or some combination thereof.

[0125] In at least one embodiment, a bias value is an integer. In at least one embodiment, as a bias value increases, one or more processor settings increase. In at least one embodiment, bias 1 increases a default maximum processor clock frequency (Fmax cap) of an energy efficiency profile from 2.6 GHz to 2.7 GHz. In at least one embodiment, bias 0 leaves processor settings of a default processor profile unchanged.

[0126] In at least one embodiment, using a bias value to adjust processor settings has advantages over modifying a job priority value because a bias value allows a RM to adjust priority of a job without having to communicate with a job scheduler to assign a new job priority, which would degrade performance of scheduling that job. In at least one embodiment, a bias value has advantages over creating additional processor profiles containing different processor settings because each processor profile requires more data to be stored than simply inputting a single bias value into a formula to adjust processor settings and / or using a single bias value to adjust one or more processor settings by a given percentage.

[0127] In at least one embodiment, the processor configuration files stored in lookup table 328 are indicated by one or more job-related information indications. In at least one embodiment, the one or more job-related information indications include one or more indications of processor performance preference, job type, job priority, processor type, or some combination thereof. In at least one embodiment, a processor performance preference (max performance) indication input as part of a job scheduler command (as described further herein) causes RM to identify a max performance configuration file in lookup table 328. In at least one embodiment, after identifying a max performance configuration file in lookup table 328, that configuration file is stored with an indication of one or more processor settings to use to implement the max performance configuration file on a processor. In at least one embodiment, the indication of one or more processor settings to use by a processor configuration file is one or more indications of one or more memory addresses storing these settings and / or related formulas.

[0128] In at least one embodiment, the type of processor is based on an operational specification of that processor. In at least one embodiment, an operational specification includes a maximum TGP, an indication of processor core type (e.g., tensor core, compute core), an amount of different types of memory (e.g., L2 cache, DRAM) on a processor, or some combination thereof. In at least one embodiment, an amount of memory is referred to as memory capacity. In at least one embodiment, if a maximum TGP of a certain type of processor is lower than a default maximum TGP of a processor setting configuration file, then that type of processor causes RM to modify the default processor setting configuration file. In at least one embodiment, if a bias factor would cause a processor setting to exceed an operational limit (e.g., maximum TGP) of that processor, then that type of processor causes RM to modify the way it incorporates bias factors into processor settings of a processor setting configuration file. In at least one embodiment, if a job type is tensor intensive and a processor assigned to execute that job has no tensor cores and limited memory, then RM determines that a maximum clock speed of the processor should be lowered so that a default clock speed is faster than needed to execute that job, and thus, if the default clock speed is not lowered, it will not speed up performance of that job and will cause wasted power consumption.

[0129] In at least one embodiment, a job to be executed using a max performance processor configuration file and another job to be executed using a tensor core intensive processor configuration file (both in Figure 1The two different jobs (delineated by the dashed line) have the same job priority 0, the same bias 0, and are scheduled to be executed concurrently at the data center. In at least one embodiment, because the bias assigned to the two different jobs that are executed at least partially concurrently are equal, the resource manager generates a new processor profile 330 by combining and / or selecting various processor settings that are estimated to optimally execute the two jobs according to one or more factors (e.g., FLOPS) for the two jobs. In at least one embodiment, the new processor profile generated by the resource manager adds the profile to the lookup table 328. In at least one embodiment, one or more existing processor profiles of the lookup table 328 are created after experimental simulations and / or benchmarking of different jobs executed on different computing resources.

[0130] In at least one embodiment, the resource manager sends an indication of the new processor profile 330 to the job scheduler. In at least one embodiment, the job scheduler utilizes the indication of the new processor profile 330 to schedule one or more jobs. In at least one embodiment, prior to execution of the one or more jobs using the new processor profile 330, the job scheduler sends an indication of the new processor profile 330 back to the resource manager, where the module checks whether the processor settings of the processors 308 have been set according to the new processor profile 330. In at least one embodiment, if the processor settings of the processors 308 have not been set accordingly, the resource manager sets the processor settings. In at least one embodiment, if the resource manager determines that the processor settings have been set according to the new processor profile 330, the module sends an indication back to the job scheduler to cause the job scheduler to initiate the one or more jobs to be executed using the new processor profile 330. In at least one embodiment, one or more operations described in connection with the new processor profile 330 apply to other processor profiles of the lookup table 328. In at least one embodiment, the processors 308 are any one or more of the processors 108 of Figure 2 or the processor group 208 of Figure 4 .

[0131] Figure 4 A block diagram of a system 400 is shown in at least one embodiment, the system 400 including a lookup table that is used at least in part to identify one or more processor profiles as part of a job scheduling process based on information about a job. In at least one embodiment, in connection with Figures 1-3 One or more aspects of one or more embodiments described herein are combinable with one or more aspects of one or more embodiments described herein, including at least in connection with Figures 5-31 and Figure 1In at least one embodiment, one or more processors perform one or more operations of system 400. In at least one embodiment, the one or more processors performing one or more operations of system 400 are any one or combination of processors described herein, including Figure 2 processor 108, Figure 3 Processor group 208, Figure 4 Processor 308, processor 408, processor group 508, Figure 38 APU 3800, combined with Figure 41A Description of the CPU 4100, combined with Figure 43A The described graphics processor 4310, combined with Figure 50 PPU 5000 as described, or Figure 49 In at least one embodiment, one or more operations of the system 400 performed by the processor are Figure 1 One or more operations of the system 100, such as the operation of the job scheduler 110. In at least one embodiment, the processor 408 executes Figure 2 One or more operations of the system 200, such as an operation for identifying a processor profile using the lookup table 228. In at least one embodiment, the processor 408 performs operations in conjunction with Figure 3 One or more operations described herein, such as generating a new performance profile 330. In at least one embodiment, the processor 408 performs the Figure 4 One or more operations described, such as the operation of the job scheduler 510 for receiving job priorities from the user processor group 508. In at least one embodiment, the processor 408 performs the operation in conjunction with Figure 5 One or more operations described, such as the operation of accessing a processor configuration file stored in a data structure using operation 604. In at least one embodiment, processor 408 performs Figure 7 In at least one embodiment, the processor 408 performs one or more operations in conjunction with the API 710. Figure 8 In at least one embodiment, processor 408 performs one or more operations described herein, such as identifying a processor setting from a database. Figures 9-31 One or more operations described.

[0132] In at least one embodiment, system 400 includes a data center, such as Figures 1-2 Data center 106. In at least one embodiment, Figure 1 The processor profile database 124 includes Figures 1-5 Lookup table 228, Figures 7-31one or more portions of the lookup table 328, the lookup table 428, or some combination thereof. In at least one embodiment, the lookup table 428 is used in conjunction with the lookup table 328. In at least one embodiment, the lookup table 428 includes processor settings for four different versions of a maximum performance profile as a function of skew. In at least one embodiment, the lookup table 428 is referred to as a profile and skew interaction table. In at least one embodiment, the processor settings for the maximum performance profile are organized by skew. In at least one embodiment, one or more processor settings for the maximum performance profile increase as skew increases. In at least one embodiment, as skew increases, a fan control algorithm does not change, but its output changes because the skew value is input into the fan control algorithm.

[0133] In at least one embodiment, the resource manager identifies a processor setting to use based on the skew value. In at least one embodiment, the identified processor setting is included in the selected processor profile 430. In at least one embodiment, the resource manager sends an indication of the selected processor profile 430 to the job scheduler. In at least one embodiment, the job scheduler uses the indication of the selected processor profile 430 to schedule one or more jobs. In at least one embodiment, prior to executing the one or more jobs using the selected processor profile 430, the job scheduler sends an indication of the selected processor profile 430 back to the resource manager, where the manager checks whether the processor settings of the processor 408 have been set according to the selected processor profile 430. In at least one embodiment, if the processor settings of the processor 408 have not been set accordingly, the resource manager sets the processor settings. In at least one embodiment, if the resource manager determines that the processor settings have been set according to the selected processor profile 430, the module sends an indication back to the job scheduler to cause the job scheduler to initiate the one or more jobs to be executed using the selected processor profile 430. In at least one embodiment, the processor 408 is Figures 1-5 the processor 108, Figures 7-31 any one or more processors of the processor group 208 or the processor 308.

[0134] In at least one embodiment, lookup table 428 illustrates how bias settings (also referred to as bias values or bias factors) are mapped to processor settings and features. In at least one embodiment, firmware of a resource manager uses lookup table 428 to map bias factors to each processor setting and feature included in a processor setting configuration file and to help bias performance of a processor using that configuration file. In at least one embodiment, bias factors and their mapping to processor settings and features are calculated, adjusted, or otherwise determined by including simulations and / or experiments of various jobs performed by various processors. In at least one embodiment, lookup table 428 includes sets of processor settings that result in optimal job execution for various jobs in various domains. In at least one embodiment, jobs are referred to as workloads or software workloads.

[0135] In at least one embodiment, techniques of mapping bias factors and / or job characteristics to processor setting configuration files can include algorithms, formulas, or models that output values as a function of bias, such as fan control algorithms of lookup table 428 described further herein. In at least one embodiment, performance adjustment coefficients of lookup table 428 are used as coefficients in performance adjustment algorithms, formulas, or models that are implemented as part of a performance estimator of a processor. In at least one embodiment, performance adjustment algorithms, formulas, or models use performance adjustment coefficients to determine processor performance at any given point when a processor is in an active state. In at least one embodiment, a resource manager automatically incorporates bias factors (when set) into one or more processor settings, features, sub-features, or models to modify each of these settings, features, sub-features, or models to cause a processor to optimally execute a job.

[0136] Figure 2 A block diagram illustrating system 500 in at least one embodiment is shown, including a job scheduler that is used, at least in part, to schedule one or more jobs based on an indication of one or more job characteristics. In at least one embodiment, one or more aspects of one or more embodiments described herein include at least in combination Figure 1 one or more aspects of one or more embodiments described herein include at least in combination Figure 2 and Figure 3 one or more aspects of one or more embodiments described herein. In at least one embodiment, one or more processors perform one or more operations of system 500. In at least one embodiment, one or more processors performing one or more operations of system 500 are any of the processors described herein or a combination of processors, including Figure 4 processor 108 of FIG. 1, Figure 1 processor group 208 of FIG. 2, Figure 2processor 308, Figure 7 Processor 408, processor group 508, Figure 1 APU 3800, combined with Figure 2 Description of the CPU 4100, combined with Figure 3 The described graphics processor 4310, combined with Figure 4 PPU 5000 as described, or Figure 5 In at least one embodiment, the processor group 508 performs one or more SMs 5114. Figure 38 In at least one embodiment, the processor group 508 performs one or more operations of the system 100, such as the operation of the job scheduler 110. Figure 41A One or more operations of the system 200, such as using the lookup table 228 to identify the operation of the processor profile. In at least one embodiment, the processor group 508 performs the operation in conjunction with Figure 43B One or more operations described herein, such as generating a new performance profile 330. In at least one embodiment, the processor group 508 performs the combined Figure 44 One or more operations described, such as operations for selecting a processor profile to be applied to processor 408. In at least one embodiment, processor group 508 performs operations in conjunction with Figure 50 One or more operations described, such as the operation of accessing a processor profile stored in a data structure using operation 604. In at least one embodiment, the processor group 508 performs Figure 1 In at least one embodiment, the processor group 508 performs one or more operations in conjunction with the API 710. Figures 1-31 In at least one embodiment, the processor group 508 performs one or more operations described in conjunction with Figure 1 One or more operations described.

[0137] In at least one embodiment, system 500 includes a data center, such as Figure 1 In at least one embodiment, the system 500 includes a job scheduler 510 for the data center. In at least one embodiment, the job scheduler 510 is Figure 1 In at least one embodiment, the job scheduler 510 schedules jobs Job0, Job1, and Job2. In at least one embodiment, Job0, Job1, and Job2 are job IDs. In at least one embodiment, the job IDs are stored in a job scheduler database. In at least one embodiment, each job ID is associated with a job priority p0, p1, and p2. In at least one embodiment, the processor executes at least one of the instructions in conjunction with Figure 8 and Figure 8One or more of the further described operations to identify processor configuration files 510.

[0138] In at least one embodiment, each of the processor configuration files 510 will be used to set processor settings for processors assigned to execute jobs submitted to the job scheduler 510. In at least one embodiment, the job scheduler 510 stores an indication of the processor configuration file to use with a job when scheduling the job. In at least one embodiment, Job0 is assigned the lowest job priority among a plurality of jobs, but the user enters a processor performance preference of maximum performance into the job scheduler 510, while other jobs are indicated as having a best-per-watt performance preference. In at least one embodiment, a processor performance preference that requires more computing resources (e.g., power) compared to other processor performance preferences causes the resource manager to assign a bias higher than the bias assigned to other jobs indicated as having those other processor performance preferences for jobs indicated as having the processor performance preference, despite and / or because the other jobs have a higher job priority.

[0139] In at least one embodiment, the system 500 includes a resource manager that allows a user to interact with and view individual processor settings for one or more GPUs through a user interface. In at least one embodiment, the system 500 allows a user to view processor setting configuration files stored in firmware of the GPUs or as part of the resource manager. In at least one embodiment, the processor setting configuration files are referred to as performance policies, profile modes, or performance profiles. In at least one embodiment, the processor setting configuration files are power policies that include processor settings that govern processor power consumption.

[0140] In at least one embodiment, the system 500 includes a resource manager that allows a user to view a list of processor performance preferences. In at least one embodiment, the list of processor performance preferences is displayed as a list of power profiles for a settings page through a user interface. In at least one embodiment, the user interface is a graphical user interface (GUI). In at least one embodiment, the system 500 allows a user to select or adjust processor performance preferences. In at least one embodiment, the system 500 allows a user to adjust processor performance preferences according to a server room, a server row, a server rack, a device, a blade server, or some combination thereof. In at least one embodiment, a processor performance preference is a more general way of identifying a processor setting configuration file that, when used to configure a processor, estimates that a threshold amount of performance is achievable. In at least one embodiment, the threshold amount of performance is a given amount of operations performed per watt of power consumed by a processor.

[0141] In at least one embodiment, system 500 allows a user to view and / or modify individual processor settings of a default processor setting configuration file stored in firmware of a resource manager. In at least one embodiment, the default processor setting configuration file stored in firmware of a resource manager is a Max Perf configuration file that includes at least the following processor settings: Vmax = 1.1V, Fmax cap = 2.6 Ghz, V939 ON, Min-TGP = 500W, Max-TGP = 750W, XBAR Ratio = 10, MCLK = 1593, Thermal policy = A, Vmax balancing feature = enabled, DLPPE = enabled. In at least one embodiment, Vmax, Fmax cap, Max-TGP, XBAR Ratio, and MCLK are described further herein. In at least one embodiment, Min-TGP is a minimum total graphics power setting for a processor that represents a minimum power consumption for that processor under normal operating conditions. In at least one embodiment, Thermal policy = A represents a specific set of settings and / or formulas used to adjust these settings to prevent one or more portions of a processor from exceeding a given temperature. In at least one embodiment, Vmax balancing feature = enabled refers to a resource manager monitoring power consumption of multiple GPUs and adjusting Vmax of each GPU as these GPUs perform workloads, thereby optimizing power consumption of these GPUs to perform a certain number of operations per watt consumed. In at least one embodiment, DLPPE = enabled refers to a deep learning parallel processing engine for performing neural network operations being enabled.

[0142] In at least one embodiment, a user can select the default Max Perf configuration file to achieve maximum performance as a base configuration file when performing a job. In at least one embodiment, a resource manager allows a user to input a bias value to further increase a priority of a job to be performed. In at least one embodiment, the bias value is referred to as a bias factor. In at least one embodiment, a resource manager automatically generates a bias value to increase a priority of a job to be performed, as described further herein. In at least one embodiment, the default Max Perf configuration file can be biased by +4, as shown in processor configuration file 510. In at least one embodiment, each setting of a processor setting configuration file will be connected to or modified by a bias factor. In at least one embodiment, firmware of a resource manager has a policy to automatically adjust each processor feature or setting to include the bias factor.

[0143] In at least one embodiment, the resource manager of system 500 includes firmware and / or drivers that utilize user-provided job priorities and adjust the characteristics of one or more circuits and the power management policy of a processor (e.g., a CPU and / or GPU). In at least one embodiment, circuit settings are changed based on a profile and a deviation factor. In at least one embodiment, the power management policy is adjusted based on the operating region of the CPU and / or GPU. In at least one embodiment, changing processor circuit settings, adjusting processor circuit characteristics, and adjusting power management policy are referred to as configuring the processor. In at least one embodiment, the resource manager dynamically adjusts voltage and frequency profiles based on the operating region of the GPU (e.g., 700W, 750W, 800W). In at least one embodiment, noise-aware frequency-locked loop (NAFLL) clock circuit settings vary based on the voltage axis, and therefore, these circuit settings are adjusted to an optimal voltage range based on the deviation. In at least one embodiment, noise settings and thermal policies are automatically changed based on the deviation of the resource manager's firmware and / or driver, so that the processor settings and various policies result in the processor operating within a given operating range. In at least one embodiment, the resource manager of system 500 dynamically adjusts the definitions of processor performance states (P-states) in the video BIOS (vBIOS) based on a bias factor applied to those definitions. In at least one embodiment, the definitions of a P-state include one or more aspects of a processor settings profile, such as one or more processor settings. In at least one embodiment, the bias factor applied to a P-state definition or processor settings profile causes the corresponding processor to be referred to as being in a bias mode. These are just a few examples; essentially, the FW / driver will have a large number of settings and policies that need to be changed based on the bias.

[0144] In at least one embodiment, the resource manager uses the indication of the processor profile stored in the job queue by the job scheduler 510 to cause the processor settings of the processor group 508 to be set according to the indicated processor profile. In at least one embodiment, the processor group 508 includes Figures 1-7 processor 108, Figures 8A-31 Processor group 208, Figure 1 Processor 308 or Figure 2 One or more processors in processor 408.

[0145] Figure 3 A block diagram of a process 600 for identifying processor settings to use based, at least in part, on one or more job characteristics (e.g., job priority) in at least one embodiment is shown. Figure 4One or more aspects of one or more embodiments described herein may be combined with one or more aspects of one or more embodiments described herein, including at least one aspect of one or more embodiments described herein. Figure 5 and Figure 38 In at least one embodiment, one or more processors perform one or more operations of process 600. In at least one embodiment, the one or more processors performing one or more operations of process 600 are any one of the processors or combinations of processors described herein, including Figure 41A processor 108, Figure 43A Processor group 208, Figure 50 processor 308, Figure 49 Processor 408, Figure 1 Processor group 508, Figure 2 APU 3800, combined with Figure 3 Description of the CPU 4100, combined with Figure 4 The described graphics processor 4310, combined with Figure 5 PPU 5000 as described, or Figure 6 In at least one embodiment, a processor performing one or more operations of process 600 executes Figure 7 One or more operations of the system 100, such as the operation of the job scheduler 110. In at least one embodiment, a processor that performs one or more operations of the process 600 executes Figures 9-31 One or more operations of the system 200, such as an operation for identifying a processor profile using the lookup table 228. In at least one embodiment, a processor performing one or more operations of the process 600 performs operations in conjunction with Figure 1 One or more operations described herein, such as generating a new performance profile 330. In at least one embodiment, a processor performing one or more operations of process 600 performs operations in conjunction with Figure 1 One or more operations described herein, such as operations for selecting a processor profile to be applied to processor 408. In at least one embodiment, a processor performing one or more operations of process 600 performs operations in conjunction with Figure 6 One or more operations described, such as identifying a processor profile 510. In at least one embodiment, a processor performing one or more operations of process 600 executes Figure 7 In at least one embodiment, a processor that performs one or more operations of process 600 performs operations in conjunction with Figure 1 In at least one embodiment, a processor that performs one or more operations of process 600 performs Figures 8A-8Cone or more operations.

[0146] In at least one embodiment, user causes processor to begin process 600 by invoking an API of a job scheduler to input an indication of one or more job characteristics (e.g., a job identifier (e.g., a job ID), a processor performance preference, a job type, a job priority, or some combination thereof) into one or more APIs, such as described herein in at least connection with Figures 8A-8C In at least one embodiment, processor performs operations of operation 602 using a single API or a single command line. In at least one embodiment, processor performs operations of operation 602 using two or more APIs or command lines. In at least one embodiment, user inputs one or more indications of one or more job characteristics into one or more APIs using a command line through a user interface (e.g., user interface 102 of FIG. 1), such as described herein in at least connection with Figures 1-8

[0147] In at least one embodiment, in response to user inputting a command line through a user interface as part of operation 602, processor executes an API to cause one or more other processors to be configured to run at one or more clock frequencies based at least in part on one or more inputs to that API. In at least one embodiment, configuring a processor to run at a clock frequency refers to an operation to set a processor setting that controls a clock frequency of a processor used to perform a job, or as described herein in at least connection with Figures 9-31 and Figures 8A-8C In at least one embodiment, processor executes an API to cause one or more other processors to be configured to run according to one or more processor settings, such as an Xbar ratio, Vmax, or as described herein in at least connection with Figures 8A-8C and Figure 1 In at least one embodiment, one or more inputs to an API of operation 602 include an indication of one or more processor target metrics to be used by one or more processors when executing one or more instructions of a job. In at least one embodiment, a processor target metric includes an amount of power to be consumed when executing a job, a time to complete execution of a job, or some combination thereof, or as described herein. In at least one embodiment, one or more inputs to an API of operation 602 include an indication of one or more processor performance preferences, such as described herein in at least connection with Figure 2 In at least one embodiment, one or more inputs to an API of operation 602 include an indication of one or more processor setting configuration files, such as described herein in at least connection with Figure 3

[0148] ​​In at least one embodiment, one or more inputs to API of operation 602 include one or more indications of a software workload for one or more processors to perform. In at least one embodiment, a software workload is a job. In at least one embodiment, a software workload is a kernel. In at least one embodiment, a software workload is a set of instructions. In at least one embodiment, a processor is used to execute API of operation 602 to cause one or more processors to be configured to run at one or more clock frequencies based at least in part on one or more observed processor performance metrics, as described further herein. In at least one embodiment, a processor is used to execute API of operation 602 to cause one or more processors to be configured to run based at least in part on a maximum operating voltage (Vmax) or other manner described herein. In at least one embodiment, a processor is used to execute API of operation 602 to cause one or more processors to be configured to run based at least in part on a crossbar ratio (Xbar ratio) or other manner described herein. In at least one embodiment, a processor is used to execute API of operation 602 to cause one or more processors to be configured to run based at least in part on a mathematical formula for controlling fan speed.

[0149] In at least one embodiment, in response to user inputting a command line through a user interface as part of operation 602, processor executes an API to indicate one or more computing resources to be used by one or more instructions based at least in part on one or more inputs to that API. In at least one embodiment, one or more inputs to API of operation 602 include one or more indications of one or more instruction types to be used by one or more computing resources, where instruction type is a job type as further described herein. In at least one embodiment, computing resources include any combination of hardware, firmware, software, or power used to execute one or more instructions. In at least one embodiment, computing resources are one or more GPUs in a set of GPUs assigned to execute one or more instructions. In at least one embodiment, processor executes API of operation 602 to indicate one or more computing resources based at least in part on an indication of floating point operations per second (FLOPS) to be performed by those one or more computing resources, or as otherwise described herein. In at least one embodiment, processor executes API of operation 602 to indicate one or more computing resources based at least in part on one or more indications of memory transfer rates of one or more computing resources, as further described herein. In at least one embodiment, one or more indications of one or more computing resources are to be used by one or more schedulers in scheduling one or more instructions, or as otherwise described herein. In at least one embodiment, one or more computing resources are one or more portions of a graphics processing unit (GPU) assigned to execute one or more instructions, or as otherwise described herein.

[0150] In at least one embodiment, in response to a user inputting a command line via a user interface as part of operation 602, the processor executes an API to indicate a priority for executing one or more instructions based at least in part on one or more inputs to the API, or as otherwise described herein. In at least one embodiment, the priority is a job priority of the one or more instructions to be executed by one or more computing resources, or as otherwise described herein. In at least one embodiment, the processor executes the API of operation 602 to indicate one or more settings for the one or more computing resources to execute the one or more instructions based at least in part on the priority of the one or more instructions. In at least one embodiment, the processor executes the API of operation 602 to indicate a priority based at least in part on an estimated time required to complete execution of the one or more instructions. In at least one embodiment, the processor executes the API of operation 602 to indicate a priority based at least in part on an estimated power consumption required to complete execution of the one or more instructions. In at least one embodiment, the priority indicated by the API of operation 602 is used by one or more schedulers when scheduling the one or more instructions. In at least one embodiment, the one or more instructions are to be executed by one or more computing resources that are one or more portions of a graphics processing unit (GPU).

[0151] In at least one embodiment, the processor continues process 600 by executing the job scheduler with operation 604 to call an API of a resource manager (RM) to send the input of operation 602 to the resource manager and cause the resource manager to access a processor profile database stored in a data structure or as otherwise described herein. In at least one embodiment, the data structure of operation 604 is a lookup table, e.g., Figure 4 Lookup table 228, Figure 5 Lookup table 328 or Figure 38 In at least one embodiment, the data structure of operation 604 is stored in a processor profile database, e.g. Figure 38 Processor profile database 124 or Figure 41A Processor profile database 224.

[0152] In at least one embodiment, processor continues process 600 by executing an RM API with operation 606 to cause the RM to select a processor configuration file from a data structure based on input of operation 602, including processor performance preferences, job type, job priority, or some combination thereof. In at least one embodiment, selecting a processor configuration file is referred to as identifying a processor configuration file, which is described further herein. In at least one embodiment, processor executes API of operation 606 to identify one or more settings to be used to configure one or more processors to run at one or more processor clock frequencies based at least in part on one or more clock frequency inputs to the API. In at least one embodiment, processor executes API of operation 606 to identify one or more settings to be used to configure one or more processors to operate according to processor settings (e.g., Xbar ratio and Vmax), or as otherwise described herein. In at least one embodiment, processor executes API of operation 606 to identify one or more settings to be used to configure one or more processors to perform one or more instructions based at least in part on one or more indications of a processor performance profile input to the API, or as otherwise described herein. In at least one embodiment, one or more settings to be used to configure one or more processors to run at one or more processor clock frequencies are based at least in part on one or more processor performance metrics observed during performance of one or more instructions by one or more processors, as described further herein. In at least one embodiment, one or more settings to be used to configure one or more processors to run at one or more processor clock frequencies include crossbar (Xbar) ratio settings. In at least one embodiment, processor executes API of operation 606 to identify one or more settings from a data structure that associates one or more indications of one or more settings with one or more clock frequency inputs, or as otherwise described herein. In at least one embodiment, processor executes API of operation 606 to identify one or more settings based at least in part on values used to bias one or more default settings used to configure one or more processors, or as otherwise described herein. In at least one embodiment, processor executes API of operation 606 to identify one or more settings based at least in part on one or more memory clock settings, or as otherwise described herein.

[0153] In at least one embodiment, processor performs API of operation 606 to identify one or more settings to be used to configure one or more processors to run at a frequency under one or more processor clocks based at least in part on a compute resource input to that API. In at least one embodiment, compute resource input is an indication of a type of compute resource that a job can require, where type of compute resource can include tensor cores, compute cores, processor memory, or some combination thereof, or as otherwise described herein. In at least one embodiment, compute resource input is an indication of a job type. In at least one embodiment, compute resource input includes an indication that one or more instruction sets to be executed by one or more processors are compute-bound, as further described herein. In at least one embodiment, compute resource input includes an indication that one or more instruction sets to be executed by one or more processors are memory-bound, as further described herein. In at least one embodiment, compute resource input includes an indication of a type of instruction to be executed by one or more processors, where such type is a job type, as further described herein. In at least one embodiment, compute resource input includes an indication of compute resources used during execution of one or more instructions by one or more processors, where such indication is based on observed processor metrics, as further described herein. In at least one embodiment, compute resource input includes an indication of a number of processors required to execute one or more instruction sets.

[0154] In at least one embodiment, a processor executes an API of operation 606 to identify one or more settings to use to configure one or more processors to run at one or more processor clock frequencies based at least in part on one or more priority inputs to the API. In at least one embodiment, a priority input is an indication of a job priority. In at least one embodiment, an indication of a job priority is referred to as a priority of one or more instructions to be executed by a processor of a data center, or as further described elsewhere herein. In at least one embodiment, an API of operation 606 receives or otherwise obtains a priority input of one or more instructions to be executed by a processor of a data center from a scheduler. In at least one embodiment, a scheduler of one or more instructions is a job scheduler as further described herein. In at least one embodiment, an API of operation 606 identifies one or more settings to use to configure one or more processors based at least in part on a priority input corresponding to one set of instructions (e.g., a particular job) and a priority input corresponding to another set of instructions (e.g., another job). In at least one embodiment, a bias value to apply to each job can be identified during execution of an API using priority inputs of multiple jobs, or as otherwise described herein. In at least one embodiment, a processor executes an API to identify one or more settings to use to configure one or more processors based at least in part on a percentage increase or decrease from one or more default values of the one or more settings, or as described elsewhere herein in connection with bias values. In at least one embodiment, a processor executes an API to identify one or more settings to use to configure one or more processors based at least in part on biasing a default value of the one or more settings by an integer, or as described elsewhere herein in connection with bias values.

[0155] In at least one embodiment, the processor executes the API of operation 606 to input a GPU handle, as further described herein. In at least one embodiment, the processor executes the API of operation 606 to identify one or more settings to be used to configure the one or more processors to run at one or more processor clock frequencies based at least in part on the one or more processors to be used. In at least one embodiment, a scheduler (e.g., a job scheduler) generates an indication of one or more processors to be used to execute one or more instructions. In at least one embodiment, the scheduler assigns one or more processors to execute a job based at least in part on processor availability in a data center. In at least one embodiment, the processor executes the API to identify settings to be used to configure the one or more processors based at least in part on the hardware specifications of the processors assigned to execute the instructions, such as hardware specifications of whether the processors include tensor cores, or other situations described herein. In at least one embodiment, the indication of the one or more processors to be used to execute the instructions is input into the API of 606. In at least one embodiment, the processor executes the API of operation 606 to identify one or more settings based at least in part on the type of processor to be used, where the types include accelerators for performing neural network operations, accelerators for using In at least one embodiment, the processor executes the API of operation 606 to identify one or more settings based at least in part on an operating specification of the one or more processors to be used, an operating specification such as a maximum total graphics power (maximum TGP). In at least one embodiment, different models of GPUs have different maximum TGPs, where one GPU may have a maximum TGP of 700W and another GPU may have a maximum TGP of 1000W. In at least one embodiment, the processor executes the API of operation 606 to identify one or more settings based at least in part on using one or more data tables that associate one or more settings with one or more processors to be used, or as otherwise described herein.

[0156] In at least one embodiment, processor continues process 600 by executing a resource manager to send an indication of a selected processor configuration file by executing an API of the resource manager. In at least one embodiment, by operation 608, processor causes a resource manager to call an API of a job scheduler to cause the job scheduler to receive an indication of a selected processor configuration file sent from the resource manager. In at least one embodiment, the resource manager sends a job ID associated with the selected processor configuration file to the job scheduler. In at least one embodiment, processor executes an API to cause execution of one or more instructions based at least in part on one or more processor setting inputs to the API. In at least one embodiment, an indication of the one or more instructions to be executed and an indication of the one or more processor settings have been stored in a queue of a scheduler (e.g., a job scheduler). In at least one embodiment, the indication of the one or more processor settings is an indication of a processor settings configuration file, as described further herein. In at least one embodiment, a processor executing a resource manager calls an API of a job scheduler and inputs a job identifier and an indication of a processor configuration file into the API. In at least one embodiment, in response to receiving these inputs, the API causes the job identifier and processor configuration file indication to be stored in a job queue, or as otherwise described herein. In at least one embodiment, within a given time prior to scheduled execution of a job, a resource manager checks whether processor settings have been set for a processor that will execute the job. In at least one embodiment, a resource manager comprises a processor management application, where an application is a software application. In at least one embodiment, a resource manager is referred to as a processor management application. In at least one embodiment, if a resource manager determines that processor settings have not been set, then the module will set the processor settings immediately. In at least one embodiment, when a resource manager identifies that processor settings have been set, the module will send an acknowledgement to a job scheduler to allow the job scheduler to launch the job.

[0157] Figure 43A A block diagram illustrating a driver and / or runtime including one or more libraries to provide one or more application programming interfaces (APIs) is shown, according to at least one embodiment. In at least one embodiment, any one processor or combination of processors executes an API 710, including Figure 50 a processor 108 of FIG. 7.08, Figure 49 a processor group 208 of FIG. 7.09, Figures 8A-8C a processor 308 of FIG. 7.10, ​ a processor 408 of FIG. 7.11, ​ a processor group 508 of FIG. 7.12, ​ an APU 3800 of FIG. 7.13, ​ a CPU 4100 of FIG. 7.14,​ Graphics processor 4340, ​ General-Purpose Graphics Processing Unit (GPGPU) 4430 and ​ In at least one embodiment, the API 710 is further described herein and includes an API for identifying a processor configuration profile based on one or more job characteristics (e.g., job priority). In at least one embodiment, the invocation of the API 710 causes one or more processors to execute any one or more components or modules described herein (e.g., ​ In at least one embodiment, the call to API 710 causes one or more processors to perform operations in conjunction with the job scheduler 110 or the data center processor management module 120. ​ In at least one embodiment, the API 710 receives as input an indication of job characteristics and causes identification of processor settings to be used when executing the job, or as otherwise described herein.

[0158] In at least one embodiment, the software program 702 is a software module. In at least one embodiment, the software program 702 includes one or more software modules. In at least one embodiment, the one or more APIs 710 are software instruction sets that, if executed, cause one or more processors to perform one or more computing operations. In at least one embodiment, the one or more APIs 710 are distributed or otherwise provided as part of one or more libraries 706, runtimes 704, drivers 704, and / or any other grouping of software and / or executable code described further herein. In at least one embodiment, the one or more APIs 710 perform one or more computing operations in response to a call by the software program 702. In at least one embodiment, the software program 702 is a collection of software codes, commands, instructions, or other text sequences that instruct a computing device to perform one or more computing operations and / or call one or more other instruction sets (e.g., APIs 710 or functions 712) for execution. In at least one embodiment, the functionality provided by one or more APIs 710 includes software functions, such as software functions that can be used to accelerate one or more portions of software program 702 using one or more parallel processing units (PPUs), such as graphics processing units (GPUs). In at least one embodiment, the software program is a compiler.

[0159] In at least one embodiment, API 710 is a hardware interface for one or more circuits to perform one or more computational operations. In at least one embodiment, one or more software APIs 710 described herein are implemented as one or more circuits to perform one or more techniques described herein. In at least one embodiment, one or more software programs 702 include instructions that, if executed, enable one or more hardware devices and / or circuits to perform one or more techniques described further herein.

[0160] In at least one embodiment, software programs 702 (e.g., user-implemented software programs) utilize one or more application programming interfaces (APIs) 710 to perform various computational operations, or any computational operations performed by parallel processing units (PPUs) (e.g., graphics processing units (GPUs)), as described further herein. In at least one embodiment, one or more APIs 710 provide a set of callable functions 712 (referred to herein as APIs, API functions, and / or functions) that each perform one or more computational operations, such as computational operations related to parallel computing. For example, in one embodiment, one or more APIs 710 provide functions 712 to cause a scheduler to schedule instructions for execution by a processor based on latency of an interconnect coupled with the processors. In at least one embodiment, API 710 provides one or more functions 712 that are one or more neural networks, e.g., trained to improve efficiency of processor usage in a rasterization process.

[0161] In at least one embodiment, one or more software programs 702 interact with or otherwise communicate with one or more APIs 710 to perform one or more computational operations using one or more PPUs (e.g., GPUs). In at least one embodiment, one or more computational operations using one or more PPUs includes at least a set or sets of computational operations that are accelerated by execution at least partially by the one or more PPUs. In at least one embodiment, one or more software programs 702 interact with one or more APIs 710 to facilitate parallel computing using a remote or local interface.

[0162] In at least one embodiment, an interface is software instructions that, if executed, provide access to one or more functions 712 provided by one or more APIs 710. In at least one embodiment, an interface is a hardware interface that provides access to one or more functions 712 provided by one or more APIs 710. ​the user interface 102. In at least one embodiment, when a software developer compiles one or more software programs 702 in conjunction with one or more libraries 706 that include one or more APIs 710 or otherwise provide access to one or more APIs 710, the software programs 702 use a native interface. In at least one embodiment, one or more software programs 702 are statically compiled with pre-compiled libraries 706 or with uncompiled source code that includes instructions for carrying out one or more APIs 710. In at least one embodiment, one or more software programs 702 are dynamically compiled and linked with a linker to one or more pre-compiled libraries 706 that include one or more APIs 710.

[0163] In at least one embodiment, when a software developer executes a software program that utilizes a library 706 that includes one or more APIs 710 or otherwise communicates with a library 706 that includes one or more APIs 710 over a network or other remote communication medium, the software program 702 uses a remote interface. In at least one embodiment, one or more libraries 706 that include one or more APIs 710 are to be executed by a remote computing service (e.g., a computing resource service provider). In another embodiment, one or more libraries 706 that include one or more APIs 710 are to be executed by any other computing host that provides the one or more software programs 702 with the one or more APIs 710.

[0164] In at least one embodiment, a processor executing or using one or more software programs 702 calls, uses, performs, or otherwise implements one or more APIs 710 to allocate and otherwise manage memory to be used by the software programs 702. In at least one embodiment, one or more software programs 702 utilize one or more APIs 710 to allocate and otherwise manage memory to be used by one or more portions of the software programs 702 to accelerate using one or more PPUs (e.g., GPUs or any other accelerator or processor further described herein). These software programs 702 can be executed by one or more processors based at least in part on latency of an interconnect coupled to the one or more processors using functions 712 provided by one or more APIs 710.

[0165] In at least one embodiment, API 710 is an API for facilitating parallel computing. In at least one embodiment, API 710 is any other API described further herein. In at least one embodiment, API 710 is provided by driver and / or runtime 704. In at least one embodiment, API 710 is provided by a CUDA user mode driver. In at least one embodiment, API 710 is provided by a CUDA runtime. In at least one embodiment, driver 704 is data values and software instructions that, if executed, perform or otherwise facilitate operations of one or more functions 712 of API 710 during loading and execution of one or more portions of software program(s) 702. In at least one embodiment, runtime 704 is data values and software instructions that, if executed, perform or otherwise facilitate operations of one or more functions 712 of API 710 during execution of software program(s) 702. In at least one embodiment, one or more software programs 702 utilize one or more APIs 710 implemented by or otherwise provided by driver and / or runtime 704 to perform combined arithmetic operations by one or more PPU(s) (e.g., GPU(s)) during execution by said one or more software programs 702.

[0166] In at least one embodiment, one or more software programs 702 utilize one or more APIs 710 provided by driver and / or runtime 704 to perform combined arithmetic operations of one or more PPU(s) (e.g., GPU(s)). In at least one embodiment, one or more APIs 710 provide combined arithmetic operations by driver and / or runtime 704, as described above. In at least one embodiment, one or more software programs 702 utilize one or more APIs 710 provided by driver and / or runtime 704 to allocate or otherwise reserve one or more blocks of memory 714 of one or more PPU(s) (e.g., GPU(s)). In at least one embodiment, one or more software programs 702 utilize one or more APIs 710 provided by driver and / or runtime 704 to allocate or otherwise reserve blocks of memory. In at least one embodiment, one or more APIs 710 are used to perform combined mathematical functions described herein.

[0167] In at least one embodiment, to improve usability of software program 702 and / or to optimize one or more portions of software program 702 for acceleration by one or more PPUs (e.g., GPUs), one or more APIs 710 provide one or more API functions 712 to perform one or more scheduling systems usable or used by one or more computing devices as described herein. In at least one embodiment, a processor executes one or more software programs to combine two or more application programming interfaces (APIs) into a single API. In at least one embodiment, a processor uses an API to cause a scheduler to select a thread selection mechanism and / or to otherwise perform operations described herein. In at least one embodiment, an API calls a scheduler to make resource allocations. In at least one embodiment, a processor uses an example API to schedule one or more instructions for execution by one or more processors based at least in part on latency of one or more interconnects coupled to the one or more processors.

[0168] In at least one embodiment, memory 714 is system memory 3904 of computing system 3900. In at least one embodiment, memory 714 is any form of hardware that stores data and is referred to as storage or data storage. In at least one embodiment, memory 714 stores data used in various operations described herein, including processing processor settings configuration files, or other operations described herein. ​ In at least one embodiment, memory 714 stores data used in various operations described herein, including default processor settings of a processing processor settings configuration file, or other operations described herein.

[0169] In at least one embodiment, memory 714 is a computer-readable storage medium and / or a code stored in the computer-readable storage medium in a form of a computer program that includes a plurality of computer-readable instructions executable by one or more processors. In at least one embodiment, a computer-readable storage medium is a non-transitory computer-readable medium. In at least one embodiment, at least some computer-readable instructions usable to perform operations described herein are not stored solely using transitory signals (e.g., propagating transitory electric or electromagnetic transmissions). In at least one embodiment, a non-transitory computer-readable medium does not necessarily include non-transitory data storage circuitry (e.g., buffers, caches, and queues) within transitory signals receivers. In at least one embodiment, memory 714 is implemented as a non-transitory computer-readable storage medium that stores executable instructions that, if executed by one or more processors of a computer system, cause the computer system to perform one or more operations of an API to identify one or more settings to use to configure one or more processors based at least in part on one or more characteristics of a job being executed by the one or more processors, as described herein with reference to, for example, FIG. 7.​ as described, or as otherwise described herein.

[0170] ​ A call flow diagram 800 is shown of a system for identifying processor settings to be used when executing a job based on job characteristics (including job priority) in at least one embodiment. ​ One or more aspects of one or more embodiments described herein may be combined with one or more aspects of one or more embodiments described herein, including at least one aspect of one or more embodiments described herein. ​ and ​ In at least one embodiment, one or more processors perform one or more operations of system 800. In at least one embodiment, the one or more processors performing one or more operations of system 800 are any one or combination of processors described herein, including ​ processor 108, ​ Processor group 208, ​ processor 308, ​ Processor 408, ​ Processor group 508, ​ APU 3800, combined with ​ Description of the CPU 4100, combined with ​ The described graphics processor 4310, combined with ​ PPU 5000 as described, or ​ In at least one embodiment, a processor that performs one or more operations of system 800 executes ​ One or more operations of the system 100, such as the operation of the job scheduler 110. In at least one embodiment, a processor that performs one or more operations of the system 800 executes ​ One or more operations of the system 200, such as operations for identifying a processor profile using the lookup table 228. In at least one embodiment, a processor that performs one or more operations of the system 800 performs operations in conjunction with ​ One or more operations described herein, such as generating a new performance profile 330. In at least one embodiment, a processor that performs one or more operations of system 800 performs operations in conjunction with ​ One or more operations described herein, such as an operation for selecting a processor profile to be applied to processor 408. In at least one embodiment, a processor performing one or more operations of system 800 performs operations in conjunction with ​One or more operations described, such as identifying processor configuration file 510. In at least one embodiment, a processor executing one or more operations of system 800 performs one or more operations described in conjunction with ​ One or more operations described, such as selecting processor setting configuration file using operation 606. In at least one embodiment, a processor executing one or more operations of system 800 performs one or more operations described in conjunction with ​ One or more APIs, such as API 710. In at least one embodiment, a processor executing one or more operations of system 800 performs one or more operations described in conjunction with ​ One or more operations described.

[0171] In at least one embodiment, invocation flowchart 800 represents at least a portion of a system for identifying processor settings for executing a job in a data center, or other as described herein. In at least one embodiment, invocation flowchart 800 includes references that indicate which components or modules of an API cause those APIs to perform actions or operations. In at least one embodiment, each action or operation performed by an API in invocation flowchart 800 is performed by a separate API. In at least one embodiment, a reference to an API that performs an action or operation refers to one or more of those APIs that perform that action or operation. In at least one embodiment, any reference to an API that performs an action or operation refers to a processor that performs that API’s action or operation.

[0172] In at least one embodiment, user 804 invokes one or more APIs of job scheduler 810 to submit a job and information about that job, which is described elsewhere herein in conjunction with, for example, ​ In at least one embodiment, job scheduler 810 is at least a portion of one or more job schedulers described herein, such as, for example, job scheduler 110 in ​ In at least one embodiment, APIs of job scheduler 810 are job scheduler APIs 830. In at least one embodiment, one or more job scheduler APIs 830 are described in conjunction with ​ or ​One or more APIs are described. In at least one embodiment, information about a job (e.g., processor performance preference, job type, job priority, job ID, or some combination thereof) is input into a job scheduler API 830, which will be otherwise described herein. In at least one embodiment, in response to receiving the information about the job, the job scheduler API 830 causes the job scheduler to store the information about the job in a job database, as further described herein. In at least one embodiment, in response to receiving the information about the job, the job scheduler API 830 causes the job scheduler to schedule the job, in part by adding an indication of the job to a job queue. In at least one embodiment, a user does not provide a priority for a job, as the job scheduler 810 generates a priority for the job. In at least one embodiment, in response to scheduling the job, the job scheduler API sends an indication of the job scheduling or queue success back to the user 804.

[0173] In at least one embodiment, in response to scheduling the job and receiving information about the job, the job scheduler 810 executes the job scheduler API 830 to invoke one or more APIs of a resource manager 820. In at least one embodiment, the resource manager 820 is a data center processor management module, such as the data center processor management module 120 in ​ In at least one embodiment, the resource manager 820 is referred to as an RM module 820. In at least one embodiment, one or more APIs of the RM module 820 are referred to as RM APIs 832. In at least one embodiment, the job scheduler API 830 sends information about the job for receipt as input by the RM APIs 832.

[0174] In at least one embodiment, in response to receiving information about a job through RM API 832, RM module 820 performs operations of the RM API to identify, from the received job information, processor settings that are best suited to perform the job. In at least one embodiment, the identified processor settings are a selected processor settings configuration file described herein. In at least one embodiment, in response to identifying the processor settings, RM API 832 invokes job scheduler API 830 to receive an indication of those processor settings sent from RM module 820 to job scheduler 810. In at least one embodiment, in response to receiving an indication of the processor settings from RM module 820, job scheduler 810 stores the indication in a job queue such that it is associated with a particular job. In at least one embodiment, prior to launching a job, job scheduler API 830 contacts RM module 820 and invokes RM API 832 to receive or otherwise obtain an indication of processor settings to use to perform the job. In at least one embodiment, in response to receiving an indication of processor settings to use to perform a job to be launched, RM module 820 checks whether the processor settings have been set on processors that will perform the job. In at least one embodiment, if the processor settings have not been set, RM module 820 sets the processor settings or configures the processors in accordance with the processor settings. In at least one embodiment, in response to determining that the indicated processor settings have been set or the processors have been configured in accordance with the processor settings, RM API 832 sends an indication to job scheduler 810 that the processors are ready to perform the job. In at least one embodiment, upon receiving an indication that the processors are ready to perform the job, job scheduler 810 launches the job and causes the processors to perform the job.

[0175] ​ An API call is shown that, in at least one embodiment, is used at least in part to identify processor settings to use when performing a job based on job characteristics, including job priority. In at least one embodiment, in conjunction with ​ one or more aspects of one or more embodiments described herein, including at least in conjunction with ​ and ​ described aspects. In at least one embodiment, one or more processors perform one or more operations described in conjunction with ​ In at least one embodiment, one or more processors that perform one or more operations described in conjunction with ​ In at least one embodiment, one or more processors that perform one or more operations described in conjunction with ​ processor 108 of FIG. 1, ​ processor group 208 of FIG. 2,​ processor 308 of FIG. 1, ​ processor 408 of FIG. 4, ​ processor group 508 of FIG. 5, ​ APU 3800 of FIG. 38, ​ APU 3800 of FIG. 38, in conjunction with ​ CPU 4100 described in conjunction with ​ graphics processor 4310 described in conjunction with ​ PPU 5000 described in conjunction with ​ one or more SMs 5114. In at least one embodiment, a processor that performs one or more operations described in conjunction with ​ performs one or more operations of system 100 of FIG. 1, such as operations of job scheduler 110. In at least one embodiment, a processor that performs one or more operations described in conjunction with FIG. 1 performs one or more operations of system 200 of FIG. 2, such as operations to identify a processor configuration file using lookup table 228. In at least one embodiment, a processor that performs one or more operations described in conjunction with FIGS. 8A-8C performs one or more operations described in conjunction with FIG. 2 performs one or more operations described in conjunction with FIGS. 8A-8C such as generating a new performance configuration file 330. In at least one embodiment, a processor that performs one or more operations described in conjunction with FIG. 3 performs one or more operations described in conjunction with FIGS. 8A-8C such as operations to select a processor configuration file to apply to processor 408. In at least one embodiment, a processor that performs one or more operations described in conjunction with FIG. 4 performs one or more operations described in conjunction with FIGS. 8A-8C such as identifying a processor configuration file 510. In at least one embodiment, a processor that performs one or more operations described in conjunction with FIG. 5 performs one or more operations described in conjunction with FIGS. 8A-8C such as selecting a processor setting configuration file using operation 606. In at least one embodiment, a processor that performs one or more operations described in conjunction with FIG. 6 performs one or more APIs of FIG. 7, such as API 710. In at least one embodiment, a processor that performs one or more operations described in conjunction with FIGS. 8A-8C performs one or more APIs of FIG. 7, such as API 710. In at least one embodiment, a processor that performs one or more operations described in conjunction with FIG. 7 performs one or more APIs of system 800 of FIG. 8, such as job scheduler API 830. In at least one embodiment, a processor that performs one or more operations of system 800 performs one or more operations described in conjunction with FIGS. 8A-8C performs one or more operations described in conjunction with FIG. 8 performs one or more operations described in conjunction with FIGS. 9-31 performs one or more operations described in conjunction with

[0176] FIG. 8A A diagram illustrating a schedule job API call 800 is shown, in accordance with at least one embodiment. In at least one embodiment, schedule job API call 800 is a call to one or more job scheduler APIs 830 in FIG. 8 In at least one embodiment, schedule job API call 800 is used to receive (e.g., by a user, application program, or library call) one or more parameters of a job ID, preference, job type, job priority, or some combination thereof, or other as described herein. In at least one embodiment, parameters received or otherwise obtained by an API are referred to as inputs. In at least one embodiment, a job ID is any type of identifier or indication of a job, one or more instructions, a software workload, or other as described herein. In at least one embodiment, a job priority is any indication of a priority of a job, one or more instructions, or a software workload. In at least one embodiment, a preference includes an indication of a processor performance preference as described herein or other described herein. In at least one embodiment, schedule job API call 800 calls an API named srun. In at least one embodiment, schedule job API call 800 calls an API of a job scheduler that schedules jobs, but has been modified to receive parameters that a resource manager uses to configure processors, as further described herein. In at least one embodiment, parameters received by a job scheduling API are referred to as hints. In at least one embodiment, parameters received by a job scheduling API are any hints that can be used to configure a data center processor. In at least one embodiment, schedule job API call 800 calls an API to schedule a job and receives a parameter that indicates a processor performance preference. In at least one embodiment, schedule job API call 800 calls an API to schedule a job and receives a parameter that indicates a job type, which is further described herein. In at least one embodiment, schedule job API call 800 calls an API to schedule a job and receives a parameter that indicates a computing resource to use, which is further described herein. In at least one embodiment, a parameter of a computing resource to use is based on an input to the API that indicates a processor type to use, such as a processor with tensor cores or a processor with a specified amount of memory. In at least one embodiment, schedule job API call 800 calls an API to schedule a job and receives a parameter that indicates a job priority, which is further described herein. In at least one embodiment, a job priority parameter is an indication of a priority of one or more instructions or a software workload. In at least one embodiment, schedule job API call 800 is one or more API calls, where each API receives one or more inputs of a job ID, preference, job type, job priority, or some combination thereof.

[0177] In at least one embodiment, response 802 to dispatch job API call 800 includes an indication of queue success. In at least one embodiment, an indication of queue success indicates that a job has been successfully queued. In at least one embodiment, an indication of queue success indicates that input received as part of dispatch job API call 800 has been stored in a job database, or the like as described herein. In at least one embodiment, response 802 to dispatch job API call 800 includes an operation to store parameters input to a dispatch job API in a storage device, such as a job scheduler database, or the like as described herein. In at least one embodiment, response 802 to dispatch job API call 800 includes an operation performed by a processor to call an API of a resource manager to identify one or more processor performance profiles. In at least one embodiment, response 802 to dispatch job API call 800 includes an operation to call an API of a resource manager, such as a data center manager (DCM), ROCm, or a data center GPU manager (DCGM).

[0178] FIG. 8B FIG. illustrates a get all available profiles API call 804, in accordance with at least one embodiment. In at least one embodiment, get all available profiles API call 804 is a call to one or more resource manager APIs 832 in FIG. 8 In at least one embodiment, get all available profiles API call 804 is used to receive or otherwise obtain one or more parameters of a job ID, preferences, job type, job priority, or some combination thereof, or the like as described herein, for example, by a user, application, or library call. In at least one embodiment, get all available profiles API call 804 causes a processor to receive or otherwise obtain parameters as described in connection with FIG. 8A In at least one embodiment, get all available profiles API call 804 is used to receive or otherwise obtain parameters indicative of one or more processor performance preferences, job type, job priority, GPU handle, or some combination thereof. In at least one embodiment, processor performance preferences are referred to as preferences. In at least one embodiment, a GPU handle is any indication of a particular processor, which identifies a particular processor.

[0179] In at least one embodiment, get all available profiles API call 804 calls an API of a resource manager to identify one or more processor performance profiles. In at least one embodiment, get all available profiles API call 804 calls an API of a resource manager, such as a data center manager (DCM), ROCm or a data center GPU manager (DCGM).

[0180] In at least one embodiment, a response 806 to the get all available profile API call 804 includes a profile ID, a profile description, or some combination thereof. In at least one embodiment, a profile ID is an indication of a particular processor settings profile, such as a processor settings profile stored in a lookup table 428 of FIG. 4 In at least one embodiment, a profile description is text describing a processor settings profile identified by a profile ID. In at least one embodiment, a profile description includes a phrase such as maximum performance, energy efficiency, tensor core intensive, compute core intensive, memory intensive, or some combination thereof. In at least one embodiment, a response 806 is sent to a job scheduler, where the job scheduler causes a profile ID to be stored such that a particular job is associated with that profile ID, or other as described herein. In at least one embodiment, a profile ID is an indication of a processor setting sent from an RM module 820 to a job scheduler 810, as described in connection with FIG. 8 .

[0181] FIG. 8C FIG. shows a processor on-schedule set specific profile API call 808, in accordance with at least one embodiment. In at least one embodiment, a processor on-schedule set specific profile API call 808 is a call to one or more job scheduler APIs 830 in FIG. 8 In at least one embodiment, a processor on-schedule set specific profile API call 808 is used (e.g., called by a user, application, or library) to receive or otherwise obtain one or more parameters in a profile ID, a profile description, a GPU handle, or some combination thereof, or other as described herein. In at least one embodiment, a processor on-schedule set specific profile API call 808 is used (e.g., called by a user, application, or library) to receive or otherwise obtain one or more parameters to set and clear a set of one or more other parameters, such as a profile ID, a profile description, a GPU handle, or some combination thereof. In at least one embodiment, a profile ID is an indication of a processor setting sent from an RM module 820 to a job scheduler 810, as described in connection with FIG. 8 .

[0182] In at least one embodiment, the response 810 to the schedule set specific profile API call 808 on the processor includes an operation performed by the processor to set a GPU handle, a profile ID, a profile description, or some combination thereof to be associated with the job ID by storing the GPU handle, profile ID, profile description, or some combination thereof in a job queue or a data structure associated with the job queue. In at least one embodiment, the response 810 to the schedule set specific profile API call 808 on the processor includes an operation performed by the processor to clear the GPU handle, profile ID, profile description, or some combination thereof from the job queue after the corresponding job is launched. In at least one embodiment, clearing parameters means deleting or otherwise removing the parameters from the set stored as a data structure. In at least one embodiment, the response 810 to the schedule set specific profile API call 808 on the processor includes an indication to the user or application that the profile setting was successful.

[0183] FIG. 9 A block diagram of a system 900 is shown for identifying processor settings to be used when executing a job based on job characteristics, including job priority, in at least one embodiment. FIG. 9 One or more aspects of one or more embodiments described herein may be combined with one or more aspects of one or more embodiments described herein, including at least one aspect of one or more embodiments described herein. FIGS. 1-8 and FIGS. 10-31 In at least one embodiment, one or more processors perform one or more operations of system 900. In at least one embodiment, the one or more processors performing one or more operations of system 900 are any one or combination of processors described herein, including FIG. 1 processor 108, FIG. 2 Processor group 208, FIG. 3 processor 308, FIG. 4 Processor 408, FIG. 5 Processor group 508, FIG. 38 APU 3800, combined with FIG. 41A Description of the CPU 4100, combined with FIG. 43A The described graphics processor 4310, combined with FIG. 50 PPU 5000 as described, or FIG. 49 In at least one embodiment, one or more operations of the system 900 are FIG. 1 One or more operations of the system 100, such as the operation of the job scheduler 110. In at least one embodiment, one or more operations of the system 900 are FIG. 2one or more operations of system 200, e.g., operations to identify a processor profile using lookup table 228. In at least one embodiment, one or more operations of system 900 are FIG. 3 one or more operations described, such as generating a new performance profile 330. In at least one embodiment, one or more operations of system 900 are FIG. 4 one or more operations described, such as operations to select a processor profile to apply to processor 408. In at least one embodiment, one or more operations of system 900 are FIG. 5 one or more operations described, such as identifying a processor profile 510. In at least one embodiment, one or more operations of system 900 are FIG. 6 one or more operations described, such as selecting a processor setting profile using operation 606. In at least one embodiment, a processor executes FIG. 7 one or more APIs 710 to perform one or more operations of system 900. In at least one embodiment, a processor executes one or more APIs of system 800 to perform one or more operations of system 900.

[0184] In at least one embodiment, FIG. 9 A system is shown that describes balancing autonomous processor performance across a number N of nodes. In at least one embodiment, a node is any computing system, such as a server. In at least one embodiment, a node is a processor, such as a GPU. In at least one embodiment, a node is a portion of a processor, such as a partition of a GPU, which will be further described herein. In at least one embodiment, a reference to a command to a setting policy refers to configuring a processor according to a processor setting. In at least one embodiment, a reference to an RM refers to a resource manager. In at least one embodiment, an RM is a resource manager, such as FIG. 1 data center processor management module 120 of FIG. 1. In at least one embodiment, a reference to a job being executed refers to a job that a processor is launching and / or executing.

[0185] FIG. 10 A block diagram of system 1000 is shown in at least one embodiment, which is used to identify a processor setting to use when executing a job based on job characteristics, including job priority. In at least one embodiment, system 1000 is used in conjunction with FIG. 10 one or more aspects of one or more embodiments described herein are combined with one or more aspects of one or more embodiments described herein, including at least in conjunction with FIGS. 1-9 and FIGS. 11-31In at least one embodiment, one or more processors perform one or more operations of system 1000. In at least one embodiment, the one or more processors performing one or more operations of system 1000 are any one of the processors or combinations of processors described herein, including FIG. 1 processor 108, FIG. 2 Processor group 208, FIG. 3 processor 308, FIG. 4 Processor 408, FIG. 5 Processor group 508, FIG. 38 APU 3800, combined with FIG. 41A Description of the CPU 4100, combined with FIG. 43A The described graphics processor 4310, combined with FIG. 50 PPU 5000 as described, or FIG. 49 In at least one embodiment, one or more operations of system 1000 are FIG. 1 One or more operations of the system 100, such as the operation of the job scheduler 110. In at least one embodiment, one or more operations of the system 1000 are FIG. 2 One or more operations of the system 200, such as operations for identifying a processor profile using the lookup table 228. In at least one embodiment, one or more operations of the system 1000 are combined with FIG. 3 One or more operations described herein, such as generating a new performance profile 330. In at least one embodiment, one or more operations of system 1000 are combined with FIG. 4 One or more operations described herein, such as an operation for selecting a processor profile to be applied to processor 408. In at least one embodiment, one or more operations of system 1000 are combined with FIG. 5 One or more operations described herein, such as identifying a processor profile 510. In at least one embodiment, one or more operations of system 1000 are combined with FIG. 6 One or more operations described, such as selecting a processor settings profile using operation 606. In at least one embodiment, the processor performs FIG. 7 The processor executes one or more APIs 710 of the system 800 to perform one or more operations of the system 1000. In at least one embodiment, the processor executes one or more APIs of the system 800 to perform one or more operations of the system 1000.

[0186] In at least one embodiment, FIG. 10A system for selecting an optimal power and / or optimal performance profile for an application based on user input is depicted. In at least one embodiment, references to a scheduler of system 1000 refer to a job scheduler, e.g. FIG. 1 In at least one embodiment, reference to a RM refers to a resource manager, which is any combination of hardware, firmware, or software that manages the configuration and / or operation of one or more processors, or otherwise as described herein. In at least one embodiment, the RM is a data center processor management module, e.g., FIG. 1 In at least one embodiment, the RM includes any combination of hardware, firmware, or software that manages communications between data center components (e.g., between a job scheduler and a processor) using a communication protocol. In at least one embodiment, reference to the RMSMI refers to the Resource Manager System Management Interface or aspects thereof, as described herein in at least some embodiments in conjunction with FIG. 1 In at least one embodiment, reference to RMMCI refers to the Resource Manager Baseboard Management Controller Interface or aspects thereof, as described herein in conjunction with at least FIG. 1 In at least one embodiment, reference to the RM driver refers to the resource manager's processor driver, as described herein at least in conjunction with FIG. 1 The data center processor management module 120 is further described.

[0187] FIG. 11 A block diagram of a system 1100 is shown for identifying processor settings to use when executing a job based on job characteristics, including job priority, in at least one embodiment. FIG. 11 One or more aspects of one or more embodiments described herein may be combined with one or more aspects of one or more embodiments described herein, including at least one aspect of one or more embodiments described herein. FIGS. 1-10 and FIGS. 12-31 In at least one embodiment, one or more processors perform one or more operations of the system 1100. In at least one embodiment, the one or more processors performing one or more operations of the system 1100 are any one or combination of processors described herein, including FIG. 1 The processor 108 in FIG. 2 Processor group 208 in FIG. 3 The processor 308 in FIG. 4 Processor 408, FIG. 5 Processor group 508 in FIG. 38 APU 3800, combined with FIG. 41AThe described CPU 4100, in combination FIG. 43A The described graphics processor 4310, in combination FIG. 50 The described PPU 5000, or FIG. 49 one or more SMs 5114. In at least one embodiment, one or more operations of system 1100 are FIG. 1 one or more operations of system 100, e.g., operations of job scheduler 110. In at least one embodiment, one or more operations of system 1100 are FIG. 2 one or more operations of system 200, e.g., operations for identifying a processor configuration file using lookup table 228. In at least one embodiment, one or more operations of system 1100 are FIG. 3 one or more operations described in combination with FIG. 4 one or more operations described in combination with FIG. 5 one or more operations described in combination with FIG. 6 one or more operations described in combination with FIG. 7 one or more APIs 710 to perform one or more operations of system 1100. In at least one embodiment, a processor executes one or more APIs of system 800 to perform one or more operations of system 1100.

[0188] In at least one embodiment, FIG. 11 depicts a system for selecting an optimal power and / or optimal performance configuration file for an application based on user input. In at least one embodiment, system 1100 is an extension of system 1000. FIG. 10 In at least one embodiment, a reference to a target refers to a processor performance preference as described herein. In at least one embodiment, a reference to a target refers to a processor target metric as described herein.

[0189] FIG. 12 shows a block diagram for system 1200 in at least one embodiment, which is for identifying a processor setting to use when executing a job based on job characteristics, including job priority. In at least one embodiment, system 1200 is in combination with FIG. 12One or more aspects of one or more embodiments described herein may be combined with one or more aspects of one or more embodiments described herein, including at least one aspect of one or more embodiments described herein. FIGS. 1-11 and FIGS. 13-31 In at least one embodiment, one or more processors perform one or more operations of the system 1200. In at least one embodiment, the one or more processors performing one or more operations of the system 1200 are any one or combination of processors described herein, including FIG. 1 processor 108, FIG. 2 Processor group 208, FIG. 3 processor 308, FIG. 4 Processor 408, FIG. 5 Processor group 508, FIG. 38 APU 3800, combined with FIG. 41A Description of the CPU 4100, combined with FIG. 43A The described graphics processor 4310, combined with FIG. 50 PPU 5000 as described, or FIG. 49 In at least one embodiment, one or more operations of the system 1200 are FIG. 1 One or more operations of the system 100, such as the operation of the job scheduler 110. In at least one embodiment, one or more operations of the system 1200 are FIG. 2 One or more operations of the system 200, such as operations for identifying a processor profile using the lookup table 228. In at least one embodiment, one or more operations of the system 200 are combined with FIG. 3 One or more operations described herein, such as generating a new performance profile 330. In at least one embodiment, one or more operations of system 1200 are combined with FIG. 4 One or more operations described herein, such as an operation for selecting a processor profile to be applied to processor 408. In at least one embodiment, one or more operations of system 1200 are combined with FIG. 5 One or more operations described herein, such as identifying a processor profile 510. In at least one embodiment, one or more operations of the system 1200 are combined with FIG. 6 One or more operations described, such as selecting a processor settings profile using operation 606. In at least one embodiment, the processor performs FIG. 7 In at least one embodiment, the processor executes one or more APIs 710 of the system 800 to perform one or more operations of the system 1200. In at least one embodiment, the processor executes one or more APIs of the system 800 to perform one or more operations of the system 1200. FIG. 12Depicts a system of tools that schedulers use to understand power. In at least one embodiment, system 1200 is FIG. 10 System 1000 and FIG. 11 An extension of system 1100.

[0190] FIG. 13 A block diagram of a system 1300 is shown for identifying processor settings to use when executing a job based on job characteristics, including job priority, in at least one embodiment. FIG. 13 One or more aspects of one or more embodiments described herein may be combined with one or more aspects of one or more embodiments described herein, including at least one combination thereof. FIGS. 1-12 and FIGS. 14-31 In at least one embodiment, one or more processors perform one or more operations of system 1300. In at least one embodiment, the one or more processors performing one or more operations of system 1300 are any one of the processors or combinations of processors described herein, including FIG. 1 processor 108, FIG. 2 Processor group 208, FIG. 3 processor 308, FIG. 4 Processor 408, FIG. 5 Processor group 508, FIG. 38 APU 3800, combined with FIG. 41A Description of the CPU 4100, combined with FIG. 43A The described graphics processor 4310, combined with FIG. 50 PPU 5000 as described, or FIG. 49 In at least one embodiment, one or more operations of the system 1300 are FIG. 1 One or more operations of the system 100, such as the operation of the job scheduler 110. In at least one embodiment, one or more operations of the system 1300 are FIG. 2 One or more operations of the system 200, such as operations for identifying a processor profile using the lookup table 228. In at least one embodiment, one or more operations of the system 1300 are combined with FIG. 3 One or more operations described herein, such as generating a new performance profile 330. In at least one embodiment, one or more operations of system 1300 are combined with FIG. 4 One or more operations described herein, such as an operation for selecting a processor profile to be applied to processor 408. In at least one embodiment, one or more operations of system 1300 are combined with FIG. 5One or more operations described herein, such as identifying a processor profile 510. In at least one embodiment, one or more operations of system 1300 are combined with FIG. 6 One or more operations described, such as selecting a processor settings profile using operation 606. In at least one embodiment, the processor performs FIG. 7 In at least one embodiment, the processor executes one or more APIs 710 of the system 800 to perform one or more operations of the system 1300. In at least one embodiment, the processor executes one or more APIs of the system 800 to perform one or more operations of the system 1300. FIG. 13 A system for identifying a power policy to apply to individual and / or shared computing resources is described.

[0191] FIG. 14 A block diagram of a system 1400 is shown for identifying processor settings to use when executing a job based on job characteristics, including job priority, in at least one embodiment. FIG. 14 One or more aspects of one or more embodiments described herein may be combined with one or more aspects of one or more embodiments described herein, including at least one combination thereof. FIGS. 1-13 and FIGS. 15-31 In at least one embodiment, one or more processors perform one or more operations of system 1400. In at least one embodiment, the one or more processors performing one or more operations of system 1400 are any one or combination of processors described herein, including FIG. 1 processor 108, FIG. 2 Processor group 208, FIG. 3 processor 308, FIG. 4 Processor 408, FIG. 5 Processor group 508, FIG. 38 APU 3800, combined with FIG. 41A Description of the CPU 4100, combined with FIG. 43A The described graphics processor 4310, combined with FIG. 50 PPU 5000 as described, or FIG. 49 In at least one embodiment, one or more operations of the system 1400 are FIG. 1 One or more operations of the system 100, such as the operation of the job scheduler 110. In at least one embodiment, one or more operations of the system 1400 are FIG. 2 One or more operations of the system 200, such as operations for identifying a processor profile using the lookup table 228. In at least one embodiment, one or more operations of the system 1400 are combined withFIG. 3 One or more operations described, such as generating a new performance profile 330. In at least one embodiment, one or more operations of system 1400 are performed by one or more processors in combination with memory module 402, in combination with graphics processing unit 406, in combination with system interconnect 404, in combination with one or more of FIG. 4 One or more operations described, such as operations for selecting a processor profile to apply to processor 408. In at least one embodiment, one or more operations of system 1400 are performed by one or more processors in combination with memory module 402, in combination with graphics processing unit 406, in combination with system interconnect 404, in combination with one or more of FIG. 5 One or more operations described, such as identifying a processor profile 510. In at least one embodiment, one or more operations of system 1400 are performed by one or more processors in combination with memory module 402, in combination with graphics processing unit 406, in combination with system interconnect 404, in combination with one or more of FIG. 6 One or more operations described, such as selecting a processor settings profile using operation 606. In at least one embodiment, one or more processors perform one or more operations of system 1400. In at least one embodiment, one or more processors perform one or more operations of system 1400 to perform one or more operations of system 1400. In at least one embodiment, one or more processors perform one or more operations of system 1400 to perform one or more operations of system 1400. In at least one embodiment, FIG. 7 One or more APIs 710 to perform one or more operations of system 1400. In at least one embodiment, one or more processors perform one or more APIs of system 800 to perform one or more operations of system 1400. In at least one embodiment, FIG. 14 A system for identifying a power policy to apply to a single and / or shared compute resource is shown. In at least one embodiment, FIG. 14 A system for identifying a power policy based on whether one or more GPUs are running in multi-instance GPU (MIG) mode is shown, which is described further herein.

[0192] FIG. 15 A block diagram of a system 1500 in at least one embodiment for identifying processor settings to use when executing a job based on job characteristics, including job priority. In at least one embodiment, one or more aspects of one or more embodiments described herein are FIG. 15 One or more aspects of one or more embodiments described are combined with one or more aspects of one or more embodiments described herein, including at least in combination with FIGS. 1-14 and FIGS. 16-31 Aspects described. In at least one embodiment, one or more processors perform one or more operations of system 1500. In at least one embodiment, one or more processors performing one or more operations of system 1500 are any of the processors described herein or a combination of processors, including FIG. 1 Processor 108 of FIG. 1, FIG. 2 Processor group 208 of FIG. 2, FIG. 3 Processor 308 of FIG. 3, FIG. 4 Processor 408 of FIG. 4, FIG. 5 Processor group 508 of FIG. 5, FIG. 38 APU 3800 of FIG. 38, in combination with FIG. 41A CPU 4100 described in combination with FIG. 43AThe described graphics processor 4310, in conjunction FIG. 50 The described PPU 5000, or FIG. 49 one or more SMs 5114. In at least one embodiment, one or more operations of system 1500 are FIG. 1 one or more operations of system 100, such as operations of job scheduler 110. In at least one embodiment, one or more operations of system 1500 are FIG. 2 one or more operations of system 200, such as operations for identifying a processor configuration file using lookup table 228. In at least one embodiment, one or more operations of system 1500 are FIG. 3 one or more operations described in conjunction with FIG. 4 one or more operations described in conjunction with FIG. 5 one or more operations described in conjunction with FIG. 6 one or more operations described in conjunction with FIG. 7 one or more APIs 710 to perform one or more operations of system 1500. In at least one embodiment, a processor executes one or more APIs of system 800 to perform one or more operations of system 1500. In at least one embodiment, FIG. 15 depicts a system for identifying a power policy to apply to single and / or shared compute resources. In at least one embodiment, FIG. 15 depicts a system for identifying a power policy based on whether one or more GPUs are running in multi-instance GPU (MIG) mode, which is described further herein. In at least one embodiment, system 1500 includes a workflow for identifying a processor configuration file based on a single exclusive access scenario using multiple nodes. In at least one embodiment, single exclusive access to multiple nodes refers to a user having exclusive access to multiple nodes to perform a job.

[0193] FIG. 16 depicts a block diagram of system 1600 in at least one embodiment for identifying a processor setting to use when executing a job based on job characteristics, including job priority. In at least one embodiment, FIG. 16One or more aspects of one or more embodiments described herein may be combined with one or more aspects of one or more embodiments described herein, including at least one aspect of one or more embodiments described herein. FIGS. 1-15 and FIGS. 17-31 In at least one embodiment, one or more processors perform one or more operations of system 1600. In at least one embodiment, the one or more processors performing one or more operations of system 1600 are any one or combination of processors described herein, including FIG. 1 processor 108, FIG. 2 Processor group 208, FIG. 3 processor 308, FIG. 4 Processor 408, FIG. 5 Processor group 508, FIG. 38 APU 3800, combined with FIG. 41A Description of the CPU 4100, combined with FIG. 43A The described graphics processor 4310, combined with FIG. 50 PPU 5000 as described, or FIG. 49 In at least one embodiment, one or more operations of system 1600 are FIG. 1 One or more operations of the system 100, such as the operation of the job scheduler 110. In at least one embodiment, one or more operations of the system 1600 are FIG. 2 One or more operations of the system 200, such as operations for identifying a processor profile using the lookup table 228. In at least one embodiment, one or more operations of the system 1600 are combined with FIG. 3 One or more operations described herein, such as generating a new performance profile 330. In at least one embodiment, one or more operations of system 1600 are combined with FIG. 4 One or more operations described herein, such as operations for selecting a processor profile to be applied to processor 408. In at least one embodiment, one or more operations of system 1600 are combined with FIG. 5 One or more operations described herein, such as identifying a processor profile 510. In at least one embodiment, one or more operations of system 1600 are combined with FIG. 6 In at least one embodiment, the processor executes one or more APIs of system 800 to perform one or more operations of system 1600. In at least one embodiment, the processor executes one or more APIs of system 800 to perform one or more operations of system 1600. FIG. 7 One or more APIs 710 to perform one or more operations of the system 1600. In at least one embodiment, FIG. 16A system for identifying power policies to apply to individual and / or shared computing resources is depicted. In at least one embodiment, system 1600 includes a workflow for, in part, identifying whether a job requires multiple GPUs or a single GPU.

[0194] FIG. 17 A block diagram of a system 1700 is shown for identifying processor settings to use when executing a job based on job characteristics, including job priority, in at least one embodiment. FIG. 17 One or more aspects of one or more embodiments described herein may be combined with one or more aspects of one or more embodiments described herein, including at least one combination thereof. FIGS. 1-16 and FIGS. 18-31 In at least one embodiment, one or more processors perform one or more operations of system 1700. In at least one embodiment, the one or more processors performing one or more operations of system 1700 are any one or combination of processors described herein, including FIG. 1 processor 108, FIG. 2 Processor group 208, FIG. 3 processor 308, FIG. 4 Processor 408, FIG. 5 Processor group 508, FIG. 38 APU 3800, combined with FIG. 41A Description of the CPU 4100, combined with FIG. 43A The described graphics processor 4310, combined with FIG. 50 PPU 5000 as described, or FIG. 49 In at least one embodiment, one or more operations of system 1700 are FIG. 1 One or more operations of the system 100, such as the operation of the job scheduler 110. In at least one embodiment, one or more operations of the system 1700 are FIG. 2 One or more operations of the system 200, such as operations for identifying a processor profile using the lookup table 228. In at least one embodiment, one or more operations of the system 1700 are combined with FIG. 3 One or more operations described herein, such as generating a new performance profile 330. In at least one embodiment, one or more operations of system 1700 are combined with FIG. 4 One or more operations described herein, such as an operation for selecting a processor profile to be applied to processor 408. In at least one embodiment, one or more operations of system 1700 are combined with FIG. 5One or more operations described herein, such as identifying a processor profile 510. In at least one embodiment, one or more operations of system 1700 are combined with FIG. 6 One or more operations described, such as selecting a processor settings profile using operation 606. In at least one embodiment, the processor performs FIG. 7 In at least one embodiment, the processor executes one or more APIs 710 of the system 800 to perform one or more operations of the system 1700. In at least one embodiment, the processor executes one or more APIs of the system 800 to perform one or more operations of the system 1700. FIG. 17 A system is shown for identifying a processor setting profile to use a power policy based on the number of GPUs required to execute a job.

[0195] FIG. 18 A block diagram of a system 1800 is shown for identifying processor settings to use when executing a job based on job characteristics, including job priority, in at least one embodiment. FIG. 18 One or more aspects of one or more embodiments described herein may be combined with one or more aspects of one or more embodiments described herein, including at least one combination thereof. FIGS. 1-17 and FIGS. 19-31 In at least one embodiment, one or more processors perform one or more operations of system 1800. In at least one embodiment, the one or more processors performing one or more operations of system 1800 are any one or combination of processors described herein, including FIG. 1 processor 108, FIG. 2 Processor group 208, FIG. 3 processor 308, FIG. 4 Processor 408, FIG. 5 Processor group 508, FIG. 38 APU 3800, combined with FIG. 41A Description of the CPU 4100, combined with FIG. 43A The described graphics processor 4310, combined with FIG. 50 PPU 5000 as described, or FIG. 49 In at least one embodiment, one or more operations of system 1800 are FIG. 1 One or more operations of the system 100, such as the operation of the job scheduler 110. In at least one embodiment, one or more operations of the system 1800 are FIG. 2One or more operations of the system 200, such as operations for identifying a processor profile using the lookup table 228. In at least one embodiment, one or more operations of the system 1800 are combined with FIG. 3 One or more operations described herein, such as generating a new performance profile 330. In at least one embodiment, one or more operations of system 1800 are combined with FIG. 4 One or more operations described herein, such as an operation for selecting a processor profile to be applied to processor 408. In at least one embodiment, one or more operations of system 1800 are combined with FIG. 5 One or more operations described herein, such as identifying a processor profile 510. In at least one embodiment, one or more operations of system 1800 are combined with FIG. 6 One or more operations described, such as selecting a processor settings profile using operation 606. In at least one embodiment, the processor performs FIG. 7 The processor executes one or more APIs 710 of system 800 to perform one or more operations of system 1800. In at least one embodiment, the processor executes one or more APIs of system 800 to perform one or more operations of system 1800. In at least one embodiment, system 1800 includes a workflow for identifying a processor setting profile based on a single exclusive access scenario using multiple nodes.

[0196] FIG. 19 A block diagram of a system 1900 is shown for identifying processor settings to use when executing a job based on job characteristics, including job priority, in at least one embodiment. FIG. 19 One or more aspects of one or more embodiments described herein may be combined with one or more aspects of one or more embodiments described herein, including at least one combination thereof. FIGS. 1-18 and FIGS. 20-31 In at least one embodiment, one or more processors perform one or more operations of system 1900. In at least one embodiment, the one or more processors performing one or more operations of system 1900 are any one of the processors or combinations of processors described herein, including FIG. 1 processor 108, FIG. 2 Processor group 208, FIG. 3 processor 308, FIG. 4 Processor 408, FIG. 5 Processor group 508, FIG. 38 APU 3800, combined with FIG. 41A Description of the CPU 4100, combined with FIG. 43A The described graphics processor 4310, combined withFIG. 50 the described PPU 5000, or FIG. 49 one or more SMs 5114. In at least one embodiment, one or more operations of system 1900 are FIG. 1 one or more operations of system 100, such as operations of job scheduler 110. In at least one embodiment, one or more operations of system 1900 are FIG. 2 one or more operations of system 200, such as operations for identifying a processor configuration file using lookup table 228. In at least one embodiment, one or more operations of system 1900 are FIG. 3 one or more operations described in conjunction with FIG. 4 one or more operations described in conjunction with FIG. 5 one or more operations described in conjunction with FIG. 6 one or more operations described in conjunction with FIG. 7 one or more APIs 710 to perform one or more operations of system 1900. In at least one embodiment, one or more processors perform one or more APIs of system 800 to perform one or more operations of system 1900. In at least one embodiment, system 1800 includes a workflow to identify a processor settings configuration file based on a single, exclusive access scenario using multiple nodes.

[0197] FIG. 20 a block diagram of system 2000 is shown in at least one embodiment, which is used to identify processor settings to use when executing a job based on job characteristics, including job priority. In at least one embodiment, one or more aspects of one or more embodiments described herein are combined with one or more aspects of one or more embodiments described herein, including at least aspects described in conjunction with FIG. 20 one or more aspects of one or more embodiments described herein are combined with one or more aspects of one or more embodiments described herein, including at least aspects described in conjunction with FIGS. 1-19 and FIGS. 21-31 In at least one embodiment, one or more processors perform one or more operations of system 2000. In at least one embodiment, one or more processors performing one or more operations of system 2000 are any of the processors or combinations of processors described herein, including FIG. 1 processor 108 of system 100,FIG. 2 the processor group 208 of FIG. 2, FIG. 3 the processor 308 of FIG. 3, FIG. 4 the processor 408 of FIG. 4, FIG. 5 the processor group 508 of FIG. 5, FIG. 38 the APU 3800 of FIG. 38, in conjunction with FIG. 41A the CPU 4100 described in conjunction with FIG. 43A the graphics processor 4310 described in conjunction with FIG. 50 the PPU 5000 described in conjunction with FIG. 49 one or more SMs 5114. In at least one embodiment, one or more operations of system 2000 are FIG. 1 one or more operations of system 100, such as operations of job scheduler 110. In at least one embodiment, one or more operations of system 2000 are FIG. 2 one or more operations of system 200, such as operations for identifying a processor configuration file using lookup table 228. In at least one embodiment, one or more operations of system 2000 are FIG. 3 one or more operations described in conjunction with FIG. 4 one or more operations described in conjunction with FIG. 5 one or more operations described in conjunction with FIG. 6 one or more operations described in conjunction with FIG. 7 one or more APIs 710 of FIG. 7 to perform one or more operations of system 2000. In at least one embodiment, a processor performs one or more APIs of system 800 to perform one or more operations of system 2000. In at least one embodiment, system 2000 includes multiple GPUs and telemetry information that is used to inform a data center administrator which jobs should be concentrated on the same busbar and which jobs should not be concentrated on the busbar. In at least one embodiment, concentrating jobs on the same busbar means connecting nodes that perform the jobs to the same busbar. In at least one embodiment, a busbar is an electrically conductive connection that distributes power to connected nodes.

[0198] FIG. 21A block diagram of a system 2100 is shown for identifying processor settings to use when executing a job based on job characteristics, including job priority, in at least one embodiment. FIG. 21 One or more aspects of one or more embodiments described herein may be combined with one or more aspects of one or more embodiments described herein, including at least one aspect of one or more embodiments described herein. FIGS. 1-20 and FIGS. 22-31 In at least one embodiment, one or more processors perform one or more operations of the system 2100. In at least one embodiment, the one or more processors performing one or more operations of the system 2100 are any one or combination of processors described herein, including FIG. 1 processor 108, FIG. 2 Processor group 208, FIG. 3 processor 308, FIG. 4 Processor 408, FIG. 5 Processor group 508, FIG. 38 APU 3800, combined with FIG. 41A Description of the CPU 4100, combined with FIG. 43A The described graphics processor 4310, combined with FIG. 50 PPU 5000 as described, or FIG. 49 In at least one embodiment, one or more operations of the system 2100 are FIG. 1 One or more operations of the system 100, such as the operation of the job scheduler 110. In at least one embodiment, one or more operations of the system 2100 are FIG. 2 One or more operations of the system 2100, such as operations for identifying a processor profile using the lookup table 228. In at least one embodiment, one or more operations of the system 2100 are combined with FIG. 3 One or more operations described herein, such as generating a new performance profile 330. In at least one embodiment, one or more operations of the system 2100 are combined with FIG. 4 One or more operations described herein, such as an operation for selecting a processor profile to be applied to processor 408. In at least one embodiment, one or more operations of system 2100 are combined with FIG. 5 One or more operations described herein, such as identifying a processor profile 510. In at least one embodiment, one or more operations of the system 2100 are combined with FIG. 6 One or more operations described, such as selecting a processor settings profile using operation 606. In at least one embodiment, the processor performs FIG. 7The processor executes one or more APIs 710 of the system 800 to perform one or more operations of the system 2100. In at least one embodiment, the processor executes one or more APIs of the system 800 to perform one or more operations of the system 2100. In at least one embodiment, the system 2100 includes multiple GPUs and telemetry information that is used to inform a data center administrator which jobs are worse in terms of current change rate over a short period of time.

[0199] FIG. 22 A block diagram of a system 2200 is shown for identifying processor settings to use when executing a job based on job characteristics, including job priority, in at least one embodiment. FIG. 22 One or more aspects of one or more embodiments described herein may be combined with one or more aspects of one or more embodiments described herein, including at least one aspect of one or more embodiments described herein. FIGS. 1-21 and FIGS. 23-31 In at least one embodiment, one or more processors perform one or more operations of the system 2200. In at least one embodiment, the one or more processors performing one or more operations of the system 2200 are any one or combination of processors described herein, including FIG. 1 processor 108, FIG. 2 Processor group 208, FIG. 3 processor 308, FIG. 4 Processor 408, FIG. 5 Processor group 508, FIG. 38 APU 3800, combined with FIG. 41A Description of the CPU 4100, combined with FIG. 43A The described graphics processor 4310, combined with FIG. 50 PPU 5000 as described, or FIG. 49 In at least one embodiment, one or more operations of the system 2200 are FIG. 1 One or more operations of the system 100, such as the operation of the job scheduler 110. In at least one embodiment, one or more operations of the system 2200 are FIG. 2 One or more operations of the system 2200, such as operations for identifying a processor profile using the lookup table 228. In at least one embodiment, one or more operations of the system 2200 are combined with FIG. 3 One or more operations described herein, such as generating a new performance profile 330. In at least one embodiment, one or more operations of the system 2200 are combined with FIG. 4One or more operations described, such as operations to select a processor configuration file to apply to processor 408. In at least one embodiment, one or more operations of system 2200 are performed by one or more processors executing one or more APIs 710 to perform one or more operations of system 2200. In at least one embodiment, one or more processors execute one or more APIs of system 800 to perform one or more operations of system 2200. In at least one embodiment, system 2200 includes a workflow to identify a processor setting configuration file based on a single, exclusive access scenario using multiple nodes. In at least one embodiment, system 2200 takes into account whether RMs can have stickiness, which refers to a functionality that allows requests from clients or users to be repeatedly routed to the same node, so that data integrity can be maintained. FIG. 5 One or more operations described, such as identifying a processor configuration file 510. In at least one embodiment, one or more operations of system 2200 are performed by one or more processors executing one or more APIs 710 to perform one or more operations of system 2200. In at least one embodiment, one or more processors execute one or more APIs of system 800 to perform one or more operations of system 2200. In at least one embodiment, system 2200 includes a workflow to identify a processor setting configuration file based on a single, exclusive access scenario using multiple nodes. In at least one embodiment, system 2200 takes into account whether RMs can have stickiness, which refers to a functionality that allows requests from clients or users to be repeatedly routed to the same node, so that data integrity can be maintained. FIG. 6 One or more operations described, such as selecting a processor setting configuration file using operation 606. In at least one embodiment, one or more processors execute one or more APIs 710 to perform one or more operations of system 2200. In at least one embodiment, one or more processors execute one or more APIs of system 800 to perform one or more operations of system 2200. In at least one embodiment, system 2200 includes a workflow to identify a processor setting configuration file based on a single, exclusive access scenario using multiple nodes. In at least one embodiment, system 2200 takes into account whether RMs can have stickiness, which refers to a functionality that allows requests from clients or users to be repeatedly routed to the same node, so that data integrity can be maintained. FIG. 7 One or more operations described, such as selecting a processor setting configuration file using operation 606. In at least one embodiment, one or more processors execute one or more APIs 710 to perform one or more operations of system 2200. In at least one embodiment, one or more processors execute one or more APIs of system 800 to perform one or more operations of system 2200. In at least one embodiment, system 2200 includes a workflow to identify a processor setting configuration file based on a single, exclusive access scenario using multiple nodes. In at least one embodiment, system 2200 takes into account whether RMs can have stickiness, which refers to a functionality that allows requests from clients or users to be repeatedly routed to the same node, so that data integrity can be maintained.

[0200] FIG. 23 A block diagram of a system 2300 is shown in at least one embodiment, which is used to identify a processor setting to use when executing a job based on job characteristics, including job priority. In at least one embodiment, system 2300 is used in conjunction with FIG. 23 One or more aspects of one or more embodiments described include one or more aspects of one or more embodiments described herein, including at least one or more aspects of one or more embodiments described in connection with FIGS. 1-22 and Figures 24-31 Aspects described. In at least one embodiment, one or more processors perform one or more operations of system 2300. In at least one embodiment, one or more processors performing one or more operations of system 2300 are any one or combination of processors described herein, including Figure 1 processor 108 of FIG. 1, Figure 2 processor group 208 of FIG. 2, Figure 3 processor 308 of FIG. 3, Figure 4 processor 408 of FIG. 4, Figure 5 processor group 508 of FIG. 5, Figure 38 APU 3800 of FIG. 38, in conjunction with Figure 41A CPU 4100 described in connection with Figure 43A graphics processor 4310 described in connection with Figure 50 PPU 5000 described in connection with Figure 49 one or more SMs 5114 of FIG. 51. In at least one embodiment, one or more operations of system 2300 areFigure 1 one or more operations of system 100, such as operations of job scheduler 110. In at least one embodiment, one or more operations of system 2300 are Figure 2 one or more operations of system 200, such as operations for identifying a processor configuration file using lookup table 228. In at least one embodiment, one or more operations of system 2300 are Figure 3 one or more operations described in conjunction with Figure 4 one or more operations described in conjunction with Figure 5 one or more operations described in conjunction with Figure 6 one or more operations described in conjunction with Figure 7 one or more APIs 710 of system 800 to perform one or more operations of system 2300. In at least one embodiment, one or more processors perform one or more APIs of system 800 to perform one or more operations of system 2300. In at least one embodiment, system 2300 includes a workflow to identify a processor settings configuration file based on a single, exclusive access scenario using multiple nodes.

[0201] Figure 24 a block diagram of a system 2400 to identify processor settings to use when executing a job based on job characteristics, including job priority, in at least one embodiment. In at least one embodiment, one or more aspects of one or more embodiments described herein are Figure 24 one or more aspects of one or more embodiments described herein, including at least in conjunction with aspects described in Figures 1-23 and Figures 25-31 In at least one embodiment, one or more processors perform one or more operations of system 2400. In at least one embodiment, one or more processors performing one or more operations of system 2400 are any of the processors or combinations of processors described herein, including Figure 1 processor 108 of system 100, Figure 2 processor group 208 of system 200, Figure 3 processor 308 of system 300, Figure 4 processor 408 of system 400, Figure 5 processor group 508 of system 500,Figure 38 APU 3800, in combination with Figure 41A CPU 4100, as described in Figure 43A graphics processor 4310, as described in Figure 50 PPU 5000, or Figure 49 one or more SMs 5114. In at least one embodiment, one or more operations of system 2400 are Figure 1 one or more operations of system 100, such as operations of job scheduler 110. In at least one embodiment, one or more operations of system 2400 are Figure 2 one or more operations of system 200, such as operations for identifying a processor configuration file using lookup table 228. In at least one embodiment, one or more operations of system 2400 are Figure 3 one or more operations described in combination with Figure 4 one or more operations described in combination with Figure 5 one or more operations described in combination with Figure 6 one or more operations described in combination with Figure 7 one or more APIs 710 to perform one or more operations of system 2400. In at least one embodiment, a processor executes one or more APIs of system 800 to perform one or more operations of system 2400. In at least one embodiment, system 2400 includes a workflow for identifying a processor setting configuration file based on a single, exclusive access scenario using multiple nodes. In at least one embodiment, system 2200 takes into account whether an RM can have stickiness, which refers to a functionality that allows requests from a client or user to be repeatedly routed to the same node, so that data integrity can be maintained.

[0202] Figure 25 illustrates a block diagram of a system 2500 for identifying processor settings to use when executing a job based on job characteristics, including job priority, in at least one embodiment. In at least one embodiment, one or more aspects of one or more embodiments described in combination with Figure 25 one or more aspects of one or more embodiments described herein, including at least in combination with Figures 1-24 andFigures 26-31 In at least one embodiment, one or more processors perform one or more operations of the system 2500. In at least one embodiment, the one or more processors performing one or more operations of the system 2500 are any one or combination of processors described herein, including Figure 1 processor 108, Figure 2 Processor group 208, Figure 3 processor 308, Figure 4 Processor 408, Figure 5 Processor group 508, Figure 38 APU 3800, combined with Figure 41A Description of the CPU 4100, combined with Figure 43A The described graphics processor 4310, combined with Figure 50 PPU 5000 as described, or Figure 49 In at least one embodiment, one or more operations of the system 2500 are Figure 1 One or more operations of the system 100, such as the operation of the job scheduler 110. In at least one embodiment, one or more operations of the system 2500 are Figure 2 One or more operations of the system 2500, such as operations for identifying a processor profile using the lookup table 228. In at least one embodiment, one or more operations of the system 2500 are combined with Figure 3 One or more operations described herein, such as generating a new performance profile 330. In at least one embodiment, one or more operations of the system 2500 are combined with Figure 4 One or more operations described herein, such as an operation for selecting a processor profile to be applied to processor 408. In at least one embodiment, one or more operations of system 2500 are combined with Figure 5 One or more operations described, such as identifying a processor profile 510. In at least one embodiment, one or more operations of the system 2500 are combined with Figure 6 One or more operations described, such as selecting a processor settings profile using operation 606. In at least one embodiment, the processor performs Figure 7 The processor executes one or more APIs 710 of system 800 to perform one or more operations of system 2500. In at least one embodiment, the processor executes one or more APIs of system 800 to perform one or more operations of system 2500. In at least one embodiment, system 2500 includes a workflow for identifying a processor setting profile based on a single exclusive access scenario using multiple nodes.

[0203] Figure 26A block diagram of a system 2600 is shown for identifying processor settings to use when executing a job based on job characteristics, including job priority, in at least one embodiment. Figure 26 One or more aspects of one or more embodiments described herein may be combined with one or more aspects of one or more embodiments described herein, including at least one aspect of one or more embodiments described herein. Figures 1-15 and Figures 27-31 In at least one embodiment, one or more processors perform one or more operations of the system 2600. In at least one embodiment, the one or more processors performing one or more operations of the system 2600 are any one or combination of processors described herein, including Figure 1 processor 108, Figure 2 Processor group 208, Figure 3 processor 308, Figure 4 Processor 408, Figure 5 Processor group 508, Figure 38 APU 3800, combined with Figure 41A Description of the CPU 4100, combined with Figure 43A The described graphics processor 4310, combined with Figure 50 PPU 5000 as described, or Figure 49 In at least one embodiment, one or more operations of the system 2600 are Figure 1 One or more operations of the system 100, such as the operation of the job scheduler 110. In at least one embodiment, one or more operations of the system 2600 are Figure 2 One or more operations of the system 2600, such as operations for identifying a processor profile using the lookup table 228. In at least one embodiment, one or more operations of the system 2600 are combined with Figure 3 One or more operations described herein, such as generating a new performance profile 330. In at least one embodiment, one or more operations of the system 2600 are combined with Figure 4 One or more operations described herein, such as an operation for selecting a processor profile to be applied to processor 408. In at least one embodiment, one or more operations of system 2600 are combined with Figure 5 One or more operations described herein, such as identifying a processor profile 510. In at least one embodiment, one or more operations of the system 2600 are combined with Figure 6 One or more operations described, such as selecting a processor settings profile using operation 606. In at least one embodiment, the processor performs Figure 7The processor executes one or more APIs 710 of system 800 to perform one or more operations of system 2600. In at least one embodiment, the processor executes one or more APIs of system 800 to perform one or more operations of system 2600. In at least one embodiment, system 2600 includes a workflow for identifying a processor setting profile based on a shared access scenario using a single node or multiple nodes.

[0204] Figure 27 A block diagram of a system 2700 is shown for identifying processor settings to use when executing a job based on job characteristics, including job priority, in at least one embodiment. Figure 27 One or more aspects of one or more embodiments described herein may be combined with one or more aspects of one or more embodiments described herein, including at least one aspect of one or more embodiments described herein. Figures 1-26 and Figures 28-31 In at least one embodiment, one or more processors perform one or more operations of the system 2700. In at least one embodiment, the one or more processors performing one or more operations of the system 2700 are any one or combination of processors described herein, including Figure 1 processor 108, Figure 2 Processor group 208, Figure 3 processor 308, Figure 4 Processor 408, Figure 5 Processor group 508, Figure 38 APU 3800, combined with Figure 41A Description of the CPU 4100, combined with Figure 43A The described graphics processor 4310, combined with Figure 50 PPU 5000 as described, or Figure 49 In at least one embodiment, one or more operations of the system 2700 are Figure 1 One or more operations of the system 100, such as the operation of the job scheduler 110. In at least one embodiment, one or more operations of the system 2700 are Figure 2 One or more operations of the system 2700, such as operations for identifying a processor profile using the lookup table 228. In at least one embodiment, one or more operations of the system 2700 are combined with Figure 3 One or more operations described herein, such as generating a new performance profile 330. In at least one embodiment, one or more operations of the system 2700 are combined with Figure 4One or more operations described, such as operations to select a processor configuration file to apply to processor 408. In at least one embodiment, one or more operations of system 2700 are performed by one or more processors executing one or more APIs 710 to perform one or more operations of system 2700. In at least one embodiment, one or more processors execute one or more APIs of system 800 to perform one or more operations of system 2700. In at least one embodiment, system 2700 includes a workflow to identify a processor setting configuration file based on GPUs running in MIG mode, where these GPUs are capable of being partitioned or other as described herein. Figure 5 One or more operations described, such as identifying processor configuration file 510. In at least one embodiment, one or more operations of system 2700 are performed by one or more processors executing one or more APIs 710 to perform one or more operations of system 2700. In at least one embodiment, one or more processors execute one or more APIs of system 800 to perform one or more operations of system 2700. In at least one embodiment, system 2700 includes a workflow to identify a processor setting configuration file based on GPUs running in MIG mode, where these GPUs are capable of being partitioned or other as described herein. Figure 6 One or more operations described, such as selecting a processor setting configuration file using operation 606. In at least one embodiment, one or more processors execute one or more APIs 710 to perform one or more operations of system 2700. In at least one embodiment, one or more processors execute one or more APIs of system 800 to perform one or more operations of system 2700. In at least one embodiment, system 2700 includes a workflow to identify a processor setting configuration file based on GPUs running in MIG mode, where these GPUs are capable of being partitioned or other as described herein. Figure 7 One or more APIs 710 to perform one or more operations of system 2700. In at least one embodiment, one or more processors execute one or more APIs of system 800 to perform one or more operations of system 2700. In at least one embodiment, system 2700 includes a workflow to identify a processor setting configuration file based on GPUs running in MIG mode, where these GPUs are capable of being partitioned or other as described herein.

[0205] Figure 28 A block diagram of system 2800 is shown in at least one embodiment, which is used to identify a processor setting to use when executing a job based on job characteristics, including job priority. In at least one embodiment, one or more operations of system 2800 are performed by one or more processors executing one or more APIs 710 to perform one or more operations of system 2800. In at least one embodiment, one or more processors execute one or more APIs of system 800 to perform one or more operations of system 2800. In at least one embodiment, system 2800 includes a workflow to identify a processor setting configuration file based on GPUs running in MIG mode, where these GPUs are capable of being partitioned or other as described herein. Figure 28 One or more aspects of one or more embodiments described herein are described in combination with one or more aspects of one or more embodiments described herein, including at least in combination with Figures 1-27 and Figures 28-31 Aspects described. In at least one embodiment, one or more processors perform one or more operations of system 2800. In at least one embodiment, one or more processors performing one or more operations of system 2800 are any one processor or combination of processors described herein, including Figure 1 Processor 108 of FIG. 1, Figure 2 Processor group 208 of FIG. 2, Figure 3 Processor 308 of FIG. 3, Figure 4 Processor 408 of FIG. 4, Figure 5 Processor group 508 of FIG. 5, Figure 38 APU 3800 of FIG. 38, in combination with Figure 41A CPU 4100 described in combination with Figure 43A Graphics processor 4310 described in combination with Figure 50 PPU 5000 described in combination with Figure 49 One or more SMs 5114 of FIG. 51. In at least one embodiment, one or more operations of system 2800 are Figure 1one or more operations of system 100, such as operations of job scheduler 110. In at least one embodiment, one or more operations of system 2800 are Figure 2 one or more operations of system 200, such as operations for identifying a processor configuration file using lookup table 228. In at least one embodiment, one or more operations of system 2800 are Figure 3 one or more operations described in conjunction with, such as generating a new performance configuration file 330. In at least one embodiment, one or more operations of system 2800 are Figure 4 one or more operations described in conjunction with, such as operations for selecting a processor configuration file to apply to processor 408. In at least one embodiment, one or more operations of system 2800 are Figure 5 one or more operations described in conjunction with, such as identifying a processor configuration file 510. In at least one embodiment, one or more operations of system 2800 are Figure 6 one or more operations described in conjunction with, such as selecting a processor setting configuration file using operation 606. In at least one embodiment, a processor executes Figure 7 one or more APIs 710 of system 2800 to perform one or more operations of system 2800. In at least one embodiment, a processor executes one or more APIs of system 800 to perform one or more operations of system 2800. In at least one embodiment, system 2800 includes a workflow for identifying a processor setting configuration file based on identifying a processor setting configuration file based on using a shared access scenario with a single node or multiple nodes, while accounting for the possibility that these nodes can be able to run in MIG mode. In at least one embodiment, the identified processor setting configuration file is a MIG-based performance configuration file.

[0206] Figure 29 illustrates a block diagram of a system 2900 for identifying processor settings to use when executing a job based on job characteristics, including job priority, in at least one embodiment. In at least one embodiment, one or more aspects of one or more embodiments described in conjunction with Figure 29 include at least aspects described in conjunction with Figures 1-28 and Figures 29-31 In at least one embodiment, one or more processors perform one or more operations of system 2900. In at least one embodiment, one or more processors performing one or more operations of system 2900 are any of the processors described herein, or a combination of processors, including Figure 1 processor 108 of system 100, Figure 2 processor group 208 of system 200, Figure 3 processor 308 of system 300,Figure 4 Processor 408, Figure 5 Processor group 508, Figure 38 APU 3800, combined with Figure 41A Description of the CPU 4100, combined with Figure 43A The described graphics processor 4310, combined with Figure 50 PPU 5000 as described, or Figure 49 In at least one embodiment, one or more operations of the system 2900 are Figure 1 One or more operations of the system 100, such as the operation of the job scheduler 110. In at least one embodiment, one or more operations of the system 2900 are Figure 2 One or more operations of the system 2900, such as operations for identifying a processor profile using the lookup table 228. In at least one embodiment, one or more operations of the system 2900 are combined with Figure 3 One or more operations described herein, such as generating a new performance profile 330. In at least one embodiment, one or more operations of the system 2900 are combined with Figure 4 One or more operations described herein, such as an operation for selecting a processor profile to be applied to processor 408. In at least one embodiment, one or more operations of system 2900 are combined with Figure 5 One or more operations described, such as identifying a processor profile 510. In at least one embodiment, one or more operations of the system 2900 are combined with Figure 6 One or more operations described, such as selecting a processor settings profile using operation 606. In at least one embodiment, the processor performs Figure 7 In at least one embodiment, the processor executes one or more APIs 710 of the system 2900 to perform one or more operations of the system 2900. In at least one embodiment, the processor executes one or more APIs of the system 800 to perform one or more operations of the system 2900. In at least one embodiment, the system 2900 includes a workflow for identifying a processor setting profile based on a shared access scenario using a single node or multiple nodes. In at least one embodiment, the scheduler searches a list of processor setting profiles and selects a processor setting profile to apply to the processor assigned to execute the job.

[0207] Figure 30 A block diagram of a system 3000 is shown for identifying processor settings to use when executing a job based on job characteristics, including job priority, in at least one embodiment. Figure 30One or more aspects of one or more embodiments described herein may be combined with one or more aspects of one or more embodiments described herein, including at least one aspect of one or more embodiments described herein. Figures 1-29 and Figure 31 In at least one embodiment, one or more processors perform one or more operations of the system 3000. In at least one embodiment, the one or more processors performing one or more operations of the system 3000 are any one or combination of processors described herein, including Figure 1 processor 108, Figure 2 Processor group 208, Figure 3 processor 308, Figure 4 Processor 408, Figure 5 Processor group 508, Figure 38 APU 3800, combined with Figure 41A Description of the CPU 4100, combined with Figure 43A The described graphics processor 4310, combined with Figure 50 PPU 5000 as described, or Figure 49 In at least one embodiment, one or more operations of the system 3000 are Figure 1 One or more operations of the system 100, such as the operation of the job scheduler 110. In at least one embodiment, one or more operations of the system 3000 are Figure 2 One or more operations of the system 200, such as operations for identifying a processor profile using the lookup table 228. In at least one embodiment, one or more operations of the system 3000 are combined with Figure 3 One or more operations described herein, such as generating a new performance profile 330. In at least one embodiment, one or more operations of system 3000 are combined with Figure 4 One or more operations described herein, such as an operation for selecting a processor profile to be applied to processor 408. In at least one embodiment, one or more operations of system 3000 are combined with Figure 5 One or more operations described herein, such as identifying a processor profile 510. In at least one embodiment, one or more operations of system 3000 are combined with Figure 6 One or more operations described herein, such as selecting a processor settings profile using operation 606. In at least one embodiment, a processor that performs one or more operations of system 3000 executes Figure 7In at least one embodiment, the processor executes one or more APIs of system 800 to perform one or more operations of system 3000. In at least one embodiment, system 3000 includes a workflow for identifying a processor setting profile based on a single exclusive access scenario using multiple nodes. In at least one embodiment, system 3000 identifies an optimal power performance profile for a job. In at least one embodiment, system 3000 is Figure 29 An expansion of System 2900.

[0208] Figure 31 A block diagram of a system 3100 is shown for identifying processor settings to use when executing a job based on job characteristics, including job priority, in at least one embodiment. Figure 31 One or more aspects of one or more embodiments described herein may be combined with one or more aspects of one or more embodiments described herein, including at least one aspect of one or more embodiments described herein. Figures 1-30 In at least one embodiment, one or more processors perform one or more operations of the system 3100. In at least one embodiment, the one or more processors performing one or more operations of the system 3100 are any one or combination of processors described herein, including Figure 1 processor 108, Figure 2 Processor group 208, Figure 3 processor 308, Figure 4 Processor 408, Figure 5 Processor group 508, Figure 38 APU 3800, combined with Figure 41A Description of the CPU 4100, combined with Figure 43A The described graphics processor 4310, combined with Figure 50 PPU 5000 as described, or Figure 49 In at least one embodiment, one or more operations of the system 3100 are Figure 1 One or more operations of the system 100, such as the operation of the job scheduler 110. In at least one embodiment, one or more operations of the system 3100 are Figure 2 One or more operations of the system 200, such as operations for identifying a processor profile using the lookup table 228. In at least one embodiment, one or more operations of the system 3100 are combined with Figure 3 One or more operations described herein, such as generating a new performance profile 330. In at least one embodiment, one or more operations of the system 3100 are combined with Figure 4One or more operations described, such as operations to select a processor configuration file to apply to processor 408. In at least one embodiment, one or more operations of system 3100 are performed by a processor executing instructions of Figure 5 One or more operations described, such as identifying processor configuration file 510. In at least one embodiment, one or more operations of system 3100 are performed by a processor executing instructions of Figure 6 One or more operations described, such as selecting a processor setting configuration file using operation 606. In at least one embodiment, a processor performing one or more operations of system 3100 executes one or more operations of API 710 of Figure 7 In at least one embodiment, a processor executing system 800 performs one or more APIs to perform one or more operations of system 3100. In at least one embodiment, system 3100 includes a tool for a scheduler to understand power. In at least one embodiment, system 3100 includes a workflow to identify node allocations in a single exclusive access scenario using multiple nodes. In at least one embodiment, system 3100 is an extension of system 2900 of Figure 29 and / or system 3000 of Figure 30 In at least one embodiment, system 3100 is an extension of system 2900 of

[0209] Data Center

[0210] Figure 32 FIG. 32 illustrates an example data center 3200, in accordance with at least one embodiment. In at least one embodiment, data center 3200 includes, without limitation, a data center infrastructure layer 3210, a framework layer 3220, a software layer 3230, and an application layer 3240.

[0211] In at least one embodiment, data center 3200 is configurable to execute a schedule job API to receive, as input, an indication of a processor performance preference, a job type, a job priority, or some combination thereof, to cause a processor to be configured to run at a clock frequency, as described in connection with Figures 8A-8C In at least one embodiment, data center 3200 is configurable to execute a schedule job API to receive, as input, an indication of a software instruction to use a computing resource. In at least one embodiment, data center 3200 is configurable to execute a schedule job API to receive, as input, an indication of a priority of an execution instruction.

[0212] In at least one embodiment, the data center 3200 may be configured to execute a Get All Available Profiles API to receive as input an indication of a processor performance preference, a frequency input, a job type, a computing resource input, a priority input, a processor identifier, or some combination thereof, thereby identifying a processor settings profile, or as otherwise described herein. In at least one embodiment, the data center 3200 may be configured to execute a Get All Available Profiles API to receive as input an indication of a processor performance preference, a frequency input, a computing resource input, a priority input, a processor identifier, or some combination thereof, thereby identifying a processor settings profile, or as otherwise described herein.

[0213] In at least one embodiment, the data center 3200 may be configured to execute a scheduling set specific profile API on a processor to receive an indication of a processor setting profile so that instructions are executed by a processor configured according to the processor setting profile, or otherwise as described herein.

[0214] In at least one embodiment, Figure 32 As shown, the data center infrastructure layer 3210 may include a resource coordinator 3212, grouped computing resources 3214, and node computing resources ("node CRs") 3216(1)-3216(N), where "N" represents any complete positive integer. In at least one embodiment, the node CRs 3216(1)-3216(N) may include, but are not limited to, any number of central processing units ("CPUs") or other processors (including accelerators, field programmable gate arrays ("FPGAs"), data processing units ("DPUs") in network devices, graphics processors, etc.), memory devices (e.g., dynamic read-only memories), storage devices (e.g., solid-state drives or disk drives), network input / output ("NW I / O") devices, network switches, virtual machines ("VMs"), power modules, and cooling modules, etc. In at least one embodiment, one or more of the node CRs 3216(1)-3216(N) may be servers having one or more of the above-mentioned computing resources.

[0215] In at least one embodiment, Figure 32 At least one component shown or described is used to implement the combination Figures 1-33 In at least one embodiment, a node computing resource ("node CR") 716(1)-716(N) performs one or more operations of an API to identify one or more settings to be used to configure one or more processors based at least in part on one or more characteristics of a job to be executed by the one or more processors, such as in conjunction with Figure 1 as otherwise described herein.

[0216] In at least one embodiment, the grouped computing resources 3214 may include separate groups of node CRs housed in one or more racks (not shown), or may be housed in many racks (also not shown) in data centers at various geographic locations. The separate groups of node CRs within the grouped computing resources 3214 may include computing, networking, memory, or storage resources that may be configured or allocated to support groupings of one or more workloads. In at least one embodiment, several node CRs including a CPU or processor may be grouped in one or more racks to provide computing resources to support one or more workloads. In at least one embodiment, one or more racks may also include any number of power modules, cooling modules, and network switches, in any combination.

[0217] In at least one embodiment, resource coordinator 3212 may configure or otherwise control one or more nodes CR 3216(1)-3216(N) and / or grouped computing resources 3214. In at least one embodiment, resource coordinator 3212 may comprise a software design infrastructure ("SDI") management entity for data center 3200. In at least one embodiment, resource coordinator 3212 may comprise hardware, software, or some combination thereof.

[0218] In at least one embodiment, Figure 32 As shown, the framework layer 3220 includes, but is not limited to, a job scheduler 3232, a configuration manager 3234, a resource manager 3236, and a distributed file system 3238. In at least one embodiment, the framework layer 3220 may include a framework that supports software 3252 of the software layer 3230 and / or one or more applications 3242 of the application layer 3240. In at least one embodiment, the software 3252 or the application 3242 may include web-based service software or applications, such as those provided by Amazon Web Services, Google Cloud, and Microsoft Azure. In at least one embodiment, the framework layer 3220 may be, but is not limited to, a free and open source software web application framework, such as Apache Spark, which may utilize the distributed file system 3238 for large-scale data processing (e.g., "big data"). TM(“Spark”). In at least one embodiment, the job scheduler 3232 can include a Spark driver to facilitate scheduling of workloads supported by various layers of the datacenter 3200. In at least one embodiment, the configuration manager 3234 can be capable of configuring different layers, such as the software layer 3230 and the framework layer 3220 including Spark and a distributed file system 3238 for supporting large-scale data processing. In at least one embodiment, the resource manager 3236 can be capable of managing clustered or grouped computing resources mapped to or allocated for supporting the distributed file system 3238 and the job scheduler 3232. In at least one embodiment, the clustered or grouped computing resources can include grouped computing resources 3214 on the datacenter infrastructure layer 3210. In at least one embodiment, the resource manager 3236 can coordinate with the resource orchestrator 3212 to manage these mapped or allocated computing resources.

[0219] In at least one embodiment, software 3252 included in the software layer 3230 can include software used by at least a portion of the node C.R.s 3216(1)- 3216(N), the grouped computing resources 3214, and / or the distributed file system 3238 of the framework layer 3220. One or more types of software can include, but are not limited to, Internet web page search software, email virus scanning software, database software, and streaming video content software.

[0220] In at least one embodiment, one or more application programs 3242 included in the application layer 3240 can include one or more types of application programs used by at least a portion of the node C.R.s 3216(1)-3216(N), the grouped computing resources 3214, and / or the distributed file system 3238 of the framework layer 3220. In at least one embodiment, one or more types of application programs can include, but are not limited to, CUDA application programs.

[0221] In at least one embodiment, any of the configuration manager 3234, the resource manager 3236, and the resource orchestrator 3212 can implement any number and type of self-modifying actions based on any number and type of data acquired in any technically feasible fashion. In at least one embodiment, self-modifying actions can mitigate datacenter operators of the datacenter 3200 making possibly poor configuration decisions and can avoid underutilized and / or poorly performing portions of a datacenter.

[0222] Computer-based system

[0223] The following figures set forth, without limitation, exemplary computer-based systems that can be used to implement at least one embodiment.

[0224] Figure 33 A processing system 3300 is shown, in accordance with at least one embodiment. In at least one embodiment, system 3300 includes one or more processor(s) 3302 and one or more graphics processor(s) 3308, and can be a single processor desktop system, a multiprocessor workstation system, or a server system having many processor(s) 3302 or processor cores 3307. In at least one embodiment, processing system 3300 is a processing platform incorporated within a system- on-a-chip (SoC) integrated circuit for use in mobile, handheld, or embedded devices. In at least one embodiment, processor cores 3307 are referred to as computational units or processing units.

[0225] In at least one embodiment, processing system 3300 is configured to execute a schedule job API to receive, as input, an indication of a processor performance preference, a job type, a job priority, or some combination thereof, to cause a processor to be configured to run at a certain clock frequency, as described in connection with Figures 8A-8C In at least one embodiment, processing system 3300 is configured to execute a schedule job API to receive, as input, to cause an indication of a computational resource to be used by software instructions. In at least one embodiment, processing system 3300 is configured to execute a schedule job API to receive, as input, an indication of a priority of an execution instruction.

[0226] In at least one embodiment, processing system 3300 is configured to execute a get all available profile API to receive, as input, an indication of a processor performance preference, a frequency input, a job type, a computational resource input, a priority input, a processor identifier, or some combination thereof, to identify a processor setup profile, or as otherwise described herein. In at least one embodiment, processing system 3300 is configured to execute a get all available profile API to receive, as input, a frequency input, a computational resource input, a priority input, a processor identifier, or some combination thereof, to identify a processor setup profile, or as otherwise described herein.

[0227] In at least one embodiment, processing system 3300 is configured to execute a schedule set specific profile on processor API to receive an indication of a processor setup profile to cause instructions to be executed by a processor configured in accordance with that processor setup profile, or as otherwise described herein.

[0228] In at least one embodiment, at least one component shown or described in relation to Figure 33 is used to implement a combination of Figures 1-33technology and / or functionality described. In at least one embodiment, processing system 3300 performs one or more operations of an API to identify one or more settings to use to configure one or more processors based at least in part on one or more characteristics of a job to be performed by the one or more processors, as described in connection with Figure 1 or as otherwise described herein.

[0229] In at least one embodiment, processing system 3300 can include or be incorporated into a server-based gaming platform, including a game console, a gaming

[0230] In at least one embodiment, one or more processor(s) 3302 each include one or more processor cores 3307 to process instructions which, when executed, perform operations for system and user software. In at least one embodiment, each of the one or more processor cores 3307 are configured to process a specific instruction set 3309. In at least one embodiment, instruction set 3309 can facilitate complex instruction set computing (“CISC”), reduced instruction set computing (“RISC”), or computing via a very long instruction word (“VLIW”). In at least one embodiment, multiple processor cores 3307 can each process a different instruction set 3309, which can include instructions to facilitate emulation of other instruction sets. In at least one embodiment, processor core 3307 can also include other processing devices, such as a digital signal processor (“DSP”).

[0231] In at least one embodiment, processor 3302 includes cache memory 3304. In at least one embodiment, processor 3302 can have a single internal cache or multiple levels of internal caches. In at least one embodiment, cache memory is shared among multiple components of processor 3302. In at least one embodiment, processor 3302 also uses an external cache, which can be shared by processor cores 3307 using known cache coherency techniques (not illustrated), e.g., a level three (L3) cache or last level cache (LLC). In at least one embodiment, additionally included in processor 3302 are register file 3306, which processor 3302 can include different types of registers to store different types of data (e.g., integer registers, floating point registers, status registers, and instruction pointer registers). In at least one embodiment, register file 3306 can include general registers or other registers.

[0232] In at least one embodiment, one or more processor(s) 3302 are coupled with one or more interface bus(es) 3310 for communicating data between processor 3302 and other components of system 3300, e.g., address, data, or control signals. In at least one embodiment, interface bus 3310 can be a version of a processor bus, such as a direct media interface (DMI) bus, in at least one embodiment. In at least one embodiment, interface bus 3310 is not limited to DMI bus, and can include one or more peripheral component interconnect buses (e.g., PCI, PCI Express (“PCIe”)), memory buses, or other types of interface buses. In at least one embodiment, processor 3302 includes integrated memory controller 3316 and platform controller hub 3330. In at least one embodiment, memory controller 3316 facilitates communication between memory devices and other components of processing system 3300, while platform controller hub 3330 provides connections to input / output (“I / O”) devices via local I / O bus. In at least one embodiment, one or more peripheral component interconnect buses include a PCIe Gen 5, which provides an interface for a processor.

[0233] In at least one embodiment, storage device 3320 can be a dynamic random access memory (DRAM) device, a static random access memory (SRAM) device, flash memory device, or a phase change memory device, among others. In at least one embodiment, storage device 3320 can be used as a main memory for processing system 3300. In at least one embodiment, storage device 3320 can be used for storage of data 3322 and instructions 3321 for use when implementing applications or processes by one or more processors 3302. In at least one embodiment, a memory controller 3316 is used to control storage device 3320. In at least one embodiment, memory controller 3316 can be a separate component from processing system 3300. In at least one embodiment, memory controller 3316 is part of the processing system 3300, such as a processor or a core of a processor. In at least one embodiment, memory controller 3316 is on a chip or a die of a processor that includes one or more cores. In at least one embodiment, memory controller 3316 is included in processing system 3300 as a set of instructions executable by one or more processors 3302 or an operating system (not shown) executable by one or more processors 3302. In at least one embodiment, memory controller 3316 is hardware.

[0234] In at least one embodiment, platform controller hub 3330 enables peripherals to connect to storage devices 3320 and processor 3302 via a high-speed I / O bus. In at least one embodiment, I / O peripherals include, without limitation, audio controller 3346, network controller 3334, firmware interface 3328, wireless transceiver 3326, touch sensors 3325, data storage device 3324 (e.g., hard disk drive, flash memory, etc.). In at least one embodiment, data storage device 3324 can be connected via a storage interface (e.g., SATA) or via a peripheral bus, such as a Peripheral Component Interconnect bus (e.g., PCI, PCIe). In at least one embodiment, touch sensors 3325 can include touch screen sensors, pressure sensors, or fingerprint sensors. In at least one embodiment, wireless transceiver 3326 can be a Wi-Fi transceiver, a Bluetooth transceiver, or a mobile network transceiver such as a 3G, 4G, or Long Term Evolution (LTE) transceiver. In at least one embodiment, firmware interface 3328 enables communication with system firmware, and can be, for example, a unified extensible firmware interface (UEFI). In at least one embodiment, network controller 3334 can enable network connectivity to one or more private or public networks. In at least one embodiment, a high-performance network controller (not shown) is coupled to interface bus 3310. In at least one embodiment, audio controller 3346 is a multi-channel high definition audio controller. In at least one embodiment, processing system 3300 includes optional legacy I / O controller 3340 for coupling legacy (e.g., Personal System 2 (PS / 2)) devices to processing system 3300. In at least one embodiment, platform controller hub 3330 can also connect to one or more Universal Serial Bus (USB) controllers 3342 that connect to input devices, such as keyboard and mouse combination 3343, camera 3344, or other USB input devices.

[0235] In at least one embodiment, memory controller 3316 and instances of platform controller hub 3330 can be integrated into a discrete external graphics processor, such as external graphics processor 3312. In at least one embodiment, platform controller hub 3330 and / or memory controller 3316 can be external to one or more processor(s) 3302. For example, in at least one embodiment, processing system 3300 can include an external memory controller 3316 and platform controller hub 3330, which can be configured as a memory controller hub and a peripheral controller hub in a system chipset that is in communication with processor(s) 3302.

[0236] Figure 34A computer system 3400 according to at least one embodiment is shown. In at least one embodiment, computer system 3400 can be a system with interconnected devices and components, a SOC, or some combination thereof. In at least one embodiment, computer system 3400 is formed from a processor 3402 that can include execution units to execute an instruction, in at least one embodiment, computer system 3400 can include, without limitation, components such as processor 3402 to employ execution units including logic to perform algorithms for process data. In at least one embodiment, computer system 3400 can include processors such as Pentium®, CoreTM, XeonTM, AtomTM, TM or Nervana TM microprocessors, although other systems (including PCs, workstations, set-top boxes, etc. with other microprocessors) can also be used. In at least one embodiment, computer system 3400 can execute a version of the WINDOWS operating system available from Microsoft Corporation of Redmond, Wash., although other operating systems (UNIX and Linux, for example), embedded software, and / or graphical user interfaces, can also be used.

[0237] In at least one embodiment, computing system 3400 can be configured to execute a schedule job API to receive, as input, an indication of a processor performance preference, a job type, a job priority, or some combination thereof, to cause a processor to be configured to run at a certain clock frequency, as described in connection with Figures 8A-8C or elsewhere described herein. In at least one embodiment, computing system 3400 can be configured to execute a schedule job API to receive input to cause an indication of a computing resource to be used by software instructions. In at least one embodiment, computing system 3400 can be configured to execute a schedule job API to receive input to indicate a priority of an execution instruction.

[0238] In at least one embodiment, computing system 3400 can be configured to perform a get all available profile API to receive, as input, an indication of a processor performance preference, a frequency input, a job type, a compute resource input, a priority input, a processor identifier, or some combination thereof, to identify a processor setup profile, or as otherwise described herein. In at least one embodiment, computing system 3400 can be configured to perform a get all available profile API to receive, as input, an indication of a frequency input, a compute resource input, a priority input, a processor identifier, or some combination thereof, to identify a processor setup profile, or as otherwise described herein.

[0239] In at least one embodiment, computing system 3400 can be configured to perform a processor on schedule set specific profile API to receive an indication of a processor setup profile to execute instructions by a processor configured according to that processor setup profile, or as otherwise described herein.

[0240] In at least one embodiment, with respect to Figure 34 At least one component shown or described as being implemented within computing system 3400 can be implemented outside of computing system 3400. Alternatively, a combination of components shown or described as being implemented within computing system 3400 can be implemented outside of computing system 3400. In at least one embodiment, any part of computing system 3400 can be implemented as software instructions. Figures 1-33 In at least one embodiment, computer system 3400 performs one or more operations of an API to identify one or more settings to use to configure one or more processors based at least in part on one or more characteristics of a job to be performed by the one or more processors, as described in connection with Figure 1 In at least one embodiment, computer system 3400 performs one or more operations of an API to identify one or more settings to use to configure one or more processors based at least in part on one or more characteristics of a job to be performed by the one or more processors, as described in connection with

[0241] In at least one embodiment, computer system 3400 can be used in other devices such as handheld devices and embedded applications. Some examples of handheld devices include cellular phones, Internet Protocol devices, digital cameras, personal digital assistants (“PDAs”), and handheld PCs. In at least one embodiment, embedded applications can include a microcontroller, a digital signal processor (“DSP”), a SoC, a network computer (“NetPC”), a set-top box, a network hub, a wide area

[0242] In at least one embodiment, computer system 3400 can include, without limitation, a processor 3402, which can include, without limitation, one or more execution units 3408, which can be configured to implement a compute unified device architecture (“CUDA”) (available from NVIDIA Corporation of Santa Clara, California), an OpenCL architecture (available from the Khronos Group), or a vector CUDA program. In at least one embodiment, a CUDA program is at least a portion of a software application written in the CUDA programming language. In at least one embodiment, computer system 3400 is a single processor desktop or server system. In at least one embodiment, computer system 3400 can be a multiprocessor system. In at least one embodiment, processor 3402 can include, without limitation, a CISC microprocessor, a RISC microprocessor, a VLIW microprocessor, a processor implementing a combination of instruction sets, or any other processor device, such as a digital signal processor. In at least one embodiment, processor 3402 can be coupled to a processor bus 3410 that can transmit data signals between processor 3402 and other components in computer system 3400.

[0243] In at least one embodiment, processor 3402 can include, without limitation, a level 1 (“L1”) internal cache memory (“cache”) 3404. In at least one embodiment, processor 3402 can have a single internal cache or more than one level of internal cache. In at least one embodiment, cache memory can reside in the processor 3402’s external. In at least one embodiment, processor 3402 can include a combination of internal and external caches. In at least one embodiment, register file 3406 can store different types of data within various registers including, without limitation, integer registers, floating point registers, status registers, and instruction pointer registers.

[0244] In at least one embodiment, execution unit 3408, including, without limitation, logic to perform integer and floating point operations, also resides in processor 3402. Processor 3402 can also include microcode (“ucode”) read-only memory (“ROM”) that stores microcode for certain macro instructions. In at least one embodiment, execution unit 3408 can include logic to handle a packed instruction set 3409. In at least one embodiment, by including the packed instruction set 3409 in the instruction set of a general-purpose processor 3402, along with associated circuitry to execute the instructions, the general-purpose processor 3402 can be used to perform the operations of many multimedia applications faster and more effectively than a general-purpose processor without such support. In at least one embodiment, by using the full width of the processor’s data bus frequently to execute operations on packed data, many multimedia applications can be accelerated compared to processors that use a method of fetching and executing one or more instructions, one data element at a time, on smaller units of data.

[0245] In at least one embodiment, execution unit 3408 can also be used in microcontrollers, embedded processors, graphics devices, DSPs, and other types of logic circuits. In at least one embodiment, computer system 3400 can include, but not limited to, memory 3420. In at least one embodiment, memory 3420 can be implemented as a DRAM device, SRAM device, flash memory device, or other memory device. Memory 3420 can store data and / or instructions (e.g., software) that can be executed by processor 3402.

[0246] In at least one embodiment, system logic chip can be coupled to processor bus 3410 and memory 3420. In at least one embodiment, system logic chip can include, without limitation, a memory controller hub (“MCH”) 3416, and processor 3402 can communicate with MCH 3416 via processor bus 3410. In at least one embodiment, MCH 3416 can provide a high bandwidth memory path 3418 to memory 3420 for instruction and data storage and for storage of graphics commands, data, and textures. In at least one embodiment, MCH 3416 can direct data signals between processor 3402, memory 3420, and other components in computer system 3400, and can

[0247] In at least one embodiment, computer system 3400 can use system I / O 3422 as a proprietary hub interface bus to couple MCH 3416 to I / O controller hub (“ICH”) 3430. In at least one embodiment, ICH 3430 can provide direct connection to some I / O devices and can be connected to other devices via a local I / O bus. In at least one embodiment, local I / O bus can include, without limitation, a high-speed I / O bus for connecting peripheral devices to memory 3420, a chipset, and processor 3402. Examples can include, without limitation, an audio controller 3429, a firmware hub (“Flash BIOS”) 3428, a wireless transceiver 3426, a data storage 3424, a legacy I / O controller 3423 containing user input 3425 and keyboard interface, a serial expansion port 3427 (e.g., USB), and a network controller 3434. Data storage 3424 can include a hard disk drive, a floppy disk drive, a CD-ROM device, a flash memory device, or other mass storage device.

[0248] In at least one embodiment, Figure 34 A system including interconnected hardware devices or “chips” is shown. In at least one embodiment, system 3400 is a server system including multiple processors 3402, memory 3420, and I / O devices interconnected through system I / O 3422. Figure 34 An exemplary SoC can be shown. In at least one embodiment, system 3400 is a server system including multiple processors 3402, memory 3420, and I / O devices interconnected through system I / O 3422. Figure 34 Devices shown in FIG. 34A can be interconnected with proprietary interconnects, standardized interconnects (e.g., PCIe), or some combination thereof. In at least one embodiment, one or more components of system 3400 are interconnected using Compute Express Link (CXL) interconnects.

[0249] Figure 35 System 3500 according to at least one embodiment is shown. In at least one embodiment, system 3500 is an electronic device that utilizes processor 3510. In at least one embodiment, system 3500 can be, for example and without limitation, a laptop personal computer, a tower server, a rack server, a blade server, an edge device communicatively coupled to one or more local or cloud service providers, a laptop computer, a desktop computer, a tablet, a mobile device, a phone, an embedded computer, or any other suitable electronic device.

[0250] In at least one embodiment, system 3500 can be configured to execute a dispatch job API to receive, as input, an indication of a processor performance preference, a job type, a job priority, or some combination thereof, to cause a processor to be configured to run at a certain clock frequency, as described in connection with Figures 8A-8CThe or other described herein. In at least one embodiment, system 3500 is configurable to execute a schedule job API to receive input indicative of a computing resource to be used by an instruction. In at least one embodiment, system 3500 is configurable to execute a schedule job API to receive input indicative of a priority of an execution instruction.

[0251] In at least one embodiment, system 3500 is configurable to execute a get all available profile API to receive as input an indication of a processor performance preference, a frequency input, a job type, a computing resource input, a priority input, a processor identifier, or some combination thereof, to identify a processor setup profile, or other described herein. In at least one embodiment, system 3500 is configurable to execute a get all available profile API to receive a frequency input, a computing resource input, a priority input, a processor identifier, or some combination thereof, to identify a processor setup profile, or other as described herein.

[0252] In at least one embodiment, system 3500 is configurable to execute a processor on schedule set specific profile API to receive an indication of a processor setup profile to cause an instruction to be executed by a processor configured according to the processor setup profile, or other as described herein.

[0253] In at least one embodiment, with respect to Figure 35 At least one component shown or described is used to implement functionality described in connection with Figures 1-33 the technology and / or functionality described. In at least one embodiment, system 3500 performs one or more operations of an API to identify one or more settings to be used to configure one or more processors based at least in part on one or more characteristics of a job to be executed by the one or more processors, as described in connection with Figure 1 the technology and / or functionality described herein.

[0254] In at least one embodiment, system 3500 can include, without limitation, a processor 3510 communicatively coupled to any suitable number or kind of components, peripherals, modules, or devices. In at least one embodiment, processor 3510 is coupled using a bus or interface, such as an I2C bus, a System Management Bus (“SMBus”), a Low Pin Count (LPC) bus, a Serial Peripheral Interface (“SPI”), a High-Definition Audio (“HDA”) bus, a Serial Advanced Technology Attachment (“SATA”) bus, a USB (versions 1, 2, 3), or a Universal Asynchronous Receiver / Transmitter (“UART”) bus. In at least one embodiment, system 3500 is communicatively coupled to a network, such as a local area network (“LAN”), a metropolitan area network (“MAN”), or a wide area network (“WAN”), network such as the Internet, or some combination thereof. Figure 35 A system is shown that includes interconnected hardware devices or “chips.” In at least one embodiment, system 3500 is a system-on-a-chip (“SoC”). In at least one embodiment, system 3500 is a system-in-package (“SiP”). In at least one embodiment, system 3500 is a multi-chip package (“MCP”). In at least one embodiment, system 3500 is a single package that includes two or more heterogeneous chips. Figure 35An example SoC can be shown. In at least one embodiment, Figure 35 The devices shown in FIG. 15A can be interconnected with proprietary interconnects, standardized interconnects (e.g., PCIe), or some combination thereof.

[0255] In at least one embodiment, Figure 35 One or more components of FIG. 15B are interconnected using compute express link (CXL) interconnects.

[0256] In at least one embodiment, Figure 35 may include a display 3524, a touch screen 3525, a touch pad 3530, a near field communication unit (“NFC”) 3545, a sensor hub 3540, a thermal sensor 3546, an express chipset (“EC”) 3535, a trusted platform module (“TPM”) 3538, a BIOS / firmware / flash memory (“BIOS, FW Flash”) 3527, a DSP 3560, a solid state disk (“SSD”) or hard drive (“HDD”) 3520, a wireless local area network unit (“WLAN”) 3550, a Bluetooth unit 3552, a wireless wide area network unit (“WW AN”) 3556, a global positioning system (GPS) 3555, a camera (“USB 3.0 camera”) 3554 (e.g., a USB 3.0 camera), or a low power double data rate (“LPDDR”) memory unit (“LPDDR3”) 3515 implemented in, for example, LPDDR3 standard. These components can each be implemented in any suitable manner.

[0257] In at least one embodiment, other components can be communicatively coupled to processor 3510 by components discussed above. In at least one embodiment, an accelerometer 3541, an ambient light sensor (“ALS”) 3542, a compass 3543, and a gyroscope 3544 can be communicatively coupled to sensor hub 3540. In at least one embodiment, a thermal sensor 3539, a fan 3537, a keyboard 35736, and a touch pad 3530 can be communicatively coupled to EC 3535. In at least one embodiment, a speaker 3563, a headphone 3564, and a microphone (“mic”) 3565 can be communicatively coupled to an audio unit (“audio codec and class D amplifier”) 3562, which can in turn be communicatively coupled to DSP 3560. In at least one embodiment, audio unit 3562 can include, for example and without limitation, an audio coder / decoder (“codec”) and a class D amplifier. In at least one embodiment, a SIM card (“SIM”) 3557 can be communicatively coupled to WWAN unit 3556. In at least one embodiment, components such as WLAN unit 3550 and Bluetooth unit 3552, as well as WWAN unit 3556, can be implemented as a next generation form factor (NGFF).

[0258] Figure 36 An exemplary integrated circuit 3600 is shown, in accordance with at least one embodiment. In at least one embodiment, the exemplary integrated circuit 3600 is a SoC, which can be fabricated using one or more IP cores. In at least one embodiment, integrated circuit 3600 includes one or more application processors 3605 (e.g., CPUs, DPUs), at least one graphics processor 3610, and can additionally include an image processor 3615 and / or a video processor 3620, any of which can be a modular IP core. In at least one embodiment, integrated circuit 3600 includes peripheral or bus logic including a USB controller 3625, a UART controller 3630, an SPI / SDIO controller 3635, and an I2S / I2C controller 3640. In at least one embodiment, integrated circuit 3600 can include a display device 3645 coupled to one or more of a high-definition multimedia interface (HDMI) controller 3650 and a mobile industry processor interface (MIPI) display interface 3655. In at least one embodiment, storage can be provided by a flash memory subsystem 3660 including flash memory and a flash memory controller. In at least one embodiment, a memory interface can be provided via a memory controller 3665 for access to SDRAM or SRAM memory devices. In at least one embodiment, some integrated circuits also include an embedded security engine 3670.

[0259] In at least one embodiment, exemplary integrated circuit 3600 can be configured to execute a schedule job API to receive, as input, an indication of a processor performance preference, a job type, a job priority, or some combination thereof, causing the processor to be configured to run at a certain clock frequency, as described in connection with Figures 8A-8C or elsewhere herein. In at least one embodiment, exemplary integrated circuit 3600 can be configured to execute a schedule job API to receive input causing an indication of a computing resource to be used by software instructions. In at least one embodiment, exemplary integrated circuit 3600 can be configured to execute a schedule job API to receive input indicating a priority of an execution instruction.

[0260] In at least one embodiment, exemplary integrated circuit 3600 can be configured to perform a get all available configuration file API to receive, as input, an indication of a processor performance preference, a frequency input, a job type, a compute resource input, a priority input, a processor identifier, or some combination thereof, to identify a processor setup configuration file, or as otherwise described herein. In at least one embodiment, exemplary integrated circuit 3600 can be configured to perform a get all available configuration file API to receive, as input, an indication of a frequency input, a compute resource input, a priority input, a processor identifier, or some combination thereof, to identify a processor setup configuration file, or as otherwise described herein.

[0261] In at least one embodiment, exemplary integrated circuit 3600 can be configured to perform a processor on-die set specific configuration file API to receive an indication of a processor setup configuration file to execute instructions by a processor configured according to that processor setup configuration file, or as otherwise described herein.

[0262] In at least one embodiment, with respect to Figure 36 At least one component shown or described is used to implement techniques and / or functionality described in conjunction with Figures 1-33 described. In at least one embodiment, IC 3600 performs one or more operations of an API to identify one or more settings to use to configure one or more processors based at least in part on one or more characteristics of a job performed by the one or more processors, as described in conjunction with Figure 1 described, or as otherwise described herein.

[0263] Figure 37A computing system 3700 is shown in accordance with at least one embodiment. In at least one embodiment, computing system 3700 includes a processing subsystem 3701 having one or more processor(s) 3702 and system memory 3704 communicating via an interconnection path 3705 that can include an on-chip interconnection medium 3706. In at least one embodiment, memory hub 3705 can be a separate component coupled with one or more processors 3702 via on-chip interconnection medium 3706. In at least one embodiment, memory hub 3705 can be integrated into one or more processors 3702. In at least one embodiment, memory hub 3705 couples with an I / O subsystem 3711 via a communication link 3706 to facilitate communication between processing subsystem 3701 and I / O subsystem 3711. In at least one embodiment, I / O subsystem 3711 includes a memory hub 3705 and I / O hub 3707. In at least one embodiment, memory hub 3705 can be a separate component coupled with one or more processors 3702 via on-chip interconnection medium 3706. In at least one embodiment, memory hub 3705 can be integrated into one or more processors 3702. In at least one embodiment, I / O hub 3707 can be a separate component coupled with memory hub 3705 and one or more processors 3702 via on-chip interconnection medium 3706. In at least one embodiment, I / O hub 3707 can be integrated into memory hub 3705 or one or more processors 3702.

[0264] In at least one embodiment, computing system 3700 can be configured to execute a schedule job API to receive as input an indication of a processor performance preference, a job type, a job priority, or some combination thereof, to cause a processor to be configured to run at a certain clock frequency, as described or described elsewhere herein. In at least one embodiment, computing system 3700 can be configured to execute a schedule job API to receive input to cause an indication of a computing resource to be used by software instructions. Figures 8A-8C In at least one embodiment, computing system 3700 can be configured to execute a schedule job API to receive input to cause an indication of a computing resource to be used by software instructions.

[0265] In at least one embodiment, computing system 3700 can be configured to execute a get all available profiles API to receive as input an indication of a processor performance preference, a frequency input, a job type, a computing resource input, a priority input, a processor identifier, or some combination thereof, to identify a processor setup profile, as described or described elsewhere herein. In at least one embodiment, computing system 3700 can be configured to execute a get all available profiles API to receive a frequency input, a computing resource input, a priority input, a processor identifier, or some combination thereof, to identify a processor setup profile, as otherwise described herein.

[0266] In at least one embodiment, computing system 3700 is configured to execute an on-processor dispatch set-specific configuration file API to receive an indication of a processor setup configuration file to cause instructions to be executed by a processor configured according to that processor setup configuration file, or as otherwise described herein.

[0267] In at least one embodiment, regarding Figure 37 At least one component shown or described is used to implement a technique and / or functionality described in connection with Figures 1-33 In at least one embodiment, computing system 3700 performs one or more operations of an API to identify one or more settings to use to configure one or more processors based at least in part on one or more characteristics of a job to be performed by the one or more processors, as described in connection with Figure 1 or as otherwise described herein.

[0268] In at least one embodiment, processing subsystem 3701 includes one or more parallel processors 3712 coupled to memory hub 3705 via bus or other communication link 3713. In at least one embodiment, communication link 3713 can be one of many based on well-known bus or communications link technologies, such as but not limited to PCI, PCI Express or a vendor-specific communications protocol, or can be a proprietary communication structure. In at least one embodiment, one or more parallel processors 3712 form a programmable graphics processing subsystem that can include one or more graphics processing cluster(s) (GPCs) that can be configured to perform processing of pixels at one or more resolutions. In at least one embodiment, one or more parallel processors 3712 form a parallel or vector processing system that can include a large number of processing cores and / or processing clusters, such as a many integrated core (MIC) processor or compute unit. In at least one embodiment, one or more parallel processors 3712 form a graphics processing subsystem that can output pixels to one of one or more display devices 3710A coupled via I / O Hub 3707. In at least one embodiment, one or more parallel processors 3712 can also include a display controller and display interface (not shown) to enable a direct connection to one or more display devices 3710B.

[0269] In at least one embodiment, system storage unit 3714 can connect to I / O hub 3707 to provide storage mechanisms for computing system 3700. In at least one embodiment, I / O switches 3716 can be used to provide interface mechanisms to enable connections between I / O hub 3707 and other components such as network adapters 3718 and / or wireless network adapters 3719 that can be integrated into a platform, as well as various other devices that can be added via one or more add-in devices 3720. In at least one embodiment, network adapters 3718 can be Ethernet adapters or another wired network adapters. In at least one embodiment, wireless network adapters 3719 can include one or more of Wi-Fi, Bluetooth, NFC, or other network devices that include one or more radios.

[0270] In at least one embodiment, computing system 3700 can include other components not explicitly shown, including USB or other port connections, optical storage drives, video capture devices, etc., that can also be connected to I / O hub 3707. In at least one embodiment, communication paths interconnecting various Figure 37 Communication paths between various components in FIG. 37 can use any suitable protocol, such as a protocol based on PCI (Peripheral Component Interconnect) such as PCI Express, or another suitable bus or point-to-point communication interface and / or protocol (e.g., NVLink high-speed interconnect, or interconnect protocols).

[0271] In at least one embodiment, parallel processor(s) 3712 include circuitry optimized for graphics and video processing, including for example video output circuitry, and are configured for general-purpose processing as well. In at least one embodiment, parallel processor(s) 3712 incorporate circuitry optimized for graphics and video processing, including for example video output circuitry, and include state of the art vector processors. In at least one embodiment, parallel processor(s) 3712 can also include, without limitation, SIMD instruction set domain specific instruction set, a memory semantics that include in-flight cache misses and memory writes stall cycle count, and scalar processing cores including single- and double-precision floating point, integer, and Boolean logic. In at least one embodiment, parallel processor(s) 3712 incorporate circuitry optimized for graphics and video processing, including for example video

[0272] Processing system

[0273] The following figures illustrate example processing systems that can be used to implement at least one embodiment, without limitation.

[0274] Figure 38 An accelerated processing unit (“APU”) 3800, in accordance with at least one embodiment, is shown. In at least one embodiment, APU 3800 is developed by AMD Corporation, located in Santa Clara, CA. In at least one embodiment, APU 3800 is configurable to execute a dispatch work API to receive as input an indication of a processor performance preference, a job type, a job priority, or some combination thereof, causing the processor to be configured to run at a clock frequency, or as otherwise described herein. In at least one embodiment, APU 3800 is configurable to execute a dispatch work API to receive as input, causing an indication of a computing resource to be used by a software instruction. In at least one embodiment, APU 3800 is configurable to execute a dispatch work API to receive as input, causing an indication of a priority of an execution instruction.

[0275] In at least one embodiment, the APU 3800 may be configured to execute a Get All Available Profiles API to receive as input an indication of a processor performance preference, a frequency input, a job type, a compute resource input, a priority input, a processor identifier, or some combination thereof, thereby identifying a processor settings profile, or as otherwise described herein. In at least one embodiment, the APU 3800 may be configured to execute a Get All Available Profiles API to receive as input an indication of a frequency input, a compute resource input, a priority input, a processor identifier, or some combination thereof, thereby identifying a processor settings profile, or as otherwise described herein.

[0276] In at least one embodiment, the APU 3800 may be configured to execute an on-processor scheduler profile API to receive an indication of a processor configuration profile, thereby causing the execution of instructions by a processor configured according to the processor configuration profile, or as otherwise described herein. In at least one embodiment, the APU 3800 includes, but is not limited to, a core complex 3810, a graphics complex 3840, a fabric 3860, an I / O interface 3870, a memory controller 3880, a display controller 3892, and a multimedia engine 3894. In at least one embodiment, the APU 3800 may include, but is not limited to, any number of core complexes 3810, any number of graphics complexes 3850, any number of display controllers 3892, and any number of multimedia engines 3894 in any combination. For purposes of illustration, reference numerals are used herein to refer to multiple instances of similar objects, where the reference numeral identifies the object and a number in parentheses identifies the desired instance.

[0277] In at least one embodiment, Figure 38 At least one component shown or described is used to implement the combination Figures 1-33 In at least one embodiment, the APU 3800 performs one or more operations of the API to identify one or more settings to be used to configure the one or more processors based at least in part on one or more characteristics of a job to be performed by the one or more processors, such as in conjunction with Figure 1 as described, or as otherwise described herein.

[0278] In at least one embodiment, core complex 3810 is a CPU and graphics complex 3840 is a GPU, and APU 3800 is a processing unit that integrates, without limitation, 3810 and 3840 onto a single chip. In at least one embodiment, some tasks can be assigned to core complex 3810 while other tasks can be assigned to graphics complex 3840. In at least one embodiment, core complex 3810 is configured to execute host executable code derived from CUDA source code and graphics complex 3840 is configured to execute device executable code derived from CUDA source code.

[0279] In at least one embodiment, core complex 3810 includes, without limitation, cores 3820(1)-3820(4) and L3 cache 3830. In at least one embodiment, core complex 3810 can include, without limitation, any number of cores 3820 and any combination and number of caches. In at least one embodiment, cores 3820 are configured to execute instructions of a particular instruction set architecture (“ISA”). In at least one embodiment, each core 3820 is a CPU core. In at least one embodiment, cores 3820 are referred to as processing units or execution units.

[0280] In at least one embodiment, each core 3820 includes, without limitation, a fetch / decode unit 3822, an integer execution engine 3824, a floating point execution engine 3826, and an L2 cache 3828. In at least one embodiment, fetch / decode unit 3822 fetches instructions, decodes them, generates micro-operations, and dispatches individual micro-instructions to integer execution engine 3824 and floating point execution engine 3826. In at least one embodiment, fetch / decode unit 3822 can dispatch one micro-instruction to integer execution engine 3824 and another micro-instruction to floating point execution engine 3826 simultaneously. In at least one embodiment, integer execution engine 3824 executes, without limitation, integer and memory operations. In at least one embodiment, floating point engine 3826 executes, without limitation, floating point and vector operations. In at least one embodiment, fetch-decode unit 3822 dispatches micro-instructions to a single execution engine in place of both integer execution engine 3824 and floating point execution engine 3826.

[0281] In at least one embodiment, each core 3820(i) has access to an L2 cache 3828(i) included in the core 3820(i), where i is an integer representing a particular instance of a core 3820. In at least one embodiment, each core 3820 included in a core complex 3810(j) is connected to the other cores 3820 included in the core complex 3810(j) via an L3 cache 3830(j) included in the core complex 3810(j), where j is an integer representing a particular instance of a core complex 3810. In at least one embodiment, a core 3820 included in a core complex 3810(j) has access to all L3 caches 3830(j) included in the core complex 3810(j), where j is an integer representing a particular instance of a core complex 3810. In at least one embodiment, an L3 cache 3830 can include, without limitation, any number of slices.

[0282] In at least one embodiment, graphics complex 3840 can be configured to perform computational operations in a highly parallel manner. In at least one embodiment, graphics complex 3840 is configured to perform graphics pipeline operations such as draw commands, pixel operations, geometric computations, and other operations associated with rendering images to a display. In at least one embodiment, graphics complex 3840 is configured to perform operations that are not graphics related. In at least one embodiment, graphics complex 3840 is configured to perform both graphics related operations and operations that are not graphics related.

[0283] In at least one embodiment, graphics complex 3840 includes, without limitation, any number of compute units 3850 and an L2 cache 3842. In at least one embodiment, compute units 3850 share L2 cache 3842. In at least one embodiment, L2 cache 3842 is partitioned. In at least one embodiment, graphics complex 3840 includes, without limitation, any number of compute units 3850 and any number (including zero) and type of cache. In at least one embodiment, graphics complex 3840 includes, without limitation, any number of specialized graphics hardware.

[0284] In at least one embodiment, each compute unit 3850 includes, without limitation, any number of SIMD units 3852 and shared memory 3854. In at least one embodiment, each SIMD unit 3852 implements a SIMD architecture and is configured to execute operations in parallel. In at least one embodiment, each compute unit 3850 can execute any number of thread blocks, but each thread block executes on a single compute unit 3850. In at least one embodiment, a thread block includes, without limitation, any number of execution threads. In at least one embodiment, a workgroup is a thread block. In at least one embodiment, each SIMD unit 3852 executes a different thread warp. In at least one embodiment, a thread warp is a group of threads (e.g., 16 threads), where each thread in a thread warp belongs to a single thread block and is configured to process a different set of data based on a single instruction set. In at least one embodiment, one or more threads in a thread warp can be disabled using predication. In at least one embodiment, a lane is a thread. In at least one embodiment, a work item is a thread. In at least one embodiment, a wavefront is a thread warp. In at least one embodiment, different wavefronts in a thread block can be synchronized together and communicate via shared memory 3854. In at least one embodiment, each compute unit 3850 includes one or more clusters of thread blocks, where a cluster of thread blocks can enable programming control of locality at a greater granularity than a single thread block of a single streaming multi-processor (SM). In at least one embodiment, a cluster of thread blocks (also referred to as a “cluster”) enables multiple thread blocks running concurrently across streaming multi-processors to synchronize and cooperatively fetch, exchange, or otherwise use data.

[0285] In at least one embodiment, fabric 3860 is a system interconnect that facilitates data and control transmissions across core complex 3810, graphics complex 3840, I / O interface 3870, memory controllers 3880, display controller 3892, and multimedia engine 3894. In at least one embodiment, APU 3800 can include, without limitation, any number and type of system interconnects in addition to or instead of fabric 3860 that facilitate data and control transmissions across any number and type of directly or indirectly linked components that can be internal or external to APU 3800. In at least one embodiment, I / O interface 3870 represents any number and type of I / O interfaces (e.g., PCI, PCI-Extended (“PCI-X”), PCIe, Gigabit Ethernet (“GBE”), USB, etc.). In at least one embodiment, various types of peripheral devices are coupled to I / O interface 3870. In at least one embodiment, peripheral devices coupled to I / O interface 3870 can include, without limitation, a keyboard, a mouse, a printer, a scanner, a joystick or other type of game controller, a media recording device, an external storage device, a network interface card, etc.

[0286] In at least one embodiment, display controller AMD92 displays images on one or more display devices, such as liquid crystal display (“LCD”) devices. In at least one embodiment, multimedia engine 3894 includes, without limitation, any number and type of multimedia-related circuitry, such as a video decoder, a video encoder, an image signal processor, etc. In at least one embodiment, memory controllers 3880 facilitate data transfers between APU 3800 and unified system memory 3890. In at least one embodiment, core complex 3810 and graphics complex 3840 share unified system memory 3890.

[0287] In at least one embodiment, APU 3800 implements a memory subsystem that includes, without limitation, any number and type of memory controllers 3880 and memory devices (e.g., shared memory 3854) that can be dedicated to one component or shared among multiple components. In at least one embodiment, APU 3800 implements a cache subsystem that includes, without limitation, one or more cache memories (e.g., L2 cache 3128, L3 cache 3830, and L2 cache 3842), each of which can be private to a component or shared among any number of components (e.g., core 3820, core complex 3810, SIMD unit 3852, compute unit 3850, and graphics complex 3840).

[0288] Figure 39A CPU 3900 is shown, in accordance with at least one embodiment. In at least one embodiment, CPU 3900 is developed by AMD Corporation of Santa Clara, California. In at least one embodiment, CPU 3900 can be configured to execute application programs. In at least one embodiment, CPU 3900 is configured to execute host executable code derived from CUDA source code, and an external GPU can be configured to execute device executable code derived from such CUDA source code. In at least one embodiment, CPU 3900 includes, without limitation, any number of core complexes 3910, fabric 3960, I / O interfaces 3970, and memory controllers 3980.

[0289] In at least one embodiment, CPU 3900 can be configured to execute a schedule job API to receive, as input, an indication of a processor performance preference, a job type, a job priority, or some combination thereof, causing a processor to be configured to run at a certain clock frequency, as described in connection with Figures 8A-8C In at least one embodiment, CPU 3900 can be configured to execute a schedule job API to receive, as input, an indication of a processor performance preference, a job type, a job priority, or some combination thereof, causing a processor to be configured to run at a certain clock frequency, as described in connection with

[0290] In at least one embodiment, CPU 3900 can be configured to execute a get all available profiles API to receive, as input, an indication of a processor performance preference, a frequency input, a job type, a compute resource input, a priority input, a processor identifier, or some combination thereof, causing a processor settings profile to be identified, or as otherwise described herein. In at least one embodiment, CPU 3900 can be configured to execute a get all available profiles API to receive, as input, a frequency input, a compute resource input, a priority input, a processor identifier, or some combination thereof, causing a processor settings profile to be identified, or as otherwise described herein.

[0291] In at least one embodiment, CPU 3900 can be configured to execute a schedule settings specific profile on processor API to receive an indication of a processor settings profile, causing instructions to be executed by a processor configured in accordance with that processor settings profile, or as otherwise described herein.

[0292] In at least one embodiment, at least one component shown or described in connection with Figure 39 is used to implement the functionality described in connection withFigures 1-33 The described techniques and / or functionality. In at least one embodiment, CPU 3900 performs one or more operations of an API to identify one or more settings to use to configure one or more processors based at least in part on one or more characteristics of a job to be performed by the one or more processors, as described in connection with Figure 1 as described herein or otherwise described herein.

[0293] In at least one embodiment, core complex 3910 includes, without limitation, cores 3920(1)-3920(4) and L3 cache 3930. In at least one embodiment, core complex 3910 can include, without limitation, any number of cores 3920 in any combination and any number and type of caches. In at least one embodiment, cores 3920 are configured to execute instructions of a particular ISA. In at least one embodiment, each core 3920 is a CPU core.

[0294] In at least one embodiment, each core 3920 includes, without limitation, fetch / decode unit 3922, integer execution engine 3924, floating point execution engine 3926, and L2 cache 3928. In at least one embodiment, fetch / decode unit 3922 fetches instructions, decodes them, generates micro-operations, and dispatches individual micro-instructions to integer execution engine 3924 and floating point execution engine 3926. In at least one embodiment, fetch / decode unit 3922 can simultaneously dispatch one micro-instruction to integer execution engine 3924 and another micro-instruction to floating point execution engine 3926. In at least one embodiment, integer execution engine 3924 executes, without limitation, integer and memory operations. In at least one embodiment, floating point engine 3926 executes, without limitation, floating point and vector operations. In at least one embodiment, fetch-decode unit 3922 dispatches micro-instructions to a single execution engine in place of both integer execution engine 3924 and floating point execution engine 3926.

[0295] In at least one embodiment, each core 3920(i) has access to an L2 cache 3928(i) included in the core 3920(i), where i is an integer representing a particular instance of a core 3920. In at least one embodiment, each core 3920 included in a core complex 3910(j) is connected to the other cores 3920 in the core complex 3910(j) via an L3 cache 3930(j) included in the core complex 3910(j), where j is an integer representing a particular instance of a core complex 3910. In at least one embodiment, a core 3920 included in a core complex 3910(j) has access to all L3 caches 3930(j) included in the core complex 3910(j), where j is an integer representing a particular instance of a core complex 3910. In at least one embodiment, an L3 cache 3930 can include, without limitation, any number of slices.

[0296] In at least one embodiment, fabric 3960 is a system interconnect that facilitates data and control transfers across core complexes 3910(1)-3910(N) (where N is an integer greater than zero), I / O interface 3970, and memory controllers 3980. In at least one embodiment, CPU 3900 can include, without limitation, any number and type of system interconnects in addition to or instead of fabric 3960 that facilitate data and control transfers across any number and type of directly or indirectly linked components that can be internal or external to CPU 3900. In at least one embodiment, I / O interface 3970 represents any number and type of I / O interface (e.g., PCI, PCI-X, PCIe, GBE, USB, etc.). In at least one embodiment, various types of peripheral devices are coupled to I / O interface 3970. In at least one embodiment, peripheral devices coupled to I / O interface 3970 can include, without limitation, a display, a keyboard, a mouse, a printer, a scanner, a joystick or other type of game controller, a media recording device, an external storage device, a network interface card, etc.

[0297] In at least one embodiment, a memory controller 3980 facilitates data transfers between the CPU 3900 and system memory 3990. In at least one embodiment, the core complex 3910 and the graphics complex 3940 share system memory 3990. In at least one embodiment, the CPU 3900 implements a memory subsystem that includes, but is not limited to, any number and type of memory controllers 3980 and memory devices that can be dedicated to a component or shared among multiple components. In at least one embodiment, the CPU 3900 implements a cache subsystem that includes, but is not limited to, one or more cache memories (e.g., an L2 cache 3928 and an L3 cache 3930), each of which can be private to a component or shared among any number of components (e.g., a core 3920 and the core complex 3910).

[0298] Figure 40 An exemplary accelerator integrated slice 4090 is shown according to at least one embodiment. As used herein, a "slice" includes a specified portion of the processing resources of an accelerator integrated circuit. In at least one embodiment, the accelerator integrated circuit provides cache management, memory access, environment management, and interrupt management services on behalf of multiple graphics processing engines in multiple graphics acceleration modules. The graphics processing engines may each include a separate GPU. Optionally, the graphics processing engines may include different types of graphics processing engines within the GPU, such as a graphics execution unit, a media processing engine (e.g., a video encoder / decoder), a sampler, and a blit engine. In at least one embodiment, the graphics acceleration module may be a GPU having multiple graphics processing engines. In at least one embodiment, the graphics processing engines may be individual GPUs integrated on a common package, line card, or chip.

[0299] In at least one embodiment, the exemplary accelerator integrated slice 4090 may be configured to execute a schedule job API to receive as input an indication of processor performance preference, job type, job priority, or some combination thereof, thereby configuring the processor to run at a particular clock frequency, such as in conjunction with Figures 8A-8C In at least one embodiment, the exemplary accelerator integration slice 4090 may be configured to execute a schedule job API to receive input indicating the computing resources to be used by the software instructions. In at least one embodiment, the exemplary accelerator integration slice 4090 may be configured to execute a schedule job API to receive input indicating the priority of the executed instructions.

[0300] In at least one embodiment, exemplary accelerator integration slice 4090 can be configured to perform a get all available profile API to receive an indication of processor performance preference, frequency input, job type, compute resource input, priority input, processor identifier, or some combination thereof as input to identify a processor setup profile, or as otherwise described herein. In at least one embodiment, exemplary accelerator integration slice 4090 can be configured to perform a get all available profile API to receive an indication of frequency input, compute resource input, priority input, processor identifier, or some combination thereof to identify a processor setup profile, or as otherwise described herein.

[0301] In at least one embodiment, exemplary accelerator integration slice 4090 can be configured to perform a processor on schedule set specific profile API to receive an indication of a processor setup profile to execute instructions by a processor configured according to that processor setup profile, or as otherwise described herein.

[0302] In at least one embodiment, with respect to Figure 40 At least one component shown or described is used to implement a technique and / or functionality described in conjunction with Figures 1-33 In at least one embodiment, accelerator integration slice 4090 performs one or more operations of an API to identify one or more settings for configuring one or more processors based at least in part on one or more characteristics of a job to be performed by the one or more processors, as described in conjunction with Figure 1 or as otherwise described herein.

[0303] Application effective address space 4082 within system memory 4014 stores process elements 4083. In one embodiment, process elements 4083 are stored in response to GPU calls 4081 from applications 4080 executing on processor 4007. Process elements 4083 contain processing state for corresponding applications 4080. Work descriptors (WDs) 4084 contained in process elements 4083 can be a single job requested by an application or can contain a pointer to a queue of jobs. In at least one embodiment, WD 4084 is a pointer to a job request queue in application effective address space 4082.

[0304] Graphics acceleration module 4046 and / or individual graphics processing engines can be shared by all or a subset of processes in a system. In at least one embodiment, an infrastructure for setting up processing state and sending WDs 4084 to graphics acceleration module 4046 to start jobs in a virtualized environment can be included.

[0305] In at least one embodiment, a dedicated process programming model is implemented. In this model, a single process owns a graphics acceleration module 4046 or individual graphics processing engines. As graphics acceleration module 4046 is owned by a single process, a hypervisor initializes the accelerator integration circuit for the owning partition and an operating system initializes the accelerator integration circuit for the owning partition when graphics acceleration module 4046 is assigned.

[0306] In operation, a WD fetch unit 4091 in accelerator integration slice 4090 fetches a next WD 4084 including an indication of work to be completed by one or more graphics processing engines of graphics acceleration module 4046. Data from WD 4084 can be stored in registers 4045 used by memory management unit (MMU) 4039, interrupt management circuit 4047, and / or environment management circuit 4048, as shown. For example, one embodiment of MMU 4039 includes segment / page walk circuitry to access segment / page tables 4086 within OS virtual address space 4085. Interrupt management circuit 4047 can handle interrupt events (“INT”) 4092 received from graphics acceleration module 4046. Effective addresses 4093 produced by graphics processing engines, when executing graphics operations, are translated to real addresses by MMU 4039.

[0307] In one embodiment, the same set of registers 4045 are replicated for each graphics processing engine and / or graphics acceleration module 4046 and can be initialized by a system hypervisor or operating system. Each of these replicated registers can be included in accelerator integration slice 4090. Exemplary registers that can be initialized by a hypervisor are shown in Table 1.

[0308] Table 1 - Hypervisor Initialized Registers

[0309]

[0310]

[0311] Exemplary registers that can be initialized by an operating system are shown in Table 2.

[0312] Table 2 - Operating System Initialized Registers

[0313] 1 Process and thread identification 2 Effective address (EA) environment save / restoration pointer 3 Virtual address (VA) accelerator utilization record pointer 4 Virtual address (VA) storage segment table pointer 5 Authority mask 6 Work descriptor

[0314] In one embodiment, each WD 4084 is specific to a particular graphics acceleration module 4046 and / or a particular graphics processing engine. It contains all information needed by the graphics processing engines to do the work or it can be a pointer to a memory location where an application has set up a command queue of work to be completed.

[0315] Figures 41A-41B An exemplary graphics processor according to at least one embodiment is shown. In at least one embodiment, any of the exemplary graphics processors can be fabricated using one or more IP cores. In addition to the illustrated embodiments, in at least one embodiment, other logic and circuits can be included, including additional graphics processor cores, peripheral interface controllers or general-purpose processor cores. In at least one embodiment, an exemplary graphics processor is used within an SoC.

[0316] Figure 41A An exemplary graphics processor 4110 of an SoC integrated circuit that can be fabricated using one or more IP cores is shown, according to at least one embodiment. Figure 41B An additional exemplary graphics processor 4140 of an SoC integrated circuit that can be fabricated using one or more IP cores is shown, according to at least one embodiment. In at least one embodiment, Figure 41A The graphics processor 4110 is a low power graphics processor core. In at least one embodiment, Figure 41B The graphics processor 3340 is a higher performance graphics processor core. In at least one embodiment, each graphics processor 4110, 4140 can be a Figure 36 variant of the graphics processor 3610.

[0317] In at least one embodiment, the exemplary graphics processor 4110 and / or graphics processor 4140 can be configured to execute a dispatch operations API to receive as input an indication of a processor performance preference, a job type, a job priority, or some combination thereof, to cause the processor to be configured to run at a particular clock frequency, as described in connection with Figures 8A-8C In at least one embodiment, the exemplary graphics processor 4110 and / or graphics processor 4140 can be configured to execute a dispatch operations API to receive as input to cause an indication of a compute resource to be used by software instructions. In at least one embodiment, the exemplary graphics processor 4110 and / or graphics processor 4140 can be configured to execute a dispatch operations API to receive as input to indicate a priority of an execution instruction.

[0318] In at least one embodiment, exemplary graphics processor 4110 and / or graphics processor 4140 can be configured to perform a get all available profile API to receive, as input, an indication of a processor performance preference, a frequency input, a job type, a compute resource input, a priority input, a processor identifier, or some combination thereof, to identify a processor setting profile, or as otherwise described herein. In at least one embodiment, exemplary graphics processor 4110 and / or graphics processor 4140 can be configured to perform a get all available profile API to receive, as input, a frequency input, a compute resource input, a priority input, a processor identifier, or some combination thereof, to identify a processor setting profile, or as otherwise described herein.

[0319] In at least one embodiment, exemplary graphics processor 4110 and / or graphics processor 4140 can be configured to perform a processor on schedule set specific profile API to receive an indication of a processor setting profile to execute instructions by a processor configured according to that processor setting profile, or as otherwise described herein.

[0320] In at least one embodiment, with respect to Figures 41A-41B At least one component shown or described is used to implement techniques and / or functionality described in connection with Figures 1-33 described. In at least one embodiment, graphics processor 4140 performs one or more operations of an API to identify one or more settings to use to configure one or more processors based at least in part on one or more characteristics of a job to be performed by the one or more processors, as described in connection with Figure 1 or otherwise described herein.

[0321] In at least one embodiment, graphics processor 4110 includes a vertex processor 4105 and one or more fragment processor(s) 4115A-4115N (e.g., 4115A, 4115B, 4115C, 4115D up to 4115N-1 and 4115N). In at least one embodiment, graphics processor 4110 can execute different shader programs via separate logic for vertex processing and / or for fragment / pixel processing. In at least one embodiment, vertex processor 4105 is optimized to execute operations for vertex shader programs, while one or more fragment processor(s) 4115A-4115N can be optimized for the execution of fragment or pixel shader programs. In at least one embodiment, vertex processor 4105 processes vertex data that is transmitted from a graphics application within a vertex processing stage. In at least one embodiment, the vertex processing stage includes a number of separate operations: vertex setup, vertex main, setup / attribute, and setup / geometry. In at least one embodiment, vertex processor 4105 executes the vertex setup and vertex main operations. In at least one embodiment, one or more of setup / attribute, setup / geometry, and a setup / hardware interface execute one or more operations.

[0322] In at least one embodiment, graphics processor 4110 additionally includes one or more MMU(s) 4120A-4120B, cache memory 4125A-4125B, and circuit interconnect 4130A-4130B. In at least one embodiment, one or more MMU(s) 4120A-4120B provide for virtual to physical address mapping for graphics processor 4110, including for vertex processor 4105 and / or for fragment processor(s) 4115A-4115N, which can reference vertex or image / texture data stored in memory, in addition to vertex or image / texture data stored in one or more cache(s) 4125A-4125B. In at least one embodiment, one or more MMU(s) 4120A-4120B can be synchronized with one or more MMUs within Figure 36 application processor(s) 3605, image processor 3615, and / or video p...

Claims

1. A processor, comprising: One or more circuits for executing an application programming interface (API) to configure one or more processors to operate at one or more clock frequencies based at least in part on one or more inputs to the API. 2 . The processor of claim 1 , wherein the one or more inputs comprise one or more indications by the one or more processors of one or more processor target metrics to be used when executing the one or more instructions. 3 . The processor of claim 1 , wherein the one or more inputs include one or more indications of software workloads to be executed by the one or more processors.

4. The processor of claim 1 , wherein the one or more circuits are to execute the API to cause the one or more processors to be configured to operate at one or more clock frequencies based at least in part on one or more observed processor performance metrics. 5 . The processor of claim 1 , wherein the one or more circuits are configured to execute an API to cause the one or more processors to be configured to operate based at least in part on a maximum operating voltage Vmax. 6 . The processor of claim 1 , wherein the one or more circuits are to execute the API to cause the one or more processors to be configured to operate based at least in part on a crossbar Xbar ratio.

7. The processor of claim 1, wherein the one or more circuits are to execute the API to cause the one or more processors to be configured to operate at one or more clock frequencies to execute one or more instructions in a data center.

8. A system comprising: One or more processors configured to execute an application programming interface (API) such that the one or more processors are configured to operate at one or more clock frequencies based at least in part on one or more inputs to the API.

9. The system of claim 8, wherein the one or more inputs include one or more indications of processor performance preferences.

10. The system of claim 8, wherein the one or more inputs include one or more instructions for a job.

11. The system of claim 8, wherein the one or more processors are configured to execute the API such that the one or more processors are configured to operate at one or more clock frequencies based at least in part on one or more processor performance metrics generated when the one or more processors execute one or more instruction sets. 12 . The system of claim 8 , wherein the one or more processors are to execute the API to cause the one or more processors to be configured to operate based at least in part on a maximum total graphics power (max TGP).

13. The system of claim 8, wherein the one or more processors are to execute the API to cause the one or more processors to be configured to operate based at least in part on a crossbar ratio.

14. The system of claim 8, wherein the one or more processors are to execute the API such that the one or more processors are configured to operate based at least in part on an algorithm for setting fan settings.

15. A method comprising: An application programming interface (API) is executed to configure one or more processors to operate at one or more clock frequencies based at least in part on one or more inputs to the API.

16. The method of claim 15, wherein the one or more inputs include one or more instructions to set up a profile for the processor. The method of claim 15 , wherein the one or more inputs include one or more indications of an instruction set.

18. The method according to claim 15, further comprising: The API is executed to configure the one or more processors to operate at one or more clock frequencies based at least in part on one or more processor performance metrics generated when the one or more processors execute one or more software workloads.

19. The method according to claim 15, further comprising: The API is executed to configure the one or more processors to operate at one or more clock frequencies based at least in part on a maximum operating voltage Vmax indicated by one or more job scheduling software programs.

20. The method of claim 15, further comprising: The API is executed to cause the one or more processors to be configured to operate at one or more clock frequencies based at least in part on one or more instruction sets scheduled to be executed by the one or more processors.