Processor clock scaling techniques

By identifying the proximity and utilization patterns of processor cores and dynamically adjusting clock frequency, this solves the problem of temperature sensors being unable to accurately measure local hotspots, achieving more efficient thermal management and improving processor performance and reliability.

CN120610602APending Publication Date: 2025-09-09NVIDIA CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510251100.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-03-06
Filing Date
2025-03-04
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

In the prior art, temperature sensors on processor chips are unable to accurately measure local hot spots, resulting in thermal overload or excessive throttling, affecting processor performance and reliability.

Method used

By identifying the proximity and utilization patterns of processor cores, it dynamically adjusts clock frequency to avoid adverse thermal conditions, combining profiles and firmware management strategies for more precise thermal management.

Benefits of technology

It improves processor performance and reliability, avoids thermal overload and excessive throttling issues, and implements a more efficient thermal management strategy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120610602A_ABST
    Figure CN120610602A_ABST
Patent Text Reader

Abstract

The invention discloses processor clock scaling techniques, and particularly discloses devices, systems, and techniques to scale a processor clock. In at least one embodiment, one or more circuits are configured to scale one or more clocks of one or more cores based at least in part on a proximity of the one or more cores to each other.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] At least one embodiment relates to scaling the clocks of one or more processor cores. For example, at least one embodiment relates to a processor including one or more circuits for scaling one or more clocks of one or more cores based at least in part on the proximity of the one or more cores to each other. Background Art

[0002] The processor can manage heat through a software thermal policy that monitors the average temperature of thermal sensors located in the set of running cores. Currently, a worst-case thermal hotspot excursion estimate based on a profiled application is added as a safety margin. BRIEF DESCRIPTION OF THE DRAWINGS

[0003] The following detailed description of exemplary, non-limiting illustrative embodiments should be read in conjunction with the accompanying drawings, in which:

[0004] Figure 1 A system including a controller and one or more processor cores according to at least one embodiment is shown;

[0005] Figure 2 A multi-core processor chip according to at least one embodiment is shown;

[0006] Figure 3 A multi-core processor chip according to at least one embodiment is shown;

[0007] Figure 4 illustrates a multi-core processor chip with a combination of processor cores in use according to at least one embodiment;

[0008] Figure 5 A block diagram illustrating a multi-core processor chip with different combinations of processor cores in use according to at least one embodiment;

[0009] Figure 6 A block diagram illustrating a multi-core processor chip having multiple different combinations of processor cores in use according to at least one embodiment;

[0010] Figure 7 A block diagram illustrating a multi-core processor chip having multiple different combinations of processor cores in use according to at least one embodiment;

[0011] Figure 8A shows a thermal map of a processor die illustrating an example of a worst-case hotspot in accordance with at least one embodiment;

[0012] Figure 8B illustrates a heat map of a processor die showing an evenly loaded core configuration with no hot spots in accordance with at least one embodiment;

[0013] Figure 9 A process flow for implementing a thermal policy management system according to at least one embodiment is shown;

[0014] Figure 10 A process for updating thermal policy offsets according to at least one embodiment is shown;

[0015] Figure 11 A process for implementing a system for thermal policy management according to at least one embodiment is shown;

[0016] Figure 12 shows a processor module according to at least one embodiment;

[0017] Figure 13 Depicting an API for scaling one or more clocks of one or more cores based at least in part on proximity of the one or more cores to each other in accordance with at least one embodiment;

[0018] Figure 14 A distributed system according to at least one embodiment is shown;

[0019] Figure 15 An exemplary data center is shown in accordance with at least one embodiment;

[0020] Figure 16 illustrates a client-server network according to at least one embodiment;

[0021] Figure 17 illustrates an example of a computer network in accordance with at least one embodiment;

[0022] Figure 18A illustrates a networked computer system according to at least one embodiment;

[0023] Figure 18B illustrates a networked computer system according to at least one embodiment;

[0024] Figure 18C illustrates a networked computer system according to at least one embodiment;

[0025] Figure 19 illustrates one or more components of a system environment in which services may be provided as third-party network services according to at least one embodiment;

[0026] Figure 20illustrates a cloud computing environment according to at least one embodiment;

[0027] Figure 21 illustrates a set of functional abstraction layers provided by a cloud computing environment in accordance with at least one embodiment;

[0028] Figure 22 shows a supercomputer at a chip level according to at least one embodiment;

[0029] Figure 23 illustrates a supercomputer at the rack module level according to at least one embodiment;

[0030] Figure 24 illustrates a supercomputer at the rack level according to at least one embodiment;

[0031] Figure 25 illustrates a supercomputer at an overall system level according to at least one embodiment;

[0032] Figure 26A Inference and / or training logic according to at least one embodiment is shown;

[0033] Figure 26B Inference and / or training logic according to at least one embodiment is shown;

[0034] Figure 27 illustrates the training and deployment of a neural network according to at least one embodiment;

[0035] Figure 28 shows the architecture of a network system according to at least one embodiment;

[0036] Figure 29 shows the architecture of a network system according to at least one embodiment;

[0037] Figure 30 illustrates a control plane protocol stack according to at least one embodiment;

[0038] Figure 31 illustrates a user plane protocol stack according to at least one embodiment;

[0039] Figure 32 Components of a core network according to at least one embodiment are shown;

[0040] Figure 33 Components of a system supporting network functions virtualization (NFV) according to at least one embodiment are shown;

[0041] Figure 34 A processing system according to at least one embodiment is shown;

[0042] Figure 35A computer system according to at least one embodiment is shown;

[0043] Figure 36 A system according to at least one embodiment is shown;

[0044] Figure 37 An exemplary integrated circuit according to at least one embodiment is shown;

[0045] Figure 38 A computing system according to at least one embodiment is shown;

[0046] Figure 39 An APU is shown according to at least one embodiment;

[0047] Figure 40 A CPU according to at least one embodiment is shown;

[0048] Figure 41 An exemplary accelerator integrated slice is shown in accordance with at least one embodiment;

[0049] Figures 42A-42B An exemplary graphics processor is shown in accordance with at least one embodiment;

[0050] Figure 43A illustrates a graphics core according to at least one embodiment;

[0051] Figure 43B GPGPU according to at least one embodiment is shown;

[0052] Figure 44A A parallel processor according to at least one embodiment is shown;

[0053] Figure 44B illustrates a processing cluster according to at least one embodiment;

[0054] Figure 44C A graphics multiprocessor is shown in accordance with at least one embodiment;

[0055] Figure 45 illustrates a software stack for a programming platform according to at least one embodiment;

[0056] Figure 46 According to at least one embodiment, Figure 45 CUDA implementation of the software stack;

[0057] Figure 47 According to at least one embodiment, Figure 45 ROCm implementation of the software stack;

[0058] Figure 48 According to at least one embodiment, Figure 45OpenCL implementation of the software stack;

[0059] Figure 49 illustrates software supported by a programming platform according to at least one embodiment; and

[0060] Figure 50 According to at least one embodiment, a method for Figures 45-48 Compiled code that is executed on a programming platform. DETAILED DESCRIPTION

[0061] Figure 1 A system 100 including a controller and one or more processor cores according to at least one embodiment is shown. In at least one embodiment, a chip 120 including multiple cores 110 may experience localized heat buildup (referred to as "hotspots"), a type of thermal overload that can lead to hardware failure. In at least one embodiment, this occurs because temperature sensors on chip 120 do not accurately measure the temperature of each processor core or chip region because temperature sensors (e.g., thermal sensors 115) are not evenly or ideally distributed throughout the chip 120. In at least one embodiment, hotspots can be caused by core processor activity when certain combinations of processor cores 110 experience sustained levels of operation, such as combinations where the proximity of the processor cores 110 causes heat buildup in corresponding regions of the chip 120. In at least one embodiment, the hotspots may not be detected by thermal sensors 115 or may be inaccurately measured, which may occur in part due to the proximity of thermal sensors 115 to the hotspots.

[0062] In at least one embodiment, when the temperature exceeds a certain value, processor activity can be scaled down to eliminate thermally induced failures. In at least one embodiment, the value is based on an estimated temperature derived in part from thermal sensor 115. In at least one embodiment, due to the lack of specific temperature data for each location, the estimated temperature is adjusted by adding a thermal offset. In at least one embodiment, processor activity can be scaled based on the utilization pattern of processor core 110, rather than using an offset value that is too high (which can help prevent overheating but also leads to excessive throttling of processor activity) or too low (which risks thermally induced failures). In at least one embodiment, the pattern includes a pattern that indicates localized heat accumulation that can be erroneously measured by thermal sensor 115.

[0063] In at least one embodiment, the use of adjustable thermal offsets for different processor core utilization patterns gives thermal management strategies greater flexibility to manage hot spots and processor performance. In at least one embodiment, records of processor core utilization and recorded or predicted temperatures can be used to identify patterns that would lead to thermally induced failures in the absence of thermal offsets or where fixed thermal offsets are too low to provide sufficient safety margins. In at least one embodiment, such patterns (which may be described as worst-case thermal combinations, worst-case combinations, worst-case thermal scenarios, etc.) can be avoided or accounted for by the processor to improve processor performance. In at least one embodiment, this is accomplished by scaling the clocks of one or more processor cores.

[0064] In at least one embodiment, a configuration file (e.g., configuration file 104) may include information indicating a combination of processor cores that indicates a worst-case scenario or other pattern as described above that may influence what thermal excursions are feasible. In at least one embodiment, the information may indicate processor core utilization patterns that may be avoided to allow for smaller thermal excursions. In at least one embodiment, the information may indicate patterns that, when identified, may be used as a basis for dynamically setting a relatively high thermal excursion.

[0065] In at least one embodiment, scaling the clock of one or more processor cores or combinations thereof includes increasing or decreasing the speed of the processor cores, thereby causing a corresponding increase or decrease in the amount of heat energy generated by the processor cores. In at least one embodiment, adjusting the operation of a first group of one or more processor cores (e.g., by temporarily removing the processor cores from operation) allows increasing the clock of another group of processor cores, where the first group is a group that includes processors whose proximity is associated with an adverse thermal condition.

[0066] In at least one embodiment, the thermal output of one or more processor cores is correlated with the proximity of the processor cores. In at least one embodiment, for example, the thermal output of a group of processor cores is correlated with the proximity of the operating processor cores relative to each other. In at least one embodiment, there is a further correlation between the proximity and activity of the processor cores. In at least one embodiment, processor cores operating in close proximity to each other (e.g., processor cores relatively close to each other on chip 120) can generate thermal activity associated with an adverse thermal condition, and the clocks of the processor cores can be scaled accordingly.

[0067] In at least one embodiment, the processor core utilization pattern includes specifying a combination of processor cores that may result in a particular thermal output, such as a thermal output associated with a hotspot or other worst-case scenario, such as a hotspot that is not accurately measured by the onboard thermal sensor 115. In at least one embodiment, the core utilization pattern may indicate that the thermal activity generated by a processor core or combination of processor cores is within the limits of acceptable thermal output and does not require intervention by a thermal management policy. In at least one embodiment, the processor core utilization pattern may indicate that the thermal activity generated by the processor core or combination of processor cores exceeds the limits of acceptable thermal output, but can be managed by scaling the activity of the processor cores in accordance with the thermal management policy. In at least one embodiment, this includes adjusting the thermal offsets described herein. In at least one embodiment, the processor core utilization pattern may indicate that the processor core or combination of processor cores will generate thermal activity that exceeds the limits of acceptable thermal output, but can be managed by temporarily removing the processor core combination from operation in accordance with the thermal management policy.

[0068] In at least one embodiment, the thermal condition corresponds to one or more of a temperature of a processor core, a combination of processor cores, a processor component other than a processor core, or a location on a chip.

[0069] In at least one embodiment, an adverse thermal condition is a condition that jeopardizes the sustained operational performance of one or more combinations of processor cores. In at least one embodiment, avoiding one or more core utilization patterns that would result in an adverse thermal condition includes implementing a thermal management strategy that removes a processor core or combination of processor cores from operation or adjusts the operation of the processor core or combination of processor cores to avoid the adverse thermal condition by scaling the clocks of the processor cores.

[0070] In at least one embodiment, the information indicative of utilization patterns of one or more processor cores includes data correlating a processor core or combination of processor cores with thermal conditions.

[0071] In at least one embodiment, system 100 includes a chip 120 that includes a controller 105 and a plurality of processor cores 110. In at least one embodiment, controller 105 manages the overall performance of processor cores 110, including implementing thermal management strategies as described herein. Figure 1 One or more aspects of one or more embodiments described herein may be combined with one or more aspects of one or more embodiments described herein, including at least one aspect of one or more embodiments described herein. Figures 2 to 13 Described embodiment.

[0072] In at least one embodiment, chip 120 includes multiple groups of processor cores. For example, in at least one embodiment, chip 120 includes two or more groups of processor cores, which are connected to chip 120 via two or more sockets or circuit systems configured to operate in this way, and each group of processors functions as an independent multi-core processor. In at least one embodiment, the groups of processor cores communicate via a communication bus (e.g., PCIe, NVLink, or Infinity Fabric xGMI). In at least one embodiment, chip 120 corresponds to a GRACE chip, an AMD Instinct MI300 series chip, or other similar chip.

[0073] In at least one embodiment, the measurements from the thermal sensor 115 provide substantially real-time thermal measurements. In at least one embodiment, the configuration file 104 includes stored information identifying processor combinations associated with pathological patterns of processor usage. In at least one embodiment, the usage pattern of a group of processors (e.g., processor 110) by an application correlates with a pattern indicated in the configuration file 104. In at least one embodiment, the pattern indicates an adverse thermal condition. In at least one embodiment, the configuration file 104 is specific to a chip, operating system, or computing device. In at least one embodiment, the processor core utilization pattern is indicated in the configuration file 104 as data or text. For example, in at least one embodiment, the configuration file 104 may include a text string indicating groups of processor cores, such as "C33, C34, C43," indicating a group of processor cores that may result in an adverse thermal condition if used at sufficiently high levels. It should be understood that this example is not intended to limit potential embodiments to only those that conform to the specific example provided.

[0074] In at least one embodiment, different hotspot offsets can be assigned to different processor configurations in order to dynamically configure the processor core to strike a balance between reducing the processor core temperature and maintaining high performance. In at least one embodiment, this is accomplished by system 100 identifying a processor core utilization pattern associated with an adverse thermal condition and scaling the clock of one or more processors based on the identification. In at least one embodiment, system 100 provides an alternative thermal strategy in response to identifying the pattern. In at least one embodiment, for example, system 100 correlates the current utilization of processor core 110 with the processor core utilization pattern indicated in configuration file 104.

[0075] In at least one embodiment, configuration file 104 corresponds to configuration file 225, such as Figure 2 In at least one embodiment, the system 100 includes an Advanced Configuration Power Interface (ACPI), such as Figure 2AFCPI 230 is shown. In at least one embodiment, system 100 includes firmware, such as firmware 220, for implementing a thermal management policy (e.g., policy 235). In at least one embodiment, the firmware monitors application usage of processor cores 110, compares the processor cores being used, and compares the processor cores to a processor core utilization pattern that may indicate an adverse thermal condition. In at least one embodiment, the pattern is indicated in configuration file 105. In at least one embodiment, the firmware and software adjust one or more of the processor cores' clocks according to the policy to increase or decrease their performance and corresponding thermal output.

[0076] Figure 2 FIG2 shows a multi-core processor chip according to at least one embodiment. In at least one embodiment, chip 205 includes multiple processor cores 210, which are Figure 2 The chip 205 is further labeled as processor cores 1-76, but the number of processor cores is for illustration only and should not be construed as limiting any particular embodiment. In at least one embodiment, the chip 205 includes thermal sensors, such as thermal sensors 215A-215D. In at least one embodiment, the chip 205 includes Figure 2 The arrangement of processor cores is similar to the arrangement of processor cores depicted in , but this arrangement is intended to be illustrative only and may vary in different embodiments.

[0077] In at least one embodiment, thermal sensors 215A-215D are located at different locations on chip 205 to measure the temperature of different areas of the chip, such as the temperature of processor core 210, other circuitry, or other materials. In at least one embodiment, various hardware and architectural constraints or design limitations prevent the temperature sensors from being ideally positioned relative to processor core 210 or other areas of the chip, and therefore, temperature sensors 215A-215D may not accurately or uniformly detect hot spots on chip 205. In at least one embodiment, sustained processor core activity near the maximum clock speed (e.g., Fmax) generates a thermal load on that particular core and may also induce thermal loads on nearby cores, components, or materials. In at least one embodiment, the likelihood of thermal overload or hot spots increases when two or more processor cores are in close proximity to each other and operate in a state of sustained activity. In at least one embodiment, other proximity relationships (e.g., a processor core near an area with poor heat transfer) may also result in hot spots.

[0078] In at least one embodiment, system 200 includes firmware 220, configuration files 225, ACPI 230, and policies 235. In at least one embodiment, one or more of firmware 220, configuration files 225, ACPI 230, and policies 235 are included in chip 205.

[0079] In at least one embodiment, firmware 220 includes software instructions to be executed by a processor, such as one or more of processor cores 210. In at least one embodiment, the instructions are stored in one or more read-only memories, such as one or more read-only memories of system 200 or chip 205.

[0080] In at least one embodiment, instructions in firmware 220 are executed to cause system 200 to verify that the most recent version of the latest worst-case thermal scenario has been read from configuration file 225; verify that the ACPI 230 table has been updated; and confirm that thermal sensors are being used to monitor thermal conditions on chip 205. In at least one embodiment, the instructions are executed to update policy 235. In at least one embodiment, policy 235 includes software and / or data for implementing thermal policies on chip 205. In at least one embodiment, the instructions update policy 235 with any newly identified processor core utilization patterns (e.g., worst-case processor core combinations).

[0081] Figure 3 In at least one embodiment, system 300 depicts a multi-core processor chip. Figure 2 The arrangement of the processor cores on the CPU. In at least one embodiment, the processor cores 310 are used in various combinations depending on the workload. In at least one embodiment, software and / or firmware (e.g., operating system software and / or firmware, such as Figure 2 The firmware 220 depicted in FIG determines which processor cores 210 should be used to execute the workload. In at least one embodiment, the workload can be distributed across the selected processor cores 210.

[0082] In at least one embodiment, the workload may include an application that includes threads that are scheduled on a subset of the total number of cores of the available processor. In at least one embodiment, for example, cores 12 and 22 may be selected to execute the threads, or in another case, cores 15 and 16 may be selected.

[0083] In at least one embodiment, some thermal sensors 315A-315D will more accurately measure the temperature of certain processing cores 210 than other thermal sensors. In at least one embodiment, thermal sensors 315A-315D more accurately measure the thermal output of the processor core 210 closest to the sensor. In at least one embodiment, for example, thermal sensor TS-1 will more accurately reflect the actual temperature at processor cores 12, 13, 21, and 22 than other thermal sensors on chip 305. In at least one embodiment, the thermal conditions of some processor cores or other circuitry, areas, or components of chip 305 are not accurately measured by any thermal sensor. For example, in at least one embodiment, a processor core (e.g., Figure 3 The processor core 34 depicted in FIG is not near any of the thermal sensors 315A-315D, so its temperature may not be accurately measured.

[0084] In at least one embodiment, system 300 measures the temperature on chip 305 and adjusts one or more clock speeds of one or more processor cores. In at least one embodiment, by reducing the clock speed, the amount of heat generated by a given processor core can be reduced because thermal energy is a byproduct of computing processing. However, in at least one embodiment, reducing the processing speed (clock speed) of one or more processor cores can reduce computing performance. In at least one embodiment, problems associated with some thermal management issues are avoided; these problems can include: the thermal management strategy being too conservative, reducing the processor clock speed by more than necessary to prevent thermal overload failures, or not being conservative enough, as it can lead to thermally induced failures.

[0085] In at least one embodiment, the thermal management strategy is implemented by dynamically collecting temperature readings from sensors during use and correlating these sensor readings with processor activity. In at least one embodiment, firmware (e.g., Figure 2Firmware 220 in the processor core monitors these temperature readings and, when the temperature reaches a given threshold, throttles back processor activity for processor cores that may be contributing to adverse thermal conditions. In at least one embodiment, throttling back processor activity includes reducing the clock speed of one or more processor cores. In at least one embodiment, the distribution of temperature sensors on chip 305 is not ideal, which may mean that the sensors inaccurately measure the temperature of certain areas of the chip. In at least one embodiment, a scalar temperature offset is added to the temperature reading obtained at the sensor location. In at least one embodiment, this value is adjusted based on the distance from the processor core to estimate the actual temperature at the core. In at least one embodiment, assigning an offset to the sensor reading is intended to more accurately reflect the actual temperature of the processor, which in turn avoids failures due to thermal overload. In at least one embodiment, using an offset in this manner can degrade performance in some cases because, for example, it can lead to overly aggressive or underaggressive thermal policies, but this consequence can be avoided by identifying and / or preventing certain processor core utilization patterns.

[0086] Figure 4 A multi-core processor chip with a combination of processor cores in use is shown in accordance with at least one embodiment. In at least one embodiment, system 400 corresponds to Figure 2 and Figure 3 2 or 300, respectively, and includes a processor core 410 and thermal sensors 415A-D, which correspond to similarly named elements in these figures. In at least one embodiment, as Figure 4 As shown using cross hatching in , certain processor cores are selected to execute threads associated with the computational workload. In at least one embodiment, thermal sensor TS-1 415A records temperature readings associated with the thermal condition of the processor cores in the local area of ​​TS-1 on chip 405. In at least one embodiment, if these processors generate hot spots, thermal sensor TS-1 may accurately measure the temperature. However, in at least one embodiment, such processor core utilization patterns may cause hot spots to appear in a portion of chip 405 that TS-1 does not accurately measure (e.g., the area near processor core 33). However, in at least one embodiment, this problem can be avoided by identifying processor core utilization patterns that contribute to such conditions. In at least one embodiment, the patterns can be identified based on the physical location of the cores on chip 405. In at least one embodiment, the patterns can be identified in a configuration file (e.g., Figure 2The pattern may be indicated in a configuration file 225 (as depicted in FIG. 1 ). In at least one embodiment, one or more thermal management strategies may be applied in response to identifying the pattern. In at least one embodiment, the response to the identification may include adjusting thermal offsets. In at least one embodiment, the response to the identification may include changing how the processor core 410 is used to avoid the pattern.

[0087] Figure 5 A block diagram of a multi-core processor chip with a combination of processor cores in use according to at least one embodiment is shown. In at least one embodiment, the system 500 corresponds to Figure 2 、 Figure 3 and Figure 4 In at least one embodiment, the system 500 includes the systems 200, 300, and / or 400 depicted in the respective embodiments. Figures 2 to 4 505 is a chip 505 that is a chip 205, 305, and / or 405 depicted in FIG, and also includes processor cores 410 that correspond to the processor cores depicted in these figures. In at least one embodiment, a group of processor cores 410 are selected to execute threads associated with a workload, such as processor cores 25, 33, 34, 35, 42, 43, 44, 52, 53, and 54 (in FIG, Figure 5 In at least one embodiment, these processor cores generate thermal conditions that are not accurately detected by any of the sensors TS-1, TS-2, TS-3, or TS-4, and continue to operate at similar levels.

[0088] In at least one embodiment, to prevent such thermal failures, an offset adjustment technique is used. In at least one embodiment, the offset compensates for potentially inaccurate temperature readings at one or more processors by adding an offset to the recorded temperature. In at least one embodiment, the firmware therefore compensates the temperature recorded by the sensor and makes this offset determination based on the number and location of processor cores in use. In at least one embodiment, this determination is based on one or more indicated core utilization patterns. In at least one embodiment, the pattern is indicated in a configuration file. In at least one embodiment, the pattern is dynamically determined. In at least one embodiment, the pattern indicates a proximity pattern of processor cores that may contribute to adverse thermal conditions (e.g., hot spots that are not accurately measured by onboard thermal sensors). In at least one embodiment, the proximity pattern includes on-chip location. In at least one embodiment, the proximity pattern includes on-chip distance. In at least one embodiment, the proximity pattern includes relative location, which may include, for example, the relationship between a processor core and other cores, circuitry, components, materials, and / or temperature sensors.

[0089] Figure 6A block diagram of a multi-core processor chip with different combinations of processor cores in use according to at least one embodiment is shown. In at least one embodiment, system 600 corresponds to Figures 2 to 5 In at least one embodiment, the potential utilization pattern that may be detected includes a cluster of processor cores located at a certain distance from a thermal center. However, it should be understood that these examples are intended to be illustrative and not limiting.

[0090] Figure 7 A block diagram of a multi-core processor chip with different combinations of processor cores in use according to at least one embodiment is shown. In at least one embodiment, system 700 corresponds to Figures 2 to 6 One or more of the systems 200, 300, 400, 500, and 600 depicted in .

[0091] In at least one embodiment, Figure 7 As depicted, the initial core utilization pattern may include processor cores 730 whose proximity to one another and / or to other components of the chip 705 may create an unfavorable thermal condition. In at least one embodiment, such a pattern is identified. In at least one embodiment, this is accomplished by software and / or firmware of the system 700 and / or chip 705, which compares the currently observed processor core utilization with a processor core utilization pattern that is indicated as potentially leading to an unfavorable thermal condition. In at least one embodiment, the software and / or firmware, when executed by the processor, causes the workload to be reassigned to other processors, thereby ending usage according to the pattern. In at least one embodiment, the software may, for example, cause the workload to be assigned to other processor cores 720 located elsewhere on the chip 705. Similarly, the software and / or firmware, when executed by the processor, may proactively schedule workloads to processor cores to prevent the unfavorable processor utilization pattern from occurring.

[0092] In at least one embodiment, rather than reassigning workloads, the clocks of the processor cores associated with the unfavorable pattern can be adjusted so that at least some of the processors operate at a temperature low enough to avoid the unfavorable thermal condition. For example, the clocks of processors 34, 43, and 53 in region 730 can be scaled down below a threshold amount to avoid an unfavorable processor core utilization pattern consisting of all processors in region 730 operating at or above the threshold.

[0093] Figure 8AA thermal map of a processor die is shown, illustrating an example of a worst-case thermal hotspot, according to at least one embodiment. In at least one embodiment, region 810 of the processor die includes multiple active processors in close proximity to each other, which can create a hotspot that, if not addressed, can lead to thermally induced failures. Other points (e.g., at 820) have active processors that are not arranged in a pattern that would result in a hotspot.

[0094] In at least one embodiment, Figure 8A As shown, the processor chip may gradually become cooler as it moves away from this region, as shown at point 805. In at least one embodiment, excessive heat accumulation may be identified by the thermal sensor at 815, but not by the thermal sensor at point 805. In at least one embodiment, the software and / or firmware controlling the operation of the chip may therefore reschedule workloads or prevent workloads from being assigned to the processor core in the configuration shown in region 810. In at least one embodiment, this may be as shown in FIG. Figure 8B As shown, Figure 8B A heat map of a processor die is shown showing an evenly loaded core configuration with no hot spots, according to at least one embodiment. In at least one embodiment, the diagram shows the area shown as 835 against a background of 840, and there are no thermally hazardous hot spots as seen at 810 or 820. In the configuration shown in this embodiment, the processor is not thermally overloaded and performance is not hampered.

[0095] In at least one embodiment, if an unfavorable core utilization pattern is identified, a thermal management strategy may be applied to adjust thermal excursions. In at least one embodiment, the pseudo-code associated with the thermal management strategy is represented as follows:

[0096]

[0097] Where 15 is a higher thermal offset value and 7 is a lower thermal offset value. It should be understood that this example is illustrative and should not be considered limiting.

[0098] In at least one embodiment, the logic of the thermal management policy distinguishes between processor cores whose temperatures are accurately measured by onboard thermal sensors and processor cores whose temperatures are not accurately measured by the thermal sensors. In at least one embodiment, processors in certain processor core combinations may require less or no thermal offset because nearby temperature sensors accurately measure their temperatures. In at least one embodiment, certain processor core combinations are identified as unfavorable processor core utilization patterns in comparison scenarios. In at least one embodiment, the thermal management policy assigns higher offset values ​​to these processors, which the firmware then uses to trigger throttling. In at least one embodiment, the performance of these processors is reduced, but the processor cores do not overheat or fail. In at least one embodiment, the processor cores in such processor core combinations are assigned lower thermal offsets to compensate for the fact that their actual temperatures are not accurately measured by the thermal sensors. In at least one embodiment, the unfavorable scenario can be identified as less than a worst-case scenario, and in such cases, a thermal offset can be used that is greater than the thermal offset that would be used in a non-unfavorable mode but less than the thermal offset that would be used in a worst-case mode.

[0099] Figure 9 1 shows a process flow for implementing a thermal policy management system according to at least one embodiment. In at least one embodiment, process 900 includes steps or operations for updating a thermal management policy. In at least one embodiment, process 900 is performed by firmware, e.g. Figure 2 The firmware 220 shown in or herein with respect to Figures 1 to 8B In at least one embodiment, the firmware at 905 monitors temperature using a thermal sensor proximate to the processor core.

[0100] In at least one embodiment, CPU temperature is monitored at 910 and correlated to processor core combinations associated with adverse thermal conditions.

[0101] In at least one embodiment, process 900 includes determining whether a core combination associated with an adverse thermal condition is identified at 915. In at least one embodiment, this includes determining whether processor cores operating at or near peak processing capacity correspond to a processor core utilization pattern associated with an adverse thermal condition (given their proximity to each other and / or other chip components).

[0102] In at least one embodiment, such a combination is identified at 915, and the thermal policy offset is updated at 925. In at least one embodiment, such a combination is not identified, and the thermal policy offset is not updated at 920.

[0103] In at least one embodiment, if a processor core combination has been identified at 915, the operation of the processor cores in the combination can be adjusted. For example, in at least one embodiment, the ACPI table can be adjusted to cause one or more processors in the combination to be used at a lower capacity or temporarily disabled (at 930).

[0104] Figure 10 1 shows a process for updating a thermal policy offset according to at least one embodiment. In at least one embodiment, process 900 includes steps or operations for updating a thermal management policy. In at least one embodiment, process 900 is performed by firmware, e.g. Figure 2 The firmware 220 shown in or herein with respect to Figures 1 to 8B In at least one embodiment, the firmware at 905 monitors temperature using a thermal sensor near the processor core.

[0105] In at least one embodiment, a combination of processor cores associated with an adverse thermal condition is identified at 1005. In at least one embodiment, one or more thermal excursions are identified based on the combination at 1010. In at least one embodiment, the combination is used to populate a configuration file (e.g., configuration file 225) to indicate the combination at 1015.

[0106] In at least one embodiment, the operations described in connection with elements 1005-1015 are performed using one or more test systems or simulations, while subsequent operations described in connection with elements 1020-1040 are performed by a system (e.g., Figures 2 to 7 Any system described in the relevant description) is executed.

[0107] In at least one embodiment, process 1000 includes reading the combination at 1020. In at least one embodiment, this is done by software and / or firmware (e.g., firmware 220 of system 200) or Figures 3 to 7 This is accomplished using similarly named components as shown in .

[0108] In at least one embodiment, process 1000 includes updating the ACPI table to indicate the thermal policy at 1025. In at least one embodiment, this is done by software and / or firmware (e.g., firmware 220 of system 200) or Figures 3 to 7 This is accomplished using similarly named components as shown in .

[0109] In at least one embodiment, at 1030, process 1000 includes monitoring the operation of the processor core and the output of the thermal sensor, determining whether any of the combinations described are occurring, and determining whether any thermal limits have been exceeded. In at least one embodiment, this is performed by software and / or firmware (e.g., firmware 220 of system 200) or Figures 3 to 7The similarly named components shown in are completed.

[0110] In at least one embodiment, process 1000 includes performing an evaluation to determine if conditions are true relative to the operations described in relation to elements 1020, 1025, and 1030. In at least one embodiment, if all of these conditions are evaluated to be true, then at 1040, the thermal management policy is updated. In at least one embodiment, the updating includes adjusting thermal offsets to account for the identified core utilization pattern. In at least one embodiment, the updating includes disabling or throttling processor cores to avoid adverse processor core utilization patterns. In at least one embodiment, this is performed by software and / or firmware (e.g., firmware 220 of system 200) or Figures 3 to 7 The similarly named components shown in are completed.

[0111] In at least one embodiment, if the condition evaluated at decision block 1035 does not evaluate to true, then at 1045 the policy is not updated.

[0112] Figure 11 The process of implementing thermal policy management by a system according to at least one embodiment is shown. In at least one embodiment, process 1100 is performed by a system (e.g., Figures 2 to 7 Any of the systems described herein) is executed.

[0113] In at least one embodiment, at 1105, the worst case thermal combination is identified and a configuration file (e.g., configuration file 225) is populated with this information. In at least one embodiment, this is performed prior to operation of the system implementing the reminder of process 1100. In at least one embodiment, the combination or other combinations utilized by the processor cores are dynamically identified. In at least one embodiment, these operations are performed by software and / or firmware (e.g., firmware 220 of system 200) or Figures 3 to 7 This is accomplished using similarly named components as shown in .

[0114] In at least one embodiment, the configuration file is read, loaded into memory, or otherwise processed at 1110. In at least one embodiment, this is done by software and / or firmware (e.g., firmware 220 of system 200) or Figures 3 to 7 This is accomplished using similarly named components as shown in .

[0115] In at least one embodiment, at 1115, the ACPI table is updated to indicate how one or more processor cores should be used. In at least one embodiment, this is done by software and / or firmware (e.g., firmware 220 of system 200) or Figures 3 to 7 This is accomplished using similarly named components as shown in .

[0116] In at least one embodiment, at 1120, the thermal sensor is monitored and checked to see if it exceeds a limit. In at least one embodiment, this is done by software and / or firmware (e.g., firmware 220 of system 200) or Figures 3 to 7 This is accomplished using similarly named components as shown in .

[0117] In at least one embodiment, at 1125, a determination is made as to whether the conditions indicated by operations 1110, 1115, and 1120 are satisfied. In at least one embodiment, if satisfied, an updated thermal management policy is determined at 1130 and set at 1135. In at least one embodiment, these operations are implemented by software and / or firmware (e.g., firmware 220 of system 200) or Figures 3 to 7 . In at least one embodiment, updating the thermal policy comprises setting a worst-case thermal offset for an application that exhibits a worst-case thermal scenario (e.g., a worst-case core utilization pattern). In at least one embodiment, updating the thermal policy comprises setting an adjusted thermal offset for an application or workload that exhibits a core utilization pattern associated with an adverse thermal condition. In at least one embodiment, updating the thermal policy comprises setting a reduced thermal offset for an application or workload that exhibits a core utilization pattern that is not associated with an adverse thermal condition. In at least one embodiment, policies for applications or workloads that are not associated with worst-case or other adverse thermal conditions are not updated. In at least one embodiment, application of one or more of the thermal policies comprises adjusting a thermal offset associated with a processor used by the application or workload.

[0118] Figure 12 1200 illustrates a processor 1205 and modules according to at least one embodiment. In at least one embodiment, the processor 1205 is configured to process a plurality of cores based at least in part on the proximity of one or more cores to each other (e.g., Figure 1 In at least one embodiment, the processor 1205 performs the following steps in conjunction with the implementation changes described in the accompanying drawings and / or related descriptions to perform one or more processes, such as the process described herein for scaling one or more clocks of one or more cores. Figure 1 In at least one embodiment, the processor 1205 performs one or more processes, such as in conjunction with Figures 1 to 11 The process described.

[0119] In at least one embodiment, processor 1205 includes one or more processors, such as a processor in conjunction with Figures 14 to 50processor described. In at least one embodiment, the processor 1205 is any suitable processing unit and / or combination of processing units, such as one or more CPUs, GPUs, GPGPUs, PPUs, and / or variants thereof. In at least one embodiment, the processor 1205 includes: a CPU temperature monitor and policy update module 1210 for monitoring the CPU temperature and updating the thermal management policy offset; a worst-case scenario identification and policy update module 1215; and a hotspot offset management module 1220 for generating one or more new hotspot offsets and / or otherwise managing hotspot offsets in the thermal management policy. In at least one embodiment, the CPU temperature monitor and policy update module 1210 for monitoring the CPU temperature and updating the thermal management policy offset, the worst-case scenario identification and policy update module 1215 for identifying worst-case scenarios and updating the thermal management policy, and the hotspot offset management module 1220 for generating and / or managing one or more new hotspot offsets are part of the processor 1205 and / or one or more other processors. In at least one embodiment, a CPU temperature monitor and policy update module 1210 for monitoring CPU temperature and updating thermal management policy offsets, a worst-case scenario identification and policy update module 1215 for identifying worst-case scenarios and updating thermal management policies, and a hotspot offset management module 1220 for generating and / or managing one or more new hotspot offsets are distributed among multiple processors that communicate via a bus, a network, by writing to a shared memory, and / or any suitable communication process (e.g., the communication process described herein).

[0120] In at least one embodiment, some or all of the functions, methods, and functionality within one or more of these modules may be reconfigured from one of these modules to another module.

[0121] In at least one embodiment, as used in any implementation described herein, unless the context clearly dictates otherwise or clearly to the contrary, a module refers to any combination of software logic, firmware logic, hardware logic, and / or circuitry configured to provide the functionality described herein. In at least one embodiment, software may be embodied as a software package, code, and / or instruction set or instructions, and "hardware" as used in any implementation described herein may include (e.g., individually or in any combination) hardwired circuits, programmable circuits, state machine circuits, fixed function circuits, execution unit circuits, and / or firmware that stores instructions executed by programmable circuits. In at least one embodiment, modules may be collectively or individually embodied as circuits that constitute part of a larger system, such as an integrated circuit (IC), a system on a chip (SoC), and the like. In at least one embodiment, a module performs one or more processes in conjunction with any suitable processing unit and / or combination of processing units (e.g., one or more CPUs, GPUs, GPGPUs, PPUs, and / or variants thereof).

[0122] In at least one embodiment, the CPU temperature monitor and policy update module 1210 is based at least in part on the proximity of one or more cores to each other (e.g., in conjunction with Figures 1 to 11 In at least one embodiment, the CPU temperature monitor and policy update module 1210 performs one or more processes (e.g., the processes described herein) by at least including or otherwise monitoring temperature data of a CPU or a processor core of the CPU and updating a thermal policy management element (e.g., via processor 1205) with respect to the temperature data. In at least one embodiment, the CPU temperature monitor and policy update module 1210 obtains or is otherwise provided with one or more neural networks (e.g., via one or more systems, such as in conjunction with Figures 1 to 11 In at least one embodiment, the CPU temperature monitor and policy update module 1210 performs processing activities related to protecting data confidentiality by encrypting data as it is processed. In at least one embodiment, the CPU temperature monitor and policy update module 1210 performs processing activities related to protecting data confidentiality by encrypting data as it is processed.

[0123] In at least one embodiment, the worst-case scenario identification and policy update module 1215 is a module that performs processing activities related to managing one or more tenant virtual machines created or managed under the control of a hypervisor. In at least one embodiment, the module may also perform related activities such as performing variable assignments using inputs, serializing and / or storing values ​​in a database or other memory location or retrieving such values ​​from storage or performing one or more processes (e.g., in conjunction with Figures 1 to 11 In at least one embodiment, the worst-case scenario identification and policy update module 1215 performs one or more processes (e.g., one or more processes described herein) by at least including or otherwise encoding instructions that cause the one or more processes to be performed or that are otherwise operable to perform the one or more processes (e.g., by the processor 1205). In at least one embodiment, the worst-case scenario identification and policy update module 1215 obtains or is otherwise provided with one or more neural networks (e.g., by one or more systems, e.g., in conjunction with Figures 1 to 11 In at least one embodiment, the worst-case scenario identification and policy update module 1215 performs one or more processes (e.g., in conjunction with Figures 1 to 11 The worst case scenario identification and policy update module 1215 performs processing activities related to identifying the worst case scenario for the processor core and updating the thermal policy management element with respect to the worst case scenario for the processor core. Figures 1 to 11 One or more of those related processing activities described in .

[0124] In at least one embodiment, the hotspot offset management module 1220 is a module that performs management and processing activities related to identifying processor cores that are deemed eligible to remain operational but whose clocks are to be throttled by assigning offsets to the processor cores. In at least one embodiment, the module may also perform related activities such as performing variable assignments using inputs, serializing and / or storing values ​​in a database or other memory location or retrieving these values ​​from storage or performing other operations through one or more processes (e.g., in conjunction with Figures 1 to 11In at least one embodiment, the hotspot excursion management module 1220 performs one or more processes (e.g., the processes described herein) by at least managing and processing activities (e.g., performed by the processor 1205) associated with identifying the processor cores deemed eligible to remain operational but whose clocks are to be throttled by assigning excursions to the processor cores. In at least one embodiment, the hotspot excursion management module 1220 obtains or is otherwise provided with one or more neural networks (e.g., by one or more systems, e.g., in conjunction with Figures 1 to 11 In at least one embodiment, the hotspot excursion management module 1220 performs one or more processes (e.g., in conjunction with Figures 1 to 11 ) performs processing activities related to encrypting and / or decrypting data transmitted to or from the parallel processors via the interconnect. In at least one embodiment, the hotspot excursion management module 1220 manages and processes at least one of the following processes by managing and processing the identification of processor cores that are deemed eligible to remain operational but whose clocks will be interrupted by sending a signal to the processor core (e.g., in conjunction with the hotspot excursion management module 1220). Figures 1 to 11 In at least one embodiment, the hotspot offset management module 1220 can delete a hotspot offset from the thermal management strategy.

[0125] API

[0126] Figure 13 An API for scaling one or more clocks of one or more cores based at least in part on the proximity of the one or more cores to each other is depicted in accordance with at least one embodiment. In at least one embodiment, a block diagram 1300 is shown illustrating a driver and / or runtime including one or more libraries for providing one or more application programming interfaces (APIs). In at least one embodiment, the software program 1302 is a software module. In at least one embodiment, the software program 1302 includes one or more software modules. In at least one embodiment, the one or more software modules, such as Figures 1 to 121302. In at least one embodiment, one or more APIs 1310 are software instruction sets that, if executed, cause one or more processors to perform one or more computing operations. In at least one embodiment, one or more APIs 1310 are distributed or otherwise provided as part of one or more libraries 1306, runtimes 1304, drivers 1304, and / or any other grouping of software and / or executable code described further herein. In at least one embodiment, one or more APIs 1310 perform one or more computing operations in response to an invocation of a software program 1302. In at least one embodiment, a software program 1302 is a collection of software code, commands, instructions, or other text sequences that instructs a computing device to perform one or more computing operations and / or invokes one or more other instruction sets (e.g., APIs 1310 or API functions 1312) for execution. In at least one embodiment, the functionality provided by one or more APIs 1310 includes software functions 1312, such as software functions that can be used to accelerate one or more portions of software program 1302 using one or more parallel processing units (PPUs), such as graphics processing units (GPUs). In at least one embodiment, the software program is a compiler.

[0127] In at least one embodiment, the API 1310 is a hardware interface to one or more circuits for performing one or more computing operations. In at least one embodiment, the one or more software APIs 1310 described herein are implemented to perform operations in conjunction with Figures 1 to 12 In at least one embodiment, one or more software programs 1302 include instructions that, if executed, cause one or more hardware devices and / or circuits to perform operations in conjunction with one or more of the techniques described herein. Figures 1 to 12 One or more techniques further described.

[0128] In at least one embodiment, a software program 1302 (e.g., a user-implemented software program) utilizes one or more application programming interfaces (APIs) 1310 to perform various computational operations, such as memory reservations, matrix multiplications, arithmetic operations, or any computational operations performed by a parallel processing unit (PPU) (e.g., a graphics processing unit (GPU)), as further described herein. In at least one embodiment, the one or more APIs 1310 provide a set of callable functions 1312 (referred to herein as APIs, API functions, and / or functions) that each perform one or more computational operations, such as computational operations associated with parallel computing. In at least one embodiment, the one or more APIs 1310 provide functions 1312 to enable 1316 to scale one or more clocks of one or more cores based at least in part on the proximity of the one or more cores to each other and / or otherwise perform operations described herein. In at least one embodiment, one or more APIs 1310 provide functions 1312 to cause 1316 a neural network to perform one or more operations, such as by returning the called function to a processor, where the processor invokes the neural network.

[0129] In at least one embodiment, one or more software programs 1302 interact with or otherwise communicate with one or more APIs 1310 to perform one or more computing operations using one or more PPUs (e.g., GPUs). In at least one embodiment, the one or more computing operations using one or more PPUs include at least one or more groups of computing operations that are accelerated by being executed at least in part by the one or more PPUs. In at least one embodiment, one or more software programs 1302 interact with one or more APIs 1310 to facilitate parallel computing using remote or local interfaces.

[0130] In at least one embodiment, the interface is software instructions that, if executed, provide access to one or more functions 1312 provided by one or more APIs 1310. In at least one embodiment, the software programs 1302 use native interfaces when a software developer compiles one or more software programs 1302 in conjunction with one or more libraries 1306 that include or otherwise provide access to one or more APIs 1310. In at least one embodiment, the one or more software programs 1302 are statically compiled in conjunction with precompiled libraries 1306 or uncompiled source code that includes instructions for executing the one or more APIs 1310. In at least one embodiment, the one or more software programs 1302 are dynamically compiled and linked to the one or more precompiled libraries 1306 that include the one or more APIs 1310 using a linker.

[0131] In at least one embodiment, a software program 1302 uses a remote interface when a software developer executes a software program that utilizes a library 1306 including one or more APIs 1310 or otherwise communicates with a library 1306 including one or more APIs 1310 over a network or other remote communication medium. In at least one embodiment, the one or more libraries 1306 including one or more APIs 1310 are executed by a remote computing service (e.g., a computing resource service provider). In another embodiment, the one or more libraries 1306 including one or more APIs 1310 are executed by any other computing host that provides the one or more APIs 1310 to the one or more software programs 1302.

[0132] In at least one embodiment, a processor executing or using one or more software programs 1302 calls, uses, executes, or otherwise implements one or more APIs 1310 to allocate and otherwise manage memory for use by the software programs 1302. In at least one embodiment, one or more software programs 1302 utilize one or more APIs 1310 to allocate and otherwise manage memory for use by one or more portions of the software programs 1302 for acceleration using one or more PPUs (e.g., GPUs or any other accelerators or processors described further herein). These software programs 1302 are used to scale one or more clocks of one or more cores based at least in part on the proximity of the one or more cores to each other.

[0133] In at least one embodiment, API 1310 is an API for facilitating parallel computing. In at least one embodiment, API 1310 is any other API described further herein. In at least one embodiment, API 1310 is provided by a driver and / or runtime 1304. In at least one embodiment, API 1310 is provided by a CUDA user-mode driver. In at least one embodiment, API 1310 is provided by a CUDA runtime. In at least one embodiment, driver 1304 is data values ​​and software instructions that, if executed, perform or otherwise facilitate the operation of one or more functions 1312 of API 1310 during the loading and execution of one or more portions of software program 1302. In at least one embodiment, runtime 1304 is data values ​​and software instructions that, if executed, perform or otherwise facilitate the operation of one or more functions 1312 of API 1310 during the execution of software program 1302. In at least one embodiment, one or more software programs 1302 utilize one or more APIs 1310 implemented or otherwise provided by a driver and / or runtime 1304 to perform combined arithmetic operations by the one or more software programs 1302 during execution by one or more PPUs (e.g., GPUs).

[0134] In at least one embodiment, one or more software programs 1302 utilize one or more APIs 1310 provided by a driver and / or runtime 1304 to perform combined arithmetic operations for one or more PPUs (e.g., GPUs). In at least one embodiment, the one or more APIs 1310 provide combined arithmetic operations through the driver and / or runtime 1304, as described above. In at least one embodiment, one or more software programs 1302 utilize one or more APIs 1310 provided by the driver and / or runtime 1304 to allocate or otherwise reserve one or more blocks of memory 1314 for one or more PPUs (e.g., GPUs). In at least one embodiment, one or more software programs 1302 utilize one or more APIs 1310 provided by the driver and / or runtime 1304 to allocate or otherwise reserve blocks of memory. In at least one embodiment, the one or more APIs 1310 are used to perform combined arithmetic operations, as described below in conjunction with Figures 1 to 12 Any of those described in .

[0135] To improve the usability of the software program 1302 and / or optimize one or more portions of the software program 1302 for acceleration by one or more PPUs (e.g., GPUs), in one embodiment, the one or more APIs 1310 provide one or more API functions 1312 to implement a scheduling system that can be used or utilized by one or more computing devices, as described above and in conjunction with Figures 1 to 12 In at least one embodiment, block diagram 1300 depicts a processor comprising one or more circuits for executing one or more software programs to combine two or more application programming interfaces (APIs) into a single API. In at least one embodiment, block diagram 1300 depicts a system comprising one or more processors for executing one or more software programs to combine two or more application programming interfaces (APIs) into a single API. In at least one embodiment, the API is configured to cause 1316 to scale one or more clocks of one or more cores based at least in part on the proximity of the one or more cores to each other.

[0136] In at least one embodiment, Figures 14 to 18C As shown, a system for scaling one or more clocks of one or more cores based at least in part on proximity of the one or more cores to each other is used in servers and data centers.

[0137] In at least one embodiment, Figures 19 to 21 As shown, a system for scaling one or more clocks of one or more cores based at least in part on the proximity of the one or more cores to each other is used in and by other cloud computing and servers.

[0138] In at least one embodiment, Figures 22 to 25 As shown, a system for scaling one or more clocks of one or more cores based at least in part on proximity of the one or more cores to each other is used as part of a supercomputing system.

[0139] In at least one embodiment, Figure 26A and / or Figure 27 As shown, a system for scaling one or more clocks of one or more cores based at least in part on proximity of the one or more cores to each other is incorporated into other artificial intelligence.

[0140] In at least one embodiment, a system for scaling one or more clocks of one or more cores based at least in part on the proximity of the one or more cores to each other comprises: Figures 28 to 33 5G network usage shown.

[0141] In at least one embodiment, a system for scaling one or more clocks of one or more cores based at least in part on the proximity of the one or more cores to each other comprises: Figures 34 to 38 The computer-based system shown is used.

[0142] In at least one embodiment, a system for scaling one or more clocks of one or more cores based at least in part on the proximity of the one or more cores to each other comprises: Figures 39 to 44C The processing system shown is used.

[0143] In at least one embodiment, a system for scaling one or more clocks of one or more cores based at least in part on the proximity of the one or more cores to each other is configured to: Figures 45 to 50 The general calculation shown.

[0144] Technical solutions to technical problems

[0145] In at least one embodiment, a technical solution to a technical problem is provided. In at least one embodiment, the technical problem addressed is overheating in computer processors with multiple cores. In at least one embodiment, overheating in computer processors is problematic because it can lead to heat-induced failures. Computer designers often overcompensate for this problem by assigning the same offset value to all processor cores whose temperatures cannot be directly measured, which results in a reduction in the temperature of the processors during use. This practice also has the side effect of excessive "over-throttling" and unnecessarily reduces computer performance.

[0146] In at least one embodiment, the multiple offset technique is one way to avoid unnecessarily reducing all processor activity and over-throttling. In at least one embodiment, excessive over-throttling can be overcome by tracking which processor core combinations are "worst case" heat generators and assigning higher offset values ​​to those worst case combinations, while assigning lower offset values ​​or no offset values ​​to other processors that are not worst case.

[0147] In at least one embodiment, the technology improves upon this by enabling temperature control within a processor and at the processor core to be adjusted at a more precise, core-specific level, thereby improving performance and eliminating thermally induced failures.

[0148] In the following description, numerous specific details are set forth to provide a more thorough understanding of at least one embodiment. However, it will be apparent to one skilled in the art that the present invention can be practiced without one or more of these specific details.

[0149] Servers and Data Centers

[0150] The following figures illustrate, but are not limited to, exemplary network server and data center based systems that may be used to implement at least one embodiment.

[0151] Figure 14 A distributed system 1400 is shown in accordance with at least one embodiment. In at least one embodiment, the distributed system 1400 includes one or more client computing devices 1402, 1404, 1406, and 1408 configured to execute and operate client applications, such as web browsers, proprietary clients, and / or variations thereof, over one or more networks 1410. In at least one embodiment, a server 1412 can be communicatively coupled to the remote client computing devices 1402, 1404, 1406, and 1408 via the network 1410.

[0152] In at least one embodiment, server 1412 may be adapted to run one or more services or software applications, such as services and applications that can manage session activity for single sign-on (SSO) access across multiple data centers. In at least one embodiment, server 1412 may also provide other services or software applications that may include both non-virtualized and virtualized environments. In at least one embodiment, these services may be provided to users of client computing devices 1402, 1404, 1406, and / or 1408 as web-based services or cloud services or under a software as a service (SaaS) model. In at least one embodiment, users operating client computing devices 1402, 1404, 1406, and / or 1408 may, in turn, utilize one or more client applications to interact with server 1412 to utilize the services provided by these components.

[0153] In at least one embodiment, the software components 1418, 1420, and 1422 of system 1400 are implemented on server 1412. In at least one embodiment, one or more components of system 1400 and / or the services provided by these components can also be implemented by one or more of client computing devices 1402, 1404, 1406, and / or 1408. In at least one embodiment, a user operating a client computing device can then utilize one or more client applications to use the services provided by these components. In at least one embodiment, these components can be implemented in hardware, firmware, software, or a combination thereof. It should be understood that a variety of different system configurations are possible that can differ from distributed system 1400. Therefore, Figure 14 The illustrated embodiment is one example of a distributed system for implementing an embodiment system and is not intended to be limiting.

[0154] In at least one embodiment, client computing devices 1402, 1404, 1406, and / or 1408 may include different types of computing systems. In at least one embodiment, client computing devices may include portable handheld devices (e.g., Cellular phones, computing tablets, personal digital assistants (PDAs), or wearable devices (e.g., Google head-mounted display), running software (such as Microsoft Windows ) and / or various mobile operating systems (such as iOS, Windows Phone, Android, BlackBerry 10, Palm OS and / or their variants). In at least one embodiment, the device can support different applications, such as different Internet-related applications, email, short message service (SMS) applications, and can use various other communication protocols. In at least one embodiment, the client computing device can also include a general-purpose personal computer, for example, including a computer running various versions of Microsoft Apple In at least one embodiment, the client computing device can be a personal computer and / or laptop computer running various commercially available or any workstation computer running any of the UNIX-like operating systems, including but not limited to various GNU / Linux operating systems, such as Google Chrome OS. In at least one embodiment, client computing devices may also include electronic devices capable of communicating over one or more networks 1410, such as thin client computers, Internet-enabled gaming systems (e.g., with or without Internet access), and gesture input devices for Microsoft Xbox gaming consoles), and / or personal messaging devices. Figure 14 The distributed system 1400 in FIG. 1 is shown as having four client computing devices, but any number of client computing devices may be supported. Other devices (such as devices with sensors, etc.) may interact with the server 1412 .

[0155] In at least one embodiment, the network 1410 in the distributed system 1400 can be any type of network capable of supporting data communications using any of a variety of available protocols, including but not limited to TCP / IP (Transmission Control Protocol / Internet Protocol), SNA (Systems Network Architecture), IPX (Internetwork Packet Exchange), AppleTalk, and / or variations thereof. In at least one embodiment, the network 1410 can be a local area network (LAN), an Ethernet-based network, a token ring, a wide area network, the Internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a public switched telephone network (PSTN), an infrared network, a wireless network (e.g., a standard implemented in the Institute of Electrical and Electronics Engineers (IEEE) 802.11 protocol suite), a wireless network, or a combination thereof. and / or any other wireless protocols), and / or any combination of these and / or other networks.

[0156] In at least one embodiment, server 1412 may be comprised of one or more general purpose computers, dedicated server computers (including, for example, PC (personal computer) servers, In at least one embodiment, server 1412 may include one or more virtual machines running virtual operating systems or other computing architectures involving virtualization. In at least one embodiment, one or more flexible logical storage device pools may be virtualized to maintain virtual storage devices for the servers. In at least one embodiment, the virtual network may be controlled by server 1412 using software-defined networking. In at least one embodiment, server 1412 may be adapted to run one or more services or software applications.

[0157] In at least one embodiment, the server 1412 can run any operating system, and any commercially available server operating system. In at least one embodiment, the server 1412 can also run any of a variety of additional server applications and / or mid-tier applications, including HTTP (Hypertext Transfer Protocol) servers, FTP (File Transfer Protocol) servers, CGI (Common Gateway Interface) servers, Server, database server and / or variants thereof.In at least one embodiment, exemplary database servers include, but are not limited to, those commercially available from Oracle, Microsoft, Sybase, IBM (International Business Machines), and / or variants thereof.

[0158] In at least one embodiment, server 1412 may include one or more applications for analyzing and consolidating data feeds and / or event updates received from users of client computing devices 1402, 1404, 1406, and 1408. In at least one embodiment, data feeds and / or event updates may include, but are not limited to, data received from one or more third-party information sources and continuous data streams. feed, Updates or real-time updates, which may include real-time events related to sensor data applications, financial quoters, network performance measurement tools (e.g., network monitoring and business management applications), clickstream analysis tools, automobile traffic monitoring, and / or changes thereto. In at least one embodiment, server 1412 may also include one or more applications for displaying data feeds and / or real-time events via one or more display devices of client computing devices 1402, 1404, 1406, and 1408.

[0159] In at least one embodiment, the distributed system 1400 may further include one or more databases 1414 and 1416. In at least one embodiment, the database may provide a mechanism for storing information (such as user interaction information, usage pattern information, adaptation rule information, and other information). In at least one embodiment, the databases 1414 and 1416 may reside in various locations. In at least one embodiment, one or more of the databases 1414 and 1416 may reside on a non-transitory storage medium local to the server 1412 (and / or residing in the server 1412). In at least one embodiment, the databases 1414 and 1416 may be remote from the server 1412 and communicate with the server 1412 via a network-based connection or a dedicated connection. In at least one embodiment, the databases 1414 and 1416 may reside in a storage area network (SAN). In at least one embodiment, any necessary files for performing the functions attributed to the server 1412 may be appropriately stored locally on the server 1412 and / or stored remotely. In at least one embodiment, databases 1414 and 1416 may comprise relational databases, such as databases suitable for storing, updating, and retrieving data in response to SQL-formatted commands.

[0160] In at least one embodiment, Figure 14 At least one component shown or described is used to implement the combination Figure 1-13 In at least one embodiment, Figure 14 At least one component of the system is configured to cause scaling of one or more clocks of one or more cores based at least in part on proximity of the one or more cores to each other. In at least one embodiment, Figure 14 At least one component performs Figure 1The data center 120, the controller 105, at least two or more of the processors 110(a) to 110(n) and one or more thermal sensors 115(a) to 115(n), Figure 2 Firmware 220, configuration files 225, ACPI 230, policies 235, one or more processors 1-n, and / or including Figure 9 Method 900, Figure 10 Method 1000 and Figure 11 At least one aspect of one or more of the methods 1100 described herein.

[0161] Figure 15 An exemplary data center 1500 is shown in accordance with at least one embodiment. In at least one embodiment, data center 1500 includes, but is not limited to, a data center infrastructure layer 1510, a framework layer 1520, a software layer 1530, and an application layer 1540.

[0162] In at least one embodiment, Figure 15 As shown, the data center infrastructure layer 1510 may include a resource coordinator 1512, grouped computing resources 1514, and node computing resources ("node CRs") 1516(1)-1516(N), where "N" represents any complete positive integer. In at least one embodiment, the node CRs 1516(1)-1516(N) may include, but are not limited to, any number of central processing units ("CPUs") or other processors (including accelerators, field programmable gate arrays ("FPGAs"), graphics processors, etc.), memory devices (e.g., dynamic read-only memories), storage devices (e.g., solid-state drives or disk drives), network input / output ("NW I / O") devices, network switches, virtual machines ("VMs"), power modules, and cooling modules, etc. In at least one embodiment, one or more of the node CRs 1516(1)-1516(N) may be a server having one or more of the above-mentioned computing resources.

[0163] In at least one embodiment, the grouped computing resources 1514 may include separate groups of node CRs housed in one or more racks (not shown), or many racks (also not shown) housed in data centers at various geographic locations. The separate groups of node CRs within the grouped computing resources 1514 may include computing, networking, memory, or storage resources that can be configured or allocated to support groupings of one or more workloads. In at least one embodiment, several node CRs including CPUs or processors may be grouped in one or more racks to provide computing resources to support one or more workloads. In at least one embodiment, one or more racks may also include any number of power modules, cooling modules, and network switches, in any combination.

[0164] In at least one embodiment, resource coordinator 1512 may configure or otherwise control one or more nodes CR 1516(1)-1516(N) and / or grouped computing resources 1514. In at least one embodiment, resource coordinator 1512 may comprise a software design infrastructure ("SDI") management entity for data center 1500. In at least one embodiment, resource coordinator 1512 may comprise hardware, software, or some combination thereof.

[0165] In at least one embodiment, Figure 15 As shown, the framework layer 1520 includes, but is not limited to, a job scheduler 1532, a configuration manager 1534, a resource manager 1536, and a distributed file system 1538. In at least one embodiment, the framework layer 1520 may include a framework that supports software 1552 of the software layer 1530 and / or one or more applications 1542 of the application layer 1540. In at least one embodiment, the software 1552 or the application 1542 may include web-based service software or applications, such as those provided by Amazon Web Services, Google Cloud, and Microsoft Azure. In at least one embodiment, the framework layer 1520 may be, but is not limited to, a free and open source software web application framework, such as Apache Spark, which may utilize the distributed file system 1538 for large-scale data processing (e.g., "big data"). TM(hereinafter referred to as "Spark"). In at least one embodiment, the job scheduler 1532 may include a Spark driver to facilitate scheduling workloads supported by the various layers of the data center 1500. In at least one embodiment, the configuration manager 1534 may be capable of configuring the various layers, such as the software layer 1530 and the framework layer 1520 including Spark and a distributed file system 1538 for supporting large-scale data processing. In at least one embodiment, the resource manager 1536 may be capable of managing the mapping or allocation of clustered or grouped computing resources used to support the distributed file system 1538 and the job scheduler 1532. In at least one embodiment, the clustered or grouped computing resources may include grouped computing resources 1514 on the data center infrastructure layer 1510. In at least one embodiment, the resource manager 1536 may coordinate with the resource coordinator 1512 to manage these mapped or allocated computing resources.

[0166] In at least one embodiment, the software 1552 included in the software layer 1530 may include software used by at least a portion of the node CRs 1516(1)-1516(N), the grouped computing resources 1514, and / or the distributed file system 1538 of the framework layer 1520. The one or more types of software may include, but are not limited to, Internet web page search software, email virus scanning software, database software, and streaming video content software.

[0167] In at least one embodiment, the one or more applications 1542 included in the application layer 1540 may include one or more types of applications used by at least a portion of the node CRs 1516(1)-1516(N), the grouped computing resources 1514, and / or the distributed file system 1538 of the framework layer 1520. The one or more types of applications may include, but are not limited to, CUDA applications, 5G network applications, artificial intelligence applications, data center applications, and / or variations thereof.

[0168] In at least one embodiment, any of configuration manager 1534, resource manager 1536, and resource coordinator 1512 can implement any number and type of self-modification actions based on any number and type of data obtained in any technically feasible manner. In at least one embodiment, the self-modification actions can relieve a data center operator of data center 1500 from making potentially poor configuration decisions and can avoid underutilized and / or poorly performing portions of the data center.

[0169] In at least one embodiment, Figure 15 At least one component shown or described is used to implement the combination Figure 1-13In at least one embodiment, Figure 15 At least one component of the system is configured to cause scaling of one or more clocks of one or more cores based at least in part on proximity of the one or more cores to each other. In at least one embodiment, Figure 15 At least one component performs Figure 1 The data center 120, the controller 105, at least two or more of the processors 110(a) to 110(n) and one or more thermal sensors 115(a) to 115(n), Figure 2 Firmware 220, configuration files 225, ACPI 230, policies 235, one or more processors 1-n, and / or including Figure 9 Method 900, Figure 10 Method 1000 and Figure 11 At least one aspect of one or more of the methods 1100 described herein.

[0170] Figure 16 A system 1600 is illustrated, according to at least one embodiment, comprising a client-server network 1604 formed by a plurality of interconnected network server computers 1602. In at least one embodiment, in system 1600, each network server computer 1602 stores data accessible to the other network server computers 1602 and client computers 1606 and networks 1608 connected to wide area network 1604. In at least one embodiment, the configuration of client-server network 1604 can change over time as client computers 1606 and one or more networks 1608 connect and disconnect from network 1604, and as one or more trunk server computers 1602 are added to or removed from network 1604. In at least one embodiment, the client-server network includes client computers 1606 and networks 1608 when such client computers 1606 and networks 1608 are connected to network server computers 1602. In at least one embodiment, the term "computer" includes any device or machine capable of accepting data, applying a prescribed process to the data, and providing a result of the process.

[0171] In at least one embodiment, client-server network 1604 stores information accessible to network server computers 1602, remote networks 1608, and client computers 1606. In at least one embodiment, network server computers 1602 are formed from mainframe computers, minicomputers, and / or microcomputers, each having one or more processors. In at least one embodiment, server computers 1602 are linked together via wired and / or wireless transmission media (such as wires, fiber optic cables), and / or microwave transmission media, satellite transmission media, or other conductive, optical, or electromagnetic wave transmission media. In at least one embodiment, client computers 1606 access network server computers 1602 via similar wired or wireless transmission media. In at least one embodiment, client computers 1606 can connect to client-server network 1604 using a modem and a standard telephone communication network. In at least one embodiment, alternative carrier systems (such as cable and satellite communication systems) can also be used to connect to client-server network 1604. In at least one embodiment, other private or time-shared carrier systems can be used. In at least one embodiment, network 1604 is a global information network, such as the Internet. In at least one embodiment, the network is a private intranet that uses similar protocols to the Internet but with added security measures and restricted access controls.In at least one embodiment, the network 1604 is a private or semi-private network that uses a proprietary communication protocol.

[0172] In at least one embodiment, client computer 1606 is any end-user computer and may also be a mainframe computer, minicomputer, or microcomputer having one or more microprocessors. In at least one embodiment, server computer 1602 may sometimes be used as a client computer to access another server computer 1602. In at least one embodiment, remote network 1608 may be a local area network, a network added to a wide area network via an independent service provider (ISP) for the Internet, or another group of computers interconnected via a wired or wireless transmission medium with a fixed or time-varying configuration. In at least one embodiment, client computer 1606 may be linked to network 1604 independently or via remote network 1608 and access network 1604.

[0173] In at least one embodiment, Figure 16 At least one component shown or described is used to implement the combination Figure 1-13 In at least one embodiment, Figure 16 At least one component of the system is configured to cause scaling of one or more clocks of one or more cores based at least in part on proximity of the one or more cores to each other. In at least one embodiment, Figure 16At least one component performs Figure 1 The data center 120, the controller 105, at least two or more of the processors 110(a) to 110(n) and one or more thermal sensors 115(a) to 115(n), Figure 2 Firmware 220, configuration files 225, ACPI 230, policies 235, one or more processors 1-n, and / or including Figure 9 Method 900, Figure 10 Method 1000 and Figure 11 At least one aspect of one or more of the methods 1100 described herein.

[0174] Figure 17An example 1700 of a computer network 1708 connecting one or more computing machines according to at least one embodiment is shown. In at least one embodiment, the network 1708 can be any type of electrically connected group of computers, including, for example, the Internet, an intranet, a local area network (LAN), a wide area network (WAN), or an interconnected combination of these network types. In at least one embodiment, the connections within the network 1708 can be remote modems, Ethernet (IEEE 802.3), Token Ring (IEEE 802.5), Fiber Distributed Data Link Interface (FDDI), Asynchronous Transfer Mode (ATM), or any other communication protocol. In at least one embodiment, the computing devices linked to the network can be desktops, servers, portables, handhelds, set-top boxes, personal digital assistants (PDAs), terminals, or any other desired type or configuration. In at least one embodiment, network-connected devices can vary widely in terms of processing power, internal memory, and other performance depending on their functionality. In at least one embodiment, communications within the network and communications to or from computing devices connected to the network can be wired or wireless. In at least one embodiment, network 1708 may at least partially comprise the worldwide public Internet, which typically connects multiple users according to a client-server model based on the Transmission Control Protocol / Internet Protocol (TCP / IP) specification. In at least one embodiment, client-server networks are the predominant model for communication between two computers. In at least one embodiment, a client computer ("client") issues one or more commands to a server computer ("server"). In at least one embodiment, the server fulfills the client's commands by accessing available network resources and returning information to the client based on the client's commands. In at least one embodiment, client computer systems and network resources residing on network servers are assigned network addresses for identification during communications between network elements. In at least one embodiment, communications from other network-connected systems to the server will include the network address of the relevant server / network resource as part of the communication, allowing the appropriate destination of the data / request to be identified as the recipient. In at least one embodiment, when network 1708 comprises the global Internet, the network address is an IP address in TCP / IP format, which can, at least in part, route data to an email account, website, or other Internet utility residing on the server. In at least one embodiment, information and services residing on a network server may be made available to a client computer's web browser via a domain name (eg, www.site.com) that maps to the network server's IP address.

[0175] In at least one embodiment, a plurality of clients 1702, 1704, and 1706 are connected to a network 1708 via respective communication links. In at least one embodiment, each of these clients can access the network 1708 via any desired form of communication, such as via a dial-up modem connection, a cable link, a digital subscriber line (DSL), a wireless or satellite link, or any other form of communication. In at least one embodiment, each client can communicate using any machine (e.g., a personal computer (PC), workstation, dedicated terminal, personal data assistant (PDA), or other similar device) that is compatible with the network 1708. In at least one embodiment, the clients 1702, 1704, and 1706 may or may not be located in the same geographic area.

[0176] In at least one embodiment, multiple servers 1710, 1712, and 1714 are connected to a network 1718 to serve clients communicating with the network 1718. In at least one embodiment, each server is typically a powerful computer or device that manages network resources and responds to client commands. In at least one embodiment, the servers include computer-readable data storage media, such as hard drives and RAM memory, that store program instructions and data. In at least one embodiment, servers 1710, 1712, and 1714 run applications that respond to client commands. In at least one embodiment, server 1710 may run a web server application that responds to client requests for HTML pages and may also run a mail server application that receives and routes emails. In at least one embodiment, other applications may also run on server 1710, such as an FTP server or media server for streaming audio / video data to clients. In at least one embodiment, different servers may be dedicated to performing different tasks. In at least one embodiment, server 1710 may be a dedicated web server that manages website-related resources for different users, while server 1712 may be dedicated to providing email management. In at least one embodiment, other servers may be dedicated to media (audio, video, etc.), file transfer protocol (FTP), or a combination of any two or more services typically available or provided over a network. In at least one embodiment, each server may be in the same or different location than the other servers. In at least one embodiment, there may be multiple servers performing mirroring tasks for users, thereby alleviating congestion or minimizing traffic directed to and from a single server. In at least one embodiment, servers 1710, 1712, 1714 are under the control of a web hosting provider in the business of maintaining and delivering third-party content over network 1718.

[0177] In at least one embodiment, a web hosting provider delivers services to two different types of clients. In at least one embodiment, one type, which may be referred to as a browser, requests content, such as web pages, email messages, video clips, etc., from servers 1710, 1712, 1714. In at least one embodiment, a second type, which may be referred to as a user, hires the web hosting provider to maintain network resources, such as a website, and make them available to the browser. In at least one embodiment, the user contracts with the web hosting provider to make available the memory space, processor capacity, and communication bandwidth required for the network resources they desire, depending on the amount of server resources they desire to utilize.

[0178] In at least one embodiment, in order for the web hosting provider to serve both clients, an application that manages the network resources hosted by the server must be appropriately configured. In at least one embodiment, the program configuration process involves defining a set of parameters that at least partially control the application's responses to browser requests and also at least partially define the server resources available to a particular user.

[0179] In one embodiment, intranet server 1716 communicates with network 1708 via a communication link. In at least one embodiment, intranet server 1716 communicates with server manager 1718. In at least one embodiment, server manager 1718 includes a database of application configuration parameters used in servers 1710, 1712, 1714. In at least one embodiment, a user modifies database 1720 via intranet 1716, and server manager 1718 interacts with servers 1710, 1712, 1714 to modify application parameters so that they match the contents of the database. In at least one embodiment, a user logs into intranet 1716 by connecting to intranet 1716 via computer 1702 and entering authentication information such as a username and password.

[0180] In at least one embodiment, when a user wishes to log in to a new service or modify an existing service, intranet server 1716 authenticates the user and provides the user with an interactive screen display / control panel that allows the user to access configuration parameters for a particular application. In at least one embodiment, the user is presented with a plurality of modifiable text boxes that describe aspects of the configuration of the user's website or other network resource. In at least one embodiment, if the user desires to increase the memory space reserved for their website on the server, the user is provided with a field in which the user specifies the desired memory space. In at least one embodiment, in response to receiving this information, intranet server 1716 updates database 1720. In at least one embodiment, server manager 1718 forwards this information to the appropriate server, and the new parameters are used during application operation. In at least one embodiment, intranet server 1716 is configured to provide the user with access to configuration parameters for hosted network resources (e.g., web pages, email, FTP sites, media sites, etc.) that the user has contracted with a web hosting service provider.

[0181] In at least one embodiment, Figure 17 At least one component shown or described is used to implement the combination Figure 1-13 In at least one embodiment, Figure 17 At least one component of the system is configured to cause scaling of one or more clocks of one or more cores based at least in part on proximity of the one or more cores to each other. In at least one embodiment, Figure 17 At least one component performs Figure 1 The data center 120, the controller 105, at least two or more of the processors 110(a) to 110(n) and one or more thermal sensors 115(a) to 115(n), Figure 2 Firmware 220, configuration files 225, ACPI 230, policies 235, one or more processors 1-n, and / or including Figure 9 Method 900, Figure 10 Method 1000 and Figure 11 At least one aspect of one or more of the methods 1100 described herein.

[0182] Figure 18AA networked computer system 1800A is shown according to at least one embodiment. In at least one embodiment, the networked computer system 1800A includes a plurality of nodes or personal computers ("PCs") 1802, 1818, 1820. In at least one embodiment, the personal computer or node 1802 includes a processor 1814, a memory 1816, a camera 1804, a microphone 1806, a mouse 1808, a speaker 1810, and a monitor 1812. In at least one embodiment, the PCs 1802, 1818, 1820 can each run one or more desktop servers, such as an internal network within a given company, or can be servers of a general-purpose network not limited to a particular environment. In at least one embodiment, each PC node of the network has one server, such that each PC node of the network represents a specific network server with a specific network URL address. In at least one embodiment, each server defaults to a default web page for users of that server, which itself can contain embedded URLs pointing to further subpages for that user on that server, or to other servers on the network or to pages on other servers.

[0183] In at least one embodiment, nodes 1802, 1818, 1820 and other nodes of the network are interconnected via a medium 1822. In at least one embodiment, the medium 1822 can be a communication channel such as an Integrated Services Digital Network ("ISDN"). In at least one embodiment, the various nodes of the networked computer system can be connected via various communication media, including a local area network ("LAN"), a plain old telephone line ("POTS") (sometimes referred to as a public switched telephone network ("PSTN")), and / or variations thereof. In at least one embodiment, the various nodes of the network can also constitute computer system users interconnected via a network such as the Internet. In at least one embodiment, each server on the network (operated from a particular node of the network at a given instance) has a unique address or identification within the network, which can be specified according to a URL.

[0184] In at least one embodiment, a plurality of multipoint conferencing units ("MCUs") can thus be used to transmit data to and from various nodes or "endpoints" of a conferencing system. In at least one embodiment, the nodes and / or MCUs can be interconnected via ISDN links or through a local area network ("LAN"), in addition to various other communication media (such as, nodes connected via the Internet). In at least one embodiment, the nodes of a conferencing system can generally be connected directly to a communication medium (such as a LAN) or through an MCU, and the conferencing system can include other nodes or elements, such as routers, servers, and / or variations thereof.

[0185] In at least one embodiment, processor 1814 is a general-purpose programmable processor. In at least one embodiment, the processor of a node of networked computer system 1800A may also be a dedicated video processor. In at least one embodiment, the various peripherals and components of a node (such as those of node 1802) may be different from those of other nodes. In at least one embodiment, node 1818 and node 1820 may be configured to be the same as or different from node 1802. In at least one embodiment, a node may be implemented on any suitable computer system other than a PC system.

[0186] Figure 18B 18. A networked computer system 1800B is shown according to at least one embodiment. In at least one embodiment, system 1800B shows a network (such as LAN 1824) that can be used to interconnect various nodes that can communicate with each other. In at least one embodiment, attached to LAN 1824 are multiple nodes, such as PC nodes 1826, 1828, 1830. In at least one embodiment, the nodes can also be connected to the LAN via a network server or other device. In at least one embodiment, system 1800B includes other types of nodes or elements, for example, routers, servers, and nodes.

[0187] Figure 18C A networked computer system 1800C is shown in accordance with at least one embodiment. In at least one embodiment, system 1800C shows a WWW system with communications across a backbone communications network, such as the Internet 1832, which can be used to interconnect various nodes of the network. In at least one embodiment, the WWW is a set of protocols that operate on top of the Internet and allow graphical interface systems to operate on top of it to access information through the Internet. In at least one embodiment, attached to the Internet 1832 in the WWW are multiple nodes, such as PCs 1840, 1842, 1844. In at least one embodiment, the nodes interface with other nodes of the WWW through WWW HTTP servers, such as servers 1834, 1836. In at least one embodiment, PC 1844 can be a PC that forms a node of the network 1832, and PC 1844 itself runs its server 1836, although for illustrative purposes only. Figure 18C PC 1844 and server 1836 are shown separately in FIG.

[0188] In at least one embodiment, the WWW is a distributed type of application characterized by WWW HTTP, the protocol of the WWW, which runs on top of the Internet's Transmission Control Protocol / Internet Protocol ("TCP / IP"). In at least one embodiment, the WWW can therefore be characterized by a set of protocols (i.e., HTTP) running on the Internet as its "backbone."

[0189] In at least one embodiment, a web browser is an application running on a node of the network in a WWW-compatible network system that allows users of a particular server or node to view such information and, therefore, to search for graphics and text-based files linked together using hypertext links embedded in documents or files available from servers on a network that understands HTTP. In at least one embodiment, when a user retrieves a given web page from a first server associated with a first node using another server on a network such as the Internet, the retrieved document may have different hypertext links embedded therein, and a local copy of the page is created locally on the retrieving user's machine. In at least one embodiment, when a user clicks on a hypertext link, the locally stored information associated with the selected hypertext link is typically sufficient to allow the user's machine to open a connection over the Internet to the server indicated by the hypertext link.

[0190] In at least one embodiment, more than one user may be coupled to each HTTP server, for example, via a LAN (such as LAN 1838, as shown with respect to WWW HTTP server 1834). In at least one embodiment, system 1800C may also include other types of nodes or elements. In at least one embodiment, the WWW HTTP server is an application running on a machine, such as a PC. In at least one embodiment, each user may be considered to have a unique "server," as shown with respect to PC 1844. In at least one embodiment, a server may be considered to be a server, such as WWW HTTP server 1834, that provides access to the network for a LAN, or for one or more nodes, or for one or more LANs. In at least one embodiment, there are multiple users, each with a desktop PC or a node on the network, each desktop PC potentially establishing a server for its user. In at least one embodiment, each server is associated with a specific network address or URL that, when accessed, provides a default web page for that user. In at least one embodiment, the web page may contain further links (embedded URLs) pointing to further subpages for that user on that server, or to other servers on the network, or to pages on other servers on the network.

[0191] Cloud computing and services

[0192] The following figures illustrate, but are not limited to, exemplary cloud-based systems that can be used to implement at least one embodiment.

[0193] In at least one embodiment, cloud computing is a style of computing in which dynamically scalable and often virtualized resources are provided as services over the Internet. In at least one embodiment, users do not need knowledge, expertise, or control over the technical infrastructure that supports them; the technical infrastructure can be referred to as being "in the cloud." In at least one embodiment, cloud computing combines infrastructure as a service, platform as a service, software as a service, and other variations with a common theme of relying on the Internet to meet users' computing needs. In at least one embodiment, a typical cloud deployment (such as in a private cloud (e.g., an enterprise network)) or a data center (DC) in a public cloud (e.g., the Internet) can consist of thousands of servers (or alternatively, VMs), hundreds of Ethernet, Fibre Channel, or Fibre Channel over Ethernet (FCoE) ports, switching and storage infrastructure, etc. In at least one embodiment, the cloud can also consist of network service infrastructure, such as IPsec VPN concentrators, firewalls, load balancers, wide area network (WAN) optimizers, etc. In at least one embodiment, remote subscribers can securely access cloud applications and services by connecting via a VPN tunnel (e.g., an IPsec VPN tunnel).

[0194] In at least one embodiment, cloud computing is a model for enabling convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, servers, storage, applications, and services) that can be quickly provisioned and released with minimal management effort or service provider interaction.

[0195] In at least one embodiment, cloud computing is characterized by on-demand self-service, where consumers can automatically and unilaterally provision computing capacity, such as server time and network storage, as needed, without requiring human interaction with each service provider. In at least one embodiment, cloud computing is characterized by widespread network access, where capacity is available over the network and accessed through standard mechanisms that facilitate use by heterogeneous thin or thick client platforms (e.g., mobile phones, laptops, and PDAs). In at least one embodiment, cloud computing is characterized by resource pooling, where a provider's computing resources are pooled to serve multiple consumers using a multi-tenant model, where different physical and virtual resources are dynamically signed up and reallocated based on consumer demand. In at least one embodiment, there is a sense of location independence, as consumers typically have no control or knowledge of the exact location of provisioned resources but may be able to specify the location at a higher level of abstraction (e.g., country, state, or data center). In at least one embodiment, examples of resources include storage, processing, memory, network bandwidth, and virtual machines. In at least one embodiment, cloud computing is characterized by rapid elasticity, where capacity can be quickly and elastically provisioned (in some cases automatically) for rapid scaling down and rapidly released for rapid scaling up. In at least one embodiment, the capacity available for provisioning generally appears unlimited to the consumer and can be purchased at any time and in any quantity. In at least one embodiment, cloud computing is characterized by metered services, where the cloud system automatically controls and optimizes resource usage by utilizing metering capabilities at some level of abstraction appropriate to the type of service (e.g., storage, processing, bandwidth, and active user accounts). In at least one embodiment, resource usage can be monitored, controlled, and reported, thereby providing transparency to both the provider and the consumer of the utilized service.

[0196] In at least one embodiment, cloud computing can be associated with a variety of services. In at least one embodiment, cloud software as a service (SaaS) can refer to a service that provides consumers with the ability to use a provider's applications running on a cloud infrastructure. In at least one embodiment, the applications can be accessed from various client devices through a thin client interface such as a web browser (e.g., web-based email). In at least one embodiment, the consumer does not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, storage, or even individual application capabilities, with the possible exception of limited user-specific application configuration settings.

[0197] In at least one embodiment, cloud platform as a service (PaaS) may refer to a service in which the capability provided to the consumer is to deploy consumer-created or acquired applications onto a cloud infrastructure, where these applications are created using programming languages ​​and tools supported by the provider. In at least one embodiment, the consumer does not manage or control the underlying cloud infrastructure, including networks, servers, operating systems, or storage, but does have control over the deployed applications and possibly the configuration of the application hosting environment.

[0198] In at least one embodiment, cloud infrastructure as a service (IaaS) can refer to a service in which the capabilities provided to the consumer are processing, storage, networking, and other basic computing resources upon which the consumer can deploy and run arbitrary software, which may include operating systems and applications. In at least one embodiment, the consumer does not manage or control the underlying cloud infrastructure, but rather has control over the operating system, storage, deployed applications, and possibly limited control over selected networking components (e.g., host firewalls).

[0199] In at least one embodiment, cloud computing can be deployed in different ways. In at least one embodiment, a private cloud may refer to cloud infrastructure that operates only for an organization. In at least one embodiment, a private cloud may be managed by an organization or a third party and may exist on-premises or off-premises. In at least one embodiment, a community cloud may refer to cloud infrastructure shared by several organizations and supporting a specific community with shared concerns (e.g., mission, security requirements, policies, and compliance considerations). In at least one embodiment, a community cloud may be managed by an organization or a third party and may exist on-premises or off-premises. In at least one embodiment, a public cloud may refer to cloud infrastructure that is available to the general public or a large industry group and owned by an organization providing cloud services. In at least one embodiment, a hybrid cloud may refer to a cloud infrastructure that is a composite of two or more clouds (private, community, or public) that remain unique entities but are bound together by standardized or proprietary technologies that enable data and application portability (e.g., cloud bursting for load balancing between clouds). In at least one embodiment, a cloud computing environment is service-oriented, focusing on statelessness, low coupling, modularity, and semantic interoperability.

[0200] In at least one embodiment, Figures 18A-18C At least one component shown or described is used to implement the combination Figure 1-13 In at least one embodiment, Figures 18A-18C At least one component of the system is configured to cause scaling of one or more clocks of one or more cores based at least in part on proximity of the one or more cores to each other. In at least one embodiment, Figures 18A-18C At least one component performs Figure 1 The data center 120, the controller 105, at least two or more of the processors 110(a) to 110(n) and one or more thermal sensors 115(a) to 115(n), Figure 2 Firmware 220, configuration files 225, ACPI 230, policies 235, one or more processors 1-n, and / or including Figure 9 Method 900, Figure 10 Method 1000 and Figure 11 At least one aspect of one or more of the methods 1100 described herein.

[0201] Figure 19 One or more components of a system environment 1900 are shown, according to at least one embodiment, in which services may be provided as third-party network services. In at least one embodiment, the third-party network may be referred to as a cloud, a cloud network, a cloud computing network, and / or variations thereof. In at least one embodiment, the system environment 1900 includes one or more client computing devices 1904, 1906, and 1908, which can be used by users to interact with a third-party network infrastructure system 1902 that provides the third-party network services (which may be referred to as cloud computing services). In at least one embodiment, the third-party network infrastructure system 1902 may include one or more computers and / or servers.

[0202] It should be understood that Figure 19 The third-party network infrastructure system 1902 depicted in FIG may have other components in addition to those depicted. Further, Figure 19 In at least one embodiment, the third party network infrastructure system 1902 may have a Figure 19 More or fewer components may be depicted, two or more components may be combined, or there may be a different configuration or arrangement of components.

[0203] In at least one embodiment, client computing devices 1904, 1906, and 1908 can be configured to operate a client application, such as a web browser, a proprietary client application, or some other application that can be used by users of the client computing devices to interact with the third-party network infrastructure system 1902 to utilize services provided by the third-party network infrastructure system 1902. Although the exemplary system environment 1900 is shown with three client computing devices, any number of client computing devices can be supported. In at least one embodiment, other devices, such as devices with sensors, can interact with the third-party network infrastructure system 1902. In at least one embodiment, one or more networks 1910 can facilitate communication and data exchange between the client computing devices 1904, 1906, and 1908 and the third-party network infrastructure system 1902.

[0204] In at least one embodiment, the services provided by the third-party network infrastructure system 1902 may include a host of services available on-demand to users of the third-party network infrastructure system. In at least one embodiment, a variety of services may also be provided, including but not limited to online data storage and backup solutions, web-based email services, hosted office suites and document collaboration services, database management and processing, managed technical support services, and / or variations thereof. In at least one embodiment, the services provided by the third-party network infrastructure system may be dynamically scalable to meet the needs of its users.

[0205] In at least one embodiment, a specific instantiation of a service provided by the third-party network infrastructure system 1902 may be referred to as a "service instance." In at least one embodiment, generally, any service available to a user from a third-party network service provider system via a communication network (such as the Internet) is referred to as a "third-party network service." In at least one embodiment, in a public third-party network environment, the servers and systems that comprise the third-party network service provider system are different from the customer's own on-premises servers and systems. In at least one embodiment, the third-party network service provider system can host applications, and users can subscribe to and use the applications on demand via a communication network (such as the Internet).

[0206] In at least one embodiment, services within a computer network third-party network infrastructure may include protected computer network access to storage, hosted databases, hosted web servers, software applications, or other services provided to users by a third-party network provider. In at least one embodiment, services may include password-protected access to remote storage devices on the third-party network via the Internet. In at least one embodiment, services may include a hosted relational database and scripting language middleware engine based on a web service for private use by networked developers. In at least one embodiment, services may include access to an email software application hosted on a website of a third-party network provider.

[0207] In at least one embodiment, the third-party network infrastructure system 1902 may include a suite of application, middleware, and database service offerings delivered to customers in a self-service, subscription-based, elastically scalable, reliable, highly available, and secure manner. In at least one embodiment, the third-party network infrastructure system 1902 may also provide computing and analytical services related to "big data." In at least one embodiment, the term "big data" is generally used to refer to extremely large data sets that can be stored and manipulated by analysts and researchers to visualize, detect trends, and / or otherwise interact with the data. In at least one embodiment, big data and related applications can be hosted and / or manipulated by the infrastructure system at many levels and at varying scales. In at least one embodiment, dozens, hundreds, or thousands of processors linked in parallel may operate on such data to render it or simulate external forces acting on the data or its representations. In at least one embodiment, these data sets may include structured data (such as structured data organized in a database or otherwise according to a structured model) and / or unstructured data (e.g., emails, images, data blobs (binary large objects), web pages, complex event processing). In at least one embodiment, by leveraging the ability of embodiments to focus more (or fewer) computing resources on a target relatively quickly, third-party network infrastructure systems may be better available to perform tasks on large data sets based on demand from businesses, government agencies, research organizations, private individuals, groups of like-minded individuals or organizations, or other entities.

[0208] In at least one embodiment, the third-party network infrastructure system 1902 can be adapted to automatically provision, manage, and track customer subscriptions to services provided by the third-party network infrastructure system 1902. In at least one embodiment, the third-party network infrastructure system 1902 can provide third-party network services via different deployment models. In at least one embodiment, the services can be provided under a public third-party network model, in which the third-party network infrastructure system 1902 is owned by the organization selling the third-party network services and makes the services available to the general public or businesses across various industries. In at least one embodiment, the services can be provided under a private third-party network model, in which the third-party network infrastructure system 1902 operates solely for a single organization and can provide services to one or more entities within the organization. In at least one embodiment, the third-party network services can also be provided under a community third-party network model, in which the third-party network infrastructure system 1902 and the services provided by the third-party network infrastructure system 1902 are shared by several organizations within a related community. In at least one embodiment, the third-party network services can also be provided under a hybrid third-party network model, which is a combination of two or more different models.

[0209] In at least one embodiment, the services provided by the third-party network infrastructure system 1902 may include one or more services provided under the Software as a Service (SaaS) category, the Platform as a Service (PaaS) category, the Infrastructure as a Service (IaaS) category, or other service categories including hybrid services. In at least one embodiment, a customer may subscribe to one or more services provided by the third-party network infrastructure system 1902 via a subscription order. In at least one embodiment, the third-party network infrastructure system 1902 then performs processing to provide the services in the customer's subscription order.

[0210] In at least one embodiment, the services provided by the third-party network infrastructure system 1902 may include, but are not limited to, application services, platform services, and infrastructure services. In at least one embodiment, application services may be provided by the third-party network infrastructure system via a SaaS platform. In at least one embodiment, the SaaS platform may be configured to provide third-party network services that fall under the SaaS category. In at least one embodiment, the SaaS platform may provide the ability to build and deliver a suite of on-demand applications on an integrated development and deployment platform. In at least one embodiment, the SaaS platform may manage and control the underlying software and infrastructure used to provide SaaS services. In at least one embodiment, by utilizing the services provided by the SaaS platform, customers may utilize applications executed on the third-party network infrastructure system. In at least one embodiment, customers may obtain application services without the need for customers to purchase separate licenses and support. In at least one embodiment, a variety of different SaaS services may be provided. In at least one embodiment, examples include, but are not limited to, services that provide solutions for sales performance management, enterprise integration, and business agility for large organizations.

[0211] In at least one embodiment, platform services may be provided by the third-party network infrastructure system 1902 via a PaaS platform. In at least one embodiment, the PaaS platform may be configured to provide third-party network services that fall under the PaaS category. In at least one embodiment, examples of platform services may include, but are not limited to, services that enable organizations to consolidate existing applications on a shared common architecture, as well as the ability to build new applications that utilize the shared services provided by the platform. In at least one embodiment, the PaaS platform may manage and control the underlying software and infrastructure used to provide PaaS services. In at least one embodiment, customers may obtain PaaS services provided by the third-party network infrastructure system 1902 without requiring the customer to purchase separate licenses and support.

[0212] In at least one embodiment, by utilizing the services provided by the PaaS platform, customers can adopt programming languages ​​and tools supported by the third-party network infrastructure system and also control the deployed services. In at least one embodiment, the platform services provided by the third-party network infrastructure system may include database third-party network services, middleware third-party network services, and third-party network services. In at least one embodiment, the database third-party network service may support a shared service deployment model that enables organizations to aggregate database resources and provide database as a service to customers in the form of a database third-party network. In at least one embodiment, in the third-party network infrastructure system, the middleware third-party network service can provide customers with a platform to develop and deploy different business applications, and the third-party network service can provide customers with a platform to deploy applications.

[0213] In at least one embodiment, a variety of infrastructure services may be provided by an IaaS platform within a third-party network infrastructure system. In at least one embodiment, the infrastructure services facilitate the management and control of underlying computing resources (such as storage, network, and other basic computing resources) by customers utilizing services provided by SaaS and PaaS platforms.

[0214] In at least one embodiment, the third-party network infrastructure system 1902 may also include infrastructure resources 1930 for providing resources for providing various services to customers of the third-party network infrastructure system. In at least one embodiment, the infrastructure resources 1930 may include a pre-integrated and optimized combination of hardware (such as servers, storage, and networking resources) for executing the services and other resources provided by the PaaS platform and the SaaS platform.

[0215] In at least one embodiment, resources in the third-party network infrastructure system 1902 can be shared by multiple users and dynamically reallocated based on demand. In at least one embodiment, resources can be allocated to users in different time zones. In at least one embodiment, the third-party network infrastructure system 1902 can enable a first group of users in a first time zone to utilize the resources of the third-party network infrastructure system for a specified number of hours, and then enable the reallocation of the same resources to another group of users in a different time zone, thereby maximizing resource utilization.

[0216] In at least one embodiment, a plurality of internal shared services 1932 shared by different components or modules of the third-party network infrastructure system 1902 may be provided to enable services provided by the third-party network infrastructure system 1902. In at least one embodiment, these internal shared services may include, but are not limited to, security and identity services, integration services, enterprise library services, enterprise manager services, virus scanning and whitelisting services, high availability, backup and recovery services, services for enabling third-party network support, email services, notification services, file transfer services, and / or variations thereof.

[0217] In at least one embodiment, the third-party network infrastructure system 1902 can provide comprehensive management of third-party network services (e.g., SaaS, PaaS, and IaaS services) within the third-party network infrastructure system. In at least one embodiment, the third-party network management functionality can include capabilities for provisioning, managing, and tracking customer subscriptions received by the third-party network infrastructure system 1902 and / or variations thereof.

[0218] In at least one embodiment, Figure 19As shown, third-party network management functionality may be provided by one or more modules, such as an order management module 1920, an order coordination module 1922, an order provisioning module 1924, an order management and monitoring module 1926, and an identity management module 1928. In at least one embodiment, these modules may include or be provided using one or more computers and / or servers, which may be general-purpose computers, dedicated server computers, server farms, server clusters, or any other suitable arrangement and / or combination.

[0219] In at least one embodiment, at step 1934, a customer using a client device (such as client computing device 1904, 1906, or 1908) can interact with third-party network infrastructure system 1902 by requesting one or more services provided by third-party network infrastructure system 1902 and placing an order for a subscription to one or more services provided by third-party network infrastructure system 1902. In at least one embodiment, the customer can access a third-party network user interface (UI), such as third-party network UI 1912, third-party network UI 1914, and / or third-party network UI 1916, and place an order via these UIs. In at least one embodiment, the order information received by third-party network infrastructure system 1902 in response to the customer placing the order can include information identifying the customer and one or more services provided by third-party network infrastructure system 1902 to which the customer wishes to subscribe.

[0220] In at least one embodiment, at step 1936, the order information received from the customer can be stored in order database 1918. In at least one embodiment, if this is a new order, a new record can be created for the order. In at least one embodiment, order database 1918 can be one of several databases operated by third-party network infrastructure system 1918 and in conjunction with other system components.

[0221] In at least one embodiment, at step 1938, the order information may be forwarded to order management module 1920, which may be configured to perform billing and accounting functions related to the order, such as verifying the order and, upon verification, booking an order.

[0222] In at least one embodiment, at step 1940, information about the order may be communicated to order coordination module 1922, which is configured to coordinate the provisioning of services and resources for the order placed by the customer. In at least one embodiment, order coordination module 1922 may utilize the services of order provisioning module 1924 for provisioning. In at least one embodiment, order coordination module 1922 enables management of the business processes associated with each order and applies business logic to determine whether the order should proceed with provisioning.

[0223] In at least one embodiment, at step 1942, upon receiving an order for a new subscription, order coordination module 1922 sends a request to order provisioning module 1924 to allocate resources and configure the resources necessary to fulfill the subscription order. In at least one embodiment, order provisioning module 1924 implements resource allocation for the services ordered by the customer. In at least one embodiment, order provisioning module 1924 provides a level of abstraction between the third-party network services provided by third-party network infrastructure system 1900 and the physical implementation layer used to provision the resources for providing the requested services. In at least one embodiment, this enables order coordination module 1922 to be isolated from implementation details, such as whether services and resources are actually provisioned in real time, or pre-provisioned and allocated / assigned only upon request.

[0224] In at least one embodiment, once the services and resources are provisioned, a notification may be sent to the subscribing client indicating that the requested service is now ready for use, step 1944. In at least one embodiment, information (e.g., a link) may be sent to the client that enables the client to begin using the requested service.

[0225] In at least one embodiment, at step 1946, the customer's subscription order may be managed and tracked by the order management and monitoring module 1926. In at least one embodiment, the order management and monitoring module 1926 may be configured to collect usage statistics regarding the customer's use of the subscription service. In at least one embodiment, statistics may be collected regarding the amount of storage used, the amount of data transferred, the number of users, and the amount and / or changes in system power-up time and system power-down time.

[0226] In at least one embodiment, the third-party network infrastructure system 1900 may include an identity management module 1928 configured to provide identity services, such as access management and authorization services within the third-party network infrastructure system 1900. In at least one embodiment, the identity management module 1928 may control information about customers who wish to utilize services provided by the third-party network infrastructure system 1902. In at least one embodiment, such information may include information authenticating the identities of such customers and information describing which actions those customers are authorized to perform with respect to various system resources (e.g., files, directories, applications, communication ports, memory segments, etc.). In at least one embodiment, the identity management module 1928 may also include managing descriptive information about each customer, as well as information about how and by whom the descriptive information may be accessed and modified.

[0227] In at least one embodiment, Figure 19 At least one component shown or described is used to implement the combination Figure 1-13 In at least one embodiment, Figure 19 At least one component of the system is configured to cause scaling of one or more clocks of one or more cores based at least in part on proximity of the one or more cores to each other. In at least one embodiment, Figure 19 At least one component performs Figure 1 The data center 120, the controller 105, at least two or more of the processors 110(a) to 110(n) and one or more thermal sensors 115(a) to 115(n), Figure 2 Firmware 220, configuration files 225, ACPI 230, policies 235, one or more processors 1-n, and / or including Figure 9 Method 900, Figure 10 Method 1000 and Figure 11 At least one aspect of one or more of the methods 1100 described herein.

[0228] Figure 20 A cloud computing environment 2002 is shown in accordance with at least one embodiment. In at least one embodiment, the cloud computing environment 2002 includes one or more computer systems / servers 2004 with which computing devices such as a personal digital assistant (PDA) or cell phone 2006A, a desktop computer 2006B, a laptop computer 2006C, and / or an automobile computer system 2006N communicate. In at least one embodiment, this allows infrastructure, platforms, and / or software to be provided as a service from the cloud computing environment 2002 so that each client does not need to individually maintain such resources. It should be understood that Figure 20 The types of computing devices 2006A-N shown are intended to be illustrative only, and the cloud computing environment 2002 may communicate with any type of computerized device over any type of network and / or network / addressable connection (eg, using a web browser).

[0229] In at least one embodiment, computer system / server 2004, which may be represented as a cloud computing node, is operable with numerous other general-purpose or special-purpose computing system environments or configurations. In at least one embodiment, examples of computing systems, environments, and / or configurations that may be suitable for use with computer system / server 2004 include, but are not limited to, personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputer systems, mainframe computer systems, and distributed cloud computing environments including any of the foregoing, and / or variations thereof.

[0230] In at least one embodiment, the computer system / server 2004 can be described in the general context of computer system-executable instructions (such as program modules) executed by a computer system. In at least one embodiment, program modules include routines, programs, objects, components, logic, data structures, etc. that perform specific tasks or implement specific abstract data types. In at least one embodiment, the exemplary computer system / server 2004 can be practiced in a distributed cloud computing environment, where tasks are performed by remote processing devices linked through a communication network. In at least one embodiment, in a distributed cloud computing environment, program modules can be located in both local and remote computer system storage media, including memory storage devices.

[0231] In at least one embodiment, Figure 20 At least one component shown or described is used to implement the combination Figure 1-13 In at least one embodiment, Figure 20 At least one component of the system is configured to cause scaling of one or more clocks of one or more cores based at least in part on proximity of the one or more cores to each other. In at least one embodiment, Figure 20 At least one component performs Figure 1 The data center 120, the controller 105, at least two or more of the processors 110(a) to 110(n) and one or more thermal sensors 115(a) to 115(n), Figure 2 Firmware 220, configuration files 225, ACPI 230, policies 235, one or more processors 1-n, and / or including Figure 9 Method 900, Figure 10 Method 1000 and Figure 11 At least one aspect of one or more of the methods 1100 described herein.

[0232] Figure 21 FIG. 2 shows a cloud computing environment 2002 ( Figure 20 ) provides a set of functional abstraction layers. It should be understood in advance that Figure 21 The components, layers, and functions shown in are intended to be illustrative only, and the components, layers, and functions may vary.

[0233] In at least one embodiment, the hardware and software layer 2102 includes hardware and software components. In at least one embodiment, examples of hardware components include mainframes, servers based on various RISC (Reduced Instruction Set Computer) architectures, various computing systems, supercomputing systems, storage devices, networks, networking components, and / or variations thereof. In at least one embodiment, examples of software components include network application server software, various application server software, various database software, and / or variations thereof.

[0234] In at least one embodiment, the virtualization layer 2104 provides an abstraction layer from which the following exemplary virtual entities can be provided: virtual servers, virtual storage, virtual networks (including virtual private networks), virtual applications, virtual clients, and / or variations thereof.

[0235] In at least one embodiment, the management layer 2106 provides various functions. In at least one embodiment, resource provisioning provides dynamic acquisition of computing resources and other resources for performing tasks within the cloud computing environment. In at least one embodiment, metering provides usage tracking when resources are utilized within the cloud computing environment, as well as billing or invoicing for the consumption of those resources. In at least one embodiment, resources may include application software licenses. In at least one embodiment, security provides authentication for users and tasks, as well as protection of data and other resources. In at least one embodiment, a user interface provides access to the cloud computing environment for both users and system administrators. In at least one embodiment, service level management provides allocation and management of cloud computing resources so that required service levels are met. In at least one embodiment, service level agreement (SLA) management provides pre-positioning and acquisition of cloud computing resources in anticipation of future demand for the cloud computing resources according to the SLA.

[0236] In at least one embodiment, the workload layer 2108 provides functionality that utilizes a cloud computing environment. In at least one embodiment, examples of workloads and functionality that can be provided from this layer include: mapping and navigation, software development and management, educational services, data analysis and processing, transaction processing, and service delivery.

[0237] Supercomputing

[0238] The following figures illustrate, but are not limited to, exemplary supercomputer-based systems that may be used to implement at least one embodiment.

[0239] In at least one embodiment, a supercomputer may refer to a hardware system that exhibits significant parallelism and includes at least one chip, wherein the chips in the system are interconnected by a network and placed in a hierarchically organized housing. In at least one embodiment, a large hardware system that fills a computer room with several racks, each rack containing several boards / rack modules, each board / rack module containing several chips all interconnected by a scalable network, is a specific example of a supercomputer. In at least one embodiment, a single rack of such a large hardware system is another example of a supercomputer. In at least one embodiment, a single chip that exhibits significant parallelism and includes several hardware components can also be considered a supercomputer because as feature sizes can decrease, the amount of hardware that can be combined in a single chip can also increase.

[0240] In at least one embodiment, Figure 21 At least one component shown or described is used to implement the combination Figure 1-13 In at least one embodiment, Figure 21 At least one component of the system is configured to cause scaling of one or more clocks of one or more cores based at least in part on proximity of the one or more cores to each other. In at least one embodiment, Figure 21 At least one component performs Figure 1 The data center 120, the controller 105, at least two or more of the processors 110(a) to 110(n) and one or more thermal sensors 115(a) to 115(n), Figure 2 Firmware 220, configuration files 225, ACPI 230, policies 235, one or more processors 1-n, and / or including Figure 9 Method 900, Figure 10 Method 1000 and Figure 11 At least one aspect of one or more of the methods 1100 described herein.

[0241] Figure 22A chip-level supercomputer according to at least one embodiment is shown. In at least one embodiment, the main calculation is performed in a finite state machine (2204) called a thread unit inside an FPGA or ASIC chip. In at least one embodiment, a task and synchronization network (2202) connects the finite state machine and is used to dispatch threads and execute operations in the correct order. In at least one embodiment, a memory network (2206, 2210) is used to access the on-chip cache hierarchy (2208, 2212) of the multi-level partition. In at least one embodiment, a memory controller (2216) and an off-chip memory network (2214) are used to access off-chip memory. In at least one embodiment, when the design is not suitable for a single logic chip, an I / O controller (2218) is used for cross-chip communication.

[0242] In at least one embodiment, Figure 22 At least one component shown or described is used to implement the combination Figure 1-13 In at least one embodiment, Figure 22 At least one component of the system is configured to cause scaling of one or more clocks of one or more cores based at least in part on proximity of the one or more cores to each other. In at least one embodiment, Figure 22 At least one component performs Figure 1 The data center 120, the controller 105, at least two or more of the processors 110(a) to 110(n) and one or more thermal sensors 115(a) to 115(n), Figure 2 Firmware 220, configuration files 225, ACPI 230, policies 235, one or more processors 1-n, and / or including Figure 9 Method 900, Figure 10 Method 1000 and Figure 11 At least one aspect of one or more of the methods 1100 described herein.

[0243] Figure 23 A supercomputer at the rack module level is shown according to at least one embodiment. In at least one embodiment, within the rack module, there are multiple FPGA or ASIC chips (2302) connected to one or more DRAM units (2304) that constitute the main accelerator memory. In at least one embodiment, each FPGA / ASIC chip is connected to its neighboring FPGA / ASIC chips using differential high-speed signaling (2306) using a wide bus on the board. In at least one embodiment, each FPGA / ASIC chip is also connected to at least one high-speed serial communication cable.

[0244] In at least one embodiment, Figure 23At least one component shown or described is used to implement the combination Figure 1-13 In at least one embodiment, Figure 23 At least one component of the system is configured to cause scaling of one or more clocks of one or more cores based at least in part on proximity of the one or more cores to each other. In at least one embodiment, Figure 23 At least one component performs Figure 1 The data center 120, the controller 105, at least two or more of the processors 110(a) to 110(n) and one or more thermal sensors 115(a) to 115(n), Figure 2 Firmware 220, configuration files 225, ACPI 230, policies 235, one or more processors 1-n, and / or including Figure 9 Method 900, Figure 10 Method 1000 and Figure 11 At least one aspect of one or more of the methods 1100 described herein.

[0245] Figure 24 A rack-scale supercomputer is shown in accordance with at least one embodiment.

[0246] In at least one embodiment, Figure 24 At least one component shown or described is used to implement the combination Figure 1-13 In at least one embodiment, Figure 24 At least one component of the system is configured to cause scaling of one or more clocks of one or more cores based at least in part on proximity of the one or more cores to each other. In at least one embodiment, Figure 24 At least one component performs Figure 1 The data center 120, the controller 105, at least two or more of the processors 110(a) to 110(n) and one or more thermal sensors 115(a) to 115(n), Figure 2 Firmware 220, configuration files 225, ACPI 230, policies 235, one or more processors 1-n, and / or including Figure 9 Method 900, Figure 10 Method 1000 and Figure 11 At least one aspect of one or more of the methods 1100 described herein.

[0247] Figure 25 An overall system-level supercomputer according to at least one embodiment is shown. In at least one embodiment, see Figure 24 and Figure 25, between rack modules in a rack and across racks throughout the system, high-speed serial optical or copper cables (2402, 2502) are used to implement a scalable, potentially incomplete, hypercube network. In at least one embodiment, one of the accelerator's FPGA / ASIC chips is connected to a host system (2504) via a PCI-Express connection. In at least one embodiment, the host system includes a host microprocessor (2508) on which the software portion of the application runs, and memory consisting of one or more host memory DRAM units (2506) that are coherent with the memory on the accelerator. In at least one embodiment, the host system can be a separate module on one of the racks, or can be integrated with one of the modules of the supercomputer. In at least one embodiment, a circular topology of cube connections provides communication links to create a hypercube network for a large supercomputer. In at least one embodiment, small groups of FPGA / ASIC chips on a rack module can act as a single hypercube node, increasing the total number of external links per group compared to a single chip. In at least one embodiment, a group includes chips A, B, C, and D on a rack module with an internal wide differential bus connecting A, B, C, and D in a ring organization. In at least one embodiment, there are 12 serial communication cables connecting the rack modules to the outside world. In at least one embodiment, chip A on the rack module connects to serial communication cables 0, 1, and 2. In at least one embodiment, chip B connects to cables 3, 4, and 5. In at least one embodiment, chip C connects to cables 6, 7, and 8. In at least one embodiment, chip D connects to cables 9, 10, and 11. In at least one embodiment, the entire group {A, B, C, D} that makes up the rack modules can form a hypercube node within a supercomputer system, with up to 212 = 4096 rack modules (16384 FPGA / ASIC chips). In at least one embodiment, in order for chip A to send a message out on link 4 of the group {A, B, C, D}, the message must first be routed to chip B using the on-board differential wide bus connection. In at least one embodiment, a message arriving on link 4 of the group {A, B, C, D} destined for chip A (i.e., to B) must also first be routed to the correct destination chip (A) within the group {A, B, C, D}. In at least one embodiment, parallel supercomputer systems of other sizes may also be implemented.

[0248] In at least one embodiment, Figure 25 At least one component shown or described is used to implement the combination Figure 1-13 In at least one embodiment, Figure 25 At least one component of the system is configured to cause scaling of one or more clocks of one or more cores based at least in part on proximity of the one or more cores to each other. In at least one embodiment, Figure 25 At least one component performs Figure 1 The data center 120, the controller 105, at least two or more of the processors 110(a) to 110(n) and one or more thermal sensors 115(a) to 115(n), Figure 2 Firmware 220, configuration files 225, ACPI 230, policies 235, one or more processors 1-n, and / or including Figure 9 Method 900, Figure 10 Method 1000 and Figure 11 At least one aspect of one or more of the methods 1100 described herein.

[0249] AI

[0250] The following figures illustrate, but are not limited to, exemplary artificial intelligence-based systems that can be used to implement at least one embodiment.

[0251] Figure 26A Inference and / or training logic 2615 is shown for performing inference and / or training operations associated with one or more embodiments. Figure 26A and / or Figure 26B Provides details about the inference and / or training logic 2615.

[0252] In at least one embodiment, the inference and / or training logic 2615 may include, but is not limited to, code and / or data storage 2601 for storing forward and / or output weights and / or input / output data, and / or other parameters used to configure neurons or layers of a neural network being trained and / or used for inference in aspects of one or more embodiments. In at least one embodiment, the training logic 2615 may include or be coupled to code and / or data storage 2601 for storing graph code or other software to control the timing and / or sequence in which weights and / or other parameter information are loaded to configure logic, including integer and / or floating point units (collectively referred to as arithmetic logic units (ALUs)). In at least one embodiment, the code (such as graph code) loads weights or other parameter information into the processor ALU based on the architecture of the neural network to which such code corresponds. In at least one embodiment, the code and / or data storage 2601 stores weight parameters and / or input / output data for each layer of a neural network that is trained or used in conjunction with one or more embodiments during forward propagation of input / output data and / or weight parameters during training and / or inference using aspects of one or more embodiments. In at least one embodiment, any portion of code and / or data storage 2601 may be included with other on-chip or off-chip data storage devices, including the processor's L1, L2, or L3 cache memory or system memory.

[0253] In at least one embodiment, any portion of code and / or data storage 2601 may be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, code and / or code and / or data storage 2601 may be cache memory, dynamic random addressable memory ("DRAM"), static random addressable memory ("SRAM"), non-volatile memory (e.g., flash memory), or other storage device. In at least one embodiment, the choice of whether code and / or code and / or data storage 2601 is internal or external to a processor, for example, or includes DRAM, SRAM, flash memory, or some other type of storage, may depend on the available storage on-chip versus off-chip, the latency requirements of the training and / or inference functions being performed, the batch size of data used in inference and / or training of a neural network, or some combination of these factors.

[0254] In at least one embodiment, the inference and / or training logic 2615 may include, but is not limited to, code and / or data storage 2605 for storing backpropagation and / or output weights and / or input / output data corresponding to neurons or layers of a neural network trained and / or used for inference in aspects of one or more embodiments. In at least one embodiment, the code and / or data storage 2605 stores weight parameters and / or input / output data for each layer of a neural network trained or used in conjunction with one or more embodiments during backpropagation of input / output data and / or weight parameters during training and / or inference using aspects of one or more embodiments. In at least one embodiment, the training logic 2615 may include or be coupled to code and / or data storage 2605 to store graph code or other software to control the timing and / or sequence in which weights and / or other parameter information are loaded to configure logic, including integer and / or floating point units (collectively referred to as arithmetic logic units (ALUs)).

[0255] In at least one embodiment, code (such as graph code) causes weights or other parameter information to be loaded into the processor ALU based on the architecture of the neural network to which such code corresponds. In at least one embodiment, any portion of code and / or data storage 2605 may be included with other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory. In at least one embodiment, any portion of code and / or data storage 2605 may be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, code and / or data storage 2605 may be cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other storage devices. In at least one embodiment, the choice of whether code and / or data storage 2605 is internal or external to the processor, for example, or includes DRAM, SRAM, flash memory, or some other type of storage, may depend on the available on-chip versus off-chip storage, the latency requirements of the training and / or inference functions being performed, the batch size of data used in inference and / or training of the neural network, or some combination of these factors.

[0256] In at least one embodiment, code and / or data store 2601 and code and / or data store 2605 may be separate storage structures. In at least one embodiment, code and / or data store 2601 and code and / or data store 2605 may be a combined storage structure. In at least one embodiment, code and / or data store 2601 and code and / or data store 2605 may be partially combined and partially separate. In at least one embodiment, any portion of code and / or data store 2601 and code and / or data store 2605 may be included with other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory.

[0257] In at least one embodiment, inference and / or training logic 2615 may include, but is not limited to, one or more arithmetic logic units (“ALUs”) 2610 , including integer and / or floating point units, for performing logical and / or mathematical operations based at least in part on or directed by training and / or inference code (e.g., graph code), the results of which may produce activations (e.g., output values ​​from a layer or neuron within a neural network) stored in activation storage 2620 , which is a function of input / output and / or weight parameter data stored in code and / or data storage 2601 and / or code and / or data storage 2605 . In at least one embodiment, the activations stored in activation storage 2620 are generated based on linear algebra and / or matrix-based math performed by ALU 2610 in response to executing instructions or other code, where weight values ​​stored in code and / or data storage 2605 and / or data storage 2601 are used as operands along with other values ​​such as bias values, gradient information, momentum values, or other parameters or hyperparameters, any or all of which may be stored in code and / or data storage 2605 or code and / or data storage 2601 or in another storage on or off-chip.

[0258] In at least one embodiment, one or more ALUs 2610 are included within one or more processors or other hardware logic devices or circuits, while in another embodiment, one or more ALUs 2610 may be external to the processor or other hardware logic devices or circuits (e.g., coprocessors) that use them. In at least one embodiment, ALUs 2610 may be included within an execution unit of a processor or otherwise within an ALU bank accessible by an execution unit of a processor, either within the same processor or distributed across different processors of different types (e.g., central processing units, graphics processing units, fixed function units, etc.). In at least one embodiment, code and / or data storage 2601, code and / or data storage 2605, and activation storage 2620 may share a processor or other hardware logic device or circuit, while in another embodiment, they may be in different processors or other hardware logic devices or circuits, or in some combination of the same and different processors or other hardware logic devices or circuits. In at least one embodiment, any portion of active storage 2620 may be included with other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory. Additionally, inference and / or training code may be stored with other code accessible to the processor or other hardware logic or circuitry and fetched and / or processed using the processor's fetch, decode, schedule, execute, retire, and / or other logic circuitry.

[0259] In at least one embodiment, activation storage 2620 can be cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other storage devices. In at least one embodiment, activation storage 2620 can be completely or partially internal or external to one or more processors or other logic circuits. In at least one embodiment, the choice of whether activation storage 2620 is internal or external to the processor, for example, or includes DRAM, SRAM, flash memory, or some other storage type, can depend on the available storage on-chip versus off-chip, the latency requirements of the training and / or inference functions being performed, the batch size of the data used in inference and / or training of the neural network, or some combination of these factors.

[0260] In at least one embodiment, Figure 26A The inference and / or training logic 2615 shown in FIG can be used in conjunction with an application specific integrated circuit (“ASIC”), such as the one from Google. Processing unit from Graphcore TM Inference Processing Unit (IPU), or from Intel (e.g., "Lake Crest") processor. In at least one embodiment, Figure 26A The inference and / or training logic 2615 shown in FIG may be used in conjunction with central processing unit (“CPU”) hardware, graphics processing unit (“GPU”) hardware, or other hardware such as a field programmable gate array (“FPGA”).

[0261] Figure 26B Inference and / or training logic 2615 is shown in accordance with at least one embodiment. In at least one embodiment, inference and / or training logic 2615 may include, but is not limited to, hardware logic in which computing resources are dedicated or otherwise used exclusively in conjunction with weight values ​​or other information corresponding to one or more neuron layers within a neural network. In at least one embodiment, Figure 26B The inference and / or training logic 2615 shown in FIG can be combined with an application specific integrated circuit (ASIC) (such as the one from Google Processing unit from Graphcore TM Inference Processing Unit (IPU), or from Intel (e.g., "Lake Crest") processor. In at least one embodiment, Figure 26BThe inference and / or training logic 2615 shown in FIG may be used in conjunction with central processing unit (CPU) hardware, graphics processing unit (GPU) hardware, or other hardware such as a field programmable gate array (FPGA). In at least one embodiment, the inference and / or training logic 2615 includes, but is not limited to, code and / or data storage 2601 and code and / or data storage 2605, which may be used to store code (e.g., graph code), weight values, and / or other information, including bias values, gradient information, momentum values, and / or other parameter or hyperparameter information. Figure 26B In at least one embodiment described in , each of code and / or data storage 2601 and code and / or data storage 2605 is associated with dedicated computing resources, such as computing hardware 2602 and computing hardware 2606, respectively. In at least one embodiment, each of computing hardware 2602 and computing hardware 2606 includes one or more ALUs that perform mathematical functions (such as linear algebraic functions) solely on the information stored in code and / or data storage 2601 and code and / or data storage 2605, respectively, with the results being stored in activation storage 2620.

[0262] In at least one embodiment, each code and / or data storage 2601 and 2605, and corresponding computational hardware 2602 and 2606, respectively, corresponds to a different layer of a neural network, such that the resulting activations from one storage / computation pair 2601 / 2602 in code and / or data storage 2601 and computational hardware 2602 are provided as input to the next storage / computation pair 2605 / 2606 in code and / or data storage 2605 and computational hardware 2606, mirroring the conceptual organization of the neural network. In at least one embodiment, each of storage / computation pairs 2601 / 2602 and 2605 / 2606 can correspond to more than one neural network layer. In at least one embodiment, additional storage / computation pairs (not shown) can be included in the inference and / or training logic 2615, either after or in parallel with storage / computation pairs 2601 / 2602 and 2605 / 2606.

[0263] In at least one embodiment, Figures 26A-26B At least one component shown or described is used to implement the combination Figure 1-13 In at least one embodiment, Figures 26A-26B At least one component of the system is configured to cause scaling of one or more clocks of one or more cores based at least in part on proximity of the one or more cores to each other. In at least one embodiment, Figures 26A-26B At least one component performs Figure 1The data center 120, the controller 105, at least two or more of the processors 110(a) to 110(n) and one or more thermal sensors 115(a) to 115(n), Figure 2 Firmware 220, configuration files 225, ACPI 230, policies 235, one or more processors 1-n, and / or including Figure 9 Method 900, Figure 10 Method 1000 and Figure 11 At least one aspect of one or more of the methods 1100 described herein.

[0264] Figure 27 The training and deployment of a deep neural network according to at least one embodiment is shown. In at least one embodiment, an untrained neural network 2706 is trained using a training dataset 2702. In at least one embodiment, the training framework 2704 is the PyTorch framework, while in other embodiments, the training framework 2704 is TensorFlow, Boost, Caffe, Microsoft Cognitive Toolkit / CNTK, MXNet, Chainer, Keras, Deeplearning4j, or other training frameworks. In at least one embodiment, the training framework 2704 trains the untrained neural network 2706 and enables it to be trained using the processing resources described herein to generate a trained neural network 2708. In at least one embodiment, the weights can be randomly selected or selected by pre-training using a deep belief network. In at least one embodiment, the training can be performed in a supervised, partially supervised, or unsupervised manner.

[0265] In at least one embodiment, untrained neural network 2706 is trained using supervised learning, where training dataset 2702 includes inputs paired with expected outputs for the inputs, or where training dataset 2702 includes inputs with known outputs, and the outputs of neural network 2706 are manually graded. In at least one embodiment, untrained neural network 2706 is trained in a supervised manner, processing inputs from training dataset 2702 and comparing the resulting outputs to a set of expected or desired outputs. In at least one embodiment, errors are then backpropagated through untrained neural network 2706. In at least one embodiment, training framework 2704 adjusts the weights that control untrained neural network 2706. In at least one embodiment, training framework 2704 includes tools for monitoring how well untrained neural network 2706 converges toward a model (such as trained neural network 2708) suitable for generating correct answers (such as results 2714) based on input data (such as new dataset 2712). In at least one embodiment, the training framework 2704 repeatedly trains the untrained neural network 2706 while adjusting the weights using a loss function and an adjustment algorithm (such as stochastic gradient descent) to refine the output of the untrained neural network 2706. In at least one embodiment, the training framework 2704 trains the untrained neural network 2706 until the untrained neural network 2706 achieves a desired accuracy. In at least one embodiment, the trained neural network 2708 can then be deployed to implement any number of machine learning operations.

[0266] In at least one embodiment, untrained neural network 2706 is trained using unsupervised learning, wherein untrained neural network 2706 attempts to train itself using unlabeled data. In at least one embodiment, unsupervised learning training dataset 2702 will include input data without any associated output data or "ground truth" data. In at least one embodiment, untrained neural network 2706 can learn groupings within training dataset 2702 and can determine how individual inputs relate to untrained dataset 2702. In at least one embodiment, unsupervised training can be used to generate a self-organizing map in trained neural network 2708 that is capable of performing operations useful in reducing the dimensionality of new dataset 2712. In at least one embodiment, unsupervised training can also be used to perform anomaly detection, which allows identification of data points in new dataset 2712 that deviate from the normal pattern of new dataset 2712.

[0267] In at least one embodiment, semi-supervised learning can be used, which is a technique in which a mixture of labeled and unlabeled data is included in the training dataset 2702. In at least one embodiment, the training framework 2704 can be used to perform incremental learning, such as through transfer learning techniques. In at least one embodiment, incremental learning enables the trained neural network 2708 to adapt to new datasets 2712 without forgetting the knowledge infused into the trained neural network 1408 during initial training.

[0268] In at least one embodiment, Figure 27 At least one component shown or described is used to implement the combination Figure 1-13 In at least one embodiment, Figure 27 At least one component of the system is configured to cause scaling of one or more clocks of one or more cores based at least in part on proximity of the one or more cores to each other. In at least one embodiment, Figure 27 At least one component performs Figure 1 The data center 120, the controller 105, at least two or more of the processors 110(a) to 110(n) and one or more thermal sensors 115(a) to 115(n), Figure 2 Firmware 220, configuration files 225, ACPI 230, policies 235, one or more processors 1-n, and / or including Figure 9 Method 900, Figure 10 Method 1000 and Figure 11 At least one aspect of one or more of the methods 1100 described herein.

[0269] 5G network

[0270] The following figures illustrate, but are not limited to, exemplary 5G network-based systems that may be used to implement at least one embodiment.

[0271] Figure 28 The architecture of a system 2800 of a network according to at least one embodiment is shown. In at least one embodiment, the system 2800 is shown as including user equipment (UE) 2802 and UE 2804. In at least one embodiment, the UEs 2802 and 2804 are shown as smartphones (e.g., handheld touchscreen mobile computing devices that can connect to one or more cellular networks), but may also include any mobile or non-mobile computing device, such as a personal digital assistant (PDA), a pager, a laptop computer, a desktop computer, a wireless handheld device, or any computing device that includes a wireless communication interface.

[0272] In at least one embodiment, any of UE 2802 and UE 2804 may comprise an Internet of Things (IoT) UE, which may include a network access layer designed for low-power IoT applications that utilize short-lived UE connections. In at least one embodiment, the IoT UE may utilize technologies such as machine-to-machine (M2M) or machine-type communication (MTC) for exchanging data with an MTC server or device via a public land mobile network (PLMN), proximity-based services (ProSe), or device-to-device (D2D) communication, a sensor network, or an IoT network. In at least one embodiment, the M2M or MTC data exchange may be machine-initiated data exchange. In at least one embodiment, the IoT network describes interconnected IoT UEs, which may include uniquely identifiable embedded computing devices (within the Internet infrastructure) with short-lived connections. In at least one embodiment, the IoT UE may execute background applications (e.g., keep-alive messages, status updates, etc.) to facilitate connectivity to the IoT network.

[0273] In at least one embodiment, UE 2802 and UE 2804 can be configured to connect (e.g., be communicatively coupled) to a radio access network (RAN) 2816. In at least one embodiment, RAN 2816 can be, for example, an evolved universal mobile telecommunications system (UMTS) terrestrial radio access network (E-UTRAN), a NextGen RAN (NG RAN), or some other type of RAN. In at least one embodiment, UE 2802 and UE 2804 utilize connection 2812 and connection 2814, respectively, each of which includes a physical communication interface or layer. In at least one embodiment, connections 2812 and 2814 are shown as air interfaces for achieving communicative coupling and can be consistent with a cellular communication protocol, such as a Global System for Mobile Communications (GSM) protocol, a Code Division Multiple Access (CDMA) network protocol, a Push-to-Talk (PTT) protocol, a PTT over Cellular (POC) protocol, a Universal Mobile Telecommunications System (UMTS) protocol, a 3GPP Long Term Evolution (LTE) protocol, a Fifth Generation (5G) protocol, a New Radio (NR) protocol, and variations thereof.

[0274] In at least one embodiment, the UEs 2802 and 2804 may also directly exchange communication data via a ProSe interface 2806. In at least one embodiment, the ProSe interface 2806 may alternatively be referred to as a side link interface, which includes one or more logical channels, including but not limited to a physical side link control channel (PSCCH), a physical side link shared channel (PSSCH), a physical side link discovery channel (PSDCH), and a physical side link broadcast channel (PSBCH).

[0275] In at least one embodiment, UE 2804 is shown as being configured to access an access point (AP) 2810 via a connection 2808. In at least one embodiment, connection 2808 may comprise a local wireless connection, such as a connection consistent with any IEEE 802.11 protocol, wherein AP 2810 would include Wi-Fi. In at least one embodiment, AP 2810 is shown connected to the Internet and not to the core network of the wireless system.

[0276] In at least one embodiment, the RAN 2816 may include one or more access nodes that enable connections 2812 and 2814. In at least one embodiment, these access nodes (ANs) may be referred to as base stations (BSs), NodeBs, evolved NodeBs (eNBs), next generation NodeBs (gNBs), RAN nodes, etc., and may include ground stations (e.g., terrestrial access points) or satellite stations that provide coverage within a geographic area (e.g., a cell). In at least one embodiment, the RAN 2816 may include one or more RAN nodes (e.g., macro RAN nodes 2818) for providing macro cells and one or more RAN nodes (e.g., low power (LP) RAN nodes 2820) for providing femto cells or pico cells (e.g., cells with smaller coverage areas, smaller user capacity, or higher bandwidth than macro cells).

[0277] In at least one embodiment, either of the RAN nodes 2818 and 2820 may terminate the air interface protocol and may be the first point of contact for the UEs 2802 and 2804. In at least one embodiment, either of the RAN nodes 2818 and 2820 may implement various logical functions of the RAN 2816, including but not limited to radio network controller (RNC) functions such as radio bearer management, uplink and downlink dynamic radio resource management, and data packet scheduling and mobility management.

[0278] In at least one embodiment, UE 2802 and UE 2804 may be configured to communicate with each other or with any of RAN node 2818 and RAN node 2820 over a multi-carrier communication channel using orthogonal frequency division multiplexing (OFDM) communication signals in accordance with various communication techniques, such as, but not limited to, orthogonal frequency division multiple access (OFDMA) communication techniques (e.g., for downlink communication) or single-carrier frequency division multiple access (SC-FDMA) communication techniques (e.g., for uplink and ProSe or sidelink communication), and / or variations thereof. In at least one embodiment, the OFDM signal may include multiple orthogonal subcarriers.

[0279] In at least one embodiment, a downlink resource grid can be used for downlink transmissions from either RAN nodes 2818 and 2820 to UEs 2802 and 2804, while uplink transmissions can utilize similar techniques. In at least one embodiment, the grid can be a time-frequency grid, referred to as a resource grid or time-frequency resource grid, representing the physical resources in the downlink in each time slot. In at least one embodiment, this time-frequency plane representation is common practice in OFDM systems, making it intuitive for radio resource allocation. In at least one embodiment, each column and row of the resource grid corresponds to an OFDM symbol and an OFDM subcarrier, respectively. In at least one embodiment, the duration of the resource grid in the time domain corresponds to a time slot in a radio frame. In at least one embodiment, the smallest time-frequency unit in the resource grid is represented as a resource element. In at least one embodiment, each resource grid includes multiple resource blocks, which describe the mapping of certain physical channels to resource elements. In at least one embodiment, each resource block includes a collection of resource elements. In at least one embodiment, in the frequency domain, this can represent the minimum number of resources that can currently be allocated. In at least one embodiment, there are several different physical downlink channels transmitted using such resource blocks.

[0280] In at least one embodiment, a physical downlink shared channel (PDSCH) may carry user data and higher layer signaling to UEs 2802 and 2804. In at least one embodiment, a physical downlink control channel (PDCCH) may carry information regarding, among other things, the transport format and resource allocation associated with the PDSCH channel. In at least one embodiment, it may also inform UEs 2802 and 2804 of the transport format, resource allocation, and HARQ (Hybrid Automatic Repeat Request) information associated with the uplink shared channel. In at least one embodiment, downlink scheduling (allocation of control and shared channel resource blocks to UEs 2802 within a cell) may generally be performed at either RAN node 2818 or 2820 based on channel quality information fed back from either UE 2802 or 2804. In at least one embodiment, downlink resource allocation information may be sent on a PDCCH for (e.g., allocated to) each of UEs 2802 and 2804.

[0281] In at least one embodiment, the PDCCH may use control channel elements (CCEs) to transmit control information. In at least one embodiment, before being mapped to resource elements, the PDCCH complex symbols may first be organized into quadruplets, which may then be permuted using a sub-block interleaver for rate matching. In at least one embodiment, each PDCCH may be transmitted using one or more of these CCEs, where each CCE may correspond to nine sets of four physical resource elements referred to as resource element groups (REGs). In at least one embodiment, four quadrature phase shift keying (QPSK) symbols may be mapped to each REG. In at least one embodiment, one or more CCEs may be used to transmit the PDCCH, depending on the size of the downlink control information (DCI) and the channel conditions. In at least one embodiment, there may be four or more different PDCCH formats (e.g., aggregation levels, L=1, 2, 4, or 8) defined in LTE with different numbers of CCEs.

[0282] In at least one embodiment, an enhanced physical downlink control channel (EPDCCH) using PDSCH resources may be used for control information transmission. In at least one embodiment, EPDCCH may be transmitted using one or more enhanced control channel elements (ECCEs). In at least one embodiment, each ECCE may correspond to nine sets of four physical resource elements referred to as enhanced resource element groups (EREGs). In at least one embodiment, ECCEs may have other numbers of EREGs in some cases.

[0283] In at least one embodiment, the RAN 2816 is shown as being communicatively coupled to a core network (CN) 2838 via an S1 interface 2822. In at least one embodiment, the CN 2838 can be an evolved packet core (EPC) network, a NextGen packet core (NPC) network, or some other type of CN. In at least one embodiment, the S1 interface 2822 is divided into two parts: an S1-U interface 2826, which carries traffic data between the RAN nodes 2818 and 2820 and the serving gateway (S-GW) 2830; and an S1-Mobility Management Entity (MME) interface 2824, which is a signaling interface between the RAN nodes 2818 and 2820 and the MME 2828.

[0284] In at least one embodiment, CN 2838 includes an MME 2828, an S-GW 2830, a Packet Data Network (PDN) Gateway (P-GW) 2834, and a Home Subscriber Server (HSS) 2832. In at least one embodiment, MME 2828 can be functionally similar to the control plane of a traditional Serving General Packet Radio Service (GPRS) Support Node (SGSN). In at least one embodiment, MME 2828 can manage mobility aspects of access, such as gateway selection and tracking area list management. In at least one embodiment, HSS 2832 can include a database for network users, including subscription-related information used to support network entities handling communication sessions. In at least one embodiment, CN 2838 can include one or more HSSs 2832, depending on the number of mobile users, device capacity, network organization, and the like. In at least one embodiment, HSS 2832 can provide support for routing / roaming, authentication, authorization, naming / addressing resolution, location dependencies, and the like.

[0285] In at least one embodiment, the S-GW 2830 may terminate the S1 interface 2822 towards the RAN 2816 and route data packets between the RAN 2816 and the CN 2838. In at least one embodiment, the S-GW 2830 may be the local mobility anchor for inter-RAN node handovers and may also provide an anchor for inter-3GPP mobility. In at least one embodiment, other responsibilities may include lawful interception, charging, and some policy enforcement.

[0286] In at least one embodiment, the P-GW 2834 can terminate the SGi interface toward the PDN. In at least one embodiment, the P-GW 2834 can route data packets between the EPC network 2838 and an external network, such as a network including an application server 2840 (or application function (AF)), via an Internet Protocol (IP) interface 2842. In at least one embodiment, the application server 2840 can be an element that provides applications using IP bearer resources using a core network (e.g., a UMTS packet service (PS) domain, an LTE PS data service, etc.). In at least one embodiment, the P-GW 2834 is shown as being communicatively coupled to the application server 2840 via an IP communication interface 2842. In at least one embodiment, the application server 2840 can also be configured to support one or more communication services (e.g., voice over Internet protocol (VoIP) sessions, PTT sessions, group communication sessions, social networking services, etc.) for UEs 2802 and 2804 via the CN 2838.

[0287] In at least one embodiment, P-GW 2834 can also be a node for policy enforcement and charging data collection. In at least one embodiment, Policy and Charging Enforcement Function (PCRF) 2836 is the policy and charging control element of CN 2838. In at least one embodiment, in a non-roaming scenario, a single PCRF can exist in the Home Public Land Mobile Network (HPLMN) associated with the UE's Internet Protocol Connectivity Access Network (IP-CAN) session. In at least one embodiment, in a roaming scenario with local traffic breakout, two PCRFs can exist associated with the UE's IP-CAN session: a Home PCRF (H-PCRF) in the HPLMN and a Visited PCRF (V-PCRF) in the Visited Public Land Mobile Network (VPLMN). In at least one embodiment, PCRF 2836 can be communicatively coupled to Application Server 2840 via P-GW 2834. In at least one embodiment, Application Server 2840 can signal PCRF 2836 to indicate a new service flow and select appropriate Quality of Service (QoS) and charging parameters. In at least one embodiment, the PCRF 2836 may supply this rule to a Policy and Charging Enforcement Function (PCEF) (not shown) with the appropriate Traffic Flow Template (TFT) and QoS Class (QCI) identifier, which initiates the QoS and charging specified by the application server 2840.

[0288] In at least one embodiment, Figure 28 At least one component shown or described is used to implement the combination Figure 1-13 In at least one embodiment, Figure 28 At least one component of the system is configured to cause scaling of one or more clocks of one or more cores based at least in part on proximity of the one or more cores to each other. In at least one embodiment, Figure 28 At least one component performs Figure 1 The data center 120, the controller 105, at least two or more of the processors 110(a) to 110(n) and one or more thermal sensors 115(a) to 115(n), Figure 2 Firmware 220, configuration files 225, ACPI 230, policies 235, one or more processors 1-n, and / or including Figure 9 Method 900, Figure 10 Method 1000 and Figure 11 At least one aspect of one or more of the methods 1100 described herein.

[0289] Figure 29The architecture of a system 2900 of a network according to some embodiments is shown. In at least one embodiment, the system 2900 is shown to include a UE 2902, a 5G access node or RAN node (shown as a (R)AN node 2908), a user plane function (shown as a UPF 2904), a data network (DN 2906), which can be, for example, an operator service, internet access, or a third-party service, and a 5G core network (5GC) (shown as a CN 2910).

[0290] In at least one embodiment, the CN 2910 includes an authentication server function (AUSF 2914); a core access and mobility management function (AMF 2912); a session management function (SMF 2918); a network exposure function (NEF 2916); a policy control function (PCF 2922); a network function (NF) repository function (NRF 2920); a unified data management (UDM 2924); and an application function (AF 2926). In at least one embodiment, the CN 2910 may also include other elements not shown, such as a structured data storage network function (SDSF), an unstructured data storage network function (UDSF), and variations thereof.

[0291] In at least one embodiment, the UPF 2904 can serve as an anchor point for intra-RAT and inter-RAT mobility, an external PDU session point interconnected to the DN 2906, and a branching point to support multi-homed PDU sessions. In at least one embodiment, the UPF 2904 can also perform packet routing and forwarding, packet inspection, user plane enforcement of policy rules, lawful interception of packets (UP collection), service usage reporting, QoS processing for the user plane (e.g., packet filtering, gating, UL / DL rate enforcement), uplink service validation (e.g., SDF to QoS flow mapping), transport-level packet marking in the uplink and downlink, downlink packet buffering, and downlink data notification triggering. In at least one embodiment, the UPF 2904 can include an uplink classifier to support routing of service flows to the data network. In at least one embodiment, the DN 2906 can represent various network operator services, internet access, or third-party services.

[0292] In at least one embodiment, the AUSF 2914 may store data used for authentication of the UE 2902 and handle authentication-related functions. In at least one embodiment, the AUSF 2914 may facilitate a common authentication framework for various access types.

[0293] In at least one embodiment, the AMF 2912 may be responsible for registration management (e.g., for registering UE 2902, etc.), connection management, reachability management, mobility management, and lawful interception of AMF-related events, as well as access authentication and authorization. In at least one embodiment, the AMF 2912 may provide transport of SM messages for the SMF 2918 and act as a transparent proxy for routing SM messages. In at least one embodiment, the AMF 2912 may also provide UE 2902 with an SMS function (SMSF) ( Figure 29 2902 and the UE 2902. In at least one embodiment, the AMF 2912 may act as a Security Anchor Function (SEA), which may include interaction with the AUSF 2914 and the UE 2902 and receiving intermediate keys established as a result of the UE 2902 authentication process. In at least one embodiment, where USIM-based authentication is used, the AMF 2912 may retrieve security material from the AUSF 2914. In at least one embodiment, the AMF 2912 may also include a Security Context Management (SCM) function that receives keys from the SEA that it uses to derive access network-specific keys. In addition, in at least one embodiment, the AMF 2912 may be the termination point for the RAN CP interface (N2 reference point), the termination point for NAS (NI) signaling, and perform NAS encryption and integrity protection.

[0294] In at least one embodiment, the AMF 2912 may also support NAS signaling with the UE 2902 over the N3 Interworking Function (IWF) interface. In at least one embodiment, the N3 IWF may be used to provide access to untrusted entities. In at least one embodiment, the N3 IWF may be the termination point for the N2 and N3 interfaces for the control plane and user plane, respectively. Thus, it may handle N2 signaling from the SMF and AMF for PDU sessions and QoS, encapsulate / decapsulate packets for IPSec and N3 tunnels, mark N3 user plane packets in the uplink, and enforce QoS corresponding to the marking of N3 packets, taking into account the QoS requirements associated with such markings received over N2. In at least one embodiment, the N3 IWF may also relay uplink and downlink control plane NAS (NI) signaling between the UE 2902 and the AMF 2912, and relay uplink and downlink user plane packets between the UE 2902 and the UPF 2904. In at least one embodiment, the N3IWF also provides a mechanism for IPsec tunnel establishment with the UE 2902.

[0295] In at least one embodiment, the SMF 2918 may be responsible for session management (e.g., session establishment, modification, and release, including tunnel maintenance between the UPF and AN nodes); UE IP address allocation and management (including optional authorization); selection and control of UP functions; configuring traffic steering at the UPF to route traffic to the appropriate destination; interface termination towards the policy control function; policy enforcement and control portion of QoS; lawful interception (for SM events and interface to the LI system); termination of the SM portion of NAS messages; downlink data notification; originator of AN-specific SM information, which is sent to the AN via the AMF on N2; determining the SSC mode for the session. In at least one embodiment, the SMF 2918 may include the following roaming functions: handling local implementation to apply QoS SLAB (VPLMN); charging data collection and charging interface (VPLMN); lawful interception (for SM events in the VPLMN and interface to the LI system); supporting interaction with external DNs to transport signaling for PDU session authorization / authentication by the external DN.

[0296] In at least one embodiment, the NEF 2916 can provide a means for securely exposing services and capabilities provided by 3GPP network functions to third parties, internal exposure / re-exposure, application functions (e.g., AF 2926), edge computing or fog computing systems, and the like. In at least one embodiment, the NEF 2916 can authenticate, authorize, and / or throttle the AF. In at least one embodiment, the NEF 2916 can also convert information exchanged with the AF 2926 and information exchanged with internal network functions. In at least one embodiment, the NEF 2916 can convert between AF service identifiers and internal 5GC information. In at least one embodiment, the NEF 2916 can also receive information from other network functions (NFs) based on their exposed capabilities. In at least one embodiment, this information can be stored as structured data in the NEF 2916 or in a data storage NF using standardized interfaces. In at least one embodiment, the stored information can then be re-exposed by the NEF 2916 to other NFs and AFs and / or used for other purposes, such as analysis.

[0297] In at least one embodiment, the NRF 2920 may support service discovery functionality, receive NF discovery requests from NF instances, and provide information about the discovered NF instances to the NF instances. In at least one embodiment, the NRF 2920 also maintains information about available NF instances and the services they support.

[0298] In at least one embodiment, the PCF 2922 can provide policy rules to the control plane functions to implement them and can also support a unified policy framework to manage network behavior. In at least one embodiment, the PCF 2922 can also implement a front end (FE) for accessing subscription information related to policy decisions in the UDR of the UDM 2924.

[0299] In at least one embodiment, the UDM 2924 can process subscription-related information to support network entities handling communication sessions and can store subscription data for the UE 2902. In at least one embodiment, the UDM 2924 can include two components: an application FE and a user data repository (UDR). In at least one embodiment, the UDM can include a UDM FE, which is responsible for handling credentials, location management, subscription management, and the like. In at least one embodiment, several different front ends can serve the same user in different transactions. In at least one embodiment, the UDM-FE accesses the sub-subscription information stored in the UDR and performs authentication credential processing; user identity processing; access authorization; registration / mobility management; and subscription management. In at least one embodiment, the UDR can interact with the PCF 2922. In at least one embodiment, the UDM 2924 can also support SMS management, where the SMS-FE implements similar application logic as described above.

[0300] In at least one embodiment, the AF 2926 can provide application influence on service routing, access to the Network Capability Exposure (NCE), and interaction with the policy framework for policy control. In at least one embodiment, the NCE can be a mechanism that allows the 5GC and AF 2926 to provide information to each other via the NEF 2916, which can be used for edge computing implementations. In at least one embodiment, network operators and third-party services can be hosted near the UE 2902's attachment access point to achieve efficient service delivery with reduced end-to-end latency and load on the transport network. In at least one embodiment, for edge computing implementations, the 5GC can select a UPF 2904 close to the UE 2902 and perform service steering from the UPF 2904 to the DN 2906 via the N6 interface. In at least one embodiment, this can be based on UE subscription data, UE location, and information provided by the AF 2926. In at least one embodiment, the AF 2926 can influence UPF (re)selection and service routing. In at least one embodiment, based on operator deployment, the network operator may allow the AF 2926 to interact directly with the relevant NFs when the AF 2926 is considered a trusted entity.

[0301] In at least one embodiment, the CN 2910 may include an SMSF, which may be responsible for SMS subscription checking and verification, and relaying SM messages to / from the UE 2902 to / from other entities, such as SMS-GMSC / IWMSC / SMS routers. In at least one embodiment, the SMS may also interact with the AMF 2912 and the UDM 2924 for notification procedures that the UE 2902 is available for SMS delivery (e.g., setting a UE unreachable flag and notifying the UDM 2924 when the UE 2902 is available for SMS).

[0302] In at least one embodiment, the system 2900 may include the following service-based interfaces: Namf: a service-based interface exposed by AMF; Nsmf: a service-based interface exposed by SMF; Nnef: a service-based interface exposed by NEF; Npcf: a service-based interface exposed by PCF; Nudm: a service-based interface exposed by UDM; Naf: a service-based interface exposed by AF; Nnrf: a service-based interface exposed by NRF; and Nausf: a service-based interface exposed by AUSF.

[0303] In at least one embodiment, system 2900 may include the following reference points: N1: a reference point between the UE and the AMF; N2: a reference point between the (R)AN and the AMF; N3: a reference point between the (R)AN and the UPF; N4: a reference point between the SMF and the UPF; and N6: a reference point between the UPF and the data network. In at least one embodiment, there may be more reference points and / or service-based interfaces between NF services within the NF; however, these interfaces and reference points have been omitted for clarity. In at least one embodiment, the NS reference point may be between the PCF and the AF; the N7 reference point may be between the PCF and the SMF; the N11 reference point may be between the AMF and the SMF, and so on. In at least one embodiment, CN 2910 may include an Nx interface, which is an inter-CN interface between the MME and the AMF 2912 to enable interworking between CN 2910 and CN 7229.

[0304] In at least one embodiment, the system 2900 may include multiple RAN nodes (such as (R)AN nodes 2908), wherein an Xn interface is defined between two or more (R)AN nodes 2908 (e.g., gNBs) connected to the 5GC 410, between an (R)AN node 2908 (e.g., gNBs) and an eNB (e.g., macro RAN node) connected to the CN 2910, and / or between two eNBs connected to the CN 2910.

[0305] In at least one embodiment, the Xn interface may include an Xn user plane (Xn-U) interface and an Xn control plane (Xn-C) interface. In at least one embodiment, the Xn-U may provide non-guaranteed delivery of user plane PDUs and support / provide data forwarding and flow control functions. In at least one embodiment, the Xn-C may provide management and error handling functions, functions for managing the Xn-C interface, mobility support for UE 2902 in connected mode (e.g., CM-CONNECTED), including functions for managing UE mobility in connected mode between one or more (R)AN nodes 2908. In at least one embodiment, mobility support may include context transfer from an old (source) serving (R)AN node 2908 to a new (target) serving (R)AN node 2908, and control of a user plane tunnel between the old (source) serving (R)AN node 2908 and the new (target) serving (R)AN node 2908.

[0306] In at least one embodiment, the protocol stack of Xn-U may include a transport network layer built on an Internet Protocol (IP) transport layer and a GTP-U layer for carrying user plane PDUs on top of UDP and / or one or more IP layers. In at least one embodiment, the Xn-C protocol stack may include an application layer signaling protocol (referred to as the Xn Application Protocol (Xn-AP)) and a transport network layer built on the SCTP layer. In at least one embodiment, the SCTP layer may be on top of the IP layer. In at least one embodiment, the SCTP layer provides guaranteed delivery of application layer messages. In at least one embodiment, in the transport IP layer, point-to-point transport is used to deliver signaling PDUs. In at least one embodiment, the Xn-U protocol stack and / or the Xn-C protocol stack may be the same or similar to the user plane and / or control plane protocol stacks shown and described herein.

[0307] In at least one embodiment, Figure 29 At least one component shown or described is used to implement the combination Figure 1-13 In at least one embodiment, Figure 29 At least one component of the system is configured to cause scaling of one or more clocks of one or more cores based at least in part on proximity of the one or more cores to each other. In at least one embodiment, Figure 29 At least one component performs Figure 1 The data center 120, the controller 105, at least two or more of the processors 110(a) to 110(n) and one or more thermal sensors 115(a) to 115(n), Figure 2Firmware 220, configuration files 225, ACPI 230, policies 235, one or more processors 1-n, and / or including Figure 9 Method 900, Figure 10 Method 1000 and Figure 11 At least one aspect of one or more of the methods 1100 described herein.

[0308] Figure 30 30 is a diagram of a control plane protocol stack according to some embodiments. In at least one embodiment, the control plane 3000 is shown as a communication protocol stack between the UE 2802 (or alternatively, the UE 2804), the RAN 2816, and the MME 2828.

[0309] In at least one embodiment, the PHY layer 3002 may send or receive information used by the MAC layer 3004 over one or more air interfaces. In at least one embodiment, the PHY layer 3002 may also perform link adaptation or adaptive modulation and coding (AMC), power control, cell search (e.g., for initial synchronization and handover purposes), and other measurements used by higher layers (e.g., the RRC layer 3010). In at least one embodiment, the PHY layer 3002 may further perform error detection on transport channels, forward error correction (FEC) encoding / decoding of transport channels, modulation / demodulation of physical channels, interleaving, rate matching, mapping to physical channels, and multiple-input multiple-output (MIMO) antenna processing.

[0310] In at least one embodiment, the MAC layer 3004 may perform mapping between logical channels and transport channels, multiplexing MAC service data units (SDUs) from one or more logical channels onto transport blocks (TBs) to be delivered to the PHY via transport channels, demultiplexing MAC SDUs from transport blocks (TBs) delivered from the PHY via transport channels to one or more logical channels, multiplexing MAC SDUs onto TBs, scheduling information reporting, error correction through hybrid automatic repeat request (HARD), and logical channel prioritization.

[0311] In at least one embodiment, the RLC layer 3006 can operate in multiple operating modes, including transparent mode (TM), unacknowledged mode (UM), and acknowledged mode (AM). In at least one embodiment, the RLC layer 3006 can perform transmission of upper layer protocol data units (PDUs), error correction through automatic repeat request (ARQ) for AM data transmission, and concatenation, segmentation, and reassembly of RLC SDUs for UM and AM data transmission. In at least one embodiment, the RLC layer 3006 can also perform re-segmentation of RLC data PDUs for AM data transmission, reordering of RLC data PDUs for UM and AM data transmission, detection of duplicate data for UM and AM data transmission, discarding RLC SDUs for UM and AM data transmission, detecting protocol errors for AM data transmission, and performing RLC reestablishment.

[0312] In at least one embodiment, the PDCP layer 3008 can perform header compression and decompression of IP data, maintain PDCP sequence numbers (SNs), perform in-sequence delivery of higher layer PDUs when reestablishing lower layers, eliminate duplication of lower layer SDUs when reestablishing lower layers for radio bearers mapped on RLC AM, encrypt and decrypt control plane data, perform integrity protection and integrity verification on control plane data, control timer-based data discard, and perform security operations (e.g., encryption, decryption, integrity protection, integrity verification, etc.).

[0313] In at least one embodiment, the main services and functions of the RRC layer 3010 may include broadcasting of system information (e.g., included in a master information block (MIB) or system information block (SIB) related to the non-access stratum (NAS)), broadcasting of system information related to the access stratum (AS), paging, establishment, maintenance, and release of an RRC connection between a UE and an E-UTRAN (e.g., RRC connection paging, RRC connection establishment, RRC connection modification, and RRC connection release), establishment, configuration, maintenance, and release of point-to-point radio bearers, security functions including key management, inter-radio access technology (RAT) mobility, and measurement configuration for UE measurement reporting. In at least one embodiment, the MIB and SIB may include one or more information elements (IEs), each of which may include a separate data field or data structure.

[0314] In at least one embodiment, the UE 2802 and the RAN 2816 may utilize a Uu interface (e.g., an LTE-Uu interface) to exchange control plane data via a protocol stack including a PHY layer 3002, a MAC layer 3004, an RLC layer 3006, a PDCP layer 3008, and an RRC layer 3010.

[0315] In at least one embodiment, the non-access stratum (NAS) protocol (NAS protocol 3012) forms the highest layer of the control plane between the UE 2802 and the MME 2828. In at least one embodiment, the NAS protocol 3012 supports the mobility and session management procedures of the UE 2802 to establish and maintain IP connectivity between the UE 2802 and the P-GW 2834.

[0316] In at least one embodiment, the Si application protocol (Si-AP) layer (Si-AP layer 3022) can support the functionality of the Si interface and include elementary procedures (EPs). In at least one embodiment, the EP is the unit of interaction between the RAN 2816 and the CN 2828. In at least one embodiment, the S1-AP layer services can include two groups: UE-associated services and non-UE-associated services. In at least one embodiment, these services perform functions including, but not limited to, E-UTRAN Radio Access Bearer (E-RAB) management, UE capability indication, mobility, NAS signaling, RAN Information Management (RIM), and configuration transfer.

[0317] In at least one embodiment, a stream control transmission protocol (SCTP) layer (alternatively referred to as a stream control transmission protocol / internet protocol (SCTP / IP) layer) (SCTP layer 3020) can ensure reliable delivery of signaling messages between the RAN 2816 and the MME 2828 based in part on the IP protocol supported by the IP layer 3018. In at least one embodiment, the L2 layer 3016 and the L1 layer 3014 can refer to communication links (e.g., wired or wireless) used by the RAN node and the MME to exchange information.

[0318] In at least one embodiment, the RAN 2816 and one or more MMEs 2828 may utilize an S1-MME interface to exchange control plane data via a protocol stack including the L1 layer 3014 , the L2 layer 3016 , the IP layer 3018 , the SCTP layer 3020 , and the Si-AP layer 3022 .

[0319] In at least one embodiment, Figure 30 At least one component shown or described is used to implement the combination Figure 1-13 In at least one embodiment, Figure 30 At least one component of the system is configured to cause scaling of one or more clocks of one or more cores based at least in part on proximity of the one or more cores to each other. In at least one embodiment, Figure 30 At least one component performs Figure 1The data center 120, the controller 105, at least two or more of the processors 110(a) to 110(n) and one or more thermal sensors 115(a) to 115(n), Figure 2 Firmware 220, configuration files 225, ACPI 230, policies 235, one or more processors 1-n, and / or including Figure 9 Method 900, Figure 10 Method 1000 and Figure 11 At least one aspect of one or more of the methods 1100 described herein.

[0320] Figure 31 31 is a diagram of a user plane protocol stack according to at least one embodiment. In at least one embodiment, the user plane 3100 is shown as a communication protocol stack between the UE 2802, the RAN 2816, the S-GW 2830, and the P-GW 2834. In at least one embodiment, the user plane 3100 can utilize the same protocol layers as the control plane 3000. In at least one embodiment, for example, the UE 2802 and the RAN 2816 can utilize a Uu interface (e.g., an LTE-Uu interface) to exchange user plane data via a protocol stack including a PHY layer 3002, a MAC layer 3004, an RLC layer 3006, and a PDCP layer 3008.

[0321] In at least one embodiment, the General Packet Radio Service (GPRS) Tunneling Protocol (GTP-U) layer for the user plane (GTP-U layer 3104) can be used to carry user data within the GPRS core network and between the radio access network and the core network. In at least one embodiment, the transmitted user data can be packets in any format, such as IPv4, IPv6, or PPP. In at least one embodiment, the UDP and IP Security (UDP / IP) layer (UDP / IP layer 3102) can provide checksums for data integrity, port numbers for addressing different functions at the source and destination, and encryption and authentication of selected data flows. In at least one embodiment, the RAN 2816 and the S-GW 2830 can utilize the S1-U interface to exchange user plane data via a protocol stack including the L1 layer 3014, the L2 layer 3016, the UDP / IP layer 3102, and the GTP-U layer 3104. In at least one embodiment, the S-GW 2830 and the P-GW 2834 may utilize an S5 / S8a interface to exchange user plane data via a protocol stack including an L1 layer 3014, an L2 layer 3016, a UDP / IP layer 3102, and a GTP-U layer 3104. In at least one embodiment, as described above with respect to Figure 30As discussed, the NAS protocol supports the mobility of UE 2802 and session management procedures to establish and maintain IP connectivity between UE 2802 and P-GW 2834 .

[0322] In at least one embodiment, Figure 31 At least one component shown or described is used to implement the combination Figure 1-13 In at least one embodiment, Figure 31 At least one component of the system is configured to cause scaling of one or more clocks of one or more cores based at least in part on proximity of the one or more cores to each other. In at least one embodiment, Figure 31 At least one component performs Figure 1 The data center 120, the controller 105, at least two or more of the processors 110(a) to 110(n) and one or more thermal sensors 115(a) to 115(n), Figure 2 Firmware 220, configuration files 225, ACPI 230, policies 235, one or more processors 1-n, and / or including Figure 9 Method 900, Figure 10 Method 1000 and Figure 11 At least one aspect of one or more of the methods 1100 described herein.

[0323] Figure 32 Components 3200 of a core network according to at least one embodiment are shown. In at least one embodiment, the components of CN 2838 can be implemented in one physical node or in separate physical nodes, the separate physical nodes including components for reading and executing instructions from a machine-readable medium or computer-readable medium (e.g., a non-transitory machine-readable storage medium). In at least one embodiment, network function virtualization (NFV) is used to virtualize any or all of the above-described network node functions via executable instructions stored in one or more computer-readable storage media (described in further detail below). In at least one embodiment, a logical instantiation of CN 2838 can be referred to as a network slice 3202 (e.g., network slice 3202 is shown as including HSS 2832, MME 2828, and S-GW 2830). In at least one embodiment, a logical instantiation of a portion of CN 2838 can be referred to as a network sub-slice 3204 (e.g., network sub-slice 3204 is shown as including P-GW 2834 and PCRF 2836).

[0324] In at least one embodiment, the NFV architecture and infrastructure can be used to virtualize one or more network functions onto physical resources including a combination of industry-standard server hardware, storage hardware, or switches, which may alternatively be performed by dedicated hardware. In at least one embodiment, the NFV system can be used to perform a virtual or reconfigurable implementation of one or more EPC components / functions.

[0325] In at least one embodiment, Figure 32 At least one component shown or described is used to implement the combination Figure 1-13 In at least one embodiment, Figure 32 At least one component of the system is configured to cause scaling of one or more clocks of one or more cores based at least in part on proximity of the one or more cores to each other. In at least one embodiment, Figure 32 At least one component performs Figure 1 The data center 120, the controller 105, at least two or more of the processors 110(a) to 110(n) and one or more thermal sensors 115(a) to 115(n), Figure 2 Firmware 220, configuration files 225, ACPI 230, policies 235, one or more processors 1-n, and / or including Figure 9 Method 900, Figure 10 Method 1000 and Figure 11 At least one aspect of one or more of the methods 1100 described herein.

[0326] Figure 33 33 is a block diagram illustrating components of a system 3300 for supporting network function virtualization (NFV) according to at least one embodiment. In at least one embodiment, the system 3300 is shown to include a virtualization infrastructure manager (shown as VIM 3302), a network function virtualization infrastructure (shown as NFVI 3304), a VNF manager (shown as VNFM 3306), virtualized network functions (shown as VNF 3308), an element manager (shown as EM 3310), an NFV orchestrator (shown as NFVO 3312), and a network manager (shown as NM 3314).

[0327] In at least one embodiment, the VIM 3302 manages resources of the NFVI 3304. In at least one embodiment, the NFVI 3304 may include physical or virtual resources and applications (including a hypervisor) used to execute the system 3300. In at least one embodiment, the VIM 3302 may utilize the NFVI 3304 to manage the lifecycle of virtual resources (e.g., the creation, maintenance, and teardown of virtual machines (VMs) associated with one or more physical resources), track VM instances, track performance, faults, and security of VM instances and associated physical resources, and expose VM instances and associated physical resources to other management systems.

[0328] In at least one embodiment, VNFM 3306 can manage VNF 3308. In at least one embodiment, VNF 3308 can be used to perform EPC components / functions. In at least one embodiment, VNFM 3306 can manage the lifecycle of VNF 3308 and track the performance, faults, and security of the virtual aspects of VNF 3308. In at least one embodiment, EM 3310 can track the performance, faults, and security of the functional aspects of VNF 3308. In at least one embodiment, the data tracked from VNFM 3306 and EM 3310 can include, for example, performance measurement (PM) data used by VIM 3302 or NFVI 3304. In at least one embodiment, both VNFM 3306 and EM 3310 can scale up / down the number of VNFs in system 3300.

[0329] In at least one embodiment, NFVO 3312 can coordinate, authorize, release, and occupy resources of NFVI 3304 in order to provide the requested service (e.g., to execute EPC functions, components, or slices). In at least one embodiment, NM 3314 can provide an end-user function package responsible for managing the network, which can include network elements with VNFs, non-virtualized network functions, or both (management of VNFs can occur via EM 3310).

[0330] In at least one embodiment, Figure 33 At least one component shown or described is used to implement the combination Figure 1-13 In at least one embodiment, Figure 33 At least one component of the system is configured to cause scaling of one or more clocks of one or more cores based at least in part on proximity of the one or more cores to each other. In at least one embodiment, Figure 33 At least one component performs Figure 1The data center 120, the controller 105, at least two or more of the processors 110(a) to 110(n) and one or more thermal sensors 115(a) to 115(n), Figure 2 Firmware 220, configuration files 225, ACPI 230, policies 235, one or more processors 1-n, and / or including Figure 9 Method 900, Figure 10 Method 1000 and Figure 11 At least one aspect of one or more of the methods 1100 described herein.

[0331] Computer-based systems

[0332] The following figures set forth, but are not limiting of, exemplary computer-based systems that can be used to implement at least one embodiment.

[0333] Figure 34 A processing system 3400 is shown in accordance with at least one embodiment. In at least one embodiment, system 3400 includes one or more processors 3402 and one or more graphics processors 3408 and can be a single-processor desktop system, a multi-processor workstation system, or a server system having a large number of processors 3402 or processor cores 3407. In at least one embodiment, processing system 3400 is a processing platform incorporated within a system-on-chip (SoC) integrated circuit for use in a mobile, handheld, or embedded device.

[0334] In at least one embodiment, the processing system 3400 may include or be incorporated into a server-based gaming platform, including a gaming console, a mobile gaming console, a handheld gaming console, or an online gaming console. In at least one embodiment, the processing system 3400 is a mobile phone, a smartphone, a tablet computing device, or a mobile internet device. In at least one embodiment, the processing system 3400 may also include a device coupled to or integrated into a wearable device, such as a smartwatch wearable device, a smart glasses device, an augmented reality device, or a virtual reality device. In at least one embodiment, the processing system 3400 is a television or set-top box device having one or more processors 3402 and a graphical interface generated by one or more graphics processors 3408.

[0335] In at least one embodiment, each of the one or more processors 3402 includes one or more processor cores 3407 to process instructions that, when executed, perform operations for system and user software. In at least one embodiment, each of the one or more processor cores 3407 is configured to process a specific instruction set 3409. In at least one embodiment, the instruction set 3409 can facilitate complex instruction set computing (CISC), reduced instruction set computing (RISC), or computing via very long instruction words (VLIW). In at least one embodiment, multiple processor cores 3407 can each process a different instruction set 3409, which can include instructions that facilitate emulating other instruction sets. In at least one embodiment, the processor cores 3407 can also include other processing devices, such as a digital signal processor (DSP).

[0336] In at least one embodiment, the processor 3402 includes a cache memory (cache) 3404. In at least one embodiment, the processor 3402 may have a single internal cache or multiple levels of internal cache. In at least one embodiment, the cache memory is shared among various components of the processor 3402. In at least one embodiment, the processor 3402 also uses an external cache (e.g., a level 3 (L3) cache or a last level cache (LLC)) (not shown), which may share this logic among the processor cores 3407 using known cache coherence techniques. In at least one embodiment, the processor 3402 further includes a register file 3406. The processor 3402 may include different types of registers (e.g., integer registers, floating point registers, status registers, and an instruction pointer register) for storing different types of data. In at least one embodiment, the register file 3406 may include general purpose registers or other registers.

[0337] In at least one embodiment, one or more processors 3402 are coupled to one or more interface buses 3410 to transmit communication signals, such as address, data, or control signals, between the processors 3402 and other components in the system 3400. In at least one embodiment, the interface bus 3410 can be a processor bus, such as a version of a Direct Media Interface (DMI) bus. In at least one embodiment, the interface bus 3410 is not limited to a DMI bus and can include one or more peripheral component interconnect buses (e.g., PCI, PCI Express), a memory bus, or other types of interface buses. In at least one embodiment, the processor 3402 includes an integrated memory controller 3416 and a platform controller hub 3430. In at least one embodiment, the memory controller 3416 facilitates communication between storage devices and other components of the processing system 3400, while the platform controller hub (PCH) 3430 provides connections to input / output (I / O) devices via a local I / O bus.

[0338] In at least one embodiment, the memory device 3420 can be a dynamic random access memory (DRAM) device, a static random access memory (SRAM) device, a flash memory device, a phase change memory device, or a device having suitable performance for use as processor memory. In at least one embodiment, the memory device 3420 can be used as system memory for the processing system 3400 to store data 3422 and instructions 3421 for use when one or more processors 3402 execute applications or processes. In at least one embodiment, the memory controller 3416 is also coupled to an optional external graphics processor 3412, which can communicate with one or more graphics processors 3408 in the processor 3402 to perform graphics and media operations. In at least one embodiment, a display device 3411 can be connected to the processor 3402. In at least one embodiment, the display device 3411 can include one or more internal display devices, such as in a mobile electronic device or portable computer device, or an external display device connected via a display interface (such as a DisplayPort). In at least one embodiment, the display device 3411 may include a head-mounted display (HMD), such as a stereoscopic display device used in virtual reality (VR) applications or augmented reality (AR) applications.

[0339] In at least one embodiment, the platform controller hub 3430 enables peripheral devices to connect to the storage device 3420 and the processor 3402 via a high-speed I / O bus. In at least one embodiment, the I / O peripherals include, but are not limited to, an audio controller 3446, a network controller 3434, a firmware interface 3428, a wireless transceiver 3426, a touch sensor 3425, and a data storage device 3424 (e.g., a hard drive, flash memory, etc.). In at least one embodiment, the data storage device 3424 can be connected via a memory interface (e.g., SATA) or via a peripheral bus, such as a peripheral component interconnect bus (e.g., PCI, PCIe). In at least one embodiment, the touch sensor 3425 can include a touch screen sensor, a pressure sensor, or a fingerprint sensor. In at least one embodiment, the wireless transceiver 3426 can be a Wi-Fi transceiver, a Bluetooth transceiver, or a mobile network transceiver, such as a 3G, 4G, or Long Term Evolution (LTE) transceiver. In at least one embodiment, the firmware interface 3428 enables communication with the system firmware, for example, and can be a unified extensible firmware interface (UEFI). In at least one embodiment, a network controller 3434 can enable network connectivity to a wired network. In at least one embodiment, a high-performance network controller (not shown) is coupled to the interface bus 3410. In at least one embodiment, the audio controller 3446 is a multi-channel high-definition audio controller. In at least one embodiment, the processing system 3400 includes an optional legacy I / O controller 3440 for coupling legacy (e.g., Personal System 2 (PS / 2)) devices to the processing system 3400. In at least one embodiment, the platform controller hub 3430 can also be connected to one or more universal serial bus (USB) controllers 3442 that connect input devices such as a keyboard and mouse 3443 combination, a camera 3444, or other USB input devices.

[0340] In at least one embodiment, instances of the memory controller 3416 and the platform controller hub 3430 may be integrated into a discrete external graphics processor, such as the external graphics processor 3412. In at least one embodiment, the platform controller hub 3430 and / or the memory controller 3416 may be external to one or more processors 3402. For example, in at least one embodiment, the processing system 3400 may include the external memory controller 3416 and the platform controller hub 3430, which may be configured as a memory controller hub and a peripheral controller hub in a system chipset that communicates with the processor 3402.

[0341] In at least one embodiment, Figure 34 At least one component shown or described is used to implement the combination Figure 1-13In at least one embodiment, Figure 34 At least one component of the system is configured to cause scaling of one or more clocks of one or more cores based at least in part on proximity of the one or more cores to each other. In at least one embodiment, Figure 34 At least one component performs Figure 1 The data center 120, the controller 105, at least two or more of the processors 110(a) to 110(n) and one or more thermal sensors 115(a) to 115(n), Figure 2 Firmware 220, configuration files 225, ACPI 230, policies 235, one or more processors 1-n, and / or including Figure 9 Method 900, Figure 10 Method 1000 and Figure 11 At least one aspect of one or more of the methods 1100 described herein.

[0342] Figure 35 A computer system 3500 is shown in accordance with at least one embodiment. In at least one embodiment, the computer system 3500 can be a system of interconnected devices and components, a SOC, or some combination thereof. In at least one embodiment, the computer system 3500 is formed by a processor 3502, which can include an execution unit for executing instructions. In at least one embodiment, the computer system 3500 can include, but is not limited to, components such as the processor 3502, which employs an execution unit including logic to execute algorithms for processing data. In at least one embodiment, the computer system 3500 can include a processor such as the Intel® processor available from Intel Corporation of Santa Clara, California. Processor family, Xeon TM 、 XScale TM and / or StrongARM TM , Core TM or Nervana TMmicroprocessor, although other systems (including PCs with other microprocessors, engineering workstations, set-top boxes, etc.) may also be used. In at least one embodiment, computer system 3500 may execute a version of the WINDOWS operating system available from Microsoft Corporation of Redmond, Wash., although other operating systems (e.g., UNIX and Linux), embedded software, and / or graphical user interfaces may also be used.

[0343] In at least one embodiment, the computer system 3500 can be used in other devices, such as handheld devices and embedded applications. Some examples of handheld devices include cellular phones, Internet Protocol (IP) devices, digital cameras, personal digital assistants ("PDAs"), and handheld PCs. In at least one embodiment, embedded applications can include microcontrollers, digital signal processors ("DSPs"), SoCs, network computers ("NetPCs"), set-top boxes, network hubs, wide area network ("WAN") switches, or any other system that can execute one or more instructions according to at least one embodiment.

[0344] In at least one embodiment, computer system 3500 may include, but is not limited to, a processor 3502, which may include, but is not limited to, one or more execution units 3508, which may be configured to execute Compute Unified Device Architecture ("CUDA") ( Developed by NVIDIA Corporation of Santa Clara, California) program. In at least one embodiment, a CUDA program is at least a portion of a software application written in the CUDA programming language. In at least one embodiment, computer system 3500 is a single-processor desktop or server system. In at least one embodiment, computer system 3500 may be a multi-processor system. In at least one embodiment, processor 3502 may include, but is not limited to, a CISC microprocessor, a RISC microprocessor, a VLIW microprocessor, a processor implementing a combination of instruction sets, or any other processor device, such as, for example, a digital signal processor. In at least one embodiment, processor 3502 may be coupled to a processor bus 3510 that may transmit data signals between processor 3502 and other components in computer system 3500.

[0345] In at least one embodiment, processor 3502 may include, but is not limited to, level 1 ("L1") internal cache memory ("cache") 3504. In at least one embodiment, processor 3502 may have a single internal cache or multiple levels of internal cache. In at least one embodiment, cache memory may reside external to processor 3502. In at least one embodiment, processor 3502 may include a combination of internal and external caches. In at least one embodiment, register file 3506 may store different types of data in various registers, including, but not limited to, integer registers, floating point registers, status registers, and an instruction pointer register.

[0346] In at least one embodiment, an execution unit 3508, including but not limited to logic for performing integer and floating-point operations, is also located in the processor 3502. The processor 3502 may also include a microcode ("ucode") read-only memory ("ROM") for storing microcode for certain macroinstructions. In at least one embodiment, the execution unit 3508 may include logic for processing a packed instruction set 3509. In at least one embodiment, by including the packed instruction set 3509 in the instruction set of the general-purpose processor 3502, along with associated circuitry to execute the instructions, operations used by many multimedia applications may be performed using packed data in the general-purpose processor 3502. In at least one embodiment, many multimedia applications may be executed faster and more efficiently by using the full width of the processor's data bus to perform operations on the packed data, which may eliminate the need to transfer smaller units of data across the processor's data bus to perform one or more operations on one data element at a time.

[0347] In at least one embodiment, execution unit 3508 may also be used in a microcontroller, embedded processor, graphics device, DSP, or other types of logic circuits. In at least one embodiment, computer system 3500 may include, but is not limited to, memory 3520. In at least one embodiment, memory 3520 may be implemented as a DRAM device, an SRAM device, a flash memory device, or other storage device. Memory 3520 may store instructions 3519 and / or data 3521 represented by data signals that may be executed by processor 3502.

[0348] In at least one embodiment, the system logic chip can be coupled to the processor bus 3510 and the memory 3520. In at least one embodiment, the system logic chip can include, but is not limited to, a memory controller hub ("MCH") 3516, and the processor 3502 can communicate with the MCH 3516 via the processor bus 3510. In at least one embodiment, the MCH 3516 can provide a high-bandwidth memory path 3518 to the memory 3520 for instruction and data storage, as well as for storage of graphics commands, data, and textures. In at least one embodiment, the MCH 3516 can initiate data signals between the processor 3502, the memory 3520, and other components in the computer system 3500, and bridge data signals between the processor bus 3510, the memory 3520, and the system I / O 3522. In at least one embodiment, the system logic chip can provide a graphics port for coupling to a graphics controller. In at least one embodiment, the MCH 3516 may be coupled to the memory 3520 via a high-bandwidth memory path 3518 , and the graphics / video card 3512 may be coupled to the MCH 3516 via an Accelerated Graphics Port (“AGP”) interconnect 3514 .

[0349] In at least one embodiment, the computer system 3500 may use the system I / O 3522 as a proprietary hub interface bus to couple the MCH 3516 to the I / O controller hub ("ICH") 3530. In at least one embodiment, the ICH 3530 may provide direct connectivity to certain I / O devices via a local I / O bus. In at least one embodiment, the local I / O bus may include, but is not limited to, a high-speed I / O bus for connecting peripheral devices to the memory 3520, chipset, and processor 3502. Examples may include, but are not limited to, an audio controller 3529, a firmware hub ("Flash BIOS") 3528, a wireless transceiver 3526, a data store 3524, a legacy I / O controller 3523 including user input 3525 and a keyboard interface, a serial expansion port 3577 (e.g., USB), and a network controller 3534. The data store 3524 may include a hard drive, a floppy disk drive, a CD-ROM device, a flash memory device, or other mass storage device.

[0350] In at least one embodiment, Figure 35 A system comprising interconnected hardware devices or "chips" is shown. In at least one embodiment, Figure 35 An exemplary SoC may be shown. In at least one embodiment, Figure 35The devices shown in can be interconnected with a proprietary interconnect, a standardized interconnect (e.g., PCIe), or some combination thereof. In at least one embodiment, one or more components of system 3500 are interconnected using a Compute Express Link (CXL) interconnect.

[0351] In at least one embodiment, Figure 35 At least one component shown or described is used to implement the combination Figure 1-13 In at least one embodiment, Figure 35 At least one component of the system is configured to cause scaling of one or more clocks of one or more cores based at least in part on proximity of the one or more cores to each other. In at least one embodiment, Figure 35 At least one component performs Figure 1 The data center 120, the controller 105, at least two or more of the processors 110(a) to 110(n) and one or more thermal sensors 115(a) to 115(n), Figure 2 Firmware 220, configuration files 225, ACPI 230, policies 235, one or more processors 1-n, and / or including Figure 9 Method 900, Figure 10 Method 1000 and Figure 11 At least one aspect of one or more of the methods 1100 described herein.

[0352] Figure 36 A system 3600 is shown in accordance with at least one embodiment. In at least one embodiment, the system 3600 is an electronic device that utilizes a processor 3610. In at least one embodiment, the system 3600 can be, for example, but not limited to, a notebook computer, a tower server, a rack server, a blade server, a laptop computer, a desktop computer, a tablet computer, a mobile device, a phone, an embedded computer, or any other suitable electronic device.

[0353] In at least one embodiment, system 3600 may include, but is not limited to, a processor 3610 communicatively coupled to any suitable number or kind of components, peripherals, modules, or devices. In at least one embodiment, processor 3610 is coupled using a bus or interface, such as an I 2 C bus, System Management Bus ("SMBus"), Low Pin Count (LPC) bus, Serial Peripheral Interface ("SPI"), High Definition Audio ("HDA") bus, Serial Advanced Technology Attachment ("SATA") bus, USB (Revisions 1, 2, 3), or Universal Asynchronous Receiver / Transmitter ("UART") bus. In at least one embodiment, Figure 36 A system is shown that includes interconnected hardware devices or "chips". In at least one embodiment, Figure 36 An exemplary SoC may be shown. In at least one embodiment, Figure 36 The devices shown in can be interconnected with proprietary interconnects, standardized interconnects (e.g., PCIe), or some combination thereof. In at least one embodiment, Figure 36 One or more components of the system are interconnected using Compute Express Link (CXL) interconnect lines.

[0354] In at least one embodiment, Figure 36 The components may include a display 3624, a touch screen 3625, a touchpad 3630, a near field communication unit ("NFC") 3645, a sensor hub 3640, a thermal sensor 3646, a fast chipset ("EC") 3635, a trusted platform module ("TPM") 3638, a BIOS / firmware / flash memory ("BIOS, FW Flash") 3622, a DSP 3660, a solid-state disk ("SSD") or a hard disk drive ("HDD") 3620, a wireless local area network unit ("WLAN") 3650, a Bluetooth unit 3652, a wireless wide area network unit ("WWAN") 3656, a global positioning system (GPS) 3655, a camera ("USB 3.0 camera") 3654 (e.g., a USB 3.0 camera), or a low-power double data rate ("LPDDR") memory unit ("LPDDR3") 3615 implemented using, for example, the LPDDR3 standard. Each of these components may be implemented in any suitable manner.

[0355] In at least one embodiment, other components may be communicatively coupled to processor 3610 through the components discussed above. In at least one embodiment, an accelerometer 3641, an ambient light sensor (“ALS”) 3642, a compass 3643, and a gyroscope 3644 may be communicatively coupled to sensor hub 3640. In at least one embodiment, a thermal sensor 3639, a fan 3637, a keyboard 3646, and a touchpad 3630 may be communicatively coupled to EC 3635. In at least one embodiment, a speaker 3663, an earpiece 3664, and a microphone (“mic”) 3665 may be communicatively coupled to an audio unit (“audio codec and class-D amplifier”) 3664, which in turn may be communicatively coupled to DSP 3660. In at least one embodiment, audio unit 3664 may include, but is not limited to, an audio codec / decoder (“codec”) and a class-D amplifier. In at least one embodiment, a SIM card (“SIM”) 3657 may be communicatively coupled to WWAN unit 3656. In at least one embodiment, components such as the WLAN unit 3650 and the Bluetooth unit 3652 and the WWAN unit 3656 may be implemented as a next generation form factor (NGFF).

[0356] In at least one embodiment, Figure 36 At least one component shown or described is used to implement the combination Figure 1-13 In at least one embodiment, Figure 36 At least one component of the system is configured to cause scaling of one or more clocks of one or more cores based at least in part on proximity of the one or more cores to each other. In at least one embodiment, Figure 36 At least one component performs Figure 1 The data center 120, the controller 105, at least two or more of the processors 110(a) to 110(n) and one or more thermal sensors 115(a) to 115(n), Figure 2 Firmware 220, configuration files 225, ACPI 230, policies 235, one or more processors 1-n, and / or including Figure 9 Method 900, Figure 10 Method 1000 and Figure 11 At least one aspect of one or more of the methods 1100 described herein.

[0357] Figure 37 An exemplary integrated circuit 3700 is shown in accordance with at least one embodiment. In at least one embodiment, the exemplary integrated circuit 3700 is a SoC, which may be manufactured using one or more IP cores. In at least one embodiment, the integrated circuit 3700 includes one or more application processors 3705 (e.g., CPUs), at least one graphics processor 3710, and may additionally include an image processor 3715 and / or a video processor 3720, any of which may be modular IP cores. In at least one embodiment, the integrated circuit 3700 includes peripheral or bus logic including a USB controller 3725, a UART controller 3730, an SPI / SDIO controller 3735, and an I / O controller. 2 S / I 2 C controller 3740. In at least one embodiment, the integrated circuit 3700 may include a display device 3745 coupled to one or more of a High Definition Multimedia Interface (HDMI) controller 3750 and a Mobile Industry Processor Interface (MIPI) display interface 3755. In at least one embodiment, storage may be provided by a flash memory subsystem 3760, including flash memory and a flash memory controller. In at least one embodiment, a memory interface may be provided via a memory controller 3765 for accessing SDRAM or SRAM memory devices. In at least one embodiment, some integrated circuits also include an embedded security engine 3770.

[0358] In at least one embodiment, Figure 37At least one component shown or described is used to implement the combination Figure 1-13 In at least one embodiment, Figure 37 At least one component of the system is configured to cause scaling of one or more clocks of one or more cores based at least in part on proximity of the one or more cores to each other. In at least one embodiment, Figure 37 At least one component performs Figure 1 The data center 120, the controller 105, at least two or more of the processors 110(a) to 110(n) and one or more thermal sensors 115(a) to 115(n), Figure 2 Firmware 220, configuration files 225, ACPI 230, policies 235, one or more processors 1-n, and / or including Figure 9 Method 900, Figure 10 Method 1000 and Figure 11 At least one aspect of one or more of the methods 1100 described herein.

[0359] Figure 38 A computing system 3800 is shown in accordance with at least one embodiment. In at least one embodiment, computing system 3800 includes a processing subsystem 3801 having one or more processors 3802 and system memory 3804 communicating via an interconnect path that may include a memory hub 3805. In at least one embodiment, memory hub 3805 may be a separate component within a chipset assembly or integrated within one or more processors 3802. In at least one embodiment, memory hub 3805 is coupled to an I / O subsystem 3811 via a communication link 3806. In at least one embodiment, I / O subsystem 3811 includes an I / O hub 3807, which enables computing system 3800 to receive input from one or more input devices 3808. In at least one embodiment, I / O hub 3807 may enable a display controller, included in one or more processors 3802, to provide output to one or more display devices 3810A. In at least one embodiment, the one or more display devices 3810A coupled to the I / O hub 3807 may include local, internal, or embedded display devices.

[0360] In at least one embodiment, the processing subsystem 3801 includes one or more parallel processors 3812 coupled to a memory hub 3805 via a bus or other communication link 3813. In at least one embodiment, the communication link 3813 can be one of many standard-based communication link technologies or protocols, such as, but not limited to, PCIe, or can be a vendor-specific communication interface or communication structure. In at least one embodiment, the one or more parallel processors 3812 form a computationally focused parallel or vector processing system that can include a large number of processing cores and / or processing clusters, such as a multi-integrated core (MIC) processor. In at least one embodiment, the one or more parallel processors 3812 form a graphics processing subsystem that can output pixels to one of one or more display devices 3810A coupled via an I / O hub 3807. In at least one embodiment, the one or more parallel processors 3812 can also include a display controller and display interface (not shown) to enable direct connection to one or more display devices 3810B.

[0361] In at least one embodiment, a system storage unit 3814 can be connected to the I / O hub 3807 to provide a storage mechanism for the computing system 3800. In at least one embodiment, an I / O switch 3816 can be used to provide an interface mechanism to enable connections between the I / O hub 3807 and other components, such as a network adapter 3818 and / or a wireless network adapter 3819 that can be integrated into the platform, as well as various other devices that can be added via one or more add-on devices 3820. In at least one embodiment, the network adapter 3818 can be an Ethernet adapter or another wired network adapter. In at least one embodiment, the wireless network adapter 3819 can include one or more of Wi-Fi, Bluetooth, NFC, or other network devices including one or more radios.

[0362] In at least one embodiment, computing system 3800 may include other components not explicitly shown, including USB or other port connections, optical storage drives, video capture devices, and / or variations thereof, which may also be connected to I / O hub 3807. Figure 38 The communication paths that interconnect the various components in the system can be implemented using any suitable protocol, such as a PCI (Peripheral Component Interconnect)-based protocol (e.g., PCIe), or other bus or point-to-point communication interfaces and / or protocols (e.g., NVLink high-speed interconnect or interconnect protocol).

[0363] In at least one embodiment, one or more parallel processors 3812 include circuitry optimized for graphics and video processing (including video output circuitry in at least one embodiment) and constitute a graphics processing unit (GPU). In at least one embodiment, one or more parallel processors 3812 include circuitry optimized for general-purpose processing. In at least one embodiment, the components of computing system 3800 can be integrated with one or more other system elements on a single integrated circuit. For example, in at least one embodiment, one or more parallel processors 3812, memory hub 3805, processor 3802, and I / O hub 3807 can be integrated into a system-on-chip (SoC) integrated circuit. In at least one embodiment, the components of computing system 3800 can be integrated into a single package to form a system-in-package (SIP) configuration. In at least one embodiment, at least a portion of the components of computing system 3800 can be integrated into a multi-chip module (MCM), which can be interconnected with other multi-chip modules to form a modular computing system. In at least one embodiment, I / O subsystem 3811 and display device 3810B are omitted from computing system 3800.

[0364] In at least one embodiment, Figure 38 At least one component shown or described is used to implement the combination Figure 1-13 In at least one embodiment, Figure 38 At least one component of the system is configured to cause scaling of one or more clocks of one or more cores based at least in part on proximity of the one or more cores to each other. In at least one embodiment, Figure 38 At least one component performs Figure 1 The data center 120, the controller 105, at least two or more of the processors 110(a) to 110(n) and one or more thermal sensors 115(a) to 115(n), Figure 2 Firmware 220, configuration files 225, ACPI 230, policies 235, one or more processors 1-n, and / or including Figure 9 Method 900, Figure 10 Method 1000 and Figure 11 At least one aspect of one or more of the methods 1100 described herein.

[0365] Processing system

[0366] The following figures illustrate, but are not limited to, exemplary processing systems that can be used to implement at least one embodiment.

[0367] Figure 39An accelerated processing unit ("APU") 3900 is shown in accordance with at least one embodiment. In at least one embodiment, the APU 3900 was developed by Advanced Micro Devices, Inc. of Santa Clara, California. In at least one embodiment, the APU 3900 can be configured to execute applications, such as CUDA programs. In at least one embodiment, the APU 3900 includes, but is not limited to, a core complex 3910, a graphics complex 3940, a fabric 3960, an I / O interface 3970, a memory controller 3980, a display controller 3992, and a multimedia engine 3994. In at least one embodiment, the APU 3900 can include, but is not limited to, any combination of any number of core complexes 3910, any number of graphics complexes 3940, any number of display controllers 3992, and any number of multimedia engines 3994. For purposes of illustration, multiple instances of similar objects are denoted herein by reference numerals, where the reference numeral identifies the object and a number in parentheses identifies the desired instance.

[0368] In at least one embodiment, core complex 3910 is a CPU, graphics complex 3940 is a GPU, and APU 3900 is a processing unit that is not limited to integrating 3910 and 3940 onto a single chip. In at least one embodiment, some tasks may be assigned to core complex 3910, while other tasks may be assigned to graphics complex 3940. In at least one embodiment, core complex 3910 is configured to execute primary control software associated with APU 3900, such as an operating system. In at least one embodiment, core complex 3910 is the main processor of APU 3900, controlling and coordinating the operations of the other processors. In at least one embodiment, core complex 3910 issues commands that control the operations of graphics complex 3940. In at least one embodiment, core complex 3910 may be configured to execute host executable code derived from CUDA source code, and graphics complex 3940 may be configured to execute device executable code derived from CUDA source code.

[0369] In at least one embodiment, core complex 3910 includes, but is not limited to, cores 3920(1)-3920(4) and L3 cache 3930. In at least one embodiment, core complex 3910 may include, but is not limited to, any number of cores 3920 and any combination of any number and type of caches. In at least one embodiment, cores 3920 are configured to execute instructions of a particular instruction set architecture ("ISA"). In at least one embodiment, each core 3920 is a CPU core.

[0370] In at least one embodiment, each core 3920 includes, but is not limited to, a fetch / decode unit 3922, an integer execution engine 3924, a floating-point execution engine 3926, and an L2 cache 3928. In at least one embodiment, the fetch / decode unit 3922 fetches instructions, decodes these instructions, generates micro-ops, and dispatches individual micro-ops to the integer execution engine 3924 and the floating-point execution engine 3926. In at least one embodiment, the fetch / decode unit 3922 can simultaneously dispatch one micro-op to the integer execution engine 3924 and another micro-op to the floating-point execution engine 3926. In at least one embodiment, the integer execution engine 3924 performs, but is not limited to, integer and memory operations. In at least one embodiment, the floating-point engine 3926 performs, but is not limited to, floating-point and vector operations. In at least one embodiment, the fetch-decode unit 3922 dispatches micro-ops to a single execution engine that replaces both the integer execution engine 3924 and the floating-point execution engine 3926.

[0371] In at least one embodiment, each core 3920(i) can access an L2 cache 3928(i) included in the core 3920(i), where i is an integer representing a specific instance of the core 3920. In at least one embodiment, each core 3920 included in a core complex 3910(j) is connected to the other cores 3920 included in the core complex 3910(j) via an L3 cache 3930(j) included in the core complex 3910(j), where j is an integer representing a specific instance of the core complex 3910. In at least one embodiment, a core 3920 included in a core complex 3910(j) can access all L3 caches 3930(j) included in the core complex 3910(j), where j is an integer representing a specific instance of the core complex 3910. In at least one embodiment, the L3 cache 3930 can include, but is not limited to, any number of slices.

[0372] In at least one embodiment, graphics complex 3940 can be configured to perform computational operations in a highly parallel manner. In at least one embodiment, graphics complex 3940 is configured to perform graphics pipeline operations, such as draw commands, pixel operations, geometry calculations, and other operations associated with rendering an image to a display. In at least one embodiment, graphics complex 3940 is configured to perform operations that are not graphics-related. In at least one embodiment, graphics complex 3940 is configured to perform both graphics-related operations and graphics-independent operations.

[0373] In at least one embodiment, graphics complex 3940 includes, but is not limited to, any number of compute units 3950 and L2 cache 3942. In at least one embodiment, compute units 3950 share L2 cache 3942. In at least one embodiment, L2 cache 3942 is partitioned. In at least one embodiment, graphics complex 3940 includes, but is not limited to, any number of compute units 3950 and any number (including zero) and type of cache. In at least one embodiment, graphics complex 3940 includes, but is not limited to, any amount of dedicated graphics hardware.

[0374] In at least one embodiment, each compute unit 3950 includes, but is not limited to, any number of SIMD units 3952 and shared memory 3954. In at least one embodiment, each SIMD unit 3952 implements a SIMD architecture and is configured to execute operations in parallel. In at least one embodiment, each compute unit 3950 can execute any number of thread blocks, but each thread block executes on a single compute unit 3950. In at least one embodiment, a thread block includes, but is not limited to, any number of threads of execution. In at least one embodiment, a workgroup is a thread block. In at least one embodiment, each SIMD unit 3952 executes a different warp. In at least one embodiment, a warp is a group of threads (e.g., 16 threads), where each thread in a warp belongs to a single thread block and is configured to process different data sets based on a single instruction set. In at least one embodiment, predication can be used to disable one or more threads in a warp. In at least one embodiment, a channel is a thread. In at least one embodiment, a work item is a thread. In at least one embodiment, a wavefront is a warp. In at least one embodiment, different wavefronts in a thread block can be synchronized and communicated via shared memory 3954.

[0375] In at least one embodiment, fabric 3960 is a system interconnect that facilitates data and control transfers across core complex 3910, graphics complex 3940, I / O interface 3970, memory controller 3980, display controller 3992, and multimedia engine 3994. In at least one embodiment, APU 3900 may include, in addition to or in lieu of fabric 3960, but is not limited to, any number and type of system interconnects that facilitate data and control transfers across any number and type of directly or indirectly linked components that may be internal or external to APU 3900. In at least one embodiment, I / O interface 3970 represents any number and type of I / O interfaces (e.g., PCI, PCI-Extended ("PCI-X"), PCIe, Gigabit Ethernet ("GBE"), USB, etc.). In at least one embodiment, various types of peripheral devices are coupled to I / O interface 3970. In at least one embodiment, peripheral devices coupled to I / O interface 3970 may include, but are not limited to, a keyboard, a mouse, a printer, a scanner, a joystick or other type of game controller, a media recording device, an external storage device, a network interface card, etc.

[0376] In at least one embodiment, display controller AMD92 displays images on one or more display devices, such as liquid crystal display (LCD) devices. In at least one embodiment, multimedia engine 3994 includes, but is not limited to, any number and type of multimedia-related circuits, such as video decoders, video encoders, image signal processors, and the like. In at least one embodiment, memory controller 3980 facilitates data transfer between APU 3900 and unified system memory 3990. In at least one embodiment, core complex 3910 and graphics complex 3940 share unified system memory 3990.

[0377] In at least one embodiment, the APU 3900 implements a memory subsystem that includes, but is not limited to, any number and type of memory controllers 3980 and memory devices (e.g., shared memory 3954) that can be dedicated to a component or shared among multiple components. In at least one embodiment, the APU 3900 implements a cache subsystem that includes, but is not limited to, one or more cache memories (e.g., L2 cache 4028, L3 cache 3930, and L2 cache 3942), each of which can be private to a component or shared among any number of components (e.g., core 3920, core complex 3910, SIMD unit 3952, compute unit 3950, and graphics complex 3940).

[0378] In at least one embodiment, Figure 39At least one component shown or described is used to implement the combination Figure 1-13 In at least one embodiment, Figure 39 At least one component of the system is configured to cause scaling of one or more clocks of one or more cores based at least in part on proximity of the one or more cores to each other. In at least one embodiment, Figure 39 At least one component performs Figure 1 The data center 120, the controller 105, at least two or more of the processors 110(a) to 110(n) and one or more thermal sensors 115(a) to 115(n), Figure 2 Firmware 220, configuration files 225, ACPI 230, policies 235, one or more processors 1-n, and / or including Figure 9 Method 900, Figure 10 Method 1000 and Figure 11 At least one aspect of one or more of the methods 1100 described herein.

[0379] Figure 40 A CPU 4000 is shown according to at least one embodiment. In at least one embodiment, the CPU 4000 is developed by Advanced Micro Devices, Inc. of Santa Clara, California. In at least one embodiment, the CPU 4000 can be configured to execute application programs. In at least one embodiment, the CPU 4000 is configured to execute host control software, such as an operating system. In at least one embodiment, the CPU 4000 issues commands that control the operation of an external GPU (not shown). In at least one embodiment, the CPU 4000 can be configured to execute host executable code derived from CUDA source code, and the external GPU can be configured to execute device executable code derived from such CUDA source code. In at least one embodiment, the CPU 4000 includes, but is not limited to, any number of core complexes 4010, fabric 4060, I / O interfaces 4070, and memory controller 4080.

[0380] In at least one embodiment, core complex 4010 includes, but is not limited to, cores 4020(1)-4020(4) and L3 cache 4030. In at least one embodiment, core complex 4010 may include, but is not limited to, any number of cores 4020 and any combination of any number and type of caches. In at least one embodiment, cores 4020 are configured to execute instructions of a specific ISA. In at least one embodiment, each core 4020 is a CPU core.

[0381] In at least one embodiment, each core 4020 includes, but is not limited to, a fetch / decode unit 4022, an integer execution engine 4024, a floating-point execution engine 4026, and an L2 cache 4028. In at least one embodiment, the fetch / decode unit 4022 fetches instructions, decodes these instructions, generates micro-ops, and dispatches individual micro-ops to the integer execution engine 4024 and the floating-point execution engine 4026. In at least one embodiment, the fetch / decode unit 4022 can simultaneously dispatch one micro-op to the integer execution engine 4024 and another micro-op to the floating-point execution engine 4026. In at least one embodiment, the integer execution engine 4024 performs, but is not limited to, integer and memory operations. In at least one embodiment, the floating-point engine 4026 performs, but is not limited to, floating-point and vector operations. In at least one embodiment, the fetch-decode unit 4022 dispatches micro-ops to a single execution engine that replaces both the integer execution engine 4024 and the floating-point execution engine 4026.

[0382] In at least one embodiment, each core 4020(i) can access an L2 cache 4028(i) included in the core 4020(i), where i is an integer representing a specific instance of the core 4020. In at least one embodiment, each core 4020 included in a core complex 4010(j) is connected to the other cores 4020 in the core complex 4010(j) via an L3 cache 4030(j) included in the core complex 4010(j), where j is an integer representing a specific instance of the core complex 4010. In at least one embodiment, a core 4020 included in a core complex 4010(j) can access all L3 caches 4030(j) included in the core complex 4010(j), where j is an integer representing a specific instance of the core complex 4010. In at least one embodiment, the L3 cache 4030 can include, but is not limited to, any number of slices.

[0383] In at least one embodiment, fabric 4060 is a system interconnect that facilitates data and control transfers across core complexes 4010(1)-4010(N) (where N is an integer greater than zero), I / O interface 4070, and memory controller 4080. In at least one embodiment, CPU 4000 may include, in addition to or in lieu of fabric 4060, but is not limited to, any number and type of system interconnects that facilitate data and control transfers across any number and type of directly or indirectly linked components that may be internal or external to CPU 4000. In at least one embodiment, I / O interface 4070 represents any number and type of I / O interface (e.g., PCI, PCI-X, PCIe, GBE, USB, etc.). In at least one embodiment, various types of peripherals are coupled to I / O interface 4070. In at least one embodiment, peripherals coupled to I / O interface 4070 may include, but are not limited to, a display, a keyboard, a mouse, a printer, a scanner, a joystick or other type of game controller, a media recording device, an external storage device, a network interface card, and the like.

[0384] In at least one embodiment, memory controller 4080 facilitates data transfers between CPU 4000 and system memory 4090. In at least one embodiment, core complex 4010 and graphics complex 4040 share system memory 4090. In at least one embodiment, CPU 4000 implements a memory subsystem that includes, but is not limited to, any number and type of memory controllers 4080 and memory devices that can be dedicated to a component or shared among multiple components. In at least one embodiment, CPU 4000 implements a cache subsystem that includes, but is not limited to, one or more cache memories (e.g., L2 cache 4028 and L3 cache 4030), each of which can be private to a component or shared among any number of components (e.g., core 4020 and core complex 4010).

[0385] In at least one embodiment, Figure 40 At least one component shown or described is used to implement the combination Figure 1-13 In at least one embodiment, Figure 40 At least one component of the system is configured to cause scaling of one or more clocks of one or more cores based at least in part on proximity of the one or more cores to each other. In at least one embodiment, Figure 40 At least one component performs Figure 1 The data center 120, the controller 105, at least two or more of the processors 110(a) to 110(n) and one or more thermal sensors 115(a) to 115(n), Figure 2 Firmware 220, configuration files 225, ACPI 230, policies 235, one or more processors 1-n, and / or including Figure 9 Method 900, Figure 10 Method 1000 and Figure 11 At least one aspect of one or more of the methods 1100 described herein.

[0386] Figure 41 An exemplary accelerator integrated slice 4190 is shown according to at least one embodiment. As used herein, a "slice" includes a specified portion of the processing resources of an accelerator integrated circuit. In at least one embodiment, the accelerator integrated circuit provides cache management, memory access, environment management, and interrupt management services on behalf of multiple graphics processing engines in multiple graphics acceleration modules. The graphics processing engines may each include a separate GPU. Optionally, the graphics processing engines may include different types of graphics processing engines within the GPU, such as a graphics execution unit, a media processing engine (e.g., a video encoder / decoder), a sampler, and a blit engine. In at least one embodiment, the graphics acceleration module may be a GPU having multiple graphics processing engines. In at least one embodiment, the graphics processing engines may be individual GPUs integrated on a common package, line card, or chip.

[0387] The application effective address space 4182 within system memory 4114 stores a process element 4183. In one embodiment, a process element 4183 is stored in response to a GPU call 4181 from an application 4180 executing on processor 4107. The process element 4183 contains the processing state of the corresponding application 4180. The work descriptor (WD) 4184 contained in the process element 4183 can be a single job requested by the application or can contain a pointer to a job queue. In at least one embodiment, the WD 4184 is a pointer to a job request queue in the application effective address space 4182.

[0388] Graphics acceleration module 4146 and / or each graphics processing engine can be shared by all or part of the processes in the system.In at least one embodiment, an infrastructure for establishing a processing state and sending WD 4184 to graphics acceleration module 4146 to start a job in a virtualized environment can be included.

[0389] In at least one embodiment, a dedicated process programming model is implemented. In this model, a single process owns the graphics acceleration module 4146 or individual graphics processing engine. Because the graphics acceleration module 4146 is owned by a single process, the hypervisor initializes the accelerator integrated circuit for the owning partition, and the operating system initializes the accelerator integrated circuit for the owning partition when the graphics acceleration module 4146 is allocated.

[0390] In operation, the WD acquisition unit 4191 in the accelerator integrated slice 4190 acquires the next WD 4184, which includes an indication of work to be completed by one or more graphics processing engines of the graphics acceleration module 4146. Data from the WD 4184 can be stored in registers 4145 for use by the memory management unit (MMU) 4139, the interrupt management circuit 4147, and / or the context management circuit 4148, as shown. For example, the MMU 4139 includes segment / page roaming circuitry for accessing the segment / page tables 4186 within the OS virtual address space 4185. The interrupt management circuit 4147 can process interrupt events (INT) 4192 received from the graphics acceleration module 4146. When executing a graph operation, the effective address 4193 generated by the graphics processing engine is converted into a real address by the MMU 4139.

[0391] In one embodiment, the same register set 4145 is replicated for each graphics processing engine and / or graphics acceleration module 4146 and can be initialized by the hypervisor or operating system. Each of these replicated registers can be included in the accelerator integration slice 4190. Table 1 shows exemplary registers that can be initialized by the hypervisor.

[0392] Table 1 - Registers initialized by the hypervisor

[0393]

[0394]

[0395] Example registers that may be initialized by the operating system are shown in Table 2.

[0396] Table 2 - Operating System Initialization Registers

[0397] 1 Process and thread identification 2 Effective Address (EA) environment save / restore pointer 3 Virtual Address (VA) Accelerator Utilization Record Pointer 4 Virtual Address (VA) stores the segment table pointer 5 Permission blocking 6 Job Descriptor

[0398] In one embodiment, each WD 4184 is specific to a particular graphics acceleration module 4146 and / or a particular graphics processing engine. It contains all the information the graphics processing engine needs to do its work or work, or it can be a pointer to a memory location where the application has set up a command queue for work to be done.

[0399] In at least one embodiment, Figure 41 At least one component shown or described is used to implement the combination Figure 1-13 In at least one embodiment, Figure 41 At least one component of the system is configured to cause scaling of one or more clocks of one or more cores based at least in part on proximity of the one or more cores to each other. In at least one embodiment, Figure 41 At least one component performs Figure 1 The data center 120, the controller 105, at least two or more of the processors 110(a) to 110(n) and one or more thermal sensors 115(a) to 115(n), Figure 2 Firmware 220, configuration files 225, ACPI 230, policies 235, one or more processors 1-n, and / or including Figure 9 Method 900, Figure 10 Method 1000 and Figure 11 At least one aspect of one or more of the methods 1100 described herein.

[0400] Figures 42A-42B An exemplary graphics processor according to at least one embodiment of the present disclosure is shown. In at least one embodiment, any exemplary graphics processor can be manufactured using one or more IP cores. In addition to the illustrated diagram, in at least one embodiment, other logic and circuitry can be included, including additional graphics processors / cores, peripheral interface controllers, or general-purpose processor cores. In at least one embodiment, the exemplary graphics processor is used within a SoC.

[0401] Figure 42A An exemplary graphics processor 4210 of a SoC integrated circuit is shown that can be fabricated using one or more IP cores in accordance with at least one embodiment. Figure 42B An additional exemplary graphics processor 4240 of a SoC integrated circuit that can be fabricated using one or more IP cores according to at least one embodiment is shown. In at least one embodiment, Figure 42A The graphics processor 4210 is a low power graphics processor core. In at least one embodiment, Figure 42B The graphics processor 4240 is a higher performance graphics processor core. In at least one embodiment, each of the graphics processors 4210, 4240 can be Figure 37 A variant of the graphics processor 3710.

[0402] In at least one embodiment, the graphics processor 4210 includes a vertex processor 4205 and one or more fragment processors 4215A-4215N (e.g., 4215A, 4215B, 4215C, 4215D through 4215N-1 and 4215N). In at least one embodiment, the graphics processor 4210 can execute different shader programs via separate logic, such that the vertex processor 4205 is optimized to perform operations for the vertex shader program, while one or more fragment processors 4215A-4215N perform fragment (e.g., pixel) shading operations for the fragment or pixel or shader program. In at least one embodiment, the vertex processor 4205 performs the vertex processing stage of the 3D graphics pipeline and generates primitives and vertex data. In at least one embodiment, the fragment processors 4215A-4215N use the primitives and vertex data generated by the vertex processor 4205 to generate a frame buffer for display on a display device. In at least one embodiment, the fragment processors 4215A-4215N are optimized to execute fragment shader programs as provided in the OpenGL API, which can be used to perform similar operations as pixel shader programs provided in the Direct 3D API.

[0403] In at least one embodiment, graphics processor 4210 additionally includes one or more MMUs 4220A-4220B, caches 4225A-4225B, and circuit interconnects 4230A-4230B. In at least one embodiment, one or more MMUs 4220A-4220B provide a mapping of virtual to physical addresses for graphics processor 4210, including for vertex processor 4205 and / or fragment processors 4215A-4215N, which may reference vertex or image / texture data stored in memory, in addition to vertex or image / texture data stored in one or more caches 4225A-4225B. In at least one embodiment, one or more MMUs 4220A-4220B may synchronize with other MMUs within the system, including with Figure 37 One or more MMUs associated with one or more application processors 3705, graphics processor 3715, and / or video processor 3720 enable each processor 3705-3720 to participate in a shared or unified virtual memory system. In at least one embodiment, one or more circuit interconnects 4230A-4230B enable the graphics processor 4210 to connect to other IP cores within the SoC via an internal bus of the SoC or via a direct connection.

[0404] In at least one embodiment, graphics processor 4240 includes Figure 42AOne or more MMUs 4220A-4220B, caches 4225A-4225B, and circuit interconnects 4230A-4230B of the graphics processor 4210. In at least one embodiment, the graphics processor 4240 includes one or more shader cores 4255A-4255N (e.g., 4255A, 4255B, 4255C, 4255D, 4255E, 4255F, through 4255N-1 and 4255N) that provide a unified shader core architecture in which a single core or type or core can execute all types of programmable shader code, including shader program code for implementing vertex shaders, fragment shaders, and / or compute shaders. In at least one embodiment, the number of shader cores can vary. In at least one embodiment, the graphics processor 4240 includes an inter-core task manager 4245 that acts as a thread dispatcher to dispatch execution threads to one or more shader cores 4255A-4255N and a tiling unit 4258 to accelerate tile-based rendering operations in which rendering operations of a scene are subdivided in image space, for example, to exploit local spatial coherence within a scene or to optimize use of internal caches.

[0405] In at least one embodiment, Figures 42A-42B At least one component shown or described is used to implement the combination Figure 1-13 In at least one embodiment, Figures 42A-42B At least one component of the system is configured to cause scaling of one or more clocks of one or more cores based at least in part on proximity of the one or more cores to each other. In at least one embodiment, Figures 42A-42B At least one component performs Figure 1 The data center 120, the controller 105, at least two or more of the processors 110(a) to 110(n) and one or more thermal sensors 115(a) to 115(n), Figure 2 Firmware 220, configuration files 225, ACPI 230, policies 235, one or more processors 1-n, and / or including Figure 9 Method 900, Figure 10 Method 1000 and Figure 11 At least one aspect of one or more of the methods 1100 described herein.

[0406] Figure 43A FIG4 shows a graphics core 4300 according to at least one embodiment. In at least one embodiment, the graphics core 4300 may include Figure 37 In at least one embodiment, the graphics core 4300 may be Figure 42B4305N. In at least one embodiment, graphics core 4300 includes a shared instruction cache 4302, texture units 4318, and cache / shared memory 4320, which are common to execution resources within graphics core 4300. In at least one embodiment, graphics core 4300 may include multiple slices 4301A-4301N or partitions of each core, and a graphics processor may include multiple instances of graphics core 4300. Slices 4301A-4301N may include support logic including local instruction caches 4304A-4304N, thread schedulers 4306A-4306N, thread dispatchers 4308A-4308N, and a set of registers 4310A-4310N. In at least one embodiment, the slices 4301A-4301N may include a set of additional function units (AFUs) 4312A-4312N, floating point units (FPUs) 4314A-4314N, integer arithmetic logic units (ALUs) 4316A-4316N, address calculation units (ACUs) 4313A-4313N, double precision floating point units (DPFPUs) 4315A-4315N, and matrix processing units (MPUs) 4317A-4317N.

[0407] In one embodiment, the FPUs 4314A-4314N can perform single-precision (32-bit) and half-precision (16-bit) floating-point operations, while the DPFPUs 4315A-4315N can perform double-precision (64-bit) floating-point operations. In at least one embodiment, the ALUs 4316A-4316N can perform variable-precision integer operations with 8-bit, 16-bit, and 32-bit precision, and can be configured for mixed-precision operations. In at least one embodiment, the MPUs 4317A-4317N can also be configured for mixed-precision matrix operations, including half-precision floating-point operations and 8-bit integer operations. In at least one embodiment, the MPUs 4317A-4317N can perform various matrix operations to accelerate CUDA programs, including enabling support for accelerated general matrix-to-matrix multiplication (GEMM). In at least one embodiment, the AFUs 4312A-4312N can perform additional logical operations not supported by the floating-point or integer units, including trigonometric operations (e.g., Sine, Cosine, etc.).

[0408] Figure 43BA general purpose graphics processing unit (GPGPU) 4330 is shown in at least one embodiment. In at least one embodiment, GPGPU 4330 is highly parallel and suitable for deployment on a multi-chip module. In at least one embodiment, GPGPU 4330 can be configured to enable highly parallel computational operations to be performed by an array of GPUs. In at least one embodiment, GPGPU 4330 can be directly linked to other instances of GPGPU 4330 to create a multi-GPU cluster to improve execution time for CUDA programs. In at least one embodiment, GPGPU 4330 includes a host interface 4332 to enable connection to a host processor. In at least one embodiment, host interface 4332 is a PCIe interface. In at least one embodiment, host interface 4332 can be a vendor-specific communication interface or communication structure. In at least one embodiment, GPGPU 4330 receives commands from the host processor and dispatches execution threads associated with those commands to a set of compute clusters 4336A-4336H using a global scheduler 4334. In at least one embodiment, compute clusters 4336A-4336H share cache memory 4338. In at least one embodiment, cache memory 4338 may serve as a higher level cache for cache memories within compute clusters 4336A-4336H.

[0409] In at least one embodiment, GPGPU 4330 includes memory 4344A-4344B coupled to a compute cluster 4336A-4336H via a set of memory controllers 4342A-4342B. In at least one embodiment, memory 4344A-4344B may include various types of memory devices, including dynamic random access memory (DRAM) or graphics random access memory, such as synchronous graphics random access memory (SGRAM), including graphics double data rate (GDDR) memory.

[0410] In at least one embodiment, computing clusters 4336A-4336H each include a set of graphics cores, such as Figure 43A The graphics core 4300, which may include multiple types of integer and floating-point logic units, may perform computational operations at various precisions, including computations suitable for use with CUDA programs. For example, in at least one embodiment, at least a subset of the floating-point units in each of the compute clusters 4336A-4336H may be configured to perform 16-bit or 32-bit floating-point operations, while a different subset of the floating-point units may be configured to perform 64-bit floating-point operations.

[0411] In at least one embodiment, multiple instances of GPGPU 4330 can be configured to operate as a compute cluster. In at least one embodiment, compute clusters 4336A-4336H can implement any technically feasible communication technology for synchronization and data exchange. In at least one embodiment, multiple instances of GPGPU 4330 communicate via host interface 4332. In at least one embodiment, GPGPU 4330 includes an I / O hub 4339 that couples GPGPU 4330 to GPU link 4340, enabling direct connection to other instances of GPGPU 4330. In at least one embodiment, GPU link 4340 is coupled to a dedicated GPU-to-GPU bridge that enables communication and synchronization between multiple instances of GPGPU 4330. In at least one embodiment, GPU link 4340 is coupled to a high-speed interconnect to send and receive data to other GPGPUs or parallel processors. In at least one embodiment, multiple instances of GPGPU 4330 are located in separate data processing systems and communicate via a network device accessible via host interface 4332. In at least one embodiment, GPU link 4340 may be configured to connect to a host processor, in addition to or in place of host interface 4332. In at least one embodiment, GPGPU 4330 may be configured to execute CUDA programs.

[0412] In at least one embodiment, Figures 43A-43B At least one component shown or described is used to implement the combination Figure 1-13 In at least one embodiment, Figures 43A-43B At least one component of the system is configured to cause scaling of one or more clocks of one or more cores based at least in part on proximity of the one or more cores to each other. In at least one embodiment, Figures 43A-43B At least one component performs Figure 1 The data center 120, the controller 105, at least two or more of the processors 110(a) to 110(n) and one or more thermal sensors 115(a) to 115(n), Figure 2 Firmware 220, configuration files 225, ACPI 230, policies 235, one or more processors 1-n, and / or including Figure 9 Method 900, Figure 10 Method 1000 and Figure 11 At least one aspect of one or more of the methods 1100 described herein.

[0413] Figure 44AA parallel processor 4400 is shown in accordance with at least one embodiment. In at least one embodiment, the various components of the parallel processor 4400 may be implemented using one or more integrated circuit devices, such as a programmable processor, an application specific integrated circuit (ASIC), or an FPGA.

[0414] In at least one embodiment, parallel processor 4400 includes parallel processing unit 4402. In at least one embodiment, parallel processing unit 4402 includes an I / O unit 4404 that enables communication with other devices, including other instances of parallel processing unit 4402. In at least one embodiment, I / O unit 4404 can be directly connected to other devices. In at least one embodiment, I / O unit 4404 connects to other devices using a hub or switch interface (e.g., memory hub 1705). In at least one embodiment, the connection between memory hub 1905 and I / O unit 4404 forms a communication link. In at least one embodiment, I / O unit 4404 is connected to a host interface 4406 and a memory crossbar switch 4416, where host interface 4406 receives commands for performing processing operations and memory crossbar switch 4416 receives commands for performing memory operations.

[0415] In at least one embodiment, when host interface 4406 receives command buffers via I / O unit 4404, host interface 4406 can direct work operations to execute those commands to front end 4408. In at least one embodiment, front end 4408 is coupled to scheduler 4410, which is configured to dispatch commands or other work items to processing array 4412. In at least one embodiment, scheduler 4410 ensures that processing array 4412 is properly configured and in a valid state before dispatching tasks to processing array 4412. In at least one embodiment, scheduler 4410 is implemented by firmware logic executing on a microcontroller. In at least one embodiment, a microcontroller-implemented scheduler 4410 can be configured to perform complex scheduling and work dispatch operations at both coarse and fine granularity, thereby enabling fast preemption and context switching of threads executing on processing array 4412. In at least one embodiment, host software can authenticate workloads for scheduling on processing array 4412 through one of multiple graphics processing doorbells. In at least one embodiment, the workload can then be automatically distributed across the processing array 4412 by scheduler 4410 logic within a microcontroller that includes scheduler 4410 .

[0416] In at least one embodiment, processing array 4412 can include up to "N" processing clusters (e.g., cluster 4414A, cluster 4414B, through cluster 4414N). In at least one embodiment, each cluster 4414A-4414N of processing array 4412 can execute a large number of concurrent threads. In at least one embodiment, scheduler 4410 can allocate work to clusters 4414A-4414N of processing array 4412 using various scheduling and / or work distribution algorithms, which can vary depending on the workload generated by each program or computation type. In at least one embodiment, scheduling can be handled dynamically by scheduler 4410 or can be partially assisted by compiler logic during the compilation of program logic configured to be executed by processing array 4412. In at least one embodiment, different clusters 4414A-4414N of processing array 4412 can be assigned to process different types of programs or to perform different types of computations.

[0417] In at least one embodiment, processing array 4412 can be configured to perform various types of parallel processing operations. In at least one embodiment, processing array 4412 is configured to perform general-purpose parallel computing operations. For example, in at least one embodiment, processing array 4412 can include logic to perform processing tasks including filtering video and / or audio data, performing modeling operations including physics operations, and performing data transformations.

[0418] In at least one embodiment, processing array 4412 is configured to perform parallel graphics processing operations. In at least one embodiment, processing array 4412 may include additional logic to support the execution of such graphics processing operations, including but not limited to texture sampling logic to perform texture operations, as well as tessellation logic and other vertex processing logic. In at least one embodiment, processing array 4412 may be configured to execute shader programs related to graphics processing, such as, but not limited to, vertex shaders, tessellation shaders, geometry shaders, and pixel shaders. In at least one embodiment, parallel processing units 4402 may transfer data from system memory via I / O units 4404 for processing. In at least one embodiment, during processing, the transferred data may be stored in on-chip memory (e.g., parallel processor memory 4422) during processing and then written back to system memory.

[0419] In at least one embodiment, when parallel processing unit 4402 is used to perform graph processing, scheduler 4410 can be configured to divide the processing workload into tasks of approximately equal size to better distribute graphics processing operations to multiple clusters 4414A-4414N of processing array 4412. In at least one embodiment, portions of processing array 4412 can be configured to perform different types of processing. For example, in at least one embodiment, a first portion can be configured to perform vertex shading and topology generation, a second portion can be configured to perform tessellation and geometry shading, and a third portion can be configured to perform pixel shading or other screen-space operations to generate a rendered image for display. In at least one embodiment, intermediate data generated by one or more of clusters 4414A-4414N can be stored in a buffer to allow the intermediate data to be transferred between clusters 4414A-4414N for further processing.

[0420] In at least one embodiment, the processing array 4412 can receive processing tasks to be executed via the scheduler 4410, which receives commands defining the processing tasks from the front end 4408. In at least one embodiment, the processing tasks can include an index of the data to be processed, which can include, for example, surface (patch) data, primitive data, vertex data, and / or pixel data, as well as state parameters and commands defining how to process the data (e.g., what program to execute). In at least one embodiment, the scheduler 4410 can be configured to obtain the index corresponding to the task, or can receive the index from the front end 4408. In at least one embodiment, the front end 4408 can be configured to ensure that the processing array 4412 is configured in a valid state before starting the workload specified by the incoming command buffer (e.g., batch buffer, push buffer, etc.).

[0421] In at least one embodiment, each of one or more instances of parallel processing unit 4402 can be coupled to parallel processor memory 4422. In at least one embodiment, parallel processor memory 4422 can be accessed via memory crossbar 4416, which can receive memory requests from processing array 4412 and I / O unit 4404. In at least one embodiment, memory crossbar 4416 can access parallel processor memory 4422 via memory interface 4418. In at least one embodiment, memory interface 4418 can include multiple partition units (e.g., partition unit 4420A, partition unit 4420B, through partition unit 4420N), which can each be coupled to a portion of parallel processor memory 4422 (e.g., a memory unit). In at least one embodiment, the plurality of partition units 4420A-4420N are configured to be equal to the number of memory cells, such that the first partition unit 4420A has a corresponding first memory cell 4424A, the second partition unit 4420B has a corresponding memory cell 4424B, and the Nth partition unit 4420N has a corresponding Nth memory cell 4424N. In at least one embodiment, the number of partition units 4420A-4420N may not be equal to the number of memory devices.

[0422] In at least one embodiment, memory units 4424A-4424N may include various types of memory devices, including dynamic random access memory (DRAM) or graphics random access memory, such as synchronous graphics random access memory (SGRAM), including graphics double data rate (GDDR) memory. In at least one embodiment, memory units 4424A-4424N may also include 3D stacked memory, including but not limited to high bandwidth memory (HBM). In at least one embodiment, render targets such as frame buffers or texture maps may be stored across memory units 4424A-4424N, allowing partition units 4420A-4420N to write portions of each render target in parallel to efficiently use the available bandwidth of parallel processor memory 4422. In at least one embodiment, local instances of parallel processor memory 4422 may be eliminated in favor of a unified memory design utilizing system memory in combination with local cache memory.

[0423] In at least one embodiment, any of the clusters 4414A-4414N in the processing array 4412 can process data to be written to any memory unit 4424A-4424N within the parallel processor memory 4422. In at least one embodiment, the memory crossbar 4416 can be configured to transmit the output of each cluster 4414A-4414N to any partition unit 4420A-4420N or to another cluster 4414A-4414N, which can perform other processing operations on the output. In at least one embodiment, each cluster 4414A-4414N can communicate with a memory interface 4418 via the memory crossbar 4416 to read from or write to various external storage devices. In at least one embodiment, memory crossbar 4416 has connections to memory interface 4418 for communicating with I / O unit 4404, as well as connections to local instances of parallel processor memory 4422, thereby enabling processing units within different processing clusters 4414A-4414N to communicate with system memory or other memory that is not local to parallel processing unit 4402. In at least one embodiment, memory crossbar 4416 may use virtual channels to separate traffic flows between clusters 4414A-4414N and partition units 4420A-4420N.

[0424] In at least one embodiment, multiple instances of parallel processing unit 4402 can be provided on a single plug-in card, or multiple plug-in cards can be interconnected. In at least one embodiment, different instances of parallel processing unit 4402 can be configured to interoperate with each other, even if the different instances have different numbers of processing cores, different amounts of local parallel processor memory, and / or other configuration differences. For example, in at least one embodiment, some instances of parallel processing unit 4402 can include higher precision floating point units relative to other instances. In at least one embodiment, a system incorporating one or more instances of parallel processing unit 4402 or parallel processor 4400 can be implemented in a variety of configurations and form factors, including but not limited to desktop, laptop, or handheld personal computers, servers, workstations, game consoles, and / or embedded systems.

[0425] Figure 44B FIG4 illustrates a processing cluster 4494 according to at least one embodiment. In at least one embodiment, the processing cluster 4494 is included within a parallel processing unit. In at least one embodiment, the processing cluster 4494 is Figure 44AIn at least one embodiment, processing cluster 4494 can be configured to execute many threads in parallel, where the term "thread" refers to an instance of a particular program executed on a particular set of input data. In at least one embodiment, single instruction multiple data (SIMD) instruction issuance technology is used to support the parallel execution of a large number of threads without providing multiple independent instruction units. In at least one embodiment, single instruction multiple thread (SIMT) technology is used to support the parallel execution of a large number of generally synchronized threads, which uses a common instruction unit that is configured to issue instructions to a group of processing engines within each processing cluster 4494.

[0426] In at least one embodiment, the operation of the processing cluster 4494 can be controlled by a pipeline manager 4432 that assigns processing tasks to SIMT parallel processors. In at least one embodiment, the pipeline manager 4432 Figure 44A Scheduler 4410 receives instructions and manages the execution of these instructions by graphics multiprocessor 4434 and / or texture unit 4436. In at least one embodiment, graphics multiprocessor 4434 is an exemplary instance of a SIMT parallel processor. However, in at least one embodiment, various types of SIMT parallel processors with different architectures may be included within processing cluster 4494. In at least one embodiment, one or more instances of graphics multiprocessor 4434 may be included within processing cluster 4494. In at least one embodiment, graphics multiprocessor 4434 may process data, and data crossbar 4440 may be used to distribute the processed data to one of multiple possible destinations (including other shader units). In at least one embodiment, pipeline manager 4432 may facilitate the distribution of processed data by specifying the destination of the processed data to be distributed via data crossbar 4440.

[0427] In at least one embodiment, each graphics multiprocessor 4434 within a processing cluster 4494 may include the same set of function execution logic (e.g., arithmetic logic unit, load store unit (LSU), etc.). In at least one embodiment, the function execution logic may be configured in a pipelined manner, where new instructions may be issued before previous instructions have completed. In at least one embodiment, the function execution logic supports a variety of operations, including integer and floating point arithmetic, comparison operations, Boolean operations, shifts, and calculations of various algebraic functions. In at least one embodiment, the same functional unit hardware may be utilized to perform different operations, and any combination of functional units may be present.

[0428] In at least one embodiment, instructions transmitted to processing cluster 4494 constitute threads. In at least one embodiment, a group of threads executed across a set of parallel processing engines is a thread group. In at least one embodiment, a thread group executes a program on different input data. In at least one embodiment, each thread within a thread group can be assigned to a different processing engine within graphics multiprocessor 4434. In at least one embodiment, a thread group can include fewer threads than the number of processing engines within graphics multiprocessor 4434. In at least one embodiment, when a thread group includes fewer threads than the number of processing engines, one or more processing engines may be idle during the processing of a loop of the thread group. In at least one embodiment, a thread group can also include more threads than the number of processing engines within graphics multiprocessor 4434. In at least one embodiment, when a thread group includes more threads than the number of processing engines within graphics multiprocessor 4434, processing can be performed in consecutive clock cycles. In at least one embodiment, multiple thread groups can be executed simultaneously on graphics multiprocessor 4434.

[0429] In at least one embodiment, graphics multiprocessor 4434 includes internal cache memory to perform load and store operations. In at least one embodiment, graphics multiprocessor 4434 can abandon the internal cache and use cache memory within processing cluster 4494 (e.g., L1 cache 4448). In at least one embodiment, each graphics multiprocessor 4434 can also access partition units (e.g., Figure 44A L2 cache within partition units 4420A-4420N) of the graphics multiprocessor 4434 is shared across all processing clusters 4494 and can be used to transfer data between threads. In at least one embodiment, graphics multiprocessor 4434 can also access off-chip global memory, which can include one or more of local parallel processor memory and / or system memory. In at least one embodiment, any memory external to parallel processing unit 4402 can be used as global memory. In at least one embodiment, processing cluster 4494 includes multiple instances of graphics multiprocessor 4434, which can share common instructions and data, which can be stored in L1 cache 4448.

[0430] In at least one embodiment, each processing cluster 4494 may include an MMU 4445 configured to map virtual addresses to physical addresses. In at least one embodiment, one or more instances of the MMU 4445 may reside in Figure 44A4418. In at least one embodiment, the MMU 4445 includes a set of page table entries (PTEs) that are used to map virtual addresses to physical addresses of tiles (more on tiles below) and optionally to cache line indices. In at least one embodiment, the MMU 4445 may include an address translation lookaside buffer (TLB) or cache that may reside within the graphics multiprocessor 4434 or L1 cache 4448 or processing cluster 4494. In at least one embodiment, the physical addresses are processed to assign surface data access locality for efficient request interleaving between partition units. In at least one embodiment, the cache line index may be used to determine whether a request for a cache line is a hit or a miss.

[0431] In at least one embodiment, the processing clusters 4494 can be configured such that each graphics multiprocessor 4434 is coupled to a texture unit 4436 to perform texture mapping operations, which may involve, for example, determining texture sample locations, reading texture data, and filtering the texture data. In at least one embodiment, texture data is read from an internal texture L1 cache (not shown) or from an L1 cache within the graphics multiprocessor 4434, and texture data is retrieved from an L2 cache, local parallel processor memory, or system memory as needed. In at least one embodiment, each graphics multiprocessor 4434 outputs processed tasks to a data crossbar 4440 to provide the processed tasks to another processing cluster 4494 for further processing or to store the processed tasks in an L2 cache, local parallel processor memory, or system memory via...

Claims

1. A processor, comprising: One or more circuits for scaling one or more clocks of the one or more cores based at least in part on proximity of the one or more cores to each other. 2 . The processor of claim 1 , wherein the one or more circuits are configured to scale the one or more clocks based at least in part on one or more core utilization patterns indicative of one or more thermal conditions.

3. The processor of claim 1, wherein the one or more circuits are configured to avoid one or more core utilization patterns that would result in one or more adverse thermal conditions.

4. The processor of claim 1 , wherein the one or more circuits are configured to select one or more cores of the processor to execute one or more threads based at least in part on avoiding one or more core utilization patterns indicative of one or more adverse thermal conditions. 5 . The processor of claim 1 , wherein the one or more circuits are configured to scale the one or more clocks based on information indicative of one or more core utilization patterns stored in one or more configuration files.

6. The processor of claim 1, wherein the one or more circuits are configured to compare core utilization of the processor to one or more core utilization patterns. 7 . The processor of claim 1 , wherein the proximity of the one or more cores corresponds to one or more locations of the one or more cores on a chip of the processor.

8. A system comprising one or more processors configured to scale one or more clocks of one or more cores based at least in part on proximity of the one or more cores to each other.

9. The system of claim 8, wherein the one or more processors are configured to scale the one or more clocks based at least in part on one or more core utilization patterns indicative of one or more thermal conditions.

10. The system of claim 8, wherein the one or more processors are configured to avoid one or more core utilization patterns that would result in one or more adverse thermal conditions.

11. The system of claim 8, wherein the one or more processors are configured to select one or more cores of the processor to execute one or more threads based at least in part on avoiding one or more core utilization patterns indicative of one or more adverse thermal conditions.

12. The system of claim 8, wherein the one or more processors are configured to scale the one or more clocks based on information indicative of one or more core utilization patterns stored in one or more configuration files.

13. The system of claim 8, wherein the one or more processors are configured to compare core utilization of the processors to one or more core utilization patterns.

14. The system of claim 8, wherein the proximity of the one or more cores corresponds to one or more locations of the one or more cores on a chip of the processor.

15. A method comprising: One or more clocks of one or more cores of a processor are scaled based at least in part on proximity of the one or more cores to each other.

16. The method of claim 15, wherein the scaling of the one or more clocks is based at least in part on one or more core utilization patterns indicative of one or more thermal conditions.

17. The method of claim 15, wherein the scaling of the one or more clocks is used to avoid one or more core utilization patterns that would result in one or more adverse thermal conditions.

18. The method of claim 15, wherein the scaling of the one or more clocks is used to select one or more cores of the processor to execute one or more threads based at least in part on avoiding one or more core utilization patterns indicative of one or more adverse thermal conditions.

19. The method of claim 15, wherein the scaling of the one or more clocks is based on information indicative of one or more core utilization patterns stored in one or more configuration files.

20. The method of claim 15, wherein the scaling of the one or more clocks compares core utilization of the processor to one or more core utilization patterns.