Processor clock scaling technology
Patent Information
- Application Number
- DE102025104401
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-06
- Filing Date
- 2025-02-06
- Publication Date
- 2025-09-11
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
TECHNICAL FIELD
[0001] At least one embodiment relates to scaling clocks of one or more processor cores. For example, at least one embodiment relates to a processor including one or more circuits to scale one or more clocks of one or more cores based at least in part on the proximity of the one or more cores to each other. BACKGROUND
[0002] Processors can manage heat through a software temperature policy that monitors the average temperature across thermal sensors located in a collection of operating cores. Currently, a worst-case offset for thermal hotspots based on profiled applications is added as a safety margin. BRIEF DESCRIPTION OF THE DRAWINGS
[0003] The following detailed description of exemplary, non-limiting embodiments should be read in conjunction with the drawings, of which Fig. 1 shows a system including a controller and one or more processor cores, according to at least one embodiment; Fig. 2 shows a multi-core processor chip according to at least one embodiment; Fig. 3 shows a multi-core processor chip according to at least one embodiment; Fig. 4 shows a multi-core processor chip with a combination of processor cores in use according to at least one embodiment; Fig. 5 shows a block diagram of a multi-core processor chip with a different combination of processor cores in use according to at least one embodiment; Fig. 6 shows a block diagram of a multi-core processor chip with different combinations of processor cores in use according to at least one embodiment; Fig. 7 shows a block diagram of a multi-core processor chip with different combinations of processor cores in use according to at least one embodiment; Fig. 8A shows a heat map of a processor chip showing an example of a worst-case thermal hotspot, according to at least one embodiment; Fig. 8B shows a thermal map of a processor chip showing a uniformly loaded core configuration without a thermal hotspot, according to at least one embodiment; Fig. 9 shows a process flow for implementing a temperature policy management system according to at least one embodiment; Fig. 10 shows a process for updating temperature policy offsets according to at least one embodiment; Fig. 11 shows a process of a system implementing temperature policy management according to at least one embodiment; Fig. 12 shows a processor module according to at least one embodiment; Fig. 13 depicts an API for scaling one or more clocks of one or more cores based, at least in part, on a proximity of the one or more cores to each other, according to at least one embodiment; Fig. 14 shows a distributed system according to at least one embodiment; Fig. 15 shows an exemplary data center according to at least one embodiment; Fig. 16 shows a client-server network according to at least one embodiment; Fig. 17 shows an example of a computer network according to at least one embodiment; Fig. 18A shows a networked computer system according to at least one embodiment; Fig. 18B shows a networked computer system according to at least one embodiment; Fig. 18C shows a networked computer system according to at least one embodiment; Fig. 19 illustrates one or more components of a system environment in which services may be offered as third-party network services, according to at least one embodiment; Fig. 20 shows a cloud computing environment according to at least one embodiment; Fig. 21 illustrates a set of functional abstraction layers provided by a cloud computing environment, according to at least one embodiment; Fig. 22 shows a chip-level supercomputer, according to at least one embodiment; Fig. 23 shows a rack module level supercomputer, according to at least one embodiment; Fig. 24 shows a rack-level supercomputer according to at least one embodiment; Fig. 25 shows a full system level supercomputer according to at least one embodiment; Fig. 26A shows inference and / or training logic according to at least one embodiment; Fig. 26B shows inference and / or training logic according to at least one embodiment; Fig. 27 shows the training and deployment of a neural network according to at least one embodiment; Fig. 28 shows an architecture of a network system according to at least one embodiment; Fig. 29 shows an architecture of a network system according to at least one embodiment; Fig. 30 shows a control protocol stack according to at least one embodiment; Fig. 31 shows a user protocol stack according to at least one embodiment; Fig. 32 shows components of a core network according to at least one embodiment; Fig. 33 illustrates components of a system for supporting network function virtualization (NFV) in accordance with at least one embodiment; Fig. 34 shows a processing system according to at least one embodiment; Fig. 35 shows a computer system according to at least one embodiment; Fig. 36 shows a system according to at least one embodiment; Fig. 37 shows an exemplary integrated circuit according to at least one embodiment; Fig. 38 shows a computer system according to at least one embodiment; Fig. 39 shows an APU according to at least one embodiment; Fig. 40 shows a CPU according to at least one embodiment; Fig. 41 shows an exemplary accelerator integration slice according to at least one embodiment; Fig. 42A and Fig. 42B illustrate example graphics processors according to at least one embodiment; Fig. 43A shows a graphics core according to at least one embodiment; Fig. 43B shows a GPGPU according to at least one embodiment; Fig. 44A shows a parallel processor according to at least one embodiment; Fig. 44B shows a processing cluster according to at least one embodiment; Fig. 44C shows a graphics multiprocessor according to at least one embodiment; Fig. 45 shows a software stack of a programming platform according to at least one embodiment; Fig. 46 shows a CUDA implementation of a software stack from Fig. 45 according to at least one embodiment; Fig. 47 shows a ROCm implementation of a software stack from Fig. 45 according to at least one embodiment; Fig. 48 shows an OpenCL implementation of a software stack from Fig. 45 according to at least one embodiment; Fig. 49 shows software supported by a programming platform according to at least one embodiment; and Fig. 50 shows compiling code for execution on programming platforms of Fig. 45 - 48 according to at least one embodiment. DETAILED DESCRIPTION
[0004] Fig. 1 shows a system 100 including a controller and one or more processor cores, according to at least one embodiment. In at least one embodiment, a chip 120 including multiple cores 110 may suffer from localized heat buildup referred to as "hotspots," a type of thermal overload that can lead to hardware failure. In at least one embodiment, this condition occurs because the temperature sensors on the chip 120 do not accurately measure the temperature for each processor core or each region of the chip because the temperature sensors, such as thermal sensors 115, are not evenly or ideally distributed across the chip 120.In at least one embodiment, hotspots may result from processor core activity when certain combinations of processor cores 110 are subject to a sustained level of operation, such as combinations where the proximity of the processor cores 110 causes heat to build up in a corresponding area of the chip 120. In at least one embodiment, these hotspots may go undetected or be inaccurately measured by the thermal sensors 115, due at least in part to the proximity of the thermal sensors 115 to that hotspot.
[0005] In at least one embodiment, processor activity may be scaled down when a temperature exceeds a certain value, thereby avoiding a thermal failure. In at least one embodiment, this value is based on an estimated temperature derived in part from thermal sensors 115. In at least one embodiment, the estimated temperature is adjusted by adding a thermal offset because specific temperature data is not available at every location. In at least one embodiment, processor activity may be scaled according to the utilization patterns of processor cores 110, rather than using an offset value that is too high (which may help prevent overheating but also excessively throttles processor activity) or a value that is too low (which risks a thermal failure).In at least one embodiment, these patterns include information about a local heat buildup that may be inaccurately measured by thermal sensors 115.
[0006] In at least one embodiment, the use of adjustable thermal offsets for different processor core utilization patterns provides a thermal management policy with greater flexibility in managing thermal hotspots and processor performance. In at least one embodiment, records of processor core utilization and recorded or predicted temperatures can be used to identify patterns that would lead to a thermally induced failure in the absence of a thermal offset or a fixed thermal offset that is too low to provide an adequate safety margin. In at least one embodiment, these patterns, which can be described as thermal worst-case combinations, worst-case combinations, thermal worst-case scenarios, etc., can be avoided or accommodated by the processor to improve processor performance.In at least one embodiment, this is achieved by scaling the clocks of one or more processor cores.
[0007] In at least one embodiment, a configuration file, such as configuration file 104, may include information indicating processor core combinations, worst-case scenarios, or other patterns mentioned above that may influence which thermal offsets are feasible. In at least one embodiment, this information may include indications of processor core utilization patterns that may be avoided to enable a smaller thermal offset. In at least one embodiment, the information may indicate patterns that, if identified, may be used as a basis for dynamically setting a relatively high thermal offset.
[0008] In at least one embodiment, scaling clocks of one or more processor cores or combinations comprises increasing or decreasing the speeds of the processor cores, thereby causing a corresponding increase or decrease in an amount of thermal energy generated by the processor cores. In at least one embodiment, adjusting the operation of a first set of one or more processor cores (e.g., by temporarily removing the processor cores from operation) allows the clocks of another set of processor cores to be increased, the first set being a set including processors whose proximity is associated with adverse thermal conditions.
[0009] In at least one embodiment, the thermal performance of one or more processor cores is correlated with the proximity of the processor cores. For example, in at least one embodiment, the thermal performance of a group of processor cores is associated with how close the operating processor cores are to each other. In at least one embodiment, there is a further relationship between proximity and the activity of the processor cores. In at least one embodiment, processor cores operating in proximity to each other, such as processor cores located relatively close to each other on chip 120, may generate thermal activity associated with an adverse thermal condition, and the processor cores' clocks may be scaled accordingly.
[0010] In at least one embodiment, a pattern of processor core utilization includes determining combinations of processor cores that may result in a particular thermal performance, such as thermal performance associated with a hotspot or other worst-case scenario, such as hotspots not accurately measured by the onboard thermal sensors 115. In at least one embodiment, a pattern of core utilization may indicate that a processor core or combination of processor cores exhibits thermal activity that is within the limits of acceptable thermal performance and does not require intervention by a thermal management policy.In at least one embodiment, a pattern of core utilization of a processor may indicate that said processor core or combinations of processor cores would generate thermal activity that exceeds a limit of acceptable thermal performance, but that can be managed by scaling the activity of said processor cores according to a thermal management policy. In at least one embodiment, this includes adjustments to thermal offsets as described herein. In at least one embodiment, a pattern of core utilization of the processor may indicate that the processor core or combinations of processor cores would generate thermal activity that exceeds a limit of acceptable thermal performance, but that can be managed by temporarily removing the processor core combinations from operation according to a thermal management policy.
[0011] In at least one embodiment, a thermal condition corresponds to one or more temperatures of a processor core, a combination of processor cores, a processor component other than a processor core, or a location on a chip.
[0012] In at least one embodiment, an adverse thermal condition is a condition that jeopardizes the ongoing operational performance of one or more combinations of processor cores. In at least one embodiment, avoiding one or more core utilization patterns that would result in an adverse thermal condition includes implementing a thermal management policy that removes a processor core or combination of processor cores from operation or adjusts the operation of the processor core or combination of processor cores by scaling a clock of the processor core(s) to avoid the adverse thermal condition.
[0013] In at least one embodiment, the information indicating one or more patterns of core utilization of the processor includes data correlating processor cores or combinations of processor cores with thermal conditions.
[0014] In at least one embodiment, the system 100 includes a chip 120 that includes a controller 105 and a number of processor cores 110. In at least one embodiment, the controller 105 manages the overall performance of the processor cores 110, which includes implementing a thermal management policy as described herein. In at least one embodiment, one or more aspects of one or more embodiments described in connection with Fig. 1, combined with one or more aspects of one or more embodiments described herein, including at least the embodiments described in connection with Fig. 2-13 are described.
[0015] In at least one embodiment, chip 120 includes multiple groups of processor cores. In at least one embodiment, chip 120 includes, for example, two or more groups of processor cores connected to chip 120 via two or more sockets or circuits configured to operate as such, and in which each group of processors functions as an independent multi-core processor. In at least one embodiment, the groups of processor cores communicate via a communication bus, such as PCIe, NVLink, or Infinity Fabric xGMI. In at least one embodiment, chip 120 corresponds to a GRACE chip, an AMD Instinct MI300 series chip, or another similar chip.
[0016] In at least one embodiment, the measurements from thermal sensors 115 provide thermal measurements in substantially real-time. In at least one embodiment, a configuration file 104 includes stored information for identifying combinations of processors associated with pathological patterns of processor usage. In at least one embodiment, an application's usage patterns for a group of processors, such as processors 110, are correlated with the patterns specified in configuration file 104. In at least one embodiment, these patterns are indicative of adverse thermal conditions. In at least one embodiment, configuration file 104 is specific to a chip, operating system, or computing device. In at least one embodiment, processor core utilization patterns are specified in configuration file 104 as data or text.For example, in at least one embodiment, configuration file 104 could contain text strings indicating groups of processor cores, such as "C33, C34, C43," which indicate a group of processor cores that, at a sufficiently high level, could cause adverse thermal conditions. It will be appreciated that this example is not intended to limit possible embodiments to those corresponding to the present example.
[0017] In at least one embodiment, different thermal hotspot offsets may be assigned to different processor configurations to dynamically configure the processor cores and balance lowering processor core temperatures with maintaining high performance. In at least one embodiment, this is accomplished by system 100 identifying processor core utilization patterns associated with adverse thermal conditions and scaling the clocks of one or more processors based on this identification. In at least one embodiment, system 100 provides alternative temperature policies in response to identifying these patterns.For example, in at least one embodiment, system 100 correlates the current core utilization of processor cores 110 with processor core utilization patterns specified in configuration file 104.
[0018] In at least one embodiment, the configuration file 104 corresponds to the configuration file 225 as shown in Fig. 2. In at least one embodiment, the system 100 includes an Advanced Configuration Power Interface (ACPI), such as AFCPI 230, as shown in Fig. 2. In at least one embodiment, system 100 includes firmware, such as firmware 220, for implementing a thermal management policy, such as policy 235. In at least one embodiment, the firmware monitors application usage of processor cores 110, compares the processor cores in use, and compares the processor cores to core utilization patterns that may indicate adverse thermal conditions. In at least one embodiment, these patterns are specified in configuration file 105. In at least one embodiment, the firmware and software adjusts one or more processor core clocks to increase or decrease its performance and corresponding thermal performance according to the policy.
[0019] Fig. Figure 2 shows a multi-core processor chip in accordance with at least one embodiment. In at least one embodiment, chip 205 includes a plurality of processor cores 210 arranged in Fig. 2 are designated as processor cores 1-76, wherein this number of processor cores is for illustrative purposes only and should not be construed as limiting any particular embodiment. In at least one embodiment, chip 205 includes thermal sensors, such as thermal sensors 215A-D. In at least one embodiment, chip 205 includes an arrangement of processor cores corresponding to the Fig. 2, although this arrangement is for illustrative purposes only and may vary in different embodiments.
[0020] In at least one embodiment, thermal sensors 215A-D are located at various locations on the chip 205 to measure the temperature of various areas of the chip, such as the processor cores 210, other circuitry, or other materials. In at least one embodiment, various hardware and architectural constraints or design limitations do not allow the temperature sensors to be ideally positioned with respect to the processor cores 210 or other areas of the chip, so the temperature sensors 215AD may not accurately or uniformly detect hot spots on the chip 205. In at least one embodiment, sustained processor core activity near the maximum clock (e.g., Fmax) may result in thermal stress on that particular core and possibly also thermal stress on nearby cores, components, or materials.In at least one embodiment, the possibility of thermal overload or hotspotting is increased when two or more processor cores are in close proximity to each other and operating in a state of sustained activity. In at least one embodiment, other proximity relationships, such as processor cores near areas that exhibit poor thermal transfer, may also lead to hotspots.
[0021] In at least one embodiment, system 200 includes firmware 220, configuration file 225, ACPI 230, and policy 235. In at least one embodiment, firmware 220, configuration file 225, ACPI 230, and policy 235 include one or more chips 205.
[0022] In at least one embodiment, firmware 220 includes software instructions to be executed by a processor, such as one or more of processor cores 210. In at least one embodiment, the instructions are stored in one or more read-only memories, such as one or more read-only memories of system 200 or chip 205.
[0023] In at least one embodiment, instructions are executed in firmware 220 to cause system 200 to validate that a current version of the most recent worst-case thermal scenarios has been read from a configuration file 225; to validate that ACPI tables 230 have been updated; and to confirm that thermal sensors are being used to monitor thermal conditions on chip 205. In at least one embodiment, these instructions are executed to update policy 235. In at least one embodiment, policy 235 includes software and / or data for implementing a temperature policy on chip 205. In at least one embodiment, the instructions update policy 235 with any newly identified processor core utilization patterns, such as worst-case combinations of processor cores.
[0024] Fig. 3 shows a processor chip with multiple processor cores in accordance with at least one embodiment. In at least one embodiment, system 300 depicts a chip 305 that is similar to chip 205 of Fig. 2. an arrangement of processor cores on a CPU. In at least one embodiment, the processor cores 310 are used in different combinations depending on the workload. In at least one embodiment, software and / or firmware, such as operating system software and / or firmware, determine how the Fig. 2, which processor cores 210 should be used to execute the workload. In at least one embodiment, this workload may be distributed among selected processor cores 210.
[0025] In at least one embodiment, the workload could include an application that includes threads scheduled on a subset of processor cores of a total number of available processors. For example, in at least one embodiment, cores 12 and 22 could be selected for thread execution, or in another case, cores 15 and 16.
[0026] In at least one embodiment, some of the thermal sensors 315A-D will measure the temperature of certain processing cores 210 more accurately than others. In at least one embodiment, the thermal sensors 315A-D will more accurately measure the heat dissipation of the processor cores 210 that are closest to that sensor. For example, in at least one embodiment, a thermal sensor TS-1 more accurately reflects the actual temperatures at processor cores 12, 13, 21, and 22 than other thermal sensors on the chip 305. In at least one embodiment, the thermal state of some processor cores or other circuits, regions, or components of the chip 305 is not accurately measured by any thermal sensor. For example, in at least one embodiment, a processor core such as the one shown in Fig. 3, is not near any of the thermal sensors 315A-D, so its temperature may not be accurately measured.
[0027] In at least one embodiment, system 300 measures the temperature on chip 305 and adjusts one or more clocks for one or more processor cores. In at least one embodiment, reducing the clock speed may reduce the heat generated by a particular processor core, as thermal energy is a byproduct of computational processing. However, in at least one embodiment, reducing the processing speed (clock) of one or more processor cores results in a reduction in computational performance.In at least one embodiment, problems associated with some thermal management issues are avoided; these problems may include a thermal management policy that is too conservative, lowering the processor clocks more than necessary to prevent a thermal over-stress failure, or that is not conservative enough, resulting in a thermally induced failure.
[0028] In at least one embodiment, a thermal management policy is implemented by dynamically collecting sensor readings of temperature during operation and correlating these sensor readings with processor activity. In at least one embodiment, firmware, such as firmware 220 of Fig. 2, these temperature measurements and, when a temperature reaches a certain threshold, throttles processor activity for processor cores that may have contributed to an adverse thermal condition. In at least one embodiment, "throttling" processor activity includes reducing the clock speed of one or more processor cores. In at least one embodiment, the temperature sensors are not ideally distributed on the chip 305, which may mean that the sensors inaccurately measure temperatures in some areas of the chip. In at least one embodiment, a scalar temperature offset is added to the temperatures measured at a sensor location. In at least one embodiment, this value is adjusted depending on the distance from a processor core to estimate an actual temperature at that core.In at least one embodiment, assigning an offset to a sensor reading is intended to more accurately reflect the actual temperature of a processor, which in turn is intended to prevent failures due to thermal overload. In at least one embodiment, using offsets in this way could potentially degrade performance, for example, by leading to an over- or under-aggressive temperature policy, but this consequence is avoided by identifying and / or preventing certain patterns of processor core utilization.
[0029] Fig. 4 shows a multi-core processor chip with a combination of processor cores in use, according to at least one embodiment. In at least one embodiment, system 400 corresponds to system 200 or 300 as shown in Fig. 2 and 3, respectively, and includes processor cores 410 and thermal sensors 415A-D corresponding to the similarly named elements in those figures. In at least one embodiment, as shown in Fig. 4, certain processor cores are selected to execute threads associated with a workload. In at least one embodiment, a thermal sensor TS-1 415A records temperature measurements associated with thermal conditions for processor cores in the local region of TS-1 on die 405. In at least one embodiment, a thermal sensor TS-1 could accurately measure the temperature when these processors create a hotspot. However, in at least one embodiment, this pattern of core utilization of the processor cores may cause hotspots in a portion of die 405 not accurately measured by TS-1, for example, in an area near processor core 33. However, in at least one embodiment, this problem could be avoided by identifying a pattern of core utilization of the processor core that contributes to this condition.In at least one embodiment, the pattern could be identified based on a physical location of the cores on the chip 405. In at least one embodiment, the pattern may be specified in a configuration file, such as the one shown in . Fig. 2. In at least one embodiment, one or more thermal management policies may be applied in response to identifying the pattern. In at least one embodiment, a response to the identification may include adjusting a thermal offset. In at least one embodiment, a response to the identification may include changing how the processor cores 410 are utilized to avoid the pattern.
[0030] Fig. 5 shows a block diagram of a multi-core processor chip with a combination of processor cores in use, according to at least one embodiment. In at least one embodiment, the system 500 conforms to the Fig. 2, Fig. 3 and 4, respectively. In at least one embodiment, the system 500 includes a chip 505 corresponding to the chips 205, 305 and / or 405 as shown in the respective Fig. 2-4, and also includes processor cores 410 corresponding to the processor cores depicted in these figures. In at least one embodiment, a set of processor cores 410 is selected to execute threads associated with a workload, such as processor cores 25, 33, 34, 35, 42, 43, 44, 52, 53, and 54 (in Fig. 5). In at least one embodiment, these processor cores generate thermal conditions that are not accurately detected by any of the sensors TS-1, TS-2, TS-3, or TS-4 and continue to operate at similar levels.
[0031] In at least one embodiment, an offset adjustment technique is used to protect against such thermal errors. In at least one embodiment, this offset compensates for a potentially inaccurate temperature measurement on one or more processors by adding an offset to a recorded temperature. In at least one embodiment, the firmware thus compensates for the temperature recording of the sensors and makes this offset determination based on the number and location of the processor cores in use. In at least one embodiment, this determination is based on one or more specified patterns of core utilization. In at least one embodiment, these patterns are specified in a configuration file. In at least one embodiment, the patterns are determined dynamically.In at least one embodiment, the patterns indicate the proximity of processor cores that may contribute to adverse thermal conditions, such as hot spots that are not accurately measured by a built-in thermal sensor. In at least one embodiment, the proximity patterns include on-chip locations. In at least one embodiment, the proximity patterns include on-chip spacing. In at least one embodiment, the proximity patterns include relative positions, which may include, for example, the relationship between processor cores and other cores, circuits, components, materials, and / or temperature sensors.
[0032] Fig. 6 shows a block diagram of a multi-core processor chip with different combinations of processor cores in use according to at least one embodiment. In at least one embodiment, the system 600 corresponds to one or more of the Fig. 2-5, respectively. In at least one embodiment, the potential usage patterns that may be detected include clusters of processor cores located at a certain distance from heat centers. However, it will be appreciated that these examples are intended to be illustrative rather than limiting.
[0033] Fig. Figure 7 shows a block diagram of a multi-core processor chip with different combinations of processor cores in use according to at least one embodiment. In at least one embodiment, system 700 corresponds to one or more of systems 200, 300, 400, 500, and 600 as shown in Fig. 2-6 are shown accordingly.
[0034] In at least one embodiment, as in Fig. 7, an initial pattern of core utilization might include processor cores 730 whose proximity to each other and / or to other components of chip 705 may create an adverse thermal condition. In at least one embodiment, this pattern is detected. In at least one embodiment, this is done by software and / or firmware of system 700 and / or chip 705 comparing the currently observed processor core utilization to processor core utilization patterns indicated as potentially causing an adverse thermal condition. In at least one embodiment, the software and / or firmware, when executed by a processor, causes the workload to be reallocated to other processors so that utilization corresponding to the pattern is terminated.For example, in at least one embodiment, the software may cause the workload to be assigned to other processor cores 720 at other locations on the chip 705. Similarly, when executed by a processor, the software and / or firmware may proactively schedule workload to processor cores to prevent the occurrence of adverse processor usage patterns.
[0035] In at least one embodiment, instead of reallocating the workload, processor cores associated with the adverse pattern may be clocked so that at least some processors operate at a temperature low enough to avoid an adverse thermal condition. For example, the clocks of processors 34, 43, and 53 in region 730 may be scaled down to be below a threshold to avoid an adverse pattern of core utilization of the processor cores in which all processors in region 730 operate at or above the threshold.
[0036] Fig. 8A shows a thermal map of a processor chip illustrating an example of a worst-case thermal hotspot in accordance with at least one embodiment. In at least one embodiment, a region 810 of a processor chip includes a number of active processors in close proximity to each other that could create a thermal hotspot that, if left unaddressed, could lead to a thermally induced failure. Other locations, such as 820, include active processors that are not arranged in a pattern that would cause a thermal hotspot.
[0037] In at least one embodiment, as in Fig. 8, the processor chip may become progressively less hot as it moves away from this region, as shown by points 805. In at least one embodiment, excessive heat buildup may be identified by a thermal sensor located at 815, but not at points 805. In at least one embodiment, the software and / or firmware controlling the operation of the chip could therefore reschedule the workload or prevent the workload from being assigned to the processor cores in a configuration as seen in region 810. In at least one embodiment, this could be done as shown in Fig. 8B, which shows a heat map of a processor chip illustrating a uniformly loaded core configuration without a thermal hotspot, according to at least one embodiment. In at least one embodiment, this figure shows areas such as 835 background 840 and the absence of thermally hazardous hotspots seen at 810 or 820. In the configuration shown in this embodiment, the processors are not thermally overloaded, and performance is not impacted.
[0038] In at least one embodiment, when an adverse core utilization pattern is identified, a thermal management policy may be applied to adjust thermal offsets. In at least one embodiment, the pseudocode associated with a thermal management policy is expressed as follows: TJmax = Tavg + TJ_HOTSPOT_OFFSET; if (WORST CASE_CORE_CONFIG_THERMAL = 1) then (TJ_HOTSPOT_OFFSET = 15) else (TJ_HOTSPOT_OFFSET = 7); where 15 is a higher thermal offset value and 7 is a lower thermal offset value. It is understood that this example is for illustrative purposes only and should not be construed as a limitation.
[0039] In at least one embodiment, the logic of a thermal management policy distinguishes between processor cores whose temperatures are accurately measured by on-board thermal sensors and processor cores that are not measured by thermal sensors. In at least one embodiment, processors in certain processor core combinations may require less or no thermal offset because a nearby thermal sensor accurately measures their temperatures. In at least one embodiment, in a contrasting scenario, certain processor core combinations are identified as having unfavorable core utilization patterns. In at least one embodiment, a thermal management policy assigns a higher offset value to these processors, and the firmware then uses these offsets to trigger the accelerators.In at least one embodiment, the performance of these processors is reduced, but the processor cores do not overheat and cause failures. In at least one embodiment, the processor cores in such a processor core combination are assigned a lower thermal offset to compensate for the fact that their actual temperature is not accurately measured by a thermal sensor. In at least one embodiment, an adverse scenario may be identified as less adverse than a worst-case scenario, and in such cases, a thermal offset may be used that is larger than that which could be used in non-adverse patterns, but smaller than that which could be used in a worst-case pattern.
[0040] Fig. 9 shows a process flow for implementing a temperature management policy system in accordance with at least one embodiment. In at least one embodiment, the process 900 includes steps or operations for updating a temperature management policy. In at least one embodiment, the process 900 is performed by firmware, such as the firmware described in Fig. 2 or referred to herein in relation to any of the Fig. 1-8. In at least one embodiment, the firmware monitors temperature at 905 using thermal sensors near processor cores.
[0041] In at least one embodiment, CPU temperatures are monitored at 910 and correlated with processor core combinations associated with adverse thermal conditions.
[0042] In at least one embodiment, process 900 includes determining, at 915, whether a core combination associated with an adverse thermal condition is identified. In at least one embodiment, this includes determining whether processor cores operating at or near peak processing capacity, given their proximity to each other and / or other chip components, correspond to a processor core utilization pattern associated with an adverse thermal condition.
[0043] In at least one embodiment, such a combination is identified at 915, whereupon the temperature policy offsets are updated at 925. In at least one embodiment, if no such combinations are identified for identification, then the temperature policy offsets are not updated (920).
[0044] In at least one embodiment, if a processor core combination is identified at 915, the operation of the processor cores in that combination may be adjusted. For example, in at least one embodiment, an ACPI table may be adjusted to cause one or more processors in that combination to be used at a lower capacity or temporarily disabled (930).
[0045] Fig. 10 illustrates a method for updating temperature policy offsets according to at least one embodiment. In at least one embodiment, the process 900 includes steps or operations for updating a temperature management policy. In at least one embodiment, the process 900 is performed by firmware, such as the firmware described in Fig. 2 or referred to herein in relation to any of the Fig. 1-8. In at least one embodiment, the firmware monitors temperature at 905 using thermal sensors near processor cores.
[0046] In at least one embodiment, processor core combinations associated with adverse thermal conditions are identified at 1005. In at least one embodiment, one or more thermal offsets are identified based on these combinations (1010). In at least one embodiment, at 1015, the combinations are used to populate a configuration file, such as configuration file 225, to indicate the combinations.
[0047] In at least one embodiment, the operations described with respect to elements 1005-1015 are performed using one or more test systems or simulations, while the subsequent operations described with respect to elements 1020-1040 are performed by a system, such as those described in the Fig. 2-7 described systems.
[0048] In at least one embodiment, the method 1000 includes reading the combinations at 1020. In at least one embodiment, this is done by software and / or firmware, such as the firmware 220 of the system 200, or similarly named components included in the Fig. 3-7 are shown.
[0049] In at least one embodiment, the process 1000 includes updating an ACPI table to specify the temperature policies at 1025. In at least one embodiment, this is done by software and / or firmware, such as firmware 220 of the system 200, or similarly named components included in Fig. 3-7 are shown.
[0050] In at least one embodiment, the process 1000 includes monitoring the operation of processor cores and the output of thermal sensors, determining whether any of the aforementioned combinations occur, and determining whether thermal limits have been exceeded at 1030. In at least one embodiment, this is done by software and / or firmware, such as firmware 220 of the system 200, or similarly named components included in Fig. 3-7 are shown.
[0051] In at least one embodiment, process 1000 includes performing an assessment to determine whether the conditions related to the operations described in connection with elements 1020, 1025, and 1030 are true. In at least one embodiment, if all of these conditions are evaluated as true, a thermal management policy is updated 1040. In at least one embodiment, the update includes adjusting a thermal offset to account for an identified pattern of core utilization. In at least one embodiment, the update includes disabling or accelerating processor cores to avoid an unfavorable pattern of core utilization. In at least one embodiment, this is performed by software and / or firmware, such as firmware 220 of system 200, or similarly named components included in Fig. 3-7 are shown.
[0052] In at least one embodiment, the policy is not updated if the conditions evaluated in decision block 1035 are evaluated as not true, in 1035.
[0053] Fig. 11 shows a process of a system implementing temperature policy management in accordance with at least one embodiment. In at least one embodiment, process 1100 is performed by a system such as one of the systems described in Fig. 2-7 described systems.
[0054] In at least one embodiment, thermal worst-case combinations are identified in 1105, and a configuration file, such as configuration file 225, is populated with this information. In at least one embodiment, this is performed prior to operation of a system implementing a memory of process 1100. In at least one embodiment, these combinations or other combinations of processor core utilizations are dynamically identified. In at least one embodiment, these operations are performed by software and / or firmware, such as firmware 220 of system 200 or similarly named components implemented in Fig. 3-7 are shown.
[0055] In at least one embodiment, the configuration file is read, loaded into memory, or otherwise processed at 1110. In at least one embodiment, this is done by software and / or firmware, such as the firmware 220 of the system 200 or similarly named components included in the Fig. 3-7 are shown.
[0056] In at least one embodiment, at 1115, the ACPI tables are updated to indicate how one or more processor cores should be used. In at least one embodiment, this is done by software and / or firmware, such as firmware 220 of system 200, or similarly named components included in Fig. 3-7 are shown.
[0057] In at least one embodiment, the thermal sensors are monitored and checked for threshold values at 1120. In at least one embodiment, this is done by software and / or firmware, such as firmware 220 of system 200, or similarly named components included in Fig. 3-7 are shown.
[0058] In at least one embodiment, it is determined at 1125 whether the conditions specified by operations 1110, 1115, and 1120 are met. In at least one embodiment, if so, an updated thermal management policy is determined at 1130 and set at 1135. In at least one embodiment, these operations are performed by software and / or firmware, such as the firmware 220 of the system 200 or similarly named components included in Fig. 3-7. In at least one embodiment, updating the temperature policy includes setting a thermal offset for applications that have a worst-case scenario, such as a worst-case pattern of core utilization. In at least one embodiment, updating the temperature policy includes setting an adjusted thermal offset for applications or workloads that have a pattern of core utilization associated with an adverse thermal condition. In at least one embodiment, updating the temperature policy includes setting a reduced thermal offset for applications or workloads that have a pattern of core utilization not associated with an adverse thermal condition.In at least one embodiment, the policies are not updated for applications or workloads that are not associated with worst-case or other adverse thermal conditions. In at least one embodiment, applying the temperature policy or policies includes adjusting the thermal offsets associated with the processors used by the applications or workloads.
[0059] Fig. 12 shows a processor module in accordance with at least one embodiment. In at least one embodiment, 1200 depicts the processor 1205 and the modules in accordance with at least one embodiment. In at least one embodiment, a processor 1205 performs one or more processes as described herein to scale one or more clocks of one or more cores based at least in part on a proximity of the one or more cores to each other, as in Fig. 1, and / or implementation variants described in the associated description. In at least one embodiment, the processor 1205 performs the active learning process as described in connection with Fig. 1. In at least one embodiment, the processor 1205 performs one or more processes as described in connection with Fig. 1 to Fig. 11 are described.
[0060] In at least one embodiment, the processor 1205 comprises one or more processors as described in connection with Fig. 14 to 50B. In at least one embodiment, the processor 1205 is any suitable processing unit and / or combination of processing units, such as one or more CPUs, GPUs, GPGPUs, PPUs, and / or variations thereof. In at least one embodiment, the processor 1205 includes a CPU temperature monitoring and policy update module 1210 for monitoring the CPU temperature and updating the thermal management policy offsets, a module 1215 for identifying worst-case scenarios and updating the policy, and a hotspot offset management module 1220 for generating one or more new hotspot offsets and / or otherwise managing hotspot offsets in a thermal management policy.In at least one embodiment, the CPU temperature monitoring and policy update module 1210 for monitoring the CPU temperature and updating the thermal management policy offsets, the worst-case scenario identification and policy update module 1215 for identifying the worst-case scenarios and updating the thermal management policy, and the hotspot offset management module 1220 for generating and / or managing one or more new hotspot offsets are part of the processor 1205 and / or one or more other processors.In at least one embodiment, the CPU temperature monitoring and policy update module 1210 for monitoring the CPU temperature and updating the thermal management policy offsets, the worst-case scenario identification and thermal management policy update module 1215, and the hotspot offset management module 1220 for generating and / or managing one or more new hotspot offsets are distributed across multiple processors communicating over a bus, a network, by writing to shared memory, and / or any suitable communication method as described herein.
[0061] In at least one embodiment, a portion or all of the functions, methods, and functionality in one or more of these modules may be refactored from one of these modules into another module.
[0062] In at least one embodiment, a module, as used in an implementation described herein, unless the context indicates otherwise or unless expressly stated to the contrary, refers to any combination of software logic, firmware logic, hardware logic, and / or circuitry configured to provide the functionality described herein. In at least one embodiment, software may be embodied as a software package, code, and / or instruction set or instructions, and "hardware" as used in any embodiment described herein may include, for example, individually or in any combination, hard-wired circuitry, programmable circuitry, state machine circuitry, fixed-function circuitry, execution units, and / or firmware storing instructions executed by programmable circuitry.In at least one embodiment, modules may be implemented collectively or individually as circuits that are part of a larger system, such as an integrated circuit (IC), a system-on-chip (SoC), and so forth. In at least one embodiment, a module performs one or more processes in conjunction with any suitable processing unit and / or combination of processing units, such as one or more CPUs, GPUs, GPGPUs, PPUs, and / or variations thereof.
[0063] In at least one embodiment, a CPU temperature monitoring and policy update module 1210 is a module that performs processing activities associated with scaling one or more clocks of one or more cores based at least in part on the proximity of the one or more cores to one another, as used in conjunction with one or more of Fig. 1 to Fig. 11. In at least one embodiment, the CPU temperature monitoring and policy update module 1210 performs one or more processes as described herein by at least comprising or otherwise monitoring temperature data of a CPU or processor cores of the CPU and updating elements of the temperature management policy with respect to the temperature data (e.g., by the processor 1205). In at least one embodiment, the CPU temperature monitoring and policy update module 1210 receives or is otherwise provided with one or more neural networks (e.g., by one or more systems as used in conjunction with one or more of the Fig. 1 to Fig. 11). In at least one embodiment, the CPU temperature monitoring and policy update module 1210 performs processing activities related to protecting the confidentiality of data by encrypting the data during its processing. In at least one embodiment, a CPU temperature monitoring and policy update module 1210 performs processing activities related to protecting the confidentiality of data by encrypting the data during its processing.
[0064] In at least one embodiment, a worst-case scenario identification and policy update module (1215) is a module that performs processing activities related to managing one or more virtual machines (tenants) created by or under the control of a hypervisor. In at least one embodiment, this module may also perform related activities, such as performing variable assignments using the inputs, serializing and / or storing values in a database or other storage location, or retrieving these values from memory, or deserializing the data by one or more processes, as may be associated with one or more of the Fig. 1 to Fig. 11. In at least one embodiment, a worst-case scenario identification and policy update module 1215 performs one or more processes as described herein by at least including or otherwise encoding instructions that effect the performance of the one or more processes or can otherwise be used to perform the one or more processes (e.g., by processor 1205). In at least one embodiment, a worst-case scenario identification and policy update module 1215 receives one or more neural networks (e.g., by one or more systems as used in conjunction with one or more of the Fig. 1 to Fig. 11) or is otherwise provided therewith. In at least one embodiment, a worst-case scenario identification and policy update module 1215 performs processing activities related to identifying the processor core worst-case scenarios and updating the elements of the temperature policy with respect to the processor core worst-case scenario by one or more processes, as described in connection with Fig. 1 to Fig. 11. In at least one embodiment, a worst-case scenario identification and policy update module 1215 performs processing operations associated with the worst-case scenarios described in connection with one or more of the Fig. 1 to Fig. 11 described processes.
[0065] In at least one embodiment, a hotspot offset management module 1220 is a module that performs management and processing activities related to identifying processor cores that are deemed suitable to remain operational, but whose clocks are to be throttled by assigning an offset to said processor core(s). In at least one embodiment, this module may also perform related activities, such as performing variable assignments using the inputs, serializing and / or storing values in a database or other storage location, or retrieving these values from memory, or deserializing this data by one or more processes, as may be associated with one or more of the Fig. 1 to Fig. 11. In at least one embodiment, a hotspot offset management module 1220 performs one or more processes as described herein by at least managing and processing activities for identifying processor cores that are deemed suitable to remain operational, but whose clocks are to be throttled by assigning an offset to the processor core(s) (e.g., by processor 1205). In at least one embodiment, a hotspot offset management module 1220 receives or is otherwise provided with one or more neural networks (e.g., by one or more systems as described in connection with one or more of the Fig. 1 to Fig. 11). In at least one embodiment, the hotspot offset management module 1220 performs processing activities related to the encryption and / or decryption of data transmitted over an interconnection to or from a parallel processor by one or more processes, as described in connection with one or more of the Fig. 1 to Fig. 11. In at least one embodiment, a hotspot offset management module 1220 performs processing activities related at least to managing and processing activities for identifying processor cores that are deemed suitable to remain operational, but whose clocks are to be throttled by assigning an offset to the processor cores, as described in connection with one or more of the Fig. 1 to Fig. 11. In at least one embodiment, the hotspot offset management module 1220 may delete a hotspot offset from the thermal management policy.
[0066] Fig. 13 depicts an API for scaling one or more clocks of one or more cores based at least in part on the proximity of the one or more cores to each other, in accordance with at least one embodiment. In at least one embodiment, 1300 depicts a block diagram illustrating a driver and / or runtime including one or more libraries to provide one or more application programming interfaces (APIs). In at least one embodiment, a software program 1302 is a software module. In at least one embodiment, a software program 1302 comprises one or more software modules. In at least one embodiment, one or more software modules are as in the Fig. 1-12 are not exhaustive. In at least one embodiment, one or more APIs 1310 are sets of software instructions that, when executed, cause one or more processors to perform one or more computational operations. In at least one embodiment, one or more APIs 1310 are distributed or otherwise provided as part of one or more libraries 1306, runtimes 1304, drivers 1304, and / or other grouping of software and / or executable code, further described herein. In at least one embodiment, one or more APIs 1310 perform one or more computational operations in response to a call by software programs 1302.In at least one embodiment, a software program 1302 is a collection of software code, commands, instructions, or other text sequences to instruct a computing device to perform one or more computational operations and / or to invoke one or more other sets of instructions, such as APs 1310 or API functions 1312, for execution. In at least one embodiment, the functionality provided by one or more APs 1310 includes software functions 1312, such as those that can be used to accelerate one or more portions of software programs 1302 using one or more parallel processing units (PPUs), such as graphics processing units (GPUs). In at least one embodiment, a software program is a compiler.
[0067] In at least one embodiment, the APIs 1310 are hardware interfaces to one or more circuits to perform one or more computational operations. In at least one embodiment, one or more software APIs 1310 described herein are implemented as one or more circuits to perform one or more operations associated with Fig. 1-12. In at least one embodiment, one or more software programs 1302 include instructions that, when executed, cause one or more hardware devices and / or circuits to perform one or more techniques associated with Fig. 1-12 are further described.
[0068] In at least one embodiment, software programs 1302, such as user-implemented software programs, utilize one or more application programming interfaces (APIs) 1310 to perform various computational operations, such as memory allocation, matrix multiplication, arithmetic operations, or any computational operation performed by parallel processing units (PPUs), such as graphics processing units (GPUs), as further described herein. In at least one embodiment, one or more APIs 1310 provide a set of callable functions 1312, referred to herein as APIs, API functions, and / or functions, that individually perform one or more computational operations, such as computational operations related to parallel computing.In at least one embodiment, one or more APls 1310 provide functions 1312 to cause 1316 to scale one or more clocks of one or more cores based at least in part on the proximity of the one or more cores to each other, and / or otherwise perform operations described herein. In at least one embodiment, one or more APls 1310 provide functions 1312 to cause 1316 a neural network to perform one or more operations, for example, by returning a called function to a processor, where the processor invokes the neural network.
[0069] In at least one embodiment, one or more software programs 1302 interact or communicate with one or more APs 1310 to perform one or more computational operations using one or more PPUs, such as GPUs. In at least one embodiment, one or more computational operations using one or more PPUs include at least one or more groups of computational operations that are accelerated by being executed at least in part by the one or more PPUs. In at least one embodiment, one or more software programs 1302 interact with one or more APs 1310 to enable parallel computing using a remote or local interface.
[0070] In at least one embodiment, an interface consists of software instructions that, when executed, provide access to one or more functions 1312 provided by one or more APs 1310. In at least one embodiment, a software program 1302 uses a local interface when a software developer compiles one or more software programs 1302 in conjunction with one or more libraries 1306 that include or otherwise provide access to one or more APs 1310. In at least one embodiment, one or more software programs 1302 are statically compiled in conjunction with precompiled libraries 1306 or uncompiled source code that includes instructions for executing one or more APs 1310.In at least one embodiment, one or more software programs 1302 are dynamically compiled, and the one or more software programs use a linker to link to one or more precompiled libraries 1306 that include one or more APIs 1310.
[0071] In at least one embodiment, a software program 1302 uses a remote interface when a software developer executes a software program that uses or otherwise communicates with a library 1306 comprising one or more APs 1310 over a network or other remote communication medium. In at least one embodiment, one or more libraries 1306 comprising one or more APs 1310 are executed by a remote computing service, such as a computing resource service provider. In another embodiment, one or more libraries 1306 comprising one or more APs 1310 are executed by another computer host that provides the one or more APs 1310 to one or more software programs 1302.
[0072] In at least one embodiment, a processor executing or using one or more software programs 1302 invokes, uses, executes, or otherwise implements one or more APls 1310 to allocate and otherwise manage memory to be used by the software programs 1302. In at least one embodiment, one or more software programs 1302 use one or more APls 1310 to allocate and otherwise manage memory used by one or more portions of the software programs 1302 to be accelerated using one or more PPUs, such as GPUs or another accelerator or processor described further herein. These software programs 1302 to scale one or more clocks of one or more cores based at least in part on a proximity of the one or more cores to each other.
[0073] In at least one embodiment, an API 1310 is an API for facilitating parallel computing. In at least one embodiment, an API 1310 is any other API described further herein. In at least one embodiment, an API 1310 is provided by a driver and / or runtime 1304. In at least one embodiment, an API 1310 is provided by a CUDA user-mode driver.
[0074] In at least one embodiment, an API 1310 is provided by a CUDA runtime. In at least one embodiment, a driver 1304 consists of data values and software instructions that, when executed, perform or otherwise facilitate the operation of one or more functions 1312 of an API 1310 during the loading and execution of one or more portions of a software program 1302. In at least one embodiment, a runtime 1304 consists of data values and software instructions that, when executed, perform or otherwise facilitate the operation of one or more functions 1312 of an API 1310 during the execution of a software program 1302.In at least one embodiment, one or more software programs 1302 utilize one or more APIs 1310 implemented or otherwise provided by a driver and / or runtime 1304 to perform combined arithmetic operations by the one or more software programs 1302 during execution by one or more PPUs, such as GPUs.
[0075] In at least one embodiment, one or more software programs 1302 utilize one or more APls 1310 provided by a driver and / or runtime 1304 to perform combined arithmetic operations of one or more PPUs, such as GPUs. In at least one embodiment, one or more APls 1310 provide combined arithmetic operations via a driver and / or runtime 1304, as described above. In at least one embodiment, one or more software programs 1302 utilize one or more APls 1310 provided by a driver and / or runtime 1304 to allocate or otherwise reserve one or more memory blocks 1314 for one or more PPUs, such as GPUs.In at least one embodiment, one or more software programs 1302 utilize one or more APls 1310 provided by a driver and / or runtime 1304 to allocate or otherwise reserve memory blocks. In at least one embodiment, one or more APls 1310 perform combined arithmetic operations, as described below in connection with FIGS. Fig. 1-12 described.
[0076] In order to improve the usability of software programs 1302 and / or the optimization of one or more portions of the software programs 1302 that are to be accelerated by one or more PPUs, such as GPUs, in one embodiment, one or more APIs 1310 provide one or more API functions 1312 to execute a scheduling system that can be or is used by one or more computing devices, as described above and in connection with one or more of the Fig. 1-12. In at least one embodiment, a block diagram 1300 depicts a processor including one or more circuits for executing one or more software programs to combine two or more application programming interfaces (APIs) into a single API. In at least one embodiment, a block diagram 1300 depicts a system including one or more processors executing one or more software programs to combine two or more application programming interfaces (APIs) into a single API. In at least one embodiment, an API is used to cause 1316 to scale one or more clocks of one or more cores based at least in part on a proximity of the one or more cores to one another.
[0077] In at least one embodiment, a system for scaling one or more clocks of one or more cores based at least in part on the proximity of the one or more cores to each other is used in servers and data centers, as described in Fig. 14A - Fig. 18B shown.
[0078] In at least one embodiment, a system for scaling one or more clocks of one or more cores based at least in part on a proximity of the one or more cores to each other is used in and from other cloud computers and servers, as in Fig. 19 - Fig. 21 is shown.
[0079] In at least one embodiment, a system for scaling one or more clocks of one or more cores based at least in part on a proximity of the one or more cores to each other is used as part of a supercomputer, as in Fig. 22 - Fig. 25 shown.
[0080] In at least one embodiment, a system for scaling one or more clocks of one or more cores based at least in part on a proximity of the one or more cores to each other is integrated with another artificial intelligence, as in Fig. 26A and / or Fig. 27 shown.
[0081] In at least one embodiment, a system for scaling one or more clocks of one or more cores based at least in part on a proximity of the one or more cores to each other is used by 5G networks as described in Fig. 28 - Fig. 33 are shown.
[0082] In at least one embodiment, a system for scaling one or more clocks of one or more cores based at least in part on a proximity of the one or more cores to each other is used by computer systems as described in Fig. 34 - Fig. 38 are shown.
[0083] In at least one embodiment, a system for scaling one or more clocks of one or more cores based at least in part on a proximity of the one or more cores to each other is used by processing systems as described in Fig. 39 - Fig. 44 are shown.
[0084] In at least one embodiment, a system for scaling one or more clocks of one or more cores based at least in part on a proximity of the one or more cores to each other is used in general computing, as described in Fig. 45 - Fig. 50 shown. TECHNICAL SOLUTION TO A TECHNICAL PROBLEM
[0085] In at least one embodiment, a technical solution to a technical problem is presented. In at least one embodiment, a technical problem that is solved is a problem of thermal overheating in multi-core computer processors. In at least one embodiment, thermal overheating in computer processors is a problem because it can lead to thermal failure. Computer designers have overcompensated for this problem by assigning the same offset value to all processor cores whose temperature could not be directly measured, which had the effect of lowering the temperature of the processors in use. This practice also has the side effect of excessive "over-throttling," unnecessarily reducing computer performance.
[0086] In at least one embodiment, a multiple offset technique provides a way to avoid unnecessary reduction of overall processor activity and over-throttling. In at least one embodiment, excessive over-throttling is overcome by tracking which processor core combinations are worst-case heat generators and assigning a higher offset value to those worst-case combinations, while assigning a lower offset or no offset value to other processors that are not in the worst-case combinations.
[0087] In at least one embodiment, this technological improvement is that the temperature control within the processor and processor cores can be regulated at a more precise, core-specific level, resulting in an increase in performance while avoiding thermally induced failures.
[0088] In the following description, numerous specific details are set forth in order to provide a more thorough understanding of at least one embodiment. However, it will be apparent to one skilled in the art that the inventive concepts may be practiced without one or more of these specific details. Servers and data centers
[0089] The following figures illustrate, without limitation, exemplary network servers and data center-based systems that may be used to implement at least one embodiment.
[0090] Fig. 14 shows a distributed system 1400 in accordance with at least one embodiment. In at least one embodiment, the distributed system 1400 includes one or more client computer systems 1402, 1404, 1406, and 1408 configured to execute and operate a client application such as a web browser, a proprietary client, and / or variations thereof over one or more networks 1410. In at least one embodiment, the server 1412 may be communicatively coupled to remote client devices 1402, 1404, 1406, and 1408 over the network 1410.
[0091] In at least one embodiment, server 1412 may be adapted to execute one or more services or software applications, such as services and applications that can manage session activity for single sign-on (SSO) access across multiple data centers. In at least one embodiment, server 1412 may also provide other services or software applications that may include non-virtual and virtual environments. In at least one embodiment, these services may be offered as web-based or cloud services or under a software-as-a-service (SaaS) model to users of computing devices 1402, 1404, 1406, and / or 1408. In at least one embodiment, users operating computing devices 1402, 1404, 1406, and / or 1408 may, in turn, use one or more client applications to interact with server 1412 and utilize the services provided by these components.
[0092] In at least one embodiment, the software components 1418, 1420, and 1422 of the system 1400 are implemented on the server 1412. In at least one embodiment, one or more components of the system 1400 and / or services provided by these components may also be implemented by one or more of the client devices 1402, 1404, 1406, and / or 1408. In at least one embodiment, users operating client computing devices may then use one or more client applications to utilize services provided by these components. In at least one embodiment, these components may be implemented in hardware, firmware, software, or combinations thereof. It should be noted that various different system configurations are possible that may differ from the distributed system 1400. The Fig. The embodiment shown in Figure 14 is therefore an example of a distributed system for implementing an execution system and is not intended to be limiting.
[0093] In at least one embodiment, client computing systems 1402, 1404, 1406, and / or 1408 may include various types of computing systems. In at least one embodiment, a client computing system device may include portable computing devices (e.g., an iPhone®, a cellular phone, an iPad®, a computer tablet, a personal digital assistant (PDA)) or wearable devices (e.g., a Google Glass® head-mounted display) running software such as Microsoft Windows Mobile® and / or a variety of mobile operating systems such as iOS, Windows Phone, Android, BlackBerry 10, Palm OS, and / or variations thereof. In at least one embodiment, the devices may support various applications, such as various internet-related applications, email, and SMS applications, and may use various other communication protocols.In at least one embodiment, client computer systems may also include general-purpose personal computers, such as personal computers and / or laptops running various versions of the Microsoft Windows®, Apple Macintosh®, and / or Linux operating systems. In at least one embodiment, client computer systems may be workstation computers running a variety of commercially available UNIX® or UNIX-like operating systems, including, without limitation, a variety of GNU / Linux operating systems, such as Google Chrome OS. In at least one embodiment, client computer systems may also include electronic devices, such as a thin client computer, an internet-enabled gaming system (e.g., a Microsoft Xbox gaming console with or without a Kinect® gesture input device), and / or a personal messaging device capable of communicating over one or more networks 1410.Although the distributed system 1400 in . Fig. 14 with four client computer systems, any number of client computer systems may be supported. Other devices, such as devices with sensors, etc., may interact with server 1412.
[0094] In at least one embodiment, the network(s) 1410 in the distributed system 1400 may be any type of network capable of supporting data communication using a variety of available protocols, including, without limitation, TCP / IP (Transmission Control Protocol / Internet Protocol), SNA (Systems Network Architecture), IPX (Internet Packet Exchange), AppleTalk, and / or variations thereof. In at least one embodiment, the network(s) 1410 may be a local area network (LAN), Ethernet-based networks, Token Ring, a wide area network, the Internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a public switched telephone network (PSTN), an infrared network, a wireless network (e.g., a network compliant with one of the Institute of Electrical and Electronics (IEEE) 802.11, Bluetooth® and / or another wireless protocol), and / or any combination of these and / or other networks.
[0095] In at least one embodiment, server 1412 may consist of one or more general-purpose computers, specialized server computers (including, by way of example, PC (Personal Computer) servers, UNIX® servers, mid-range servers, mainframes, rack-mounted servers, etc.), server farms, server clusters, or any other suitable arrangement and / or combination. In at least one embodiment, server 1412 may include one or more virtual machines running virtual operating systems or other computer architectures that incorporate virtualization. In at least one embodiment, one or more flexible pools of logical storage devices may be virtualized to maintain virtual storage devices for a server. In at least one embodiment, virtual networks may be controlled by server 1412 using software-defined networking.In at least one embodiment, server 1412 may be adapted to execute one or more services or software applications.
[0096] In at least one embodiment, server 1412 may run any operating system or commercially available server operating system. In at least one embodiment, server 1412 may also run a variety of additional server applications and / or mid-tier applications, including Hypertext Transport Protocol (HTTP) servers, File Transfer Protocol (FTP) servers, Common Gateway Interface (CGI) servers, JAVA® servers, database servers, and / or variations thereof. In at least one embodiment, example database servers include, without limitation, those commercially available from Oracle, Microsoft, Sybase, IBM (International Business Machines), and / or variations thereof.
[0097] In at least one embodiment, server 1412 may include one or more applications to analyze and consolidate data feeds and / or event updates received from users of client computing devices 1402, 1404, 1406, and 1408. In at least one embodiment, data feeds and / or event updates may include, but are not limited to, Twitter® feeds, Facebook® updates, or real-time updates received from one or more third-party information sources, as well as continuous data streams that may include real-time events associated with sensor data applications, financial tickers, network performance measurement tools (e.g., network monitoring and traffic management applications), clickstream analysis tools, car traffic monitoring, and / or variations thereof.In at least one embodiment, server 1412 may also include one or more applications for displaying data feeds and / or real-time events via one or more display devices of client computing devices 1402, 1404, 1406, and 1408.
[0098] In at least one embodiment, distributed system 1400 may also include one or more databases 1414 and 1416. In at least one embodiment, databases may provide a mechanism for storing information such as information about interactions between users, information about usage patterns, information about customization rules, and other information. In at least one embodiment, databases 1414 and 1416 may be located in different locations. In at least one embodiment, one or more of databases 1414 and 1416 may be located on a non-transitory storage medium located on (and / or resident in) server 1412. In at least one embodiment, databases 1414 and 1416 may be remote from server 1412 and communicate with server 1412 via a network-based or dedicated connection.In at least one embodiment, databases 1414 and 1416 may be located on a storage area network (SAN). In at least one embodiment, all files required to perform the functions assigned to server 1412 may be stored locally on server 1412 and / or at a remote location, as needed. In at least one embodiment, databases 1414 and 1416 may comprise relational databases, such as databases capable of storing, updating, and retrieving data in response to SQL-formatted commands.
[0099] In at least one embodiment, at least one in Fig. 14 is used to implement techniques and / or functions associated with Fig. 1-13. In at least one embodiment, at least one component of Fig. 14 to effect scaling of one or more clocks of one or more cores based at least in part on the proximity of the one or more cores to each other. In at least one embodiment, at least one component of Fig. 14 at least one aspect relating to the data center 120, the controller 105, at least two or more of processor 110(a) to processor 110(n) and one or more thermal sensors 115(a) to thermal sensor 115(n) of Fig. 1, Firmware 220, Configuration file 225, ACPI 230, Policy 235 one or more processors 1-n of Fig. 2, and / or comprising one or more of: Method 900 of Fig. 9, Procedure 1000 of Fig. 10 and Procedure 1100 of Fig. 11.
[0100] Fig. 15 illustrates an exemplary data center 1500 in accordance with at least one embodiment. In at least one embodiment, the data center 1500 includes, without limitation, a data center infrastructure layer 1510, a framework layer 1520, a software layer 1530, and an application layer 1540.
[0101] In at least one embodiment, as in Fig. 15, the data center infrastructure layer 1510 may include a resource orchestrator 1512, clustered compute resources 1514, and node compute resources ("Node CRs") 1516(1)-1516(N), where "N" represents any integer positive number. In at least one embodiment, the Node CRs 1516(1)-1516(N) may include any number of central processing units ("CPUs") or other processors (including accelerators, field-programmable gate arrays ("FPGAs"), graphics processors, etc.), memory devices (e.g., dynamic read-only memory), storage devices (e.g., solid-state or hard disk drives), network input / output devices ("NW I / O"), network switches, virtual machines ("VMs"), power modules, and cooling modules, etc. In at least one embodiment, one or more Node CRs among the Node CRs 1516(1)-1516(N) may be a server that has one or more of the computing resources listed above.
[0102] In at least one embodiment, the grouped computing resources 1514 may include separate groupings of node CRs housed in one or more racks (not shown), or multiple racks housed in data centers in different geographic locations (also not shown). Separate groupings of node CRs within the grouped computing resources 1514 may include grouped computing, networking, memory, or storage resources that may be configured or allocated to support one or more workloads. In at least one embodiment, multiple node CRs, including CPUs or processors, may be grouped in one or more racks to provide computing resources to support one or more workloads.In at least one embodiment, one or more racks may also include any number of power modules, cooling modules, and network switches in any combination.
[0103] In at least one embodiment, resource orchestrator 1512 may configure or otherwise control one or more node CRs 1516(1)-1516(N) and / or grouped computing resources 1514. In at least one embodiment, resource orchestrator 1512 may include a software design infrastructure ("SDI") management entity for data center 1500. In at least one embodiment, resource orchestrator 1512 may include hardware, software, or a combination thereof.
[0104] In at least one embodiment, as in Fig. 15, the framework layer 1520 includes, without limitation, a job scheduler 1532, a configuration manager 1534, a resource manager 1536, and a distributed file system 1538. In at least one embodiment, the framework layer 1520 may include a framework for supporting the software 1552 of the software layer 1530 and / or one or more applications 1542 of the application layer 1540. In at least one embodiment, the software 1552 or the application(s) 1542 may include web-based service software or applications, such as those provided by Amazon Web Services, Google Cloud, and Microsoft Azure.In at least one embodiment, the framework layer 1520 may be some type of free and open-source software web application framework such as, but not limited to, Apache Spark™ (hereinafter "Spark"), which may utilize a distributed file system 1538 for processing large amounts of data (e.g., "Big Data"). In at least one embodiment, the job scheduler 1532 may include a Spark driver to facilitate the scheduling of workloads supported by different layers of the data center 1500. In at least one embodiment, the configuration manager 1534 may be capable of configuring different layers such as the software layer 1530 and the framework layer 1520, which include Spark and the distributed file system 1538 to support the processing of big data.In at least one embodiment, resource manager 1536 may be capable of managing clustered or grouped computing resources allocated to support distributed file system 1538 and job scheduler 1532. In at least one embodiment, the clustered or grouped computing resources may include clustered computing resource 1514 in data center infrastructure layer 1510. In at least one embodiment, resource manager 1536 may be coordinated with resource orchestrator 1512 to manage these allocated or assigned computing resources.
[0105] In at least one embodiment, the software 1552 included in software layer 1530 may include software used by at least portions of node CRs 1516(1)-1516(N), clustered computing resources 1514, and / or distributed file system 1538 of framework layer 1520. One or more types of software may include, but are not limited to, Internet website search software, email virus scanning software, database software, and streaming video content software.
[0106] In at least one embodiment, the application(s) 1542 included in the application layer 1540 may include one or more types of applications used by at least portions of the node CRs 1516(1)-1516(N), clustered computing resources 1514, and / or distributed file systems 1538 of the framework layer 1520. At least one or more types of applications may include, without limitation, CUDA applications, 5G networking applications, artificial intelligence applications, data center applications, and / or variations thereof.
[0107] In at least one embodiment, any configuration manager 1534, resource manager 1536, and resource orchestrator 1512 may implement any number and type of self-modifying actions based on any amount and type of data collected in any technically feasible manner. In at least one embodiment, self-modifying actions may relieve an operator of a data center 1500 from potentially making poor configuration decisions and potentially avoiding underutilized and / or poorly performing sections of a data center.
[0108] In at least one embodiment, at least one in Fig. 15 is used to implement techniques and / or functions associated with Fig. 1-13. In at least one embodiment, at least one component of Fig. 15 to effect scaling of one or more clocks of one or more cores based at least in part on the proximity of the one or more cores to each other. In at least one embodiment, at least one component of Fig. 15 at least one aspect relating to the data center 120, the controller 105, at least two or more of processor 110(a) to processor 110(n) and one or more thermal sensors 115(a) to thermal sensor 115(n) of Fig. 1, Firmware 220, Configuration file 225, ACPI 230, Policy 235 one or more processors 1-n of Fig. 2, and / or comprising one or more of: Method 900 of Fig. 9, Procedure 1000 of Fig. 10, and Procedure 1100 of Fig. 11.
[0109] Fig. 16 shows a client-server network 1604 formed by a plurality of network server computers 1602 interconnected according to at least one embodiment. In at least one embodiment, in a system 1600, each network server computer 1602 stores data accessible to other network server computers 1602 and to client computers 1606 and networks 1608 included in a wide area network 1604. In at least one embodiment, the configuration of a client-server network 1604 may change over time as client computers 1606 and one or more networks 1608 connect to and disconnect from a network 1604, and as one or more long-distance server computers 1602 are added to or removed from a network 1604.In at least one embodiment, when a client computer 1606 and a network 1608 are connected to network server computers 1602, the client-server network includes such client computer 1606 and a network 1608. In at least one embodiment, the term computer includes any device or machine capable of accepting data, applying prescribed processes to data, and providing results of processes.
[0110] In at least one embodiment, client-server network 1604 stores information accessible by network server computers 1602, remote networks 1608, and client computers 1606. In at least one embodiment, network server computers 1602 are comprised of host computers, minicomputers, and / or microcomputers, each having one or more processors. In at least one embodiment, server computers 1602 are interconnected by wired and / or wireless transmission media, such as conductive wires, fiber optic cables, and / or microwave transmission media, satellite transmission media, or other conductive, optical, or electromagnetic wave transmission media. In at least one embodiment, client computers 1606 access a network server computer 1602 via a similar wired or wireless transmission medium.In at least one embodiment, a client computer 1606 may be connected to a client-server network 1604 via a modem and a standard telephone network. In at least one embodiment, alternative carrier systems such as cable and satellite communication systems may also be used to connect to the client-server network 1604. In at least one embodiment, other private or time-shared carrier systems may also be used. In at least one embodiment, the network 1604 is a global information network, such as the Internet. In at least one embodiment, the network is a private intranet that uses similar protocols to the Internet, but with additional security measures and limited access controls. In at least one embodiment, the network 1604 is a private or semi-private network that uses proprietary communication protocols.
[0111] In at least one embodiment, client computer 1606 is any end-user computer and may also be a mainframe, minicomputer, or microcomputer having one or more microprocessors. In at least one embodiment, server computer 1602 may temporarily act as a client computer accessing another server computer 1602. In at least one embodiment, remote network 1608 may be a local area network, a network added to a wide area network through an independent Internet service provider (ISP), or another group of computers interconnected via wired or wireless transmission media and having a fixed or changing configuration over time. In at least one embodiment, client computers 1606 may connect to and access network 1604 independently or through remote network 1608.
[0112] In at least one embodiment, at least one in Fig. 16 is used to implement techniques and / or functions associated with Fig. 1-13. In at least one embodiment, at least one component of Fig. 16 to effect scaling of one or more clocks of one or more cores based at least in part on the proximity of the one or more cores to each other. In at least one embodiment, at least one component of Fig. 16 at least one aspect relating to the data center 120, the controller 105, at least two or more of processor 110(a) to processor 110(n) and one or more thermal sensors 115(a) to thermal sensor 115(n) of Fig. 1, Firmware 220, Configuration file 225, ACPI 230, Policy 235 one or more processors 1-n of Fig. 2, and / or comprising one or more of: Method 900 of Fig. 9, Procedure 1000 of Fig. 10 and Procedure 1100 of Fig. 11.
[0113] Fig. 17 shows an example 1700 of a computer network 1708 connecting one or more computers according to at least one embodiment. In at least one embodiment, the network 1708 may be any type of electronically connected group of computers, including, for example, the following networks: Internet, intranet, local area networks (LAN), wide area networks (WAN), or an interconnected combination of these network types. In at least one embodiment, the connectivity within a network 1708 may be a long-distance modem, Ethernet (IEEE 802.3), Token Ring (IEEE 802.5), Fiber Distributed Datalink Interface (FDDI), Asynchronous Transfer Mode (ATM), or another communication protocol.In at least one embodiment, computing devices connected to a network may be a desktop, a server, a portable device, a handheld device, a set-top box, a personal digital assistant (PDA), a terminal, or any other desired type or configuration. In at least one embodiment, network-attached devices may vary greatly depending on their functionality in terms of processing power, internal memory, and other performance aspects. In at least one embodiment, communication within a network and to or from computing devices connected to a network may be either wired or wireless.In at least one embodiment, network 1708 may comprise, at least in part, the worldwide public Internet, which generally connects a plurality of users according to a client-server model in accordance with a Transmission Control Protocol / Internet Protocol (TCP / IP) specification. In at least one embodiment, the client-server network is the predominant model for communication between two computers. In at least one embodiment, a client computer ("client") issues one or more commands to a server computer ("server"). In at least one embodiment, the server executes the client's commands by accessing available network resources and returning information to the client according to the client's commands.In at least one embodiment, the client computer systems and the network resources located on network servers are assigned a network address to identify them in communication between the elements of a network. In at least one embodiment, communication from other systems connected to the network to servers includes a network address of a corresponding server / network resource as part of the communication so that an appropriate destination for a data / request is identified as the recipient. In at least one embodiment, when a network 1708 comprises the global Internet, a network address is an IP address in a TCP / IP format, which can at least partially direct data to an email account, website, or other Internet tool on a server.In at least one embodiment, information and services located on network servers may be available to a web browser of a client computer via a domain name (e.g., www.site.com) associated with an IP address of a network server.
[0114] In at least one embodiment, a plurality of clients 1702, 1704, and 1706 are connected to a network 1708 via respective communication links. In at least one embodiment, each of these clients may access a network 1708 via any form of communication, such as a dial-up modem connection, a cable connection, a digital subscriber line (DSL), a wireless or satellite connection, or another form of communication. In at least one embodiment, each client may communicate via any device compatible with the network 1708, such as a personal computer (PC), a workstation, a special-purpose terminal, a personal data assistant (PDA), or a similar device. In at least one embodiment, the clients 1702, 1704, and 1706 may or may not be located in the same geographic area.
[0115] In at least one embodiment, a plurality of servers 1710, 1712, and 1714 are connected to a network 1708 to serve clients communicating with a network 1708. In at least one embodiment, each server is typically a powerful computer or device that manages network resources and responds to client commands. In at least one embodiment, the servers include computer-readable data storage such as hard disk drives and random access memory that stores program instructions and data. In at least one embodiment, the servers 1710, 1712, and 1714 execute application programs that respond to client commands. In at least one embodiment, the server 1710 may run a web server application to respond to client requests for HTML pages and may also run a mail server application for receiving and routing electronic mail.In at least one embodiment, a server 1710 may also run other application programs, such as an FTP server or a media server for streaming audio / video data to clients. In at least one embodiment, different servers may be assigned different tasks. In at least one embodiment, server 1710 may be a dedicated web server that manages resources associated with websites for various users, while a server 1712 may be responsible for managing electronic mail (e-mail). In at least one embodiment, other servers may be provided for media (audio, video, etc.), the File Transfer Protocol (FTP), or a combination of two or more services typically available or provided over a network.In at least one embodiment, each server may be located at a location identical to or different from the other servers. In at least one embodiment, there may be multiple servers performing mirrored tasks for users, thereby avoiding congestion or minimizing traffic to and from a single server. In at least one embodiment, servers 1710, 1712, 1714 are under the control of a web hosting provider engaged in maintaining and delivering third-party content over a network 1708.
[0116] In at least one embodiment, web hosting providers provide services to two different types of customers. In at least one embodiment, one type, which may be referred to as a browser, requests content from servers 1710, 1712, 1714, such as web pages, email messages, video clips, etc. In at least one embodiment, a second type, which may be referred to as a user, engages a web hosting provider to maintain a network resource, such as a website, and make it available to browsers. In at least one embodiment, users contract with a web hosting provider to provide storage space, processor capacity, and communication bandwidth for their desired network resource in accordance with an amount of server resources a user wishes to utilize.
[0117] In at least one embodiment, application programs that manage network resources hosted by servers must be properly configured so that a web hosting provider can provide services to these two customers. In at least one embodiment, the program configuration process includes defining a set of parameters that, at least in part, control an application program's response to browser requests and also, at least in part, define the server resources available to a particular user.
[0118] In one embodiment, an intranet server 1716 is connected to a network 1708 via a communications link. In at least one embodiment, the intranet server 1716 is in communication with a server manager 1718. In at least one embodiment, the server manager 1718 includes a database with configuration parameters of an application program used in the servers 1710, 1712, 1714. In at least one embodiment, users modify a database 1720 via an intranet 1716, and a server manager 1718 interacts with the servers 1710, 1712, 1714 to modify the application program parameters to conform to the contents of a database. In at least one embodiment, a user logs on to an intranet server 1716 by connecting to an intranet 1716 via computer 1702 and entering authentication information, such as a user name and password.
[0119] In at least one embodiment, when a user wishes to sign up for a new service or modify an existing service, an intranet server 1716 authenticates a user and provides a user with an interactive display / control that allows a user to access configuration parameters for a particular application program. In at least one embodiment, the user is presented with a series of editable text fields describing aspects of the configuration of the user's website or other network resource. In at least one embodiment, if the user wishes to increase the amount of space reserved on a server for their website, a field is provided for the user to specify the desired amount of space. In at least one embodiment, an intranet server 1716 updates a database 1720 in response to receiving this information.In at least one embodiment, server manager 1718 forwards this information to an appropriate server, and a new parameter is used during operation of the application program. In at least one embodiment, an intranet server 1716 is configured to provide users with access to configuration parameters of hosted network resources (e.g., web pages, email, FTP sites, media sites, etc.) for which a user has contracted with a web hosting service provider.
[0120] In at least one embodiment, at least one in Fig. 17 is used to control the component shown or described in connection with Fig. 1-13. In at least one embodiment, at least one component of Fig. 17 to effect scaling of one or more clocks of one or more cores based at least in part on the proximity of the one or more cores to each other. In at least one embodiment, at least one component of Fig. 17 at least one aspect relating to the data center 120, the controller 105, at least two or more of processor 110(a) to processor 110(n) and one or more thermal sensors 115(a) to thermal sensor 115(n) of Fig. 1, Firmware 220, Configuration file 225, ACPI 230, Policy 235 one or more processors 1-n of Fig. 2, and / or comprising one or more of: Method 900 of Fig. 9, Procedure 1000 of Fig. 10 and Procedure 1100 of Fig. 11.
[0121] Fig. 18A shows a networked computer system 1800A in accordance with at least one embodiment. In at least one embodiment, the networked computer system 1800A includes a plurality of nodes or personal computers ("PCs") 1802, 1818, 1820. In at least one embodiment, the PC or node 1802 includes a processor 1814, memory 1816, a video camera 1804, a microphone 1806, a mouse 1808, speakers 1810, and a monitor 1812. In at least one embodiment, the PCs 1802, 1818, 1820 may each run one or more desktop servers of an internal network within a particular enterprise or may be servers of a general network that is not limited to a particular environment. In at least one embodiment, there is one server per PC node of a network, such that each PC node of a network represents a particular network server having a particular network URL address.In at least one embodiment, each server displays, by default, a default web page for the user of that server, which in turn may contain embedding URLs that point to further subpages of that user on that server or to other servers or pages on other servers in a network.
[0122] In at least one embodiment, nodes 1802, 1818, 1820, and other nodes of a network are interconnected via medium 1822. In at least one embodiment, medium 1822 may be a communications channel such as an Integrated Services Digital Network ("ISDN"). In at least one embodiment, various nodes of a networked computer system may be connected via a variety of communications media, including local area networks ("LANs"), plain old telephone lines ("POTS") sometimes referred to as public switched telephone networks ("PSTN"), and / or variations thereof. In at least one embodiment, various nodes of a network may also represent computer system users interconnected via a network such as the Internet.In at least one embodiment, each server in a network (operating at a particular time from a particular node of a network) has a unique address or identification within a network, which may be specified in the form of a URL.
[0123] Therefore, in at least one embodiment, a plurality of Multi-Point Conferencing Units ("MCUs") may be used to transmit data to and from various nodes or "endpoints" of a conferencing system. In at least one embodiment, the nodes and / or MCUs may be interconnected via an ISDN connection or via a local area network ("LAN"), in addition to various other communication media such as Internet-connected nodes. In at least one embodiment, the nodes of a conferencing system may generally be connected directly to a communication medium such as a LAN or via an MCU, and a conferencing system may include other nodes or elements such as routers, servers, and / or variations thereof.
[0124] In at least one embodiment, processor 1814 is a programmable general-purpose processor. In at least one embodiment, the processors of the nodes of networked computer system 1800A may also be dedicated video processors. In at least one embodiment, various peripherals and components of a node, such as those of node 1802, may differ from those of other nodes. In at least one embodiment, node 1818 and node 1820 may be configured identically or differently from node 1802. In at least one embodiment, a node may be implemented on any suitable computer system in addition to PC systems.
[0125] Fig. 18B illustrates a networked computer system 1800B in accordance with at least one embodiment. In at least one embodiment, system 1800B illustrates a network, such as LAN 1824, that may be used to interconnect a plurality of nodes capable of communicating with one another. In at least one embodiment, a plurality of nodes, such as PC nodes 1826, 1828, 1830, are connected to LAN 1824. In at least one embodiment, a node may also be connected to the LAN via a network server or other means. In at least one embodiment, system 1800B includes other types of nodes or elements, for example, routers, servers, and nodes.
[0126] Fig. 18C illustrates a networked computer system 1800C in accordance with at least one embodiment. In at least one embodiment, system 1800C illustrates a WWW system having communication over a backbone communications network such as the Internet 1832, which may be used to interconnect a plurality of nodes of a network. In at least one embodiment, the WWW is a set of protocols that operate on the Internet and enable a graphical interface system to operate thereon to access information over the Internet. In at least one embodiment, a plurality of nodes, such as PCs 1840, 1842, 1844, are connected to the Internet 1832 on the WWW. In at least one embodiment, a node is connected to other nodes of the WWW via a WWW HTTP server such as servers 1834, 1836.In at least one embodiment, the PC 1844 may be a PC that forms a node of the network 1832 and itself runs its server 1836, although the PC 1844 and the server 1836 in . Fig. 18C are shown separately for illustrative purposes.
[0127] In at least one embodiment, the WWW is a distributed application characterized by the WWW protocol HTTP, which is built on top of the Internet Transmission Control Protocol / Internet Protocol ("TCP / IP"). Thus, in at least one embodiment, the WWW can be characterized by a set of protocols (i.e., HTTP) running on top of the Internet as a "backbone."
[0128] In at least one embodiment, a web browser is an application running on a node of a network and, in WWW-compatible network systems, enabling users of a particular server or node to view such information, thus allowing a user to search for graphical and text-based files interconnected by hypertext links embedded in documents or files available from servers on a network that understand HTTP. In at least one embodiment, when a particular web page of a first server connected to a first node is retrieved by a user via another server on a network such as the Internet, a retrieved document may have various hypertext links embedded therein, and a local copy of a page is created locally for a retrieving user.In at least one embodiment, when a user clicks on a hypertext link, the locally stored information relating to a selected hypertext link is typically sufficient to enable a user's computer to open a connection over the Internet to a server specified by a hypertext link.
[0129] In at least one embodiment, more than one user may be connected to each HTTP server, for example, via a LAN such as LAN 1838, as shown with respect to WWW HTTP server 1834. In at least one embodiment, system 1800C may also include other types of nodes or elements. In at least one embodiment, a WWW HTTP server is an application running on a computer, such as a PC. In at least one embodiment, each user may be considered to have their own "server," as shown with respect to PC 1844. In at least one embodiment, a server may be considered to be a server such as WWW HTTP server 1834 that provides access to a network for a LAN or a plurality of nodes or a plurality of LANs.In at least one embodiment, there are a plurality of users, each having a desktop PC or a node of a network, with each desktop PC potentially hosting a server for one of its users. In at least one embodiment, each server is associated with a particular network address or URL that, when accessed, provides a default web page for that user. In at least one embodiment, a web page may contain further links (embedded URLs) that point to further subpages of that user on that server or to other servers on a network or to pages on other servers on a network. Cloud computing and services
[0130] The following figures illustrate, without limitation, exemplary cloud-based systems that may be used to implement at least one embodiment.
[0131] In at least one embodiment, cloud computing is a type of computing in which dynamically scalable and often virtualized resources are delivered as a service over the internet. In at least one embodiment, users are not required to have knowledge of or control over the technological infrastructure supporting them, which may be referred to as "in the cloud." In at least one embodiment, cloud computing encompasses infrastructure as a service, platform as a service, software as a service, and other variations that share a common theme of reliance on the internet to satisfy users' computing needs.In at least one embodiment, a typical cloud deployment, such as in a private cloud (e.g., an enterprise network) or a data center (DC) in a public cloud (e.g., on the Internet), may consist of thousands of servers (or alternatively, VMs), hundreds of Ethernet, Fibre Channel, or Fibre Channel over Ethernet (FCoE) ports, switching and storage infrastructure, etc. In at least one embodiment, the cloud may also consist of infrastructure for network services such as IPsec VPN hubs, firewalls, load balancers, wide area network (WAN) optimizers, etc. In at least one embodiment, remote participants can securely access cloud applications and services by connecting through a VPN tunnel, such as an IPsec VPN tunnel.
[0132] In at least one embodiment, cloud computing is a model for enabling convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, servers, storage, applications, and services) that can be rapidly provisioned and released with minimal management effort or interaction with the service provider.
[0133] In at least one embodiment, cloud computing is characterized by on-demand self-service, where a consumer can unilaterally and automatically provision computing resources such as server time and network storage as needed, without requiring human interaction with the service provider. In at least one embodiment, cloud computing is characterized by broad network access, where functions are available over a network and accessed via standard mechanisms that facilitate use by heterogeneous thin- or thick-client platforms (e.g., mobile phones, laptops, and PDAs).In at least one embodiment, cloud computing is characterized by resource pooling, in which a provider's computing resources are pooled to serve multiple consumers using a multi-tenant model, with different physical and virtual resources dynamically allocated and reassigned based on consumer demand. In at least one embodiment, there is some location independence, as a customer generally has no control or knowledge over the exact location of the provided resources, but may be able to specify the location at a higher level of abstraction (e.g., country, state, or data center). In at least one embodiment, the resources include, for example, storage, processing, memory, network bandwidth, and virtual machines.In at least one embodiment, cloud computing is characterized by rapid elasticity, where capabilities can be quickly and elastically, in some cases automatically, provisioned to scale quickly and released quickly to scale quickly. In at least one embodiment, the capabilities available for provisioning often appear unlimited to the consumer and can be purchased in any quantity and at any time. In at least one embodiment, cloud computing is characterized by a metered service, where cloud systems automatically control and optimize resource usage by leveraging a metering function at a particular level of abstraction appropriate to the nature of the service (e.g., storage, processing, bandwidth, and active user accounts).In at least one embodiment, resource usage may be monitored, controlled, and reported, thereby providing transparency to both the provider and the user of a service being used.
[0134] In at least one embodiment, cloud computing may be associated with various services. In at least one embodiment, cloud Software as a Service (SaaS) may refer to a service in which a consumer is offered the ability to use a provider's applications running on a cloud infrastructure. In at least one embodiment, the applications are accessible from various client devices via a thin client interface, such as a web browser (e.g., web-based email). In at least one embodiment, the consumer does not manage or control the underlying cloud infrastructure, including network, servers, operating systems, storage, or even individual application features, with the possible exception of limited user-specific configuration settings for applications.
[0135] In at least one embodiment, Cloud Platform as a Service (PaaS) may refer to a service where a capability provided to a consumer is to deploy consumer-created or acquired applications, built with programming languages and tools supported by a provider, on the cloud infrastructure. In at least one embodiment, the consumer does not manage or control the underlying cloud infrastructure, including networks, servers, operating systems, or storage, but has control over the deployed applications and possibly the configurations of the application hosting environment.
[0136] In at least one embodiment, Cloud Infrastructure as a Service (IaaS) may refer to a service where a capability provided to the consumer is to provision processing, storage, networking, and other basic computing resources on which the consumer can deploy and run arbitrary software, which may include operating systems and applications. In at least one embodiment, the consumer does not manage or control the underlying cloud infrastructure, but rather has control over operating systems, storage, deployed applications, and possibly limited control over select network components (e.g., host firewalls).
[0137] In at least one embodiment, cloud computing can be used in a variety of ways. In at least one embodiment, a private cloud can refer to a cloud infrastructure operated exclusively for an organization. In at least one embodiment, a private cloud can be managed by an organization or a third party and can exist on-premises or off-site. In at least one embodiment, a community cloud can refer to a cloud infrastructure shared by multiple organizations that supports a particular community that has common concerns (e.g., mission, security requirements, policy, and compliance considerations). In at least one embodiment, a community cloud can be managed by organizations or a third party and can exist both on-premises and off-site.In at least one embodiment, a public cloud may refer to a cloud infrastructure that is made available to the general public or a large industry group and is owned by an organization that provides cloud services. In at least one embodiment, a hybrid cloud may refer to a cloud infrastructure consisting of two or more (private, community, or public) clouds that remain independent entities but are interconnected by standardized or proprietary technologies that enable data and application portability (e.g., cloud bursting for load balancing between clouds). In at least one embodiment, a computing environment in the cloud is service-oriented, emphasizing statelessness, low coupling, modularity, and semantic interoperability.
[0138] In at least one embodiment, at least one in Fig. 18 is used to implement techniques and / or functions associated with Fig. 1-13. In at least one embodiment, at least one component of Fig. 18 to effect scaling of one or more clocks of one or more cores based at least in part on the proximity of the one or more cores to each other. In at least one embodiment, at least one component of Fig. 18 at least one aspect relating to the data center 120, the controller 105, at least two or more of processor 110(a) to processor 110(n) and one or more thermal sensors 115(a) to thermal sensor 115(n) of Fig. 1, Firmware 220, Configuration file 225, ACPI 230, Policy 235 one or more processors 1-n of Fig. 2, and / or comprising one or more of: Method 900 of Fig. 9, Procedure 1000 of Fig. 10 and Procedure 1100 of Fig. 11.
[0139] Fig. 19 shows one or more components of a system environment 1900 in which services may be offered as network services by third parties, according to at least one embodiment. In at least one embodiment, a third-party network may be referred to as a cloud, cloud network, cloud computing network, and / or variations thereof. In at least one embodiment, the system environment 1900 includes one or more client computer systems 1904, 1906, and 1908 that can be used by users to interact with a third-party network infrastructure system 1902 that provides third-party network services that may be referred to as cloud computing services. In at least one embodiment, the third-party network infrastructure system 1902 may include one or more computers and / or servers.
[0140] It should be appreciated that the Fig. 19 may include other components than those shown. Furthermore, Fig. 19 depicts an embodiment of a third-party network infrastructure system. In at least one embodiment, the third-party network infrastructure system 1902 may include more or fewer components than in Fig. 19, combine two or more components or have a different configuration or arrangement of components.
[0141] In at least one embodiment, client computer systems 1904, 1906, and 1908 may be configured to run a client application, such as a web browser, a proprietary client application, or other application that may be used by a user of a client computer system device to interact with a third-party network infrastructure system 1902 to utilize services provided by a third-party network infrastructure system 1902. Although the example system environment 1900 is shown with three client computer systems, any number of client computer systems may be supported. In at least one embodiment, other devices, such as devices with sensors, etc., may interact with the third-party network infrastructure system 1902.In at least one embodiment, network(s) 1910 may facilitate communication and data exchange between client computer systems 1904, 1906, and 1908 and the third party network infrastructure system 1902.
[0142] In at least one embodiment, the services provided by the third-party network infrastructure system 1902 may include a variety of services made available upon request to users of a third-party network infrastructure system. In at least one embodiment, various services may also be offered, including, without limitation, online data storage and backup solutions, web-based email services, hosted office suites and document collaboration services, database management and processing, managed technical support services, and / or variations thereof. In at least one embodiment, the services provided by a third-party network infrastructure system may be dynamically scaled to meet the needs of its users.
[0143] In at least one embodiment, a specific instantiation of a service provided by a third-party network infrastructure system (1902) may be referred to as a "service instance." In at least one embodiment, any service provided to a user over a communications network, such as the Internet, by a third-party network service provider's system is generally referred to as a "third-party network service." In at least one embodiment, in a third-party public network environment, the servers and systems that comprise the third-party system are different from the customer's own on-premises servers and systems. In at least one embodiment, the third-party network service provider's system may host an application, and a user may order and use an application on demand over a communications network such as the Internet.
[0144] In at least one embodiment, a service in a third-party computer network infrastructure may comprise protected computer network access to storage, a hosted database, a hosted web server, a software application, or other service provided to a user by a third-party network provider. In at least one embodiment, a service may comprise password-protected access to remote storage in a third-party network via the Internet. In at least one embodiment, a service may comprise a web service-based, hosted relational database and a scripting language middleware engine for private use by a networked developer. In at least one embodiment, a service may comprise access to an email software application hosted on a third-party website.
[0145] In at least one embodiment, the third-party network infrastructure system 1902 may include a suite of applications, middleware, and database service offerings delivered to a customer in a self-service, subscription-based, elastically scalable, reliable, highly available, and secure manner. In at least one embodiment, the third-party network infrastructure system 1902 may also provide "Big Data"-related computation and analytics services. In at least one embodiment, the term "Big Data" is generally used to refer to extremely large data sets that can be stored and manipulated by analysts and researchers to visualize large amounts of data, identify trends, and / or otherwise interact with data.In at least one embodiment, big data and related applications may be hosted and / or processed by an infrastructure system at many levels and at different scales. In at least one embodiment, dozens, hundreds, or thousands of parallel processors may act on such data to present it or to simulate external forces on the data or what it represents. In at least one embodiment, these data sets may include structured data, such as data organized in a database or otherwise according to a structured model, and / or unstructured data (e.g., emails, images, data blocks (binary large objects), web pages, complex event processing).In at least one embodiment, by exploiting the ability of an embodiment to relatively quickly focus more (or fewer) computing resources on a target, a third-party network infrastructure system may be more available to perform tasks on large data sets based on the demand of a business, government agency, research organization, private individual, group of like-minded individuals or organizations, or other entity.
[0146] In at least one embodiment, the third-party network infrastructure system 1902 may be adapted to automatically provision, manage, and track a customer's subscription to the services offered by the third-party network infrastructure system 1902. In at least one embodiment, the third-party network infrastructure system 1902 may provide third-party network services through various deployment models. In at least one embodiment, the services may be provided under a public third-party network model, where the third-party network infrastructure system 1902 is owned by an organization that sells third-party network services, and the services are made available to the general public or various industries.In at least one embodiment, the services may be provided under a private third-party network model, where the third-party network infrastructure system 1902 operates exclusively for a single organization and may provide services to one or more entities within an organization. In at least one embodiment, third-party network services may also be provided under a collaborative third-party network model, where the third-party network infrastructure system 1902 and the services provided by the third-party network infrastructure system 1902 are shared by multiple organizations in a connected community. In at least one embodiment, third-party network services may also be provided under a hybrid third-party network model, which is a combination of two or more different models.
[0147] In at least one embodiment, the services provided by the third-party network infrastructure system 1902 may include one or more services provided under the Software as a Service (SaaS) category, the Platform as a Service (PaaS) category, the Infrastructure as a Service (IaaS) category, or other categories of services, including hybrid services. In at least one embodiment, a customer may order one or more services provided by the third-party network infrastructure system 1902 via a subscription order. In at least one embodiment, the third-party network infrastructure system 1902 then performs processing to provide the services in a customer's subscription order.
[0148] In at least one embodiment, the services provided by the third-party network infrastructure system 1902 may include, without limitation, application services, platform services, and infrastructure services. In at least one embodiment, application services may be provided by a third-party network infrastructure system via a SaaS platform. In at least one embodiment, the SaaS platform may be configured to provide third-party network services that fall under a SaaS category. In at least one embodiment, the SaaS platform may provide capabilities for building and deploying a range of applications on demand on an integrated development and deployment platform. In at least one embodiment, the SaaS platform may manage and control the underlying software and infrastructure for delivering SaaS services.In at least one embodiment, by using services provided by a SaaS platform, customers can utilize applications running on a third-party network infrastructure system. In at least one embodiment, customers can acquire an application service without having to purchase separate licenses and support. In at least one embodiment, various different SaaS services may be offered. In at least one embodiment, examples include, without limitation, services that provide solutions for sales performance management, enterprise integration, and business agility for large organizations.
[0149] In at least one embodiment, platform services may be provided by a third-party network infrastructure system 1902 via a PaaS platform. In at least one embodiment, the PaaS platform may be configured to provide third-party network services that fall under a PaaS category. In at least one embodiment, examples of platform services may include, without limitation, services that enable organizations to consolidate existing applications onto a common, shared architecture, as well as the ability to create new applications that utilize the shared services provided by a platform. In at least one embodiment, the PaaS platform may manage and control the underlying software and infrastructure for providing PaaS services.In at least one embodiment, customers may purchase PaaS services provided by a third-party network infrastructure system 1902 without requiring customers to purchase separate licenses and support.
[0150] In at least one embodiment, by leveraging services provided by a PaaS platform, customers can use programming languages and tools supported by a third-party network infrastructure system and also control the provided services. In at least one embodiment, the platform services provided by a third-party network infrastructure system can include third-party database network services, third-party middleware network services, and third-party network services. In at least one embodiment, third-party database network services can support shared service delivery models that enable organizations to aggregate database resources and offer customers a database as a service in the form of a third-party database network.In at least one embodiment, middleware third-party network services may provide a platform for customers to develop and deploy various business applications, and third-party network services may provide a platform for customers to deploy applications in a third-party network infrastructure system.
[0151] In at least one embodiment, various different infrastructure services may be provided by an IaaS platform in a third-party network infrastructure system. In at least one embodiment, infrastructure services facilitate the management and control of underlying computing resources, such as storage, networks, and other core computing resources, for customers using the services offered by a SaaS platform and a PaaS platform.
[0152] In at least one embodiment, the third-party network infrastructure system 1902 may also include infrastructure resources 1930 for providing resources used to provide various services to customers of a third-party network infrastructure system. In at least one embodiment, the infrastructure resources 1930 may include pre-integrated and optimized combinations of hardware, such as server, storage, and network resources for executing services provided by a PaaS platform and a SaaS platform, as well as other resources.
[0153] In at least one embodiment, the resources in the third-party network infrastructure system 1902 may be shared among multiple users and dynamically reallocated as needed. In at least one embodiment, the resources may be allocated to users in different time zones. In at least one embodiment, the third-party network infrastructure system 1902 may enable a first group of users in a first time zone to use the resources of a third-party network infrastructure system for a certain number of hours and then enable reallocation of the same resources to another group of users in a different time zone, thereby maximizing the utilization of the resources.
[0154] In at least one embodiment, a number of internal shared services 1932 may be provided that are shared by different components or modules of the third-party network infrastructure system 1902 to enable the provision of services by the third-party network infrastructure system 1902. In at least one embodiment, these internal shared services may include, without limitation, a security and identity service, an integration service, an enterprise repository service, an enterprise manager service, a virus scanning and whitelisting service, a high availability, backup, and recovery service, a service for enabling third-party network support, an email service, a notification service, a file transfer service, and / or variations thereof.
[0155] In at least one embodiment, the third-party network infrastructure system 1902 may provide comprehensive management of third-party network services (e.g., SaaS, PaaS, and IaaS services) within a third-party network infrastructure system. In at least one embodiment, the third-party network management functionality may include capabilities for provisioning, managing, and tracking a customer subscription received by the third-party network infrastructure system 1902, and / or variations thereof.
[0156] In at least one embodiment, as in Fig. 19, the third-party network management functionality may be provided by one or more modules, such as a job management module 1920, a job orchestration module 1922, a job provisioning module 1924, a job management and monitoring module 1926, and an identity management module 1928. In at least one embodiment, these modules may include or be provided using one or more computers and / or servers, which may be general-purpose computers, specialized server computers, server farms, server clusters, or any other suitable arrangement and / or combination.
[0157] In at least one embodiment, at step 1934, a customer using a client device, such as client computer systems 1904, 1906, or 1908, may interact with the third-party network infrastructure system 1902 by requesting one or more services provided by the third-party network infrastructure system 1902 and placing an order for a subscription to one or more services offered by the third-party network infrastructure system 1902. In at least one embodiment, a customer may access a third-party network interface (UI), such as the third-party network interface 1912, the third-party network interface 1914, and / or the third-party network interface 1916, and place a subscription order through those UIs.In at least one embodiment, the order information that third-party network infrastructure system 1902 receives in response to a customer's order may include information identifying a customer and one or more services offered by third-party network infrastructure system 1902 to which the customer intends to subscribe.
[0158] In at least one embodiment, in step 1936, order information received from a customer may be stored in an order database 1918. In at least one embodiment, if the order is a new order, a new record for the order may be created. In at least one embodiment, the order database 1918 may be one of a plurality of databases operated by the third-party network infrastructure system 1918 and operated in conjunction with other system elements.
[0159] In at least one embodiment, in step 1938, order information may be forwarded to an order management module 1920, which may be configured to perform billing and accounting functions with respect to an order, such as verifying an order and, upon verification, posting an order.
[0160] In at least one embodiment, at step 1940, information about an order may be communicated to an order orchestration module 1922 configured to orchestrate the provisioning of services and resources for an order placed by a customer. In at least one embodiment, the order organization module 1922 may utilize services of the order provisioning module 1924 for provisioning. In at least one embodiment, the order organization module 1922 enables the management of business processes associated with each order and applies business logic to determine whether an order should proceed to provisioning.
[0161] In at least one embodiment, upon receiving an order for a new subscription, the order organization module 1922 sends a request to the order provisioning module 1924 in step 1942 to allocate and configure resources required to fulfill a subscription order. In at least one embodiment, the order provisioning module 1924 enables allocation of resources for the services ordered by a customer. In at least one embodiment, the order provisioning module 1924 provides a level of abstraction between the third-party network services provided by the third-party network infrastructure system 1900 and a physical implementation layer used to provision resources for providing requested services.In at least one embodiment, this allows the job orchestration module 1922 to be isolated from implementation details, such as whether services and resources are actually provisioned in real time or in advance and allocated / assigned only on request.
[0162] In at least one embodiment, in step 1944, once the services and resources are provisioned, a notification may be sent to the subscribing customers indicating that a requested service is now ready for use. In at least one embodiment, information (e.g., a link) may be sent to a customer enabling them to begin using requested services.
[0163] In at least one embodiment, in step 1946, a customer's subscription order may be managed and tracked by an order management and monitoring module 1926. In at least one embodiment, the order management and monitoring module 1926 may be configured to collect usage statistics about the customer's use of the subscribed services. In at least one embodiment, statistics may be collected about the amount of storage used, the amount of data transferred, the number of users, and the amount of system uptime and downtime, and / or variations thereof.
[0164] In at least one embodiment, the third-party network infrastructure system 1900 may include an identity management module 1928 configured to provide identity services, such as access management and authorization services, within the third-party network infrastructure system 1900. In at least one embodiment, the identity management module 1928 may control information about customers seeking to use services provided by the third-party network infrastructure system 1902. In at least one embodiment, such information may include information authenticating the identities of such customers and information describing what actions these customers are permitted to perform with respect to various system resources (e.g., files, directories, applications, communication ports, memory segments, etc.).In at least one embodiment, the identity management module 1928 may also include managing descriptive information about each customer and how and by whom that descriptive information may be accessed and modified.
[0165] In at least one embodiment, at least one in Fig. 19 is used to implement techniques and / or functions associated with Fig. 1-13. In at least one embodiment, at least one component of Fig. 19 to effect scaling of one or more clocks of one or more cores based at least in part on the proximity of the one or more cores to each other. In at least one embodiment, at least one component of Fig. 19 at least one aspect relating to the data center 120, the controller 105, at least two or more of processor 110(a) to processor 110(n) and one or more thermal sensors 115(a) to thermal sensor 115(n) of Fig. 1, Firmware 220, Configuration file 225, ACPI 230, Policy 235 one or more processors 1-n of Fig. 2, and / or comprising one or more of: Method 900 of Fig. 9, Procedure 1000 of Fig. 10, and Procedure 1100 of Fig. 11.
[0166] Fig. 20 shows a computing environment 2002 in accordance with at least one embodiment. In at least one embodiment, the computing environment 2002 includes one or more computer systems / servers 2004 with which computer systems such as the personal digital assistant (PDA) or mobile phone 2006A, the desktop computer 2006B, the laptop computer 2006C, and / or the in-car computer system 2006N communicate. In at least one embodiment, this allows infrastructure, platforms, and / or software to be offered as services by the computing environment 2002 in the cloud, so that each customer does not have to maintain these resources separately. It is understood that the Fig. 20 are for illustrative purposes only, and that the cloud computing environment 2002 may communicate with any type of computing device over any type of network and / or network / addressing connection (e.g., with a web browser).
[0167] In at least one embodiment, a computer system / server 2004, which may be referred to as a cloud computing node, is operable with numerous other general-purpose or special-purpose computing environments or configurations. In at least one embodiment, examples of computing systems, environments, and / or configurations that may be suitable for use with the computer system / server 2004 include, but are not limited to, personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputer systems, mainframe systems, and distributed cloud computing environments that include any of the above-mentioned systems or devices and / or variations thereof.
[0168] In at least one embodiment, the computer system / server 2004 may be described in a general context of computer system-executable instructions, such as program modules, that are executed by a computer system. In at least one embodiment, the program modules include routines, programs, objects, components, logic, data structures, etc., that perform particular tasks or implement particular abstract data types. In at least one embodiment, an exemplary computer system / server 2004 may be deployed in distributed noisy computing environments in which tasks are executed by remote processing devices connected via a communications network. In at least one embodiment, in a distributed cloud computing environment, program modules may be located on both local and remote computer system storage media, including storage devices.
[0169] In at least one embodiment, at least one in Fig. 20 is used to implement techniques and / or functions associated with Fig. 1-13. In at least one embodiment, at least one component of Fig. 20 to effect scaling of one or more clocks of one or more cores based at least in part on the proximity of the one or more cores to each other. In at least one embodiment, at least one component of Fig. 20 at least one aspect relating to the data center 120, the controller 105, at least two or more of processor 110(a) to processor 110(n) and one or more thermal sensors 115(a) to thermal sensor 115(n) of Fig. 1, Firmware 220, Configuration file 225, ACPI 230, Policy 235 one or more processors 1-n of Fig. 2, and / or comprising one or more of: Method 900 of Fig. 9, Procedure 1000 of Fig. 10, and Procedure 1100 of Fig. 11.
[0170] Fig. 21 shows a set of functional abstraction layers used by the 2002 computer environment ( Fig. 20) in accordance with at least one embodiment. It should be understood in advance that the Fig. 21 The components, layers and functions shown are for illustrative purposes only and the components, layers and functions may vary.
[0171] In at least one embodiment, the hardware and software layer 2102 includes hardware and software components. In at least one embodiment, examples of hardware components include mainframe computers, various Reduced Instruction Set Computer (RISC) architecture servers, various computer systems, supercomputer systems, storage devices, networks, network components, and / or variations thereof. In at least one embodiment, examples of software components include network application server software, various application server software, various database software, and / or variations thereof.
[0172] In at least one embodiment, the virtualization layer 2104 provides an abstraction layer from which the following example virtual entities may be deployed: virtual servers, virtual storage, virtual networks, including virtual private networks, virtual applications, virtual clients, and / or variations thereof.
[0173] In at least one embodiment, the management layer 2106 provides various functions. In at least one embodiment, resource provisioning provides dynamic procurement of computing resources and other resources used to perform tasks within a cloud computing environment. In at least one embodiment, metering enables tracking the usage of resources within a cloud computing environment and billing or invoicing for the consumption of those resources. In at least one embodiment, the resources may include licenses for application software. In at least one embodiment, security provides identity verification for users and tasks, as well as protection for data and other resources.In at least one embodiment, the user interface provides both users and system administrators with access to a computing environment in the cloud. In at least one embodiment, service level management ensures the allocation and management of cloud computing resources so that the required service levels are met. In at least one embodiment, service level agreement (SLA) management enables the pre-arrangement and procurement of cloud computing resources for which future demand is expected according to an SLA.
[0174] In at least one embodiment, workload layer 2108 provides functions that utilize a cloud computing environment. In at least one embodiment, examples of workloads and functions that may be provided from this layer include: mapping and navigation, software development and management, educational services, data analysis and processing, transaction processing, and service provisioning. Supercomputing
[0175] The following figures illustrate, without limitation, exemplary supercomputer-based systems that may be used to implement at least one embodiment.
[0176] In at least one embodiment, a supercomputer may refer to a hardware system having significant parallelism and comprising at least one chip, wherein the chips in a system are interconnected by a network and housed in hierarchically organized enclosures. In at least one embodiment, a large, machine-room-filling hardware system having multiple racks, each containing multiple board / rack modules, each containing multiple chips, all interconnected by a scalable network, is a particular example of a supercomputer. In at least one embodiment, a single rack of such a large hardware system is another example of a supercomputer.In at least one embodiment, a single chip that has significant parallelism and includes multiple hardware components may also be considered a supercomputer, since as the size of features decreases, the amount of hardware that can be integrated into a single chip may increase.
[0177] In at least one embodiment, at least one in Fig. 21 is used to implement techniques and / or functions associated with Fig. 1-13. In at least one embodiment, at least one component of Fig. 21 to effect scaling of one or more clocks of one or more cores based at least in part on the proximity of the one or more cores to each other. In at least one embodiment, at least one component of Fig. 21 at least one aspect relating to the data center 120, the controller 105, at least two or more of processor 110(a) to processor 110(n) and one or more thermal sensors 115(a) to thermal sensor 115(n) of Fig. 1, Firmware 220, Configuration file 225, ACPI 230, Policy 235 one or more processors 1-n of Fig. 2, and / or comprising one or more of: Method 900 of Fig. 9, Procedure 1000 of Fig. 10 and Procedure 1100 of Fig. 11.
[0178] Fig. 22 shows a chip-level supercomputer in accordance with at least one embodiment. In at least one embodiment, within an FPGA or ASIC chip, the main computations are performed in finite state machines (2204), called thread units. In at least one embodiment, task and synchronization networks (2202) connect the finite state machines and serve to distribute threads and execute operations in the correct order. In at least one embodiment, a partitioned, multi-level on-chip cache hierarchy (2208, 2212) is accessed via memory networks (2206, 2210). In at least one embodiment, off-chip memory is accessed via memory controllers (2216) and an off-chip memory network (2214).In at least one embodiment, the I / O controller (2218) is used for cross-chip communication when a design does not fit on a single logic chip.
[0179] In at least one embodiment, at least one in Fig. 22 is used to implement techniques and / or functions associated with Fig. 1-13. In at least one embodiment, at least one component of Fig. 22 to effect scaling of one or more clocks of one or more cores based at least in part on the proximity of the one or more cores to each other. In at least one embodiment, at least one component of Fig. 22 at least one aspect relating to the data center 120, the controller 105, at least two or more of processor 110(a) to processor 110(n) and one or more thermal sensors 115(a) to thermal sensor 115(n) of Fig. 1, Firmware 220, Configuration file 225, ACPI 230, Policy 235 one or more processors 1-n of Fig. 2, and / or comprising one or more of: Method 900 of Fig. 9, Procedure 1000 of Fig. 10 and Procedure 1100 of Fig. 11.
[0180] Fig. Figure 23 shows a supercomputer at the level of a rock module according to at least one embodiment. In at least one embodiment, within a rack module, there are multiple FPGA or ASIC chips (2302) connected to one or more DRAM units (2304) that form the main memory of the accelerator. In at least one embodiment, each FPGA / ASIC chip is connected to its neighboring FPGA / ASIC chip via wide buses on a board with differential high-speed signaling (2306). In at least one embodiment, each FPGA / ASIC chip is also connected to at least one high-speed serial communication cable.
[0181] In at least one embodiment, at least one in Fig. 23 is used to implement techniques and / or functions associated with Fig. 1-13. In at least one embodiment, at least one component of Fig. 23 to effect scaling of one or more clocks of one or more cores based at least in part on the proximity of the one or more cores to each other. In at least one embodiment, at least one component of Fig. 23 at least one aspect relating to the data center 120, the controller 105, at least two or more of processor 110(a) to processor 110(n) and one or more thermal sensors 115(a) to thermal sensor 115(n) of Fig. 1, Firmware 220, Configuration file 225, ACPI 230, Policy 235 one or more processors 1-n of Fig. 2, and / or comprising one or more of: Method 900 of Fig. 9, Procedure 1000 of Fig. 10 and Procedure 1100 of Fig. 11.
[0182] Fig. 24 shows a rack level supercomputer in accordance with at least one embodiment.
[0183] In at least one embodiment, at least one in Fig. 24 is used to implement techniques and / or functions associated with Fig. 1-13. In at least one embodiment, at least one component of Fig. 24 to effect scaling of one or more clocks of one or more cores based at least in part on the proximity of the one or more cores to each other. In at least one embodiment, at least one component of Fig. 24 at least one aspect relating to the data center 120, the controller 105, at least two or more of processor 110(a) to processor 110(n) and one or more thermal sensors 115(a) to thermal sensor 115(n) of Fig. 1, Firmware 220, Configuration file 225, ACPI 230, Policy 235 one or more processors 1-n of Fig. 2, and / or comprising one or more of: Method 900 of Fig. 9, Procedure 1000 of Fig. 10 and Procedure 1100 of Fig. 11.
[0184] Fig. Figure 25 shows a supercomputer at an entire system level, according to at least one embodiment. In at least one embodiment, relating to Fig. 24 and Fig. 25, high-speed serial optical or copper cables (2402, 2502) are used between rack modules within a rack and across racks throughout the system to realize a scalable, possibly incomplete hypercube network. In at least one embodiment, one of the FPGA / ASIC chips of an accelerator is connected to a host system via a PCI Express connection (2504). In at least one embodiment, the host system includes a host microprocessor (2508) running a software portion of an application and memory comprised of one or more host memory DRAM units (2506) maintained coherent with the memory of an accelerator. In at least one embodiment, the host system may be a separate module on one of the racks or integrated into one of the modules of a supercomputer.In at least one embodiment, cube-connected cycles provide communication connections to create a hypercube network for a large supercomputer. In at least one embodiment, a small group of FPGA / ASIC chips on a rack module can act as a single hypercube node, thus increasing the total number of external connections of each group compared to a single chip. In at least one embodiment, a group includes chips A, B, C, and D on a rack module with internal wide differential buses connecting A, B, C, and D in a torus organization. In at least one embodiment, there are 12 serial communication cables connecting a rack module to the outside world. In at least one embodiment, chip A on a rack module is connected to serial communication cables 0, 1, 2. In at least one embodiment, chip B is connected to cables 3, 4, 5.In at least one embodiment, chip C is connected to cables 6, 7, 8. In at least one embodiment, chip D is connected to 9, 10, 11. In at least one embodiment, an entire group {A, B, C, D} forming a rack module can form a hypercube node within a supercomputer system with up to 212 = 4096 rack modules (16384 FPGA / ASIC chips). In at least one embodiment, a message to be sent from chip A via link 4 of group {A, B, C, D} must first be routed to chip B via an on-board differential wide bus connection. In at least one embodiment, a message arriving in a group {A, B, C, D} on link 4 (i.e., arriving at B) and intended for chip A must also first be routed internally within a group {A, B, C, D} to a correct destination chip (A).In at least one embodiment, parallel supercomputer systems of other sizes can also be realized.
[0185] In at least one embodiment, at least one in Fig. 25 is used to implement techniques and / or functions associated with Fig. 1-13. In at least one embodiment, at least one component of Fig. 25 to effect scaling of one or more clocks of one or more cores based at least in part on the proximity of the one or more cores to each other. In at least one embodiment, at least one component of Fig. 25 at least one aspect relating to the data center 120, the controller 105, at least two or more of processor 110(a) to processor 110(n) and one or more thermal sensors 115(a) to thermal sensor 115(n) of Fig. 1, Firmware 220, Configuration file 225, ACPI 230, Policy 235 one or more processors 1-n of Fig. 2, and / or comprising one or more of: Method 900 of Fig. 9, Procedure 1000 of Fig. 10, and Procedure 1100 of Fig. 11. Artificial intelligence
[0186] The following figures illustrate, without limitation, exemplary artificial intelligence-based systems that may be used to implement at least one embodiment.
[0187] Fig. 26A illustrates the inference and / or training logic 2615 used to perform inference and / or training operations in connection with one or more embodiments. Details of the inference and / or training logic 2615 are described below in connection with Fig. 26A and / or 26B.
[0188] In at least one embodiment, the inference and / or training logic 2615 may include, without limitation, code and / or data storage 2601 to store feedforward and / or output weights and / or input / output data and / or other parameters to configure neurons or layers of a neural network trained and / or used for inference in aspects of one or more embodiments. In at least one embodiment, the training logic 2615 may include or be coupled to graph code and / or data storage 2601 to store graph code or other software that controls the timing and / or order in which information about weights and / or other parameters is loaded to configure the logic, including integer and / or floating-point units (collectively referred to as arithmetic logic units (ALUs)).In at least one embodiment, code, such as graph code, loads information about weights or other parameters into the processor's ALUs based on the architecture of a neural network to which that code corresponds. In at least one embodiment, the code and / or data storage 2601 stores weight parameters and / or input / output data of each layer of a neural network being trained using aspects of one or more embodiments or used in connection with one or more embodiments during forward propagation of input / output data and / or weight parameters during training and / or inferencing. In at least one embodiment, each portion of the code and / or data storage 2601 may include other on-chip or off-chip data stores, including a processor's L1, L2, or L3 cache or system memory.
[0189] In at least one embodiment, any portion of the code and / or data storage 2601 may be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, the code and / or data storage 2601 may be a cache memory, dynamic random addressable memory ("DRAM"), static random addressable memory ("SRAM"), non-volatile memory (e.g., flash memory), or other memory.In at least one embodiment, the choice of whether the code and / or code and / or data memory 2601 is, for example, internal or external to a processor or comprises DRAM, SRAM, Flash, or another memory type may depend on whether on-chip or off-chip memory is available, the latency requirements of the training and / or inference functions being performed, the batch size of the data used in inferencing and / or training a neural network, or a combination of these factors.
[0190] In at least one embodiment, the inference and / or training logic 2615 may include, without limitation, a code and / or data storage 2605 to store backward and / or output weight and / or input / output data corresponding to neurons or layers of a neural network being trained and / or used for inference in aspects of one or more embodiments. In at least one embodiment, the code and / or data storage 2605 stores weight parameters and / or input / output data of each layer of a neural network being trained or used in connection with one or more embodiments during backpropagation of input / output data and / or weight parameters during training and / or inference using aspects of one or more embodiments.In at least one embodiment, the training logic 2615 may include or be coupled to a code and / or data memory 2605 to store graph code or other software to control the timing and / or order in which weight and / or other parameter information is to be loaded to configure the logic, including integer and / or floating point units (collectively: arithmetic logic units (ALUs)).
[0191] In at least one embodiment, code, such as graph code, causes information about weights or other parameters to be loaded into processor ALUs based on a neural network architecture to which that code corresponds. In at least one embodiment, any portion of the code and / or data storage 2605 may include other on-chip or off-chip data storage, including the L1, L2, or L3 cache or system memory of a processor. In at least one embodiment, any portion of the code and / or data storage 2605 may be internal or external to one or more processors or other hardware logic devices or circuitry. In at least one embodiment, the code and / or data storage 2605 may be cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other memory.In at least one embodiment, the choice of whether the code and / or data memory 2605 is, for example, internal or external to a processor, or comprises DRAM, SRAM, flash memory, or another type of memory, may depend on the available memory on-chip versus off-chip, the latency requirements of the training and / or inference functions being performed, the batch size of the data used in inferencing and / or training a neural network, or a combination of these factors.
[0192] In at least one embodiment, code and / or data memory 2601 and code and / or data memory 2605 may be separate memory structures. In at least one embodiment, code and / or data memory 2601 and code and / or data memory 2605 may be a combined memory structure. In at least one embodiment, code and / or data memory 2601 and code and / or data memory 2605 may be partially combined and partially separate. In at least one embodiment, each portion of code and / or data memory 2601 and code and / or data memory 2605 may include other on-chip or off-chip data storage, including a processor's L1, L2, or L3 cache or system memory.
[0193] In at least one embodiment, inference and / or training logic 2615 may include, without limitation, one or more arithmetic logic unit(s) ("ALU(s)") 2610, including integer and / or floating-point units, to perform logical and / or mathematical operations based at least in part on or specified by training and / or inference code (e.g., graph code), the result of which may produce activations stored in activation memory 2620 (e.g., output data from layers or neurons within a neural network) that are functions of input / output data and / or weight parameters stored in code and / or data memory 2601 and / or code and / or data memory 2605.In at least one embodiment, activations stored in activation memory 2620 are generated according to linear algebraic and / or matrix-based mathematics performed by ALU(s) 2610 in response to execution instructions or other code, using weight values stored in code and / or data memory 2605 and / or data memory 2601 as operands along with other values, such as bias values, gradient information, momentum values, or other parameters or hyperparameters, any or all of which may be stored in code and / or data memory 2605 or code and / or data memory 2601 or other on-chip or off-chip memory.
[0194] In at least one embodiment, ALU(s) 2610 are included in one or more processors or other logical hardware devices or circuits, while in another embodiment, ALU(s) 2610 may be external to a processor or other logical hardware device or circuit that uses them (e.g., a coprocessor). In at least one embodiment, ALU(s) 2610 may be included in the execution units of a processor or otherwise in a bank of ALUs accessible by the execution units of a processor, either within the same processor or distributed among different processors of different types (e.g., central processing units, graphics processing units, fixed functional units, etc.).In at least one embodiment, code and / or data memory 2601, code and / or data memory 2605, and enable memory 2620 may share a processor or other logical hardware device or circuitry, while in another embodiment, they may be located in different processors or other logical hardware devices or circuitry, or in a combination of the same and different processors or other logical hardware devices or circuitry. In at least one embodiment, each portion of enable memory 2620 may include other on-chip or off-chip data stores, including a processor's L1, L2, or L3 cache or system memory.Furthermore, the inference and / or training code may be stored along with other code accessible by a processor or other hardware logic or circuitry and retrieved and / or processed using a processor's fetch, decode, schedule, execute, retire, and / or other logic circuitry.
[0195] In at least one embodiment, the activation memory 2620 may be a cache memory, a DRAM, an SRAM, a non-volatile memory (e.g., flash memory), or other memory. In at least one embodiment, the activation memory 2620 may be located entirely or partially inside or outside one or more processors or other logic circuits. In at least one embodiment, the choice of whether the activation memory 2620 is, for example, inside or outside a processor, or includes DRAM, SRAM, flash memory, or another type of memory, may depend on the available on-chip versus off-chip memory, the latency requirements of the training and / or inference functions performed, the batch size of the data used in inferencing and / or training a neural network, or a combination of these factors.
[0196] In at least one embodiment, the Fig. 26A may be used in conjunction with an application-specific integrated circuit (“ASIC”), such as a TensorFlow® Processing Unit from Google, an inference processing unit (IPU) from Graphcore™, or a Nervana® processor (e.g., “Lake Crest”) from Intel Corp. In at least one embodiment, the inference and / or training logic 2615 shown in Fig. 26A may be used in conjunction with central processing unit (“CPU”) hardware, graphics processing unit (“GPU”) hardware, or other hardware such as field programmable gate arrays (“FPGAs”).
[0197] Fig. 26B shows the inference and / or training logic 2615 according to at least one embodiment. In at least one embodiment, the inference and / or training logic 2615 may include, without limitation, hardware logic in which computational resources associated with weight values or other information corresponding to one or more layers of neurons within a neural network are dedicated or otherwise exclusively used. In at least one embodiment, the inference and / or training logic 2615 may Fig. 26B may be used in conjunction with an application-specific integrated circuit (ASIC), such as Google's TensorFlow® Processing Unit, a Graphcore™ Inference Processing Unit (IPU), or a Nervana® (e.g., "Lake Crest") processor from Intel Corp. In at least one embodiment, the inference and / or training logic 2615 shown in Fig. 26B may be used in conjunction with central processing unit (CPU) hardware, graphics processing unit (GPU), or other hardware, such as field programmable gate arrays (FPGAs). In at least one embodiment, the inference and / or training logic 2615 includes, without limitation, code and / or data storage 2601 and code and / or data storage 2605, which may be used to store code (e.g., graph code), weight values, and / or other information, including bias values, gradient information, momentum values, and / or other parameter or hyperparameter information. In at least one embodiment shown in Fig. 26B, each code and / or data memory 2601 and each code and / or data memory 2605 is coupled to a dedicated computing resource, such as computer hardware 2602 and computer hardware 2606, respectively. In at least one embodiment, each of computer hardware 2602 and computer hardware 2606 includes one or more ALUs that perform mathematical functions, such as linear algebraic functions, only on information stored in code and / or data memory 2601 and code and / or data memory 2605, respectively, and the result of which is stored in activation memory 2620.
[0198] In at least one embodiment, each of the code and / or data memories 2601 and 2605 and the corresponding computer hardware 2602 and 2606 correspond to different layers of a neural network, such that the resulting activation from one memory / computing pair 2601 / 2602 of code and / or data memory 2601 and computer hardware 2602 is provided as input to a next memory / computing pair 2605 / 2606 of code and / or data memory 2605 and computer hardware 2606 to reflect a conceptual organization of a neural network. In at least one embodiment, each of the memory / computing pairs 2601 / 2602 and 2605 / 2606 may correspond to more than one layer of the neural network. In at least one embodiment, additional memory / compute pairs (not shown) may be included in the inference and / or training logic 2615 subsequent to or in parallel with the memory / compute pairs 2601 / 2602 and 2605 / 2606.
[0199] In at least one embodiment, at least one in Fig. 26 is used to implement techniques and / or functions associated with Fig. 1-13. In at least one embodiment, at least one component of Fig. 26 to effect scaling of one or more clocks of one or more cores based at least in part on the proximity of the one or more cores to each other. In at least one embodiment, at least one component of Fig. 26 at least one aspect relating to the data center 120, the controller 105, at least two or more of processor 110(a) to processor 110(n) and one or more thermal sensors 115(a) to thermal sensor 115(n) of Fig. 1, Firmware 220, Configuration file 225, ACPI 230, Policy 235 one or more processors 1-n of Fig. 2, and / or comprising one or more of: Method 900 of Fig. 9, Procedure 1000 of Fig. 10 and Procedure 1100 of Fig. 11.
[0200] Fig. 27 illustrates the training and deployment of a deep neural network according to at least one embodiment. In at least one embodiment, an untrained neural network 2706 is trained using a training dataset 2702. In at least one embodiment, the training framework 2704 is a PyTorch framework, while in other embodiments, the training framework 2704 is a TensorFlow, Boost, Caffe, Microsoft Cognitive Toolkit / CNTK, MXNet, Chainer, Keras, Deeplearning4j, or another training framework. In at least one embodiment, the training framework 2704 trains an untrained neural network 2706 and facilitates its training using the processing resources described herein to generate a trained neural network 2708. In at least one embodiment, the weights may be selected randomly or by pre-training using a deep belief network.In at least one embodiment, the training may be performed in either a supervised, semi-supervised, or unsupervised manner.
[0201] In at least one embodiment, the untrained neural network 2706 is trained using supervised learning, where the training data set 2702 includes an input paired with a desired output for an input, or where the training data set 2702 includes an input with a known output and an output of the neural network 2706 is manually ranked. In at least one embodiment, the untrained neural network 2706 is trained in a supervised manner and processes inputs from the training data set 2702 and compares the resulting outputs to a set of expected or desired outputs. In at least one embodiment, the errors are then backtracked through the untrained neural network 2706. In at least one embodiment, the training framework 2704 adjusts the weights that govern the untrained neural network 2706.In at least one embodiment, the training framework 2704 includes tools for monitoring the convergence of the untrained neural network 2706 toward a model, e.g., the trained neural network 2708, capable of generating correct answers, e.g., in the output 2714, based on input data, e.g., a new data set 2712. In at least one embodiment, the training framework 2704 repeatedly trains the untrained neural network 2706 while adjusting the weights to refine an output of the untrained neural network 2706 using a loss function and an adaptation algorithm, such as stochastic gradient descent. In at least one embodiment, the training framework 2704 trains the untrained neural network 2706 until the untrained neural network 2706 achieves the desired accuracy.In at least one embodiment, the trained neural network 2708 may then be used to implement any number of machine learning operations.
[0202] In at least one embodiment, the untrained neural network 2706 is trained using unsupervised learning, where the untrained neural network 2706 attempts to train itself on unlabeled data. In at least one embodiment, the training dataset 2702 for unsupervised learning comprises input data without associated output data or ground truth data. In at least one embodiment, the untrained neural network 2706 can learn groupings within the training dataset 2702 and determine how individual inputs are related to the untrained dataset 2702. In at least one embodiment, unsupervised training can be used to generate a self-organizing map in a trained neural network 2708 capable of performing operations useful in reducing the dimensionality of the new dataset 2712.In at least one embodiment, unsupervised training may also be used for anomaly detection, enabling the identification of data points in the new data set 2712 that deviate from normal patterns of the new data set 2712.
[0203] In at least one embodiment, semi-supervised learning may be used, i.e., a technique in which the training dataset 2702 includes a mixture of labeled and unlabeled data. In at least one embodiment, the training framework 2704 may be used to perform incremental learning, for example, through transfer learning techniques. In at least one embodiment, incremental learning allows the trained neural network 2708 to adapt to a new dataset 2712 without forgetting the knowledge instilled in the trained neural network 2708 during initial training.
[0204] In at least one embodiment, at least one in Fig. 27 is used to implement techniques and / or functions associated with Fig. 1-13. In at least one embodiment, at least one component of Fig. 27 to effect scaling of one or more clocks of one or more cores based at least in part on the proximity of the one or more cores to each other. In at least one embodiment, at least one component of Fig. 27 at least one aspect relating to the data center 120, the controller 105, at least two or more of processor 110(a) to processor 110(n) and one or more thermal sensors 115(a) to thermal sensor 115(n) of Fig. 1, Firmware 220, Configuration file 225, ACPI 230, Policy 235 one or more processors 1-n of Fig. 2, and / or comprising one or more of: Method 900 of Fig. 9, Procedure 1000 of Fig. 10 and Procedure 1100 of Fig. 11. 5G networks
[0205] The following figures illustrate, without limitation, exemplary 5G network-based systems that may be used to implement at least one embodiment.
[0206] Fig. 28 shows an embodiment of a system 2800 of a network according to at least one embodiment. In at least one embodiment, the system 2800 is illustrated as including a user equipment (UE) 2802 and a UE 2804. In at least one embodiment, the UEs 2802 and 2804 are shown as smartphones (e.g., portable touchscreen mobile computing devices connectable to one or more cellular networks), but may also include any other mobile or non-mobile computing device, such as personal data assistants (PDAs), pagers, laptop computers, desktop computers, wireless handsets, or any other computing device including a wireless communication interface.
[0207] In at least one embodiment, each of the UEs 2802 and 2804 may comprise an Internet of Things (IoT) UE, which may include a network layer designed for low-power IoT applications that utilize short-lived UE connections. In at least one embodiment, an IoT UE may utilize technologies such as machine-to-machine (M2M) or machine-type communications (MTC) to exchange data with an MTC server or device via a public mobile network (PLMN), proximity-based service (ProSe) or device-to-device (D2D) communications, sensor networks, or IoT networks. In at least one embodiment, an M2M or MTC data exchange may be a machine-initiated exchange of data.In at least one embodiment, an IoT network describes the interconnection of IoT UEs, which may include uniquely identifiable embedded computing devices (within the Internet infrastructure), with short-lived connections. In at least one embodiment, IoT UEs may execute background applications (e.g., keep-alive messages, status updates, etc.) to facilitate IoT network connections.
[0208] In at least one embodiment, UEs 2802 and 2804 may be configured to connect, for example, communicatively couple, to a radio access network (RAN) 2816. In at least one embodiment, RAN 2816 may be, for example, an Evolved Universal Mobile Telecommunications System (UMTS) Terrestrial Radio Access Network (E-UTRAN), a NextGen RAN (NG RAN), or another type of RAN. In at least one embodiment, UEs 2802 and 2804 utilize links 2812 and 2814, respectively, each comprising a physical communications interface or layer.In at least one embodiment, connections 2812 and 2814 are shown as an air interface to enable communicative coupling and may conform to cellular communication protocols such as a Global System for Mobile Communications (GSM) protocol, a Code-Division Multiple Access (CDMA) network protocol, a Push-to-Talk (PTT) protocol, a PTT over Cellular (POC) protocol, a Universal Mobile Telecommunications System (UMTS) protocol, a 3GPP Long Term Evolution (LTE) protocol, a fifth generation (5G) protocol, a New Radio (NR) protocol, and variants thereof.
[0209] In at least one embodiment, UEs 2802 and 2804 may further directly exchange communication data via a ProSe interface 2806. In at least one embodiment, ProSe interface 2806 may alternatively be referred to as a sidelink interface comprising one or more logical channels, including, but not limited to, a physical sidelink control channel (PSCCH), a physical sidelink shared channel (PSSCH), a physical sidelink discovery channel (PSDCH), and a physical sidelink broadcast channel (PSBCH).
[0210] In at least one embodiment, UE 2804 is configured to access an access point (AP) 2810 via connection 2808. In at least one embodiment, connection 2808 may comprise a local wireless connection, such as a connection compliant with an IEEE 802.11 protocol, where AP 2810 would comprise a Wireless Fidelity (WiFi®) router. In at least one embodiment, AP 2810 is illustrated as being connected to the Internet without connecting to a core network of a wireless system.
[0211] In at least one embodiment, the RAN 2816 may include one or more access nodes that enable connections 2812 and 2814. In at least one embodiment, these access nodes (ANs) may be referred to as base stations (BSs), NodeBs, evolved NodeBs (eNBs), next generation NodeBs (gNBs), RAN nodes, etc., and may include ground stations (e.g., terrestrial access points) or satellite stations that provide coverage within a geographic area (e.g., a cell). In at least one embodiment, the RAN 2816 may include one or more RAN nodes for deploying macrocells, such as macro RAN nodes 2818, and one or more RAN nodes for deploying femtocells or picocells (e.g., cells with smaller coverage areas, lower user capacity, or higher bandwidth compared to macrocells), such as low power (LP) RAN nodes 2820.
[0212] In at least one embodiment, each of RAN nodes 2818 and 2820 may terminate an air interface protocol and be a first point of contact for UEs 2802 and 2804. In at least one embodiment, each of RAN nodes 2818 and 2820 may perform various logical functions for RAN 2816, including, but not limited to, radio network control (RNC) functions such as radio user management, dynamic uplink and downlink radio resource management, data packet scheduling, and mobility management.
[0213] In at least one embodiment, the UEs 2802 and 2804 may be configured to communicate with each other or with one of the RAN nodes 2818 and 2820 using orthogonal frequency-division multiplexing (OFDM) communication signals over a multi-carrier communication channel in accordance with various communication techniques, such as, but not limited to, an orthogonal frequency division multiple access (OFDMA) communication technique (e.g., for downlink communications) or a single carrier frequency division multiple access (SC-FDMA) communication technique (e.g., for uplink and ProSe or sidelink communications), and / or variations thereof. In at least one embodiment, OFDM signals may include a plurality of orthogonal subcarriers.
[0214] In at least one embodiment, a downlink resource grid may be used for downlink transmissions from one of the RAN nodes 2818 and 2820 to the UEs 2802 and 2804, while similar techniques may be employed for uplink transmissions. In at least one embodiment, a grid may be a time-frequency grid, referred to as a resource grid or time-frequency resource grid, which is a physical resource in a downlink in each slot. In at least one embodiment, such a time-frequency plane representation is a common practice for OFDM systems, enabling intuitive allocation of radio resources. In at least one embodiment, each column and each row of a resource grid corresponds to an OFDM symbol and an OFDM subcarrier, respectively. In at least one embodiment, the duration of a resource grid in a time domain corresponds to a time slot in a radio frame.In at least one embodiment, the smallest time / frequency unit in a resource grid is referred to as a resource element. In at least one embodiment, each resource grid comprises a number of resource blocks that describe a mapping of specific physical channels to resource elements. In at least one embodiment, each resource block comprises a collection of resource elements. In at least one embodiment, this may represent, in a frequency domain, a smallest set of resources that can currently be allocated. In at least one embodiment, there are multiple different physical downlink channels transmitted over such resource blocks.
[0215] In at least one embodiment, a shared physical downlink channel (PDSCH) may transmit payload data and higher-layer signaling to UEs 2802 and 2804. In at least one embodiment, a physical downlink control channel (PDCCH) may transmit, among other things, information about a transport format and resource allocations related to the PDSCH channel. In at least one embodiment, it may also inform UEs 2802 and 2804 about a transport format, resource allocation, and hybrid automatic repeat request (HARQ) information related to a shared uplink channel. In at least one embodiment, downlink scheduling (allocation of control and shared channel resource blocks to UE 2802 within a cell) may typically be performed at one of the RAN nodes 2818 and 2820 based on channel quality information returned by one of the UEs 2802 and 2804.In at least one embodiment, information about the allocation of downlink resources may be transmitted on a PDCCH used (e.g., assigned) for each of the UEs 2802 and 2804.
[0216] In at least one embodiment, a PDCCH may use control channel elements (CCEs) to transmit control information. In at least one embodiment, complex-valued PDCCH symbols may first be organized into quadruplets prior to being assigned to resource elements, which may then be permuted using a sub-block interleaver for rate adaptation. In at least one embodiment, each PDCCH may be transmitted using one or more of these CCEs, where each CCE may correspond to nine sets of four physical resource elements, called resource element groups (REGs). In at least one embodiment, each REG may be assigned four quadrature phase shift keying (QPSK) symbols. In at least one embodiment, each PDCCH may be transmitted using one or more CCEs, depending on the size of a downlink control information (DCI) and a channel condition.In at least one embodiment, there may be four or more different PDCCH formats defined in LTE with different numbers of CCEs (e.g., aggregation levels, L=1, 2, 4, or 8).
[0217] In at least one embodiment, an enhanced physical downlink control channel (EPDCCH) utilizing PDSCH resources may be used for transmitting control information. In at least one embodiment, the EPDCCH may be transmitted using one or more enhanced control channel elements (ECCEs). In at least one embodiment, each ECCE may correspond to nine sets of four physical resource elements, referred to as enhanced resource element groups (EREGs). In at least one embodiment, an ECCE may have a different number of EREGs in some situations.
[0218] In at least one embodiment, the RAN 2816 is communicatively coupled to a core network (CN) 2838 via an S1 interface 2822. In at least one embodiment, the CN 2838 may be an Evolved Packet Core (EPC) network, a NextGen Packet Core (NPC) network, or another type of CN. In at least one embodiment, the S1 interface 2822 is split into two parts: an S1-U interface 2826, which carries traffic data between RAN nodes 2818 and 2820 and a Serving Gateway (S-GW) 2830, and an S1 Mobility Management Entity (MME) interface 2824, which is a signaling interface between RAN nodes 2818 and 2820 and MMEs 2828.
[0219] In at least one embodiment, CN 2838 includes MMEs 2828, S-GW 2830, Packet Data Network (PDN) Gateway (P-GW) 2834, and a Home Subscriber Server (HSS) 2832. In at least one embodiment, MMEs 2828 may have a similar function to the control plane of legacy Serving General Packet Radio Service (GPRS) Support Nodes (SGSN). In at least one embodiment, MMEs 2828 may manage mobility aspects of access, such as gateway selection and tracking area list management. In at least one embodiment, HSS 2832 may include a database for network users, including subscription information, to support network operators' handling of communication sessions. In at least one embodiment, CN 2838 may include one or more HSS 2832, depending on the number of mobile subscribers, device capacity, network organization, etc.In at least one embodiment, HSS 2832 may provide support for routing / roaming, authentication, authorization, name / address resolution, location dependencies, etc.
[0220] In at least one embodiment, S-GW 2830 may terminate an S1 interface 2822 toward RAN 2816 and route data packets between RAN 2816 and CN 2838. In at least one embodiment, SGW 2830 may be a local mobility anchor point for inter-RAN node handovers and may also provide an anchor for inter-3GPP mobility. In at least one embodiment, other responsibilities may include lawful interception, charging, and enforcement of some policies.
[0221] In at least one embodiment, the P-GW 2834 may terminate an SGi interface toward a PDN. In at least one embodiment, the P-GW 2834 may route data packets between an EPC network 2838 and external networks, such as a network including an application server 2840 (alternatively referred to as an application function (AF)), via an Internet Protocol (IP) interface 2842. In at least one embodiment, the application server 2840 may be an element that offers applications that utilize IP bearer resources with a core network (e.g., UMTS Packet Services (PS) domain, LTE PS data services, etc.). In at least one embodiment, the P-GW 2834 is communicatively coupled to an application server 2840 via an IP communication interface 2842.In at least one embodiment, application server 2840 may also be configured to support one or more communication services (e.g., Voice over Internet Protocol (VoIP) sessions, PTT sessions, group communication sessions, social networking services, etc.) for UEs 2802 and 2804 via CN 2838.
[0222] In at least one embodiment, the P-GW 2834 may further be a node for enforcing policies and collecting charging data. In at least one embodiment, the Policy and Charging Enforcement Function (PCRF) 2836 is a policy and charging control element of the CN 2838. In at least one embodiment, in a non-roaming scenario, there may be a single PCRF in a Home Public Land Mobile Network (HPLMN) associated with a UE's Internet Protocol Connectivity Access Network (IP-CAN) session. In at least one embodiment, in a roaming scenario with local traffic splitting, there may be two PCRFs associated with a UE's IP-CAN session: a Home PCRF (H-PCRF) in an HPLMN and a Visited PCRF (V-PCRF) in a Visited Public Land Mobile Network (VPLMN). In at least one embodiment, the PCRF 2836 may be communicatively coupled to the application server 2840 via the P-GW 2834.In at least one embodiment, application server 2840 may signal PCRF 2836 to specify a new service flow and select appropriate quality of service (QoS) and charging parameters. In at least one embodiment, PCRF 2836 may provide this policy in a Policy and Charging Enforcement Function (PCEF) (not shown) with an appropriate traffic flow template (TFT) and a QoS class identifier (QCI), which initiates QoS and charging according to the application server 2840's specifications.
[0223] In at least one embodiment, at least one in Fig. 28 is used to implement techniques and / or functions associated with Fig. 1-13. In at least one embodiment, at least one component of Fig. 28 to effect scaling of one or more clocks of one or more cores based at least in part on the proximity of the one or more cores to each other. In at least one embodiment, at least one component of Fig. 28 at least one aspect relating to the data center 120, the controller 105, at least two or more of processor 110(a) to processor 110(n) and one or more thermal sensors 115(a) to thermal sensor 115(n) of Fig. 1, Firmware 220, Configuration file 225, ACPI 230, Policy 235 one or more processors 1-n of Fig. 2, and / or comprising one or more of: Method 900 of Fig. 9, Procedure 1000 of Fig. 10, and Procedure 1100 of Fig. 11.
[0224] Fig. 29 illustrates an architecture of a system 2900 of a network in accordance with some embodiments. In at least one embodiment, the system 2900 is shown to include a UE 2902, a 5G access node or RAN node (represented as (R)AN node 2908), a user plane function (represented as UPF 2904), a data network (DN 2906), which may be, for example, operator services, internet access, or third-party services, and a 5G core network (5GC) (represented as CN 2910).
[0225] In at least one embodiment, CN 2910 includes an authentication server function (AUSF 2914), a core access and mobility management function (AMF 2912), a session management function (SMF 2918), a network exposure function (NEF 2916), a policy control function (PCF 2922), a network function (NF) repository function (NRF 2920), a unified data management function (UDM 2924), and an application function (AF 2926). In at least one embodiment, CN 2910 may also include other elements not shown, such as a network function for structured data stores (SDSF), a network function for unstructured data stores (UDSF), and variations thereof.
[0226] In at least one embodiment, UPF 2904 may serve as an anchor point for intra-RAT and inter-RAT mobility, as an external PDU session point for interconnection with DN 2906, and as a branch point for supporting multi-homed PDU sessions. In at least one embodiment, UPF 2904 may also perform packet routing and forwarding, packet inspection, enforcement of a subset of user-plane policy rules, legitimate packet interception (UP collection), traffic usage reporting, user-plane QoS handling (e.g., packet filtering, gating, UL / DL rate enforcement), uplink traffic inspection (e.g., SDF-to-QoS flow mapping), transport-level packet marking in the uplink and downlink, downlink packet buffering, and triggering of downlink data notifications. In at least one embodiment, UPF 2904 may include an uplink classifier to support forwarding traffic flows to a data network.In at least one embodiment, DN 2906 may represent various network operator services, Internet access, or third-party services.
[0227] In at least one embodiment, AUSF 2914 may store data for authenticating UE 2902 and handle authentication-related functions. In at least one embodiment, AUSF 2914 may enable a common authentication framework for different access types.
[0228] In at least one embodiment, AMF 2912 may be responsible for registration management (e.g., for the registration of UE 2902, etc.), connection management, reachability management, mobility management, and legitimate interception of AMF-related events, as well as access authentication and authorization. In at least one embodiment, AMF 2912 may provide the transport of SM messages for SMF 2918 and act as a transparent proxy for routing SM messages. In at least one embodiment, AMF 2912 may also provide the transport of SMS (Short Message Service) messages between UE 2902 and an SMS function (SMSF) (not in Fig. 29). In at least one embodiment, AMF 2912 may act as a Security Anchor Function (SEA), which may include interacting with AUSF 2914 and UE 2902 and receiving an intermediate key generated as a result of the UE 2902 authentication process. In at least one embodiment using USIM-based authentication, AMF 2912 may retrieve security material from AUSF 2914. In at least one embodiment, AMF 2912 may also include a Security Context Management (SCM) function that receives a key from SEA, which it uses to derive access network-specific keys. In at least one embodiment, AMF 2912 may also be a termination point of the RAN CP interface (N2 reference point) and a termination point of NAS (NI) signaling, and may perform NAS encryption and integrity protection.
[0229] In at least one embodiment, AMF 2912 may also support NAS signaling with a UE 2902 over an N3 Interworking Function (IWF) interface. In at least one embodiment, N3IWF may be used to enable access to untrusted entities. In at least one embodiment, N3IWF may be a termination point for N2 and N3 interfaces for the control plane and user plane, respectively, and as such, may handle N2 signaling from SMF and AMF for PDU sessions and QoS, encapsulate / decapsulate packets for IPSec and N3 tunneling, mark N3 user plane packets in the uplink, and enforce QoS according to the N3 packet marking, taking into account QoS requirements associated with such marking received over N2.In at least one embodiment, N3IWF may also forward uplink and downlink control plane NAS (NI) signaling between UE 2902 and AMF 2912 and forward uplink and downlink user plane packets between UE 2902 and UPF 2904. In at least one embodiment, N3IWF also provides mechanisms for IPsec tunnel establishment with UE 2902.
[0230] In at least one embodiment, SMF 2918 may be responsible for session management (e.g., session establishment, modification, and release, including maintaining the tunnel between UPF and AN nodes); allocation and management of UE IP addresses (including optional authorization); selection and control of the UPF function; configuration of traffic routing at UPF to direct traffic to the correct destination; termination of interfaces to policy control functions; control of parts of policy enforcement and QoS; lawful interception (for SM events and the interface to the LI system); termination of SM parts of NAS messages; data notification in the downlink; initiator of AN-specific SM information sent to AN via AMF over N2; determination of the SSC mode of a session.In at least one embodiment, SMF 2918 may include the following roaming functionality: Handle Local Enforcement for applying QoS SLAB (VPLMN); Charging Data Collection and Charging Interface (VPLMN); Lawful Interception (in VPLMN for SM events and interface to LI system); External DN interaction support for transporting PDU session authorization / authentication signals by external DN.
[0231] In at least one embodiment, NEF 2916 may provide means for securely exposing services and capabilities provided by 3GPP network functions to third parties, internal exposing / exposing, application functions (e.g., AF 2926), edge computing or fog computing systems, etc. In at least one embodiment, NEF 2916 may authenticate, authorize, and / or throttle AFs. In at least one embodiment, NEF 2916 may also translate information exchanged with AF 2926 and information exchanged with internal network functions. In at least one embodiment, NEF 2916 may translate between an AF service identifier and internal 5GC information. In at least one embodiment, NEF 2916 may also receive information from other network functions (NFs) based on the disclosed capabilities of other network functions.In at least one embodiment, this information may be stored at the NEF 2916 as structured data or at a data storage NF using a standardized interface. In at least one embodiment, the stored information may then be shared by the NEF 2916 with other NFs and AFs and / or used for other purposes, such as analytics.
[0232] In at least one embodiment, NRF 2920 can support service discovery functions, receive NF discovery requests from NF instances, and communicate information about discovered NF instances to NF instances. In at least one embodiment, NRF 2920 also manages information about available NF instances and the services they support.
[0233] In at least one embodiment, PCF 2922 may provide policy rules for the control plane function(s) to enforce them, and may also support a unified policy framework to govern network behavior. In at least one embodiment, PCF 2922 may also implement a front-end (FE) to access subscription information relevant to policy decisions in a UDR of UDM 2924.
[0234] In at least one embodiment, UDM 2924 may handle subscription-related information to support the handling of communication sessions by network entities and may store subscription data of UE 2902. In at least one embodiment, UDM 2924 may comprise two parts: an application FE and a user data repository (UDR). In at least one embodiment, the UDM may comprise a UDM FE responsible for credential processing, location management, subscription management, etc. In at least one embodiment, multiple different front ends may serve the same user for different transactions. In at least one embodiment, the UDM FE accesses the subscription information stored in a UDR and performs authentication credential processing, user identification handling, access authorization, registration / mobility management, and subscription management.In at least one embodiment, the UDR may interact with the PCF 2922. In at least one embodiment, the UDM 2924 may also support SMS management, with an SMS FE implementing similar application logic as previously described.
[0235] In at least one embodiment, AF 2926 may enable application influence on traffic routing, access to a Network Capability Exposure (NCE), and interaction with a policy framework for policy control. In at least one embodiment, NCE may be a mechanism that enables a 5GC and AF 2926 to communicate information with each other via NEF 2916, which may be used for edge computing implementations. In at least one embodiment, network operator and third-party services may be hosted near the point of attachment of UE 2902 to achieve efficient service provisioning through lower end-to-end latency and load on the transport network. In at least one embodiment, for edge computing implementations, 5GC may select a UPF 2904 near UE 2902 and perform traffic routing from UPF 2904 to DN 2906 via the N6 interface.In at least one embodiment, this may be done based on the UE subscription data, the UE location, and the information provided by AF 2926. In at least one embodiment, AF 2926 may influence UPF (re)selection and traffic routing. In at least one embodiment, based on the operator's deployment, if AF 2926 is considered a trusted entity, a network operator may allow AF 2926 to directly switch with relevant NFs.
[0236] In at least one embodiment, CN 2910 may include an SMSF, which may be responsible for checking and verifying SMS subscriptions and forwarding SM messages to / from UE 2902 to / from other entities, such as an SMS GMSC / IWMSC / SMS router. In at least one embodiment, SMS may also interact with AMF 2912 and UDM 2924 for the notification procedure that UE 2902 is available for SMS transmission (e.g., setting a UE unreachable flag and notifying UDM 2924 when UE 2902 is available for SMS).
[0237] In at least one embodiment, system 2900 may include the following service-based interfaces: Namf: Service-based interface issued by AMF; Nsmf: Service-based interface issued by SMF; Nnef: Service-based interface issued by NEF; Npcf: Service-based interface issued by PCF; Nudm: Service-based interface issued by UDM; Naf: Service-based interface issued by AF; Nnrf: Service-based interface issued by NRF; and Nausf: Service-based interface represented by AUSF.
[0238] In at least one embodiment, system 2900 may include the following reference points: N1: Reference point between UE and AMF; N2: Reference point between (R)AN and AMF; N3: Reference point between (R)AN and UPF; N4: Reference point between SMF and UPF; and N6: Reference point between UPF and a data network. In at least one embodiment, there may be many other reference points and / or service-based interfaces between NF services in NFs; however, these interfaces and reference points have been omitted for clarity. In at least one embodiment, an NS reference point may be between a PCF and an AF; an N7 reference point may be between PCF and SMF; an N11 reference point between AMF and SMF; etc. In at least one embodiment, CN 2910 may include an Nx interface that is an inter-CN interface between MME and AMF 2912 to enable interworking between CN 2910 and CN 7229.
[0239] In at least one embodiment, system 2900 may include multiple RAN nodes (such as (R)AN node 2908), wherein an Xn interface is defined between two or more (R)AN nodes 2908 (e.g., gNBs) connecting to 5GC 410, between an (R)AN node 2908 (e.g., gNB) connecting to CN 2910 and an eNB (e.g., a macro RAN node), and / or between two eNBs connecting to CN 2910.
[0240] In at least one embodiment, the Xn interface may comprise an Xn user plane interface (Xn-U) and an Xn control plane interface (Xn-C). In at least one embodiment, Xn-U may provide non-guaranteed delivery of user plane PDUs and support / provide data forwarding and flow control functions. In at least one embodiment, Xn-C may provide management and error handling functions, functions for managing an Xn-C interface, mobility support for UE 2902 in a connected mode (e.g., CM-CONNECTED), including functions for managing UE mobility for the connected mode between one or more (R)AN nodes 2908.In at least one embodiment, mobility support may include context transfer from an old (source) serving (R)AN node 2908 to a new (destination) serving (R)AN node 2908; and controlling user-plane tunnels between the old (source) serving (R)AN node 2908 and the new (destination) serving (R)AN node 2908.
[0241] In at least one embodiment, an Xn-U protocol stack may include a network layer built on top of the Internet Protocol (IP) transport layer and a GTP-U layer on top of one or more UDP and / or IP layers to carry user plane PDUs. In at least one embodiment, the Xn-C protocol stack may include an application layer signaling protocol (referred to as Xn Application Protocol (XnAP)) and a transport network layer built on top of an SCTP layer. In at least one embodiment, the SCTP layer may be positioned above an IP layer. In at least one embodiment, the SCTP layer provides guaranteed delivery of application layer messages. In at least one embodiment, a point-to-point transmission is used in an IP transport layer to carry signaling PDUs.In at least one embodiment, an Xn-U protocol stack and / or an Xn-C protocol stack may be the same or similar to a user plane and / or control plane protocol stack(s) shown and described herein.
[0242] In at least one embodiment, at least one in Fig. 29 is used to implement techniques and / or functions associated with Fig. 1-13. In at least one embodiment, at least one component of Fig. 29 to effect scaling of one or more clocks of one or more cores based at least in part on the proximity of the one or more cores to each other. In at least one embodiment, at least one component of Fig. 29 at least one aspect relating to the data center 120, the controller 105, at least two or more of processor 110(a) to processor 110(n) and one or more thermal sensors 115(a) to thermal sensor 115(n) of Fig. 1, Firmware 220, Configuration file 225, ACPI 230, Policy 235 one or more processors 1-n of Fig. 2, and / or comprising one or more of: Method 900 of Fig. 9, Procedure 1000 of Fig. 10, and Procedure 1100 of Fig. 11.
[0243] Fig. Figure 30 illustrates a control plane protocol stack according to some embodiments. In at least one embodiment, a control plane 3000 is depicted as a communication protocol stack between UE 2802 (or alternatively, UE 2804), RAN 2816, and MME(s) 2828.
[0244] In at least one embodiment, PHY layer 3002 may transmit or receive information used by MAC layer 3004 over one or more air interfaces. In at least one embodiment, PHY layer 3002 may further perform link adaptation or adaptive modulation and coding (AMC), power control, cell search (e.g., for initial synchronization and handover purposes), and other measurements used by higher layers, such as RRC layer 3010. In at least one embodiment, PHY layer 3002 may further perform transport channel error detection, forward error correction (FEC), transport channel encoding / decoding, physical channel modulation / demodulation, interleaving, rate adaptation, physical channel mapping, and multiple-input multiple-output (MIMO) antenna processing.
[0245] In at least one embodiment, the MAC layer 3004 may perform mapping between logical channels and transport channels, multiplexing MAC service data units (SDUs) from one or more logical channels to transport blocks (TBs) to be delivered to the PHY over transport channels, demultiplexing MAC SDUs onto one or more logical channels from transport blocks (TBs) delivered by the PHY over transport channels, multiplexing MAC SDUs onto TBs, reporting scheduling information, hybrid automatic retry request (HARD) error correction, and prioritizing logical channels.
[0246] In at least one embodiment, the RLC layer 3006 may operate in a variety of operating modes, including Transparent Mode (TM), Unacknowledged Mode (UM), and Acknowledged Mode (AM). In at least one embodiment, the RLC layer 3006 may perform transmission of upper-layer protocol data units (PDUs), automatic repeat request (ARQ) error correction for AM data transmissions, and concatenation, segmentation, and reassembly of RLC SDUs for UM and AM data transmissions. In at least one embodiment, the RLC layer 3006 may also perform resegmentation of RLC data PDUs for AM data transmissions, reorder RLC data PDUs for UM and AM data transmissions, detect duplicate data for UM and AM data transmissions, discard RLC SDUs for UM and AM data transmissions, detect protocol errors for AM data transmissions, and perform RLC reconstruction.
[0247] In at least one embodiment, the PDCP layer 3008 may perform header compression and decompression of IP data, maintain PDCP sequence numbers (SNs), perform sequence-accurate delivery of upper layer PDUs during lower layer recovery, eliminate duplicate lower layer SDUs during lower layer recovery for radio bearers mapped to RLC AM, encrypt and decrypt control plane data, integrity protection and integrity checking of control plane data, control timer-based data discard control, and perform security operations (zg, encryption, decryption, integrity protection, integrity checking, etc.).
[0248] In at least one embodiment, the main services and functions of an RRC layer 3010 may include the transmission of system information (e.g., contained in Master Information Blocks (MIBs) or System Information Blocks (SIBs) related to a Non-Access Layer (NAS)), the transmission of system information related to an Access Layer (AS), paging, establishment, maintenance, and teardown of an RRC connection between a UE and E-UTRAN (e.g., RRC connection paging, RRC connection setup, RRC connection modification, and RRC connection release), establishment, configuration, maintenance, and release of point-to-point radio bearers, security functions including key management, mobility between Radio Access Technologies (RAT), and measurement configuration for UE measurement reports.In at least one embodiment, the MIBs and SIBs may include one or more information elements (IEs), each of which may include individual data fields or data structures.
[0249] In at least one embodiment, UE 2802 and RAN 2816 may use a Uu interface (e.g., an LTE Uu interface) to exchange control plane data via a protocol stack including PHY layer 3002, MAC layer 3004, RLC layer 3006, PDCP layer 3008, and RRC layer 3010.
[0250] In at least one embodiment, non-access layer (NAS) protocols (NAS protocols 3012) form a top-level control plane between UE 2802 and MME(s) 2828. In at least one embodiment, NAS protocols 3012 support mobility of UE 2802 and session management methods for establishing and maintaining IP connectivity between UE 2802 and P-GW 2834.
[0251] In at least one embodiment, the Si Application Protocol (S1-AP) layer (Si-AP layer 3022) may support functions of an Si interface and include elementary procedures (EPs). In at least one embodiment, an EP is a unit of interaction between RAN 2816 and CN 2828. In at least one embodiment, the S1-AP layer services may comprise two groups: UE-associated services and non-UE-associated services. In at least one embodiment, these services perform functions including, but not limited to, E-UTRAN Radio Access Bearer (E-RAB) management, UE capability indication, mobility, NAS signal transport, RAN Information Management (RIM), and configuration transfer.
[0252] In at least one embodiment, the Stream Control Transmission Protocol (SCTP) layer (alternatively referred to as Stream Control Transmission Protocol / Internet Protocol (SCTP / IP) layer) (SCTP layer 3020) may ensure reliable delivery of signaling messages between RAN 2816 and MME(s) 2828 based, at least in part, on an IP protocol supported by an IP layer 3018. In at least one embodiment, the L2 layer 3016 and an L1 layer 3014 may refer to communication links (e.g., wired or wireless) used by a RAN node and an MME to exchange information.
[0253] In at least one embodiment, RAN 2816 and MME(s) 2828 may use an S1 MME interface to exchange control plane data over a protocol stack including L1 layer 3014, L2 layer 3016, IP layer 3018, SCTP layer 3020, and Si-AP layer 3022.
[0254] In at least one embodiment, at least one in Fig. 30 is used to implement techniques and / or functions associated with Fig. 1-13. In at least one embodiment, at least one component of Fig. 30 to effect scaling of one or more clocks of one or more cores based at least in part on the proximity of the one or more cores to each other. In at least one embodiment, at least one component of Fig. 30 at least one aspect relating to the data center 120, the controller 105, at least two or more of processor 110(a) to processor 110(n) and one or more thermal sensors 115(a) to thermal sensor 115(n) of Fig. 1, Firmware 220, Configuration file 225, ACPI 230, Policy 235 one or more processors 1-n of Fig. 2, and / or comprising one or more of: Method 900 of Fig. 9, Procedure 1000 of Fig. 10, and Procedure 1100 of Fig. 11.
[0255] Fig. 31 shows a representation of a user plane protocol architecture according to at least one embodiment. In at least one embodiment, a user plane 3100 is depicted as a communication protocol stack between a UE 2802, RAN 2816, S-GW 2830, and P-GW 2834. In at least one embodiment, the user plane 3100 may use the same protocol layers as the control plane 3000. In at least one embodiment, the UE 2802 and RAN 2816 may, for example, use a Uu interface (e.g., an LTE Uu interface) to exchange user plane data over a protocol stack including the PHY layer 3002, the MAC layer 3004, the RLC layer 3006, and the PDCP layer 3008.
[0256] In at least one embodiment, the General Packet Radio Service (GPRS) Tunneling Protocol for a User Plane (GTP-U) layer (GTP-U layer 3104) may be used to transmit user data within a GPRS core network and between a radio access network and a core network. In at least one embodiment, the transported payload may be packets in one of the IPv4, IPv6, or PPP formats. In at least one embodiment, the UDP and IP security layer (UDP / IP layer 3102) may provide checksums for data integrity, port numbers for addressing different functions at a source and destination, and encryption and authentication for selected data streams.In at least one embodiment, RAN 2816 and S-GW 2830 may use an S1-U interface to exchange user plane data over a protocol stack including L1 layer 3014, L2 layer 3016, UDP / IP layer 3102, and GTPU layer 3104. In at least one embodiment, S-GW 2830 and P-GW 2834 may use an S5 / S8a interface to exchange user plane data over a protocol stack including L1 layer 3014, L2 layer 3016, UDP / IP layer 3102, and GTPU layer 3104. In at least one embodiment, as described above with respect to FIG. Fig. 30, NAS protocols support mobility of the UE 2802 and session management procedures to establish and maintain IP connectivity between the UE 2802 and the P-GW 2834.
[0257] In at least one embodiment, at least one in Fig. 31 is used to implement techniques and / or functions associated with Fig. 1-13. In at least one embodiment, at least one component of Fig. 31 to effect scaling of one or more clocks of one or more cores based at least in part on the proximity of the one or more cores to each other. In at least one embodiment, at least one component of Fig. 31 at least one aspect relating to the data center 120, the controller 105, at least two or more of processor 110(a) to processor 110(n) and one or more thermal sensors 115(a) to thermal sensor 115(n) of Fig. 1, Firmware 220, Configuration file 225, ACPI 230, Policy 235 one or more processors 1-n of Fig. 2, and / or comprising one or more of: Method 900 of Fig. 9, Procedure 1000 of Fig. 10 and Procedure 1100 of Fig. 11.
[0258] Fig. 32 shows components 3200 of a core network in accordance with at least one embodiment. In at least one embodiment, the components of CN 2838 may be implemented in one physical node or in separate physical nodes that include components for reading and executing instructions from a machine-readable or computer-readable medium (e.g., a non-transitory machine-readable storage medium). In at least one embodiment, network function virtualization (NFV) is used to virtualize any or all of the network node functions described above via executable instructions stored in one or more computer-readable storage media (described in detail below). In at least one embodiment, a logical instantiation of CN 2838 may be referred to as a network slice 3202 (e.g., network slice 3202 includes HSS 2832, MME(s) 2828, and S-GW 2830).In at least one embodiment, a logical instantiation of a portion of CN 2838 may be referred to as network slice 3204 (e.g., network slice 3204 includes P-GW 2834 and PCRF 2836).
[0259] In at least one embodiment, NFV architectures and infrastructures can be used to virtualize one or more network functions, alternatively performed by proprietary hardware, on physical resources that include a combination of industry-standard server hardware, storage hardware, or switches. In at least one embodiment, NFV systems can be used to execute virtual or reconfigurable implementations of one or more EPC components / functions.
[0260] In at least one embodiment, at least one in Fig. 32 is used to implement techniques and / or functions associated with Fig. 1-13. In at least one embodiment, at least one component of Fig. 32 to effect scaling of one or more clocks of one or more cores based at least in part on the proximity of the one or more cores to each other. In at least one embodiment, at least one component of Fig. 32 at least one aspect relating to the data center 120, the controller 105, at least two or more of processor 110(a) to processor 110(n) and one or more thermal sensors 115(a) to thermal sensor 115(n) of Fig. 1, Firmware 220, Configuration file 225, ACPI 230, Policy 235 one or more processors 1-n of Fig. 2, and / or comprising one or more of: Method 900 of Fig. 9, Procedure 1000 of Fig. 10, and Procedure 1100 of Fig. 11.
[0261] Fig. 33 is a block diagram illustrating components according to at least one embodiment of a system 3300 for supporting network functions virtualization (NFV). In at least one embodiment, system 3300 is shown to include a virtualized infrastructure manager (illustrated as VIM 3302), a network functions virtualization infrastructure (illustrated as NFVI 3304), a VNF manager (illustrated as VNFM 3306), virtualized network functions (illustrated as VNF 3308), an element manager (illustrated as EM 3310), an NFV orchestrator (illustrated as NFVO 3312), and a network manager (illustrated as NM 3314).
[0262] In at least one embodiment, VIM 3302 manages the resources of NFVI 3304. In at least one embodiment, NFVI 3304 may include physical or virtual resources and applications (including hypervisors) used to execute system 3300. In at least one embodiment, VIM 3302 may manage a lifecycle of virtual resources with NFVI 3304 (e.g., creation, maintenance, and teardown of virtual machines (VMs) associated with one or more physical resources), track VM instances, track performance, faults, and security of VM instances and associated physical resources, and expose VM instances and associated physical resources to other management systems.
[0263] In at least one embodiment, VNFM 3306 may manage VNF 3308. In at least one embodiment, VNF 3308 may be used to execute EPC components / functions. In at least one embodiment, VNFM 3306 may manage a lifecycle of VNF 3308 and track performance, faults, and security of the virtual aspects of VNF 3308. In at least one embodiment, EM 3310 may track performance, faults, and security of functional aspects of VNF 3308. In at least one embodiment, the tracking data from VNFM 3306 and EM 3310 may include, for example, performance metrics (PM) used by VIM 3302 or NFVI 3304. In at least one embodiment, both VNFM 3306 and EM 3310 may scale up / down a set of VNFs of system 3300.
[0264] In at least one embodiment, NFVO 3312 may coordinate, authorize, release, and consume resources from NFVI 3304 to provide a requested service (e.g., to execute an EPC function, component, or slice). In at least one embodiment, NM 3314 may provide a suite of end-user functions responsible for managing a network, which may include network elements with VNFs, non-virtualized network functions, or both (management of VNFs may be performed via an EM 3310).
[0265] In at least one embodiment, at least one in Fig. 33 is used to implement techniques and / or functions associated with Fig. 1-13. In at least one embodiment, at least one component of Fig. 33 to effect scaling of one or more clocks of one or more cores based at least in part on the proximity of the one or more cores to each other. In at least one embodiment, at least one component of Fig. 33 at least one aspect relating to the data center 120, the controller 105, at least two or more of processor 110(a) to processor 110(n) and one or more thermal sensors 115(a) to thermal sensor 115(n) of Fig. 1, Firmware 220, Configuration file 225, ACPI 230, Policy 235 one or more processors 1-n of Fig. 2, and / or comprising one or more of: Method 900 of Fig. 9, Procedure 1000 of Fig. 10, and Procedure 1100 of Fig. 11. Computer-aided systems
[0266] The following figures illustrate, without limitation, exemplary computer systems that may be used to implement at least one embodiment.
[0267] Fig. 34 shows a processing system 3400 in accordance with at least one embodiment. In at least one embodiment, the processing system 3400 includes one or more processors 3402 and one or more graphics processors 3408 and may be a single-processor desktop system, a multiprocessor workstation system, or a server system with a large number of processors 3402 or processor cores 3407. In at least one embodiment, the processing system 3400 is a processing platform integrated into a system-on-a-chip ("SoC") integrated circuit for use in mobile, wearable, or embedded devices.
[0268] In at least one embodiment, processing system 3400 may include or be integrated with a server-based gaming platform, a game console, a media console, a mobile game console, a handheld game console, or an online game console. In at least one embodiment, processing system 3400 is a mobile phone, a smartphone, a tablet computing device, or a mobile internet device. In at least one embodiment, processing device 3400 may also include, be coupled to, or integrated with a wearable device, such as a wearable device, a smart watch, smart glasses, an augmented reality device, or a virtual reality device.In at least one embodiment, processing system 3400 is a device for a television or set-top box that includes one or more processors 3402 and a graphical interface generated by one or more graphics processors 3408.
[0269] In at least one embodiment, one or more processors 3402 each include one or more processor cores 3407 for processing instructions that, when executed, perform operations for system and user software. In at least one embodiment, each of one or more processor cores 3407 is configured to process a particular instruction set 3409. In at least one embodiment, the instruction set 3409 may enable Complex Instruction Set Computing ("CISC"), Reduced Instruction Set Computing ("RISC"), or Very Long Instruction Word ("VLIW") computing. In at least one embodiment, the processor cores 3407 may each process a different instruction set 3409, which may include instructions that facilitate emulation of other instruction sets.In at least one embodiment, the processor core 3407 may also include other processing devices, such as a digital signal processor ("DSP").
[0270] In at least one embodiment, processor 3402 includes a cache 3404. In at least one embodiment, processor 3402 may have a single internal cache or multiple levels of internal cache. In at least one embodiment, the cache is shared by various components of processor 3402. In at least one embodiment, processor 3402 also uses an external cache (e.g., a Level 3 ("L3") cache or Last Level Cache ("LLC")) (not shown), which may be shared by processor cores 3407 using known cache coherence techniques. In at least one embodiment, processor 3402 additionally includes a register file 3406, which may include different types of registers for storing different types of data (e.g., integer registers, floating-point registers, status registers, and an instruction pointer register).In at least one embodiment, register file 3406 may include general-purpose registers or other registers.
[0271] In at least one embodiment, one or more processors 3402 are coupled to one or more interface buses 3410 to communicate communication signals, such as address, data, or control signals, between processor 3402 and other components in processing system 3400. In at least one embodiment, interface bus 3410 may be a processor bus, such as a version of a Direct Media Interface ("DMI" bus). In at least one embodiment, interface bus 3410 is not limited to a DMI bus and may include one or more Peripheral Component Interconnect (e.g., "PCI," PCI Express ("PCIe")) buses, memory interconnects, or other types of interface buses. In at least one embodiment, processor(s) 3402 include an integrated memory controller 3416 and a memory controller hub 3430.In at least one embodiment, the memory controller 3416 facilitates communication between a memory device and other components of the processing system 3400, while the platform control hub ("PCH") 3430 provides connections to input / output ("I / O") devices via a local I / O bus.
[0272] In at least one embodiment, memory device 3420 may be a dynamic random access memory ("DRAM") device, a static random access memory ("SRAM") device, a flash memory device, a phase-change memory device, or other memory device having suitable performance to serve as processor memory. In at least one embodiment, memory device 3420 may function as system memory for processing system 3400 to store data 3422 and instructions 3421 for use when one or more processors 3402 are executing an application or process. In at least one embodiment, memory controller 3416 is also coupled to an optional external graphics processor 3412 that can communicate with one or more graphics processors 3408 in processors 3402 to perform graphics and media operations.In at least one embodiment, a display device 3411 may be connected to the processor(s) 3402. In at least one embodiment, the display device 3411 may include one or more internal display devices, such as in a mobile electronic device or laptop, or an external display device connected via an interface (e.g., DisplayPort, etc.). In at least one embodiment, the display device 3411 may include a head-mounted display ("HMD"), such as a stereoscopic display device for use in virtual reality ("VR") or augmented reality ("AR") applications.
[0273] In at least one embodiment, the storage control hub 3430 enables the connection of peripherals to the storage device 3420 and the processor 3402 via a high-speed I / O bus. In at least one embodiment, the I / O peripherals include, among others, an audio controller 3446, a network interface 3434, a firmware interface 3428, a wireless transceiver 3426, touch sensors 3425, and a storage device 3424 (e.g., hard disk drive, flash memory, etc.). In at least one embodiment, the data storage device 3424 can be connected via a storage interface (e.g., SATA) or via a peripheral bus such as PCI or PCIe. In at least one embodiment, the touch sensors 3425 can include touchscreen sensors, pressure sensors, or fingerprint sensors.In at least one embodiment, wireless transceiver 3426 may be a Wi-Fi transceiver, a Bluetooth transceiver, or a cellular network transceiver such as a 3G, 4G, or Long Term Evolution ("LTE") transceiver. In at least one embodiment, firmware interface 3428 enables communication with system firmware and may, for example, be a Unified Extensible Firmware Interface ("UEFI"). In at least one embodiment, network controller 3434 may enable network connection to a wired network. In at least one embodiment, a high-performance network controller (not shown) is coupled to interface bus 3410. In at least one embodiment, audio controller 3446 is a multi-channel high-definition audio controller.In at least one embodiment, the processing system 3400 includes an optional I / O controller 3440 for connecting legacy devices (e.g., Personal System 2 ("PS / 2")) to the processing system 3400. In at least one embodiment, the platform control hub 3430 may also be connected to one or more Universal Serial Bus ("USB") controllers 3442 that connect input devices, such as keyboard and mouse combinations 3443, a camera 3444, or other USB input devices.
[0274] In at least one embodiment, an instance of the memory controller 3416 and the memory control hub 3430 may be integrated into a discrete external graphics processor, such as the external graphics processor 3412. In at least one embodiment, the platform control hub 3430 and / or the memory controller 3416 may be external to one or more processors 3402. In at least one embodiment, the processing system 3400 may, for example, include an external memory controller 3416 and a platform control hub 3430, which may be configured as a memory control hub and a peripheral control hub within a system chipset in communication with the processor(s) 3402. In at least one embodiment, at least one in Fig. 34 is used to implement techniques and / or functions associated with Fig. 1-13. In at least one embodiment, at least one component of Fig. 34 to effect scaling of one or more clocks of one or more cores based at least in part on the proximity of the one or more cores to each other. In at least one embodiment, at least one component of Fig. 34 at least one aspect relating to the data center 120, the controller 105, at least two or more of processor 110(a) to processor 110(n) and one or more thermal sensors 115(a) to thermal sensor 115(n) of Fig. 1, Firmware 220, Configuration file 225, ACPI 230, Policy 235 one or more processors 1-n of Fig. 2, and / or comprising one or more of: Method 900 of Fig. 9, Procedure 1000 of Fig. 10 and Procedure 1100 of Fig. 11.
[0275] Fig. 35 illustrates a computer system 3500 in accordance with at least one embodiment. In at least one embodiment, computer system 3500 may be a system of interconnected devices and components, a System of Operation (SOC), or a combination thereof. In at least one embodiment, computer system 3500 is configured with a processor 3502, which may include execution units for executing an instruction. In at least one embodiment, computer system 3500 may include, without limitation, a component such as processor 3502 for employing execution units including logic for performing algorithms for processing data.In at least one embodiment, computer system 3500 may include processors such as the PENTIUM® processor family, Xeon™, Itanium®, XScale™ and / or StrongARM™, Intel® Core™, or Intel® Nervana™ microprocessors available from Intel Corporation of Santa Clara, California, although other systems (including personal computers with other microprocessors, technical workstations, set-top boxes, and the like) may be used. In at least one embodiment, computer system 3500 may run a version of the WINDOWS operating system available from Microsoft Corporation of Redmond, Washington, although other operating systems (e.g., UNIX and Linux), embedding software, and / or graphical user interfaces may also be used.
[0276] In at least one embodiment, computer system 3500 may be used in other devices such as handheld devices and embedded applications. Some examples of portable devices include cellular phones, Internet Protocol devices, digital cameras, personal digital assistants ("PDAs"), and handheld PCs. In at least one embodiment, embedded applications may include a microcontroller, a digital signal processor (DSP), an SoC, network computers ("NetPCs"), set-top boxes, network hubs, wide area network ("WAN") switches, or any other system capable of executing one or more instructions.
[0277] In at least one embodiment, computer system 3500 may include, without limitation, a processor 3502, which may include, without limitation, one or more execution units 3508 that may be configured to execute a Compute Unified Device Architecture ("CUDA") program (CUDA® is developed by NVIDIA Corporation of Santa Clara, CA). In at least one embodiment, a CUDA program is at least a portion of a software application written in a CUDA programming language. In at least one embodiment, computer system 3500 is a desktop or server system having a processor. In at least one embodiment, computer system 3500 may be a multiprocessor system.In at least one embodiment, processor 3502 may include, without limitation, a CISC microprocessor, a RISC microprocessor, a VLIW microprocessor, a processor implementing a combination of instruction sets, or any other device, such as a digital signal processor. In at least one embodiment, processor 3502 may be connected to a processor bus 3510 that may transmit data signals between processor 3502 and other components of computer system 3500.
[0278] In at least one embodiment, processor 3502 may include, without limitation, an internal Level 1 ("L1") cache memory ("cache") 3504. In at least one embodiment, processor 3502 may have a single internal cache or multiple levels of internal cache. In at least one embodiment, the cache memory may be external to processor 3502. In at least one embodiment, processor 3502 may also include a combination of both internal and external caches. In at least one embodiment, a register file 3506 may store different data types in various registers, including, without limitation, integer registers, floating-point registers, status registers, and instruction pointer registers.
[0279] In at least one embodiment, execution unit 3508, which includes, without limitation, logic for performing integer and floating-point operations, is also located in processor 3502. Processor 3502 may also include microcode read-only memory ("ROM") ("ucode") that stores microcode for certain macroinstructions. In at least one embodiment, execution unit 3508 may include logic for handling a packed instruction set 3509. In at least one embodiment, by including packed instruction set 3509 in an instruction set of a general-purpose processor 3502, along with associated instruction execution circuitry, operations used by many multimedia applications may be performed using packed data in a general-purpose processor 3502.In at least one embodiment, many multimedia applications can be accelerated and executed more efficiently by utilizing the full width of a processor's data bus to perform operations on packed data, thereby eliminating the need to transfer smaller units of data across a processor's data bus to perform one or more operations on one data element at a time.
[0280] In at least one embodiment, execution unit 3508 may also be used in microcontrollers, embedded processors, graphics devices, DSPs, and other types of logic circuits. In at least one embodiment, computer system 3500 may include, without limitation, a memory 3520. In at least one embodiment, memory 3520 may be implemented as a DRAM device, an SRAM device, a flash memory device, or other storage device. Memory 3520 may store instructions 3519 and / or data 3521 represented by data signals that may be executed by processor 3502.
[0281] In at least one embodiment, a system logic chip may be connected to the processor bus 3510 and the memory 3520. In at least one embodiment, a system logic chip may include, without limitation, a memory control hub ("MCH") 3516, and the processor 3502 may communicate with the MCH 3516 via the processor bus 3510. In at least one embodiment, the MCH 3516 may provide a high-bandwidth memory path 3518 to the memory 3520 for storing instructions and data, as well as for storing graphics instructions, data, and textures. In at least one embodiment, the MCH 3516 may command data signals between the processor 3502, the memory 3520, and other components of the computer system 3500, and may bridge data signals between the processor bus 3510, the memory 3520, and a system I / O 3522.In at least one embodiment, the system logic chip may provide a graphics port for connection to a graphics controller. In at least one embodiment, the MCH 3516 may be coupled to memory 3520 via a high-bandwidth memory path 3518, and the graphics / video card 3512 may be coupled to the MCH 3516 via an Accelerated Graphics Port ("AGP") interconnect 3514.
[0282] In at least one embodiment, computer system 3500 may use system I / O 3522, which is a proprietary hub interface, to connect MCH 3516 to the I / O controller ("ICH") hub 3530. In at least one embodiment, ICH 3530 may provide direct connections to some I / O devices via a local I / O bus. In at least one embodiment, the local I / O bus may include, without limitation, a high-speed I / O bus for connecting peripherals to memory 3520, a chipset, and processor 3502. Examples may include, without limitation, an audio controller 3529, a firmware hub (“flash BIOS”) 3528, a wireless transceiver 3526, a data store 3524, a legacy I / O controller 3523 with a user input interface 3525 and a keyboard interface, a serial expansion port 3527, such as a USB, and a network controller 3534.The data storage 3524 may include a hard disk drive, a floppy disk drive, a CD-ROM device, a flash memory device, or other mass storage device.
[0283] In at least one embodiment, Fig. 35 a system comprising interconnected hardware devices or "chips." In at least one embodiment, Fig. 35 show an exemplary SoC. In at least one embodiment, the Fig. 35 may be interconnected using proprietary interconnects, standardized interconnects (e.g., PCIe), or a combination thereof. In at least one embodiment, one or more components of system 3500 are interconnected using Compute Express Link ("CXL") interconnects.
[0284] In at least one embodiment, at least one in Fig. 35 is used to implement techniques and / or functions associated with Fig. 1-13. In at least one embodiment, at least one component of Fig. 35 to effect scaling of one or more clocks of one or more cores based at least in part on the proximity of the one or more cores to each other. In at least one embodiment, at least one component of Fig. 35 at least one aspect relating to the data center 120, the controller 105, at least two or more of processor 110(a) to processor 110(n) and one or more thermal sensors 115(a) to thermal sensor 115(n) of Fig. 1, Firmware 220, Configuration file 225, ACPI 230, Policy 235 one or more processors 1-n of Fig. 2, and / or comprising one or more of: Method 900 of Fig. 9, Procedure 1000 of Fig. 10 and Procedure 1100 of Fig. 11.
[0285] Fig. 36 shows a system 3600 in accordance with at least one embodiment. In at least one embodiment, the system 3600 is an electronic device that uses a processor 3610. In at least one embodiment, the system 3600 may be, for example and without limitation, a notebook, a tower server, a rack server, a blade server, a laptop, a desktop, a tablet, a mobile device, a phone, an embedded computer, or any other suitable electronic device.
[0286] In at least one embodiment, system 3600 may include, without limitation, a processor 3610 communicatively coupled to any number or type of components, peripherals, modules, or devices. In at least one embodiment, processor 3610 is coupled via a bus or interface, such as an I2C bus, a System Management Bus ("SMBus"), a Low Pin Count Bus ("LPC"), a Serial Peripheral Interface ("SPI"), a High Definition Audio Bus ("HDA"), a Serial Advance Technology Attachment Bus ("SATA"), a USB bus (versions 1, 2, 3), or a Universal Asynchronous Receiver / Transmitter Bus ("UART"). In at least one embodiment, Fig. 36 a system comprising interconnected hardware devices or "chips." In at least one embodiment, Fig. 36 show an exemplary SoC. In at least one embodiment, the Fig. 36 may be interconnected with proprietary interconnects, standardized interconnects (e.g., PCIe), or a combination thereof. In at least one embodiment, one or more components of Fig. 36 interconnected using CXL connections.
[0287] In at least one embodiment, Fig. 36 a display 3624, a touchscreen 3625, a touchpad 3630, a near-field communication unit (“NFC”) 3645, a sensor hub 3640, a thermal sensor 3646, an express chipset (“EC”) 3635, a trusted platform module (“TPM”) 3638, BIOS / firmware / flash memory (“BIOS, FW Flash”) 3622, a DSP 3660, a solid-state disk (“SSD”) or hard disk (“HDD”) 3620, a wireless local area network unit (“WLAN”) 3650, a Bluetooth unit 3652, a wireless wide area network unit (“WWAN”) 3656, a global positioning system (“GPS”) 3655, a camera (“USB 3.0 camera”) 3654, such as a USB 3.0 camera, or a low-power double data rate ("LPDDR") memory unit ("LPDDR3") 3615, for example, implemented in the LPDDR3 standard. These components may each be implemented in any suitable manner.
[0288] In at least one embodiment, other components may be communicatively coupled to the processor 3610 via the components described above. In at least one embodiment, an accelerometer 3641, an ambient light sensor ("ALS") 3642, a compass 3643, and a gyroscope 3644 may be communicatively coupled to the sensor hub 3640. In at least one embodiment, a thermal sensor 3639, a fan 3637, a keyboard 3646, and a touchpad 3630 may be communicatively coupled to the EC 3635. In at least one embodiment, a speaker 3663, a headset 3664, and a microphone ("mic") 3665 may be communicatively coupled to an audio unit ("audio codec and class d amp") 3664, which in turn may be communicatively coupled to the DSP 3660. In at least one embodiment, the audio unit 3664 may include, for example and without limitation, an audio encoder / decoder ("codec") and a Class D amplifier.In at least one embodiment, a SIM card ("SIM") 3657 may be communicatively coupled to the WWAN unit 3656. In at least one embodiment, components such as the WLAN unit 3650 and the Bluetooth unit 3652, as well as the WWAN unit 3656, may be implemented in a Next Generation Form Factor ("NGFF").
[0289] In at least one embodiment, at least one in Fig. 36 is used to implement techniques and / or functions associated with Fig. 1-13. In at least one embodiment, at least one component of Fig. 36 to effect scaling of one or more clocks of one or more cores based at least in part on the proximity of the one or more cores to each other. In at least one embodiment, at least one component of Fig. 36 at least one aspect relating to the data center 120, the controller 105, at least two or more of processor 110(a) to processor 110(n) and one or more thermal sensors 115(a) to thermal sensor 115(n) of Fig. 1, Firmware 220, Configuration file 225, ACPI 230, Policy 235 one or more processors 1-n of Fig. 2, and / or comprising one or more of: Method 900 of Fig. 9, Procedure 1000 of Fig. 10 and Procedure 1100 of Fig. 11.
[0290] Fig. 37 shows an example integrated circuit 3700 in accordance with at least one embodiment. In at least one embodiment, the example integrated circuit 3700 is a SoC that can be manufactured using one or more IP cores. In at least one embodiment, the integrated circuit 3700 includes one or more application processors 3705 (e.g., CPUs), at least one graphics processor 3710, and may additionally include an image processor 3715 and / or a video processor 3720, each of which may be a modular IP core. In at least one embodiment, the integrated circuit 3700 includes peripheral or bus logic, including a USB controller 3725, a UART controller 3730, an SPI / SDIO controller 3735, and an I2S / I2C controller 3740.In at least one embodiment, integrated circuit 3700 may include a display device 3745 coupled to one or more of the following interfaces: a High-Definition Multimedia Interface ("HDMI") controller 3750 and a Mobile Industry Processor Interface ("MIPI") display interface 3755. In at least one embodiment, memory may be provided by a flash memory subsystem 3760 comprising flash memory and a flash memory controller. In at least one embodiment, a memory interface may be provided via a memory controller 3765 for accessing SDRAM or SRAM devices. In at least one embodiment, some integrated circuits additionally include an embedded security engine 3770.
[0291] In at least one embodiment, at least one in Fig. 37 is used to implement techniques and / or functions associated with Fig. 1-13. In at least one embodiment, at least one component of Fig. 37 to effect scaling of one or more clocks of one or more cores based at least in part on the proximity of the one or more cores to each other. In at least one embodiment, at least one component of Fig. 37 at least one aspect relating to the data center 120, the controller 105, at least two or more of processor 110(a) to processor 110(n) and one or more thermal sensors 115(a) to thermal sensor 115(n) of Fig. 1, Firmware 220, Configuration file 225, ACPI 230, Policy 235 one or more processors 1-n of Fig. 2, and / or comprising one or more of: Method 900 of Fig. 9, Procedure 1000 of Fig. 10 and Procedure 1100 of Fig. 11.
[0292] Fig. 38 shows a computer system 3800 according to at least one embodiment; in at least one embodiment, the computer system 3800 includes a processing subsystem 3801 having one or more processors 3802 and a system memory 3804 communicating via an interconnect connection that may include a memory hub 3805. In at least one embodiment, the memory hub 3805 may be a separate component within a chipset component or integrated with one or more processors 3802. In at least one embodiment, the memory hub 3805 is coupled to an I / O subsystem 3811 via a communication connection 3806. In at least one embodiment, the I / O subsystem 3811 includes an I / O hub 3807 that may enable the computer system 3800 to receive input from one or more input devices 3808.In at least one embodiment, the I / O hub 3807 may enable a display controller, which may be included in one or more processors 3802, to provide outputs to one or more display devices 3810A. In at least one embodiment, one or more display devices 3810A coupled to the I / O hub 3807 may comprise a local, internal, or embedded display device.
[0293] In at least one embodiment, the processing subsystem 3801 includes one or more parallel processors 3812 connected to the storage hub 3805 via a bus or other communication link 3813. In at least one embodiment, the communication link 3813 may be any number of standards-based communication link technologies or protocols, such as, but not limited to, PCIe, or a vendor-specific communication interface or communication structure. In at least one embodiment, one or more parallel processors 3812 form a compute-intensive parallel or vector processing system, which may include a large number of processing cores and / or processing clusters, such as a processor with many integrated cores.In at least one embodiment, one or more parallel processors 3812 form a graphics processing subsystem that can output pixels to one or more display devices 3810A coupled via the I / O hub 3807. In at least one embodiment, one or more parallel processors 3812 may also include a display controller and a display interface (not shown) to enable direct connection to one or more display devices 3810B.
[0294] In at least one embodiment, a system storage unit 3814 may be connected to the I / O hub 3807 to provide a storage mechanism for the computer system 3800. In at least one embodiment, an I / O switch 3816 may be used to provide an interface enabling connections between the I / O hub 3807 and other components, such as a network adapter 3818 and / or a wireless network adapter 3819 that may be integrated into a platform, and various other devices that may be added via one or more add-in devices 3820. In at least one embodiment, the network adapter 3818 may be an Ethernet adapter or other wired network adapter.In at least one embodiment, the wireless network adapter 3819 may include one or more Wi-Fi, Bluetooth, NFC, or other network devices that include one or more wireless radios.
[0295] In at least one embodiment, computer system 3800 may include other components not explicitly shown, including USB or other connectors, optical storage devices, video capture devices, and / or variations thereof, which may also be connected to I / O hub 3807. In at least one embodiment, communication paths connecting various components in Fig. 38 interconnection may be implemented using any suitable protocols, such as PCI-based protocols (e.g., PCIe) or other bus or point-to-point communication interfaces and / or protocols, such as NVLink high-speed interconnection or interconnection protocols.
[0296] In at least one embodiment, one or more parallel processors 3812 include circuitry optimized for graphics and video processing, including, for example, video output circuitry, and form a graphics processing unit ("GPU"). In at least one embodiment, one or more parallel processors 3812 include circuitry optimized for general processing. In at least one embodiment, components of computer system 3800 may be integrated with one or more other system elements on a single integrated circuit. For example, in at least one embodiment, one or more parallel processors 3812, memory hub 3805, processor(s) 3802, and I / O hub 3807 may be integrated into an SoC integrated circuit.In at least one embodiment, the components of computer system 3800 may be integrated into a single package to form a system-in-package ("SIP") configuration. In at least one embodiment, at least a portion of the components of computer system 3800 may be integrated into a multi-chip module ("MCM") that may be interconnected with other multi-chip modules to form a modular computer system. In at least one embodiment, I / O subsystem 3811 and display devices 3810B are not included in computer system 3800.
[0297] In at least one embodiment, at least one in Fig. 38 is used to implement techniques and / or functions associated with Fig. 1-13. In at least one embodiment, at least one component of Fig. 38 to effect scaling of one or more clocks of one or more cores based at least in part on the proximity of the one or more cores to each other. In at least one embodiment, at least one component of Fig. 38 at least one aspect relating to the data center 120, the controller 105, at least two or more of processor 110(a) to processor 110(n) and one or more thermal sensors 115(a) to thermal sensor 115(n) of Fig. 1, Firmware 220, Configuration file 225, ACPI 230, Policy 235 one or more processors 1-n of Fig. 2, and / or comprising one or more of: Method 900 of Fig. 9, Procedure 1000 of Fig. 10 and Procedure 1100 of Fig. 11. Processing systems
[0298] The following figures illustrate, without limitation, exemplary processing systems that may be used to implement at least one embodiment.
[0299] Fig. 39 shows an accelerated processing unit ("APU") 3900 in accordance with at least one embodiment. In at least one embodiment, the APU 3900 is developed by AMD Corporation of Santa Clara, CA. In at least one embodiment, the APU 3900 may be configured to execute an application program, such as a CUDA program. In at least one embodiment, the APU 3900 includes, without limitation, a core complex 3910, a graphics complex 3940, a fabric 3960, I / O interfaces 3970, memory controllers 3980, a display controller 3992, and a multimedia engine 3994. In at least one embodiment, the APU 3900 may include, without limitation, any number of core complexes 3910, any number of graphics complexes 3940, any number of display controllers 3992, and any number of multimedia engines 3994 in any combination.For explanatory purposes, multiple instances of the same object are referred to here with reference numbers that identify an object and with parentheses that identify an instance where necessary.
[0300] In at least one embodiment, core complex 3910 is a CPU, graphics complex 3940 is a GPU, and APU 3900 is a processing unit that integrates, without limitation, 3910 and 3940 on a single chip. In at least one embodiment, some tasks may be assigned to core complex 3910 and other tasks to graphics complex 3940. In at least one embodiment, core complex 3910 is configured to execute main control software associated with APU 3900, such as an operating system. In at least one embodiment, core complex 3910 is a main processor of APU 3900 that controls and coordinates the operations of the other processors. In at least one embodiment, core complex 3910 issues instructions that control operation of graphics complex 3940.In at least one embodiment, core complex 3910 may be configured to execute host executable code derived from CUDA source code, and graphics complex 3940 may be configured to execute device executable code derived from CUDA source code.
[0301] In at least one embodiment, core complex 3910 includes, without limitation, cores 3920(1)-3920(4) and an L3 cache 3930. In at least one embodiment, core complex 3910 may include, without limitation, any number of cores 3920 and any number and type of caches in any combination. In at least one embodiment, cores 3920 are configured to execute instructions of a particular instruction set architecture ("ISA"). In at least one embodiment, each core 3920 is a CPU core.
[0302] In at least one embodiment, each core 3920 includes, without limitation, a fetch / decode unit 3922, an integer execution engine 3924, a floating-point execution engine 3926, and an L2 cache 3928. In at least one embodiment, the fetch unit 3922 fetches instructions, decodes such instructions, generates micro-operations, and sends separate micro-instructions to the integer execution engine 3924 and the floating-point execution engine 3926. In at least one embodiment, the fetch unit 3922 can concurrently send one micro-instruction to the integer execution engine 3924 and another micro-instruction to the floating-point execution engine 3926. In at least one embodiment, the integer execution engine 3924 performs, without limitation, integer and memory operations. In at least one embodiment, the floating-point engine 3926 performs, without limitation, floating-point and vector operations.In at least one embodiment, fetch unit 3922 forwards microinstructions to a single execution engine that replaces both integer execution engine 3924 and floating point execution engine 3926.
[0303] In at least one embodiment, each core 3920(i), where i is an integer representing a particular instance of core 3920, can access L2 cache 3928(i) that includes core 3920(i). In at least one embodiment, each core 3920 included in core complex 3910(j), where j is an integer representing a particular instance of core complex 3910, is connected to other cores 3920 included in core complex 3910(j) via L3 cache 3930(j) included in core complex 3910(j). In at least one embodiment, the cores 3920 included in core complex 3910(j), where j is an integer representing a particular instance of core complex 3910, may access the entire L3 cache 3930(j) included in core complex 3910(j). In at least one embodiment, L3 cache 3930 may include any number of slices, without limitation.
[0304] In at least one embodiment, graphics complex 3940 can be configured to perform computational operations in a highly parallel manner. In at least one embodiment, graphics complex 3940 is configured to perform graphics pipeline operations such as drawing instructions, pixel operations, geometric calculations, and other operations related to rendering an image on a display. In at least one embodiment, graphics complex 3940 is configured to perform non-graphics operations. In at least one embodiment, graphics complex 3940 is configured to perform both graphics-related and non-graphics operations.
[0305] In at least one embodiment, the graphics complex 3940 includes, without limitation, any number of compute units 3950 and an L2 cache 3942. In at least one embodiment, the compute units 3950 share the L2 cache 3942. In at least one embodiment, the L2 cache 3942 is partitioned. In at least one embodiment, the graphics complex 3940 includes, without limitation, any number of compute units 3950 and any number (including zero) and type of caches. In at least one embodiment, the graphics complex 3940 includes, without limitation, any amount of dedicated graphics hardware.
[0306] In at least one embodiment, each compute unit 3950 includes, without limitation, any number of SIMD units 3952 and a shared memory 3954. In at least one embodiment, each SIMD unit 3952 implements a SIMD architecture and is configured to execute operations in parallel. In at least one embodiment, each compute unit 3950 can execute any number of thread blocks, but each thread block executes on a single compute unit 3950. In at least one embodiment, a thread block includes, without limitation, any number of threads of execution. In at least one embodiment, a workgroup is a thread block. In at least one embodiment, each SIMD unit 3952 executes a different warp.In at least one embodiment, a warp is a group of threads (e.g., 16 threads), where each thread in a warp belongs to a single thread block and is configured to process a different set of instructions based on a single set of instructions. In at least one embodiment, predication can be used to deactivate one or more threads in a warp. In at least one embodiment, a lane is a thread. In at least one embodiment, a work item is a thread. In at least one embodiment, a wavefront is a warp thread. In at least one embodiment, different wavefronts in a thread block can synchronize with each other and communicate via a shared memory 3954.
[0307] In at least one embodiment, structure 3960 is a system interconnect that facilitates data and control transfers between core complex 3910, graphics complex 3940, I / O interfaces 3970, memory interfaces 3980, display controller 3992, and multimedia engine 3994. In at least one embodiment, APU 3900 may include, without limitation, any number and type of system interconnects in addition to or in place of structure 3960 that facilitate data and control transfers across any number and type of directly or indirectly connected components that may be internal or external to APU 3900. In at least one embodiment, I / O interfaces 3970 represent any number and type of I / O interfaces (e.g., PCI, PCI-Extended ("PCI-X"), PCIe, Gigabit Ethernet ("GBE"), USB, etc.).In at least one embodiment, various types of peripheral devices are coupled to I / O interfaces 3970. In at least one embodiment, peripheral devices coupled to I / O interfaces 3970 may include, without limitation, keyboards, mice, printers, scanners, joysticks or other types of gaming controllers, media recording devices, external storage devices, network interface cards, and so on.
[0308] In at least one embodiment, the display controller AMD92 displays images on one or more display devices, such as a liquid crystal display ("LCD") device. In at least one embodiment, the multimedia engine 3994 includes, without limitation, any number and type of multimedia-related circuitry, such as a video decoder, a video processor, an image signal processor, etc. In at least one embodiment, memory controllers 3980 facilitate data transfer between the APU 3900 and a unified system memory 3990. In at least one embodiment, the core complex 3910 and the graphics complex 3940 share the unified system memory 3990.
[0309] In at least one embodiment, the APU 3900 implements a memory subsystem, including, without limitation, any number and type of memory controllers 3980 and memory devices (e.g., shared memory 3954), which may be dedicated to a component or shared among multiple components. In at least one embodiment, the APU 3900 implements a cache subsystem, including, without limitation, one or more caches (e.g., L2 caches 4028, L3 cache 3930, and L2 cache 3942), each of which may be private or shared among any number of components (e.g., cores 3920, core complex 3910, SIMD units 3952, compute units 3950, and graphics complex 3940).
[0310] In at least one embodiment, at least one in Fig. 39 is used to implement techniques and / or functions associated with Fig. 1-13. In at least one embodiment, at least one component of Fig. 39 to effect scaling of one or more clocks of one or more cores based at least in part on the proximity of the one or more cores to each other. In at least one embodiment, at least one component of Fig. 39 at least one aspect relating to the data center 120, the controller 105, at least two or more of processor 110(a) to processor 110(n) and one or more thermal sensors 115(a) to thermal sensor 115(n) of Fig. 1, Firmware 220, Configuration file 225, ACPI 230, Policy 235 one or more processors 1-n of Fig. 2, and / or comprising one or more of: Method 900 of Fig. 9, Procedure 1000 of Fig. 10 and Procedure 1100 of Fig. 11.
[0311] Fig. 40 shows a CPU 4000 in accordance with at least one embodiment. In at least one embodiment, the CPU 4000 is developed by AMD Corporation of Santa Clara, CA. In at least one embodiment, the CPU 4000 may be configured to execute an application program. In at least one embodiment, the CPU 4000 is configured to execute main control software, such as an operating system. In at least one embodiment, the CPU 4000 issues instructions that control the operation of an external GPU (not shown). In at least one embodiment, the CPU 4000 may be configured to execute executable code on the host derived from CUDA source code, and an external GPU may be configured to execute executable code on the device derived from such CUDA source code.In at least one embodiment, CPU 4000 includes, without limitation, any number of core complexes 4010, fabric 4060, I / O interfaces 4070, and memory controllers 4080.
[0312] In at least one embodiment, core complex 4010 includes, without limitation, cores 4020(1)-4020(4) and an L3 cache 4030. In at least one embodiment, core complex 4010 may include, without limitation, any number of cores 4020 and any number and type of caches in any combination. In at least one embodiment, cores 4020 are configured to execute instructions of a particular ISA. In at least one embodiment, each core 4020 is a CPU core.
[0313] In at least one embodiment, each core 4020 includes, without limitation, a fetch / decode unit 4022, an integer execution engine 4024, a floating-point execution engine 4026, and an L2 cache 4028. In at least one embodiment, the fetch unit 4022 fetches instructions, decodes such instructions, generates micro-operations, and sends separate micro-instructions to the integer execution engine 4024 and the floating-point execution engine 4026. In at least one embodiment, the fetch unit 4022 may concurrently send one micro-instruction to the integer execution engine 4024 and another micro-instruction to the floating-point execution engine 4026. In at least one embodiment, the integer execution engine 4024 performs, without limitation, integer and memory operations. In at least one embodiment, the floating-point engine 4026 performs, without limitation, floating-point and vector operations.In at least one embodiment, fetch unit 4022 forwards microinstructions to a single execution engine that replaces both integer execution engine 4024 and floating point execution engine 4026.
[0314] In at least one embodiment, each core 4020(i), where i is an integer representing a particular instance of core 4020, can access L2 cache 4028(i) that includes core 4020(i). In at least one embodiment, each core 4020 included in core complex 4010(j), where j is an integer representing a particular instance of core complex 4010, is connected to other cores 4020 in core complex 4010(j) via L3 cache 4030(j) included in core complex 4010(j). In at least one embodiment, the cores 4020 included in core complex 4010(j), where j is an integer representing a particular instance of core complex 4010, may access the entire L3 cache 4030(j) included in core complex 4010(j). In at least one embodiment, L3 cache 4030 may include any number of slices, without limitation.
[0315] In at least one embodiment, structure 4060 is a system interconnect that enables data and control transfers between core complexes 4010(1)-4010(N) (where N is an integer greater than zero), I / O interfaces 4070, and memory controllers 4080. In at least one embodiment, CPU 4000 may include, without limitation, any number and type of system interconnects in addition to or in place of structure 4060 that facilitate data and control transfers via any number and type of directly or indirectly linked components that may be internal or external to CPU 4000. In at least one embodiment, I / O interfaces 4070 are representative of any number and type of I / O interfaces (e.g., PCI, PCI-X, PCIe, GBE, USB, etc.). In at least one embodiment, various types of peripheral devices are coupled to I / O interfaces 4070.In at least one embodiment, peripheral devices coupled to I / O interfaces 4070 may include, without limitation, displays, keyboards, mice, printers, scanners, joysticks or other types of gaming controllers, media recording devices, external storage devices, network interface cards, and so forth.
[0316] In at least one embodiment, memory controllers 4080 facilitate data transfer between CPU 4000 and system memory 4090. In at least one embodiment, core complex 4010 and graphics complex 4040 share system memory 4090. In at least one embodiment, CPU 4000 implements a memory subsystem, including, without limitation, any number and type of memory controllers 4080 and memory devices that may be dedicated to a component or shared by multiple components. In at least one embodiment, CPU 4000 implements a cache subsystem, including, without limitation, one or more caches (e.g., L2 caches 4028 and L3 caches 4030), each of which may be dedicated to or shared by any number of components (e.g., cores 4020 and core complex 4010).
[0317] In at least one embodiment, at least one in Fig. 40 is used to implement techniques and / or functions associated with Fig. 1-13. In at least one embodiment, at least one component of Fig. 40 is used to effect scaling of one or more clocks of one or more cores based at least in part on the proximity of the one or more cores to each other. In at least one embodiment, at least one component of Fig. 40 at least one aspect relating to the data center 120, the controller 105, at least two or more of processor 110(a) to processor 110(n) and one or more thermal sensors 115(a) to thermal sensor 115(n) of Fig. 1, Firmware 220, Configuration file 225, ACPI 230, Policy 235 one or more processors 1-n of Fig. 2, and / or comprising one or more of: Method 900 of Fig. 9, Procedure 1000 of Fig. 10 and Procedure 1100 of Fig. 11.
[0318] Fig. 41 shows an exemplary accelerator integration slice 4190 in accordance with at least one embodiment. As used herein, a "slice" comprises a particular portion of the processing resources of an accelerator integration circuit. In at least one embodiment, an accelerator integration circuit provides cache management, memory access, context management, and interrupt management services for multiple graphics processing engines included in a graphics acceleration module. The graphics processing engines may each comprise their own GPU. Alternatively, graphics processing engines may comprise different types of graphics processing engines within a GPU, such as graphics execution units, media processing engines (e.g., video encoders / decoders), samplers, and blit engines.In at least one embodiment, a graphics acceleration module may be a graphics processor with multiple graphics processing engines. In at least one embodiment, the graphics processing engines may be individual GPUs integrated on a common package, line card, or die.
[0319] An effective address space 4182 within system memory 4114 stores process elements 4183. In one embodiment, process elements 4183 are stored in response to GPU calls 4181 from applications 4180 executing on processor 4107. A process element 4183 contains the state of the process for the corresponding application 4180. A work description ("WD") 4184 contained in process element 4183 may be a single job requested by an application or may contain a pointer to a queue of jobs. In at least one embodiment, the WD 4184 is a pointer to a job request queue in the application's effective address space 4182.
[0320] The graphics acceleration module 4146 and / or individual graphics processing engines may be shared by all or a subset of the processes in a system. In at least one embodiment, an infrastructure for establishing process state and sending WD 4184 to the graphics acceleration module 4146 to start a job in a virtualized environment may be included.
[0321] In at least one embodiment, a dedicated process programming model is implementation-specific. In this model, a single process owns the graphics acceleration module 4146 or an individual graphics processing engine. Because the graphics acceleration module 4146 is owned by a single process, a hypervisor initializes an accelerator integration circuit for an owning partition, and an operating system initializes an accelerator integration circuit for an owning process when the graphics acceleration module 4146 is allocated.
[0322] In operation, a WD fetch unit 4191 in accelerator integration slice 4190 fetches the next WD 4184, which includes an indication of the work to be performed by one or more graphics processing engines of graphics acceleration module 4146. The data from WD 4184 may be stored in registers 4145 and used by a memory management unit ("MMU") 4139, interrupt management circuitry 4147, and / or context management circuitry 4148 (see figure). For example, one embodiment of MMU 4139 includes segment / page walkup circuitry for accessing segment / page tables 4186 within operating system virtual address space 4185. Circuitry 4147 may process interrupt events ("INT") 4192 received from graphics acceleration module 4146. When performing graphics operations, an effective address 4193 generated by a graphics processing engine is translated into a real address by the MMU 4139.
[0323] In one embodiment, the same set of registers 4145 is duplicated for each graphics processing engine and / or graphics acceleration module 4146 and may be initialized by a hypervisor or operating system. Each of these duplicated registers may comprise the accelerator integration slice 4190. Example registers that may be initialized by a hypervisor are listed in Table 1. Table 1 - Registers initialized by the hypervisor 1 Slice-Steuerregister 2 Reale Adresse (RA) Bereich Zeiger für geplante Prozesse 3 Autoritätsmasken-Überschreibungsregister 4 Interrupt vector table entry offset 5 Interrupt vector table entry boundary 6 Condition register 7 Logical partition ID 8 Real Address (RA) Pointer for the Accelerator Workload Set (Hypervisor Pointer for the Accelerator Workload Set) 9 Memory description register
[0324] Example registers that can be initialized by an operating system are listed in Table 2. Table 2 - Initialized operating system registers 1 Process and thread identification 2 Effective Address (EA) Context Store / Restore Pointer 3 Virtual Address (VA) Pointer for the Accelerator Workload Set (Pointer for the Accelerator Workload Set) 4 Virtual Address (VA) Pointer to the memory segment table 5 Authority mask 6 Job description
[0325] In one embodiment, each WD 4184 is specific to a particular graphics acceleration module 4146 and / or a particular graphics processing engine. It contains all the information a graphics processing engine needs to perform its work, or it may be a pointer to a memory location where an application has set up a command queue for the work to be performed.
[0326] In at least one embodiment, at least one in Fig. 41 is used to implement techniques and / or functions associated with Fig. 1-13. In at least one embodiment, at least one component of Fig.41 to effect scaling of one or more clocks of one or more cores based at least in part on the proximity of the one or more cores to each other. In at least one embodiment, at least one component of Fig. 41 at least one aspect relating to the data center 120, the controller 105, at least two or more of processor 110(a) to processor 110(n) and one or more thermal sensors 115(a) to thermal sensor 115(n) of Fig. 1, Firmware 220, Configuration file 225, ACPI 230, Policy 235 one or more processors 1-n of Fig. 2, and / or comprising one or more of: Method 900 of Fig. 9, Procedure 1000 of Fig. 10 and Procedure 1100 of Fig. 11.
[0327] Fig.42A-42B illustrate example graphics processors in accordance with at least one embodiment. In at least one embodiment, each of the example graphics processors may be fabricated using one or more IP cores. In addition to the embodiments shown, other logic and circuitry may also be included in at least one embodiment, including additional graphics processors / cores, peripheral interface controllers, or general processor cores. In at least one embodiment, the example graphics processors are intended for use within a SoC.
[0328] Fig. 42A shows an exemplary graphics processor 4210 of an SoC integrated circuit that may be manufactured with one or more IP cores in accordance with at least one embodiment. Fig.42B shows another exemplary graphics processor 4240 of an integrated circuit (SoC) that may be manufactured using one or more IP cores in accordance with at least one embodiment. In at least one embodiment, the graphics processor 4210 is Fig. 42A is a low-power graphics processor core. In at least one embodiment, the graphics processor 4240 is Fig. 42B, a higher performance graphics processor core. In at least one embodiment, each of the graphics processors 4210, 4240 may be a variant of the graphics processor 1810 of Fig. be 18.
[0329] In at least one embodiment, graphics processor 4210 includes a vertex processor 4205 and one or more fragment processors 4215A-4215N (e.g., 4215A, 4215B, 4215C, 4215D, 4215N-1, and 4215N). In at least one embodiment, graphics processor 4210 may execute different shader programs via separate logic, such that vertex processor 4205 is optimized to perform operations for vertex shader programs, while one or more fragment processors 4215A-4215N perform fragment (e.g., pixel) shading operations for fragment or pixel shader programs. In at least one embodiment, the vertex processor 4205 executes a vertex processing stage of a 3D graphics pipeline and generates primitives and vertex data.In at least one embodiment, fragment processor(s) 4215A-4215N use the primitive and vertex data generated by vertex processor 4205 to generate a framebuffer displayed on a display device. In at least one embodiment, fragment processor(s) 4215A-4215N are optimized for executing fragment shader programs such as those provided in an OpenGL API, which can be used to perform similar operations as a pixel shader program such as those provided in a Direct 3D API.
[0330] In at least one embodiment, graphics processor 4210 additionally includes one or more MMU(s) 4220A-4220B, cache(s) 4225A-4225B, and circuit interconnect(s) 4230A-4230B. In at least one embodiment, one or more MMU(s) 4220A-4220B provide virtual-to-physical address mapping for graphics processor 4210, including vertex processor 4205 and / or fragment processor(s) 4215A-4215N, which may reference vertex or image / texture data stored in memory in addition to the vertex or image / texture data stored in one or more cache(s) 4225A-4225B. In at least one embodiment, one or more MMU(s) 4220A-4220B may be synchronized with other MMUs within a system, including one or more MMUs associated with one or more application processors 1805, image processors 1815, and / or video processors 1820 of Fig.18, so that each processor 1805-1820 can participate in a shared or unified virtual memory system. In at least one embodiment, one or more circuit interconnects 4230A-4230B enable the graphics processor 4210 to interface with other IP cores within an SoC, either via an internal bus of an SoC or via a direct connection.
[0331] In at least one embodiment, the graphics processor 4240 includes one or more MMU(s) 4220A-4220B, caches 4225A-4225B, and circuit interconnects 4230A-4230B of the graphics processor 4210 of Fig.42A. In at least one embodiment, the graphics processor 4240 includes one or more shader cores 4255A-4255N (e.g., 4255A, 4255B, 4255C, 4255D, 4255E, 4255F, through 4255N-1 and 4255N) that provide a unified shader core architecture in which a single core or type of core can execute all types of programmable shader code, including shader code implementing vertex shaders, fragment shaders, and / or compute shaders. In at least one embodiment, the number of shader cores may vary.In at least one embodiment, the graphics processor 4240 includes an inter-core task manager 4245 acting as a thread dispatcher to distribute execution threads to one or more shader cores 4255A-4255N and a tiling unit 4258 to accelerate tiling operations for tile-based rendering, in which rendering operations for a scene are partitioned into image space, for example, to exploit local spatial coherence within a scene or to optimize the use of internal caches.
[0332] In at least one embodiment, at least one in Fig. 42 is used to implement techniques and / or functions associated with Fig. 1-13. In at least one embodiment, at least one component of Fig.42 is used to effect scaling of one or more clocks of one or more cores based at least in part on the proximity of the one or more cores to each other. In at least one embodiment, at least one component of Fig. 42 at least one aspect relating to the data center 120, the controller 105, at least two or more of processor 110(a) to processor 110(n) and one or more thermal sensors 115(a) to thermal sensor 115(n) of Fig. 1, Firmware 220, Configuration file 225, ACPI 230, Policy 235 one or more processors 1-n of Fig. 2, and / or comprising one or more of: Method 900 of Fig. 9, Procedure 1000 of Fig. 10 and Procedure 1100 of Fig. 11.
[0333] Fig.43A shows a graphics core 4300 in accordance with at least one embodiment. In at least one embodiment, the graphics core 4300 may be included in the graphics processor 3710 of Fig. 37. In at least one embodiment, the graphics core 4300 may be a unified shader core 4255A-4255N as shown in Fig.42B. In at least one embodiment, the graphics core 4300 includes a shared instruction cache 4302, a texture unit 4318, and a cache / shared memory 4320 common to the execution resources within the graphics core 4300. In at least one embodiment, the graphics core 4300 may include multiple slices 4301A-4301N or partitions for each core, and a graphics processor may include multiple instances of the graphics core 4300. The slices 4301A-4301N may include support logic including a local instruction cache 4304A-4304N, a thread scheduler 4306A-4306N, a thread dispatcher 4308A-4308N, and a set of registers 4310A-4310N.In at least one embodiment, slices 4301A-4301N may include a set of additional functional units ("AFUs") 4312A-4312N, floating point units ("FPUs") 4314A-4314N, integer arithmetic logic units ("ALUs") 4316-4316N, address calculation units ("ACUs") 4313A-4313N, double precision floating point units ("DPFPUs") 4315A-4315N, and matrix processing units ("MPUs") 4317A-4317N.
[0334] In at least one embodiment, the FPUs 4314A-4314N can perform single-precision (32-bit) and half-precision (16-bit) floating-point operations, while the DPFPUs 4315A-4315N can perform double-precision (64-bit) floating-point operations. In at least one embodiment, the ALUs 4316A-4316N can perform variable-precision integer operations at 8-bit, 16-bit, and 32-bit precision and can be configured for mixed-precision operations. In at least one embodiment, the MPUs 4317A-4317N can also be configured for mixed-precision matrix operations, including half-precision floating-point and 8-bit integer operations. In at least one embodiment, MPUs 4317-4317N may perform a variety of matrix operations to accelerate CUDA programs, including support for accelerated general purpose matrix-matrix multiplication ("GEMM").In at least one embodiment, the AFUs 4312A-4312N may perform additional logical operations not supported by floating point or integer units, including trigonometric operations (e.g., sine, cosine, etc.).
[0335] Fig.43B illustrates a general purpose graphics processing unit (“SGPGPU”) 4330 in accordance with at least one embodiment. In at least one embodiment, the GPGPU 4330 is highly parallel and suitable for deployment on a multi-chip module. In at least one embodiment, the GPGPU 4330 can be configured to allow highly parallel computational operations to be performed by an array of GPUs. In at least one embodiment, the GPGPU 4330 can be directly connected to other instances of the GPGPU 4330 to form a multi-GPU cluster and improve execution time for CUDA programs. In at least one embodiment, the GPGPU 4330 includes a host interface 4332 to enable connection to a host processor. In at least one embodiment, the host interface 4332 is a PCIe interface.In at least one embodiment, host interface 4332 may be a vendor-specific communications interface or communications structure. In at least one embodiment, GPGPU 4330 receives instructions from a host processor and uses a global scheduler 4334 to distribute the execution threads associated with those instructions among a number of compute clusters 4336A-4336H. In at least one embodiment, compute clusters 4336A-4336H share a cache 4338. In at least one embodiment, cache 4338 may serve as a higher-level cache for caches within compute clusters 4336A-4336H.
[0336] In at least one embodiment, GPGPU 4330 includes memory 4344A-4344B coupled to compute clusters 4336A-4336H via a series of memory controllers 4342A-4342B. In at least one embodiment, memory 4344A-4344B may include various types of memory devices, including DRAM or graphics random access memory, such as synchronous graphics random access memory ("SGRAM"), including graphics double data rate memory ("GDDR").
[0337] In at least one embodiment, the compute clusters 4336A-4336H each include a set of graphics cores, such as the graphics core 4300 of Fig.43A, which may include multiple types of integer and floating-point logic units capable of performing computational operations with a range of precisions, also suitable for computations associated with CUDA programs. For example, in at least one embodiment, at least a subset of floating-point units in each of compute clusters 4336A-4336H may be configured to perform 16-bit or 32-bit floating-point operations, while a different subset of floating-point units may be configured to perform 64-bit floating-point operations.
[0338] In at least one embodiment, multiple instances of the GPGPU 4330 can be configured to operate as a compute cluster. In at least one embodiment, the compute clusters 4336A-4336H can use any technically feasible communication techniques for synchronization and data exchange. In at least one embodiment, multiple instances of the GPGPU 4330 communicate via the host interface 4332. In at least one embodiment, the GPGPU 4330 includes an I / O hub 4339 that couples the GPGPU 4330 to a GPU interconnect 4340 that enables direct connection to other instances of the GPGPU 4330. In at least one embodiment, the GPU interconnect 4340 is coupled to a dedicated GPU-to-GPU bridge that enables communication and synchronization between multiple instances of the GPGPU 4330.In at least one embodiment, GPU link 4340 is coupled to a high-speed interconnect to send and receive data to other GPGPUs 4330 or parallel processors. In at least one embodiment, multiple GPGPU instances 4330 are located in separate computing systems and communicate via a network interface accessible via host interface 4332. In at least one embodiment, GPU link 4340 may be configured to enable connection to a processor in addition to, or alternatively to, host interface 4332. In at least one embodiment, GPGPU 4330 may be configured to execute a CUDA program.
[0339] In at least one embodiment, at least one in Fig. 43 is used to implement techniques and / or functions associated with Fig.1-13. In at least one embodiment, at least one component of Fig. 43 to effect scaling of one or more clocks of one or more cores based at least in part on the proximity of the one or more cores to each other. In at least one embodiment, at least one component of Fig. 43 at least one aspect relating to the data center 120, the controller 105, at least two or more of processor 110(a) to processor 110(n) and one or more thermal sensors 115(a) to thermal sensor 115(n) of Fig. 1, Firmware 220, Configuration file 225, ACPI 230, Policy 235 one or more processors 1-n of Fig. 2, and / or comprising one or more of: Method 900 of Fig. 9, Procedure 1000 of Fig. 10 and Procedure 1100 of Fig. 11.
[0340] Fig.Figure 44A illustrates a parallel processor 4400 in accordance with at least one embodiment. In at least one embodiment, various components of the parallel processor 4400 may be implemented using one or more integrated circuit devices, such as programmable processors, application-specific integrated circuits ("ASICs"), or FPGAs.
[0341] In at least one embodiment, parallel processor 4400 includes a parallel processing unit 4402. In at least one embodiment, parallel processing unit 4402 includes an I / O unit 4404 that enables communication with other devices, including other instances of parallel processing unit 4402. In at least one embodiment, I / O unit 4404 may be directly connected to other devices. In at least one embodiment, I / O unit 4404 is connected to other devices via a hub or switch interface, such as storage hub 1905. In at least one embodiment, the connections between storage hub 1905 and I / O unit 4404 form a communication link.In at least one embodiment, the I / O unit 4404 is coupled to a host interface 4406 and a memory crossbar 4416, where the host interface 4406 receives commands directed to perform processing operations and the memory crossbar 4416 receives commands directed to perform memory operations.
[0342] In at least one embodiment, when host interface 4406 receives a command buffer via I / O unit 4404, host interface 4406 may direct work operations to execute those commands to a front end 4408. In at least one embodiment, front end 4408 is coupled to a scheduler 4410 configured to dispatch commands or other work items to a processing array 4412. In at least one embodiment, scheduler 4410 ensures that processing array 4412 is properly configured and in a valid state before dispatching tasks to processing array 4412. In at least one embodiment, scheduler 4410 is implemented via firmware logic executing on a microcontroller.In at least one embodiment, the microcontroller-implemented scheduler 4410 is configurable to perform complex scheduling and work distribution operations at coarse and fine granularity, enabling rapid preemption and context switching of threads executing on the processing array 4412. In at least one embodiment, host software may announce workloads for scheduling on the processing array 4412 via one of several graphics processing doorbells. In at least one embodiment, the workloads may then be automatically distributed across the processing array 4412 by the logic of the scheduler 4410 within a microcontroller that includes the scheduler 4410.
[0343] In at least one embodiment, the processing array 4412 may include up to "N" clusters (e.g., cluster 4414A, cluster 4414B, through cluster 4414N). In at least one embodiment, each cluster 4414A-4414N of the processing array 4412 may execute a large number of concurrent threads. In at least one embodiment, the scheduler 4410 may allocate work to the clusters 4414A-4414N of the processing array 4412 using various scheduling and / or work distribution algorithms, which may vary depending on the workload associated with each type of program or computation. In at least one embodiment, scheduling may be handled dynamically by the scheduler 4410 or may be partially assisted by compiler logic during compilation of the program logic configured for execution by the processing array 4412.In at least one embodiment, different clusters 4414A-4414N of the processing array 4412 may be allocated for processing different types of programs or for performing different types of computations.
[0344] In at least one embodiment, processing array 4412 may be configured to perform various types of parallel processing operations. In at least one embodiment, processing array 4412 is configured to perform general-purpose parallel computing operations. For example, in at least one embodiment, processing array 4412 may include logic for performing processing tasks, including filtering video and / or audio data, performing modeling operations, including physics operations, and performing data transformations.
[0345] In at least one embodiment, processing array 4412 is configured to perform parallel graphics processing operations. In at least one embodiment, processing array 4412 may include additional logic to support the execution of such graphics processing operations, including, but not limited to, texture sampling logic to perform texture operations, as well as tessellation logic and other vertex processing logic. In at least one embodiment, processing array 4412 may be configured to execute graphics processing-related shader programs, such as vertex shaders, tessellation shaders, geometry shaders, and pixel shaders. In at least one embodiment, parallel processing unit 4402 may transfer data from system memory via I / O unit 4404 for processing.In at least one embodiment, the transferred data may be stored in on-chip memory (e.g., parallel processor memory 4422) during processing and then written back to system memory.
[0346] In at least one embodiment, when the graphics processing unit 4402 is used to perform graphics processing, the scheduler 4410 may be configured to divide a workload into approximately equal-sized tasks to enable better distribution of graphics processing operations across multiple clusters 4414A-4414N of the processing array 4412. In at least one embodiment, portions of the processing array 4412 may be configured to perform different types of processing. For example, in at least one embodiment, a first portion may be configured to perform vertex shading and topology generation, a second portion may be configured to perform tessellation and geometry shading, and a third portion may be configured to perform pixel shading or other screen operations to generate a rendered image for display.In at least one embodiment, intermediate data generated by one or more of clusters 4414A-4414N may be stored in buffers to facilitate transfer of intermediate data between clusters 4414A-4414N for further processing.
[0347] In at least one embodiment, processing array 4412 may receive processing tasks to be executed via scheduler 4410, which receives commands defining processing tasks from frontend 4408. In at least one embodiment, the processing tasks may include indices of the data to be processed, such as surface (patch) data, primitive data, vertex data, and / or pixel data, as well as state parameters and commands that define how the data is to be processed (e.g., which program is to be executed). In at least one embodiment, scheduler 4410 may be configured to retrieve or receive indices corresponding to the tasks from frontend 4408.In at least one embodiment, the front end 4408 may be configured to ensure that the processing array 4412 is placed in a valid state before initiating a workload specified by incoming instruction buffers (e.g., stack buffers, push buffers, etc.).
[0348] In at least one embodiment, each of one or more instances of parallel processing unit 4402 may be coupled to parallel processor memory 4422. In at least one embodiment, parallel processor memory 4422 may be accessed via memory crossbar 4416, which may receive memory requests from processing array 4412 as well as from I / O unit 4404. In at least one embodiment, memory crossbar 4416 may access parallel processor memory 4422 via a memory interface 4418. In at least one embodiment, memory interface 4418 may include multiple partition units (e.g., partition unit 4420A, partition unit 4420B, through partition unit 4420N), each of which may be coupled to a portion (e.g., memory unit) of parallel processor memory 4422.In at least one embodiment, a number of partition units 4420A-4420N is configured to be equal to a number of storage units, such that a first partition unit 4420A has a corresponding first storage unit 4424A, a second partition unit 4420B has a corresponding storage unit 4424B, and an Nth partition unit 4420N has a corresponding Nth storage unit 4424N. In at least one embodiment, a number of partition units 4420A-4420N may not be equal to a number of storage devices.
[0349] In at least one embodiment, the memory units 4424A-4424N may include various types of memory devices, including DRAM or graphics random access memory, such as SGRAM, including GDDR memory. In at least one embodiment, the memory units 4424A-4424N may also include 3D stacks, including but not limited to high-width memories ("HBM"). In at least one embodiment, rendering targets, such as frame buffers or texture maps, may be stored across the memory units 4424A-4424N so that the partition units 4420A-4420N can write portions of each rendering target in parallel to efficiently utilize the available bandwidth of the parallel processor memory 4422.In at least one embodiment, a local instance of parallel processor memory 4422 may be eliminated in favor of a unified memory design that utilizes system memory in conjunction with the local cache memory.
[0350] In at least one embodiment, each of the clusters 4414A-4414N of the processing array 4412 can process data written to each of the memory units 4424A-4424N within the parallel processor memory 4422. In at least one embodiment, the memory crossbar 4416 can be configured to transfer an output of each cluster 4414A-4414N to any partition unit 4420A-4420N or to another cluster 4414A-4414N that can perform additional processing operations on an output. In at least one embodiment, each cluster 4414A-4414N can communicate with the memory interface 4418 via the memory crossbar 4416 to read from or write to various external devices.In at least one embodiment, the memory crossbar 4416 includes a connection to the memory interface 4418 to communicate with the I / O unit 4404, as well as a connection to a local instance of the parallel processor memory 4422, which enables the processing units in the different clusters 4414A-4414N to communicate with system memory or other memory not local to the parallel processing unit 4402. In at least one embodiment, the memory crossbar 4416 may use virtual channels to separate traffic flows between clusters 4414A-4414N and partition units 4420A-4420N.
[0351] In at least one embodiment, multiple instances of the parallel processing unit 4402 may be provided on a single add-in card, or multiple add-in cards may be interconnected. In at least one embodiment, different instances of the parallel processing unit 4402 may be configured to interoperate, even if different instances have different numbers of processor cores, different amounts of local parallel processor memory, and / or other configuration differences. For example, in at least one embodiment, some instances of the parallel processing unit 4402 may include higher-precision floating-point units than other instances.In at least one embodiment, systems including one or more instances of the parallel processing unit 4402 or the parallel processor 4400 may be implemented in a variety of configurations and form factors, including, but not limited to, desktop, laptop, or handheld personal computers, servers, workstations, game consoles, and / or embedding systems.
[0352] Fig. 44B shows a processing cluster 4494 in accordance with at least one embodiment. In at least one embodiment, the processing cluster 4494 is included in a parallel processing unit. In at least one embodiment, the processing cluster 4494 is one of the processing clusters 4414A-4414N of Fig.44. In at least one embodiment, the processing cluster 4494 may be configured to execute many threads in parallel, where the term "thread" refers to an instance of a particular program executing on a particular set of input data. In at least one embodiment, single instruction, multiple data ("SIMD") instruction issuance techniques are used to support the parallel execution of a large number of threads without providing multiple independent instruction units. In at least one embodiment, single instruction multiple thread ("SIMT") techniques are used to support the parallel execution of a large number of generally synchronized threads using a common instruction unit configured to issue instructions to a set of processing engines within each processing cluster 4494.
[0353] In at least one embodiment, the operation of the processing cluster 4494 may be controlled by a pipeline manager 4432, which distributes the processing tasks to the parallel SIMT processors. In at least one embodiment, the pipeline manager 4432 receives instructions from the scheduler 4410 of the Fig.44 and manages the execution of these instructions via a graphics multiprocessor 4434 and / or a texture unit 4436. In at least one embodiment, the graphics multiprocessor 4434 is an exemplary example of a SIMT parallel processor. However, in at least one embodiment, various types of SIMT parallel processors with different architectures may be included in the processing cluster 4494. In at least one embodiment, the processing cluster 4494 may include one or more instances of the graphics multiprocessor 4434. In at least one embodiment, the graphics multiprocessor 4434 may process data, and a data crossbar 4440 may be used to distribute the processed data to one of several possible destinations, including other shader units.In at least one embodiment, the pipeline manager 4432 may facilitate the distribution of processed data by specifying destinations for processed data to be distributed across the data crossbar 4440.
[0354] In at least one embodiment, each graphics multiprocessor 4434 within the processing cluster 4494 may include an identical set of functional execution logic (e.g., arithmetic logic units, load / store units ("LSUs"), etc.). In at least one embodiment, the functional execution logic may be configured in a pipeline in which new instructions may be issued before the previous instructions complete. In at least one embodiment, the functional execution logic supports a variety of operations, including integer and floating-point arithmetic, comparison operations, Boolean operations, bit shifting, and the computation of various algebraic functions. In at least one embodiment, the same hardware with functional units may be used to perform different operations, and any combination of functional units may be present.
[0355] In at least one embodiment, the instructions transferred to the processing cluster 4494 form a thread. In at least one embodiment, a set of threads executing across a set of parallel processing engines is a thread group. In at least one embodiment, a thread group executes a program with different input data. In at least one embodiment, each thread within a thread group may be assigned to a different engine within the graphics multiprocessor 4434. In at least one embodiment, a thread group may include fewer threads than the number of processing engines within the graphics multiprocessor 4434. In at least one embodiment, when a thread group includes fewer threads than a number of processing engines, one or more of the processing engines may be idle during the cycles in which that thread group is processing.In at least one embodiment, a thread group may also include more threads than a number of processing engines within the graphics multiprocessor 4434. In at least one embodiment, if a thread group includes more threads than a number of processing engines in the graphics multiprocessor 4434, processing may occur in consecutive clocks. In at least one embodiment, multiple thread groups may execute concurrently on the graphics multiprocessor 4434.
[0356] In at least one embodiment, the graphics multiprocessor 4434 includes an internal cache for performing load and store operations. In at least one embodiment, the graphics multiprocessor 4434 may forgo an internal cache and utilize a cache (e.g., L1 cache 4448) within the processing cluster 4494. In at least one embodiment, each graphics multiprocessor 4434 also has access to Level 2 ("L2") caches within partition units (e.g., partition units 4420A-4420N of Fig.44A) that are shared by all processing clusters 4494 and can be used to transfer data between threads. In at least one embodiment, graphics multiprocessor 4434 can also access off-chip global memory, which can include one or more of the parallel processor local memories and / or system memories. In at least one embodiment, any memory external to parallel processing unit 4402 can be used as global memory. In at least one embodiment, processing cluster 4494 includes multiple instances of graphics multiprocessor 4434, which can share common instructions and data, which can be stored in L1 cache 4448.
[0357] In at least one embodiment, each processing cluster 4494 may include an MMU 4445 configured to translate virtual addresses into physical addresses. In at least one embodiment, one or more instances of the MMU 4445 may reside in the memory interface 4418 of Fig.44. In at least one embodiment, the MMU 4445 includes a set of page table entries ("PTEs") used to map a virtual address to a physical address of a tile, and optionally a cache line index. In at least one embodiment, the MMU 4445 may include address translation lookaside buffers ("TLBs") or caches, which may be located in the graphics multiprocessor 4434 or the L1 cache 4448 or the processing cluster 4494. In at least one embodiment, a physical address is processed to distribute access locality to surface data to enable efficient interleaving of requests between partition units. In at least one embodiment, a cache line index may be used to determine whether a request for a cache line is a hit or miss.
[0358] In at least one embodiment, processing cluster 4494 may be configured such that each graphics multiprocessor 4434 is coupled to a texture unit 4436 to perform texture mapping operations, such as determining texture pattern positions, reading texture data, and filtering texture data. In at least one embodiment, the texture data is read from an internal texture L1 cache (not shown) or from an L1 cache within graphics multiprocessor 4434 and retrieved from an L2 cache, local parallel processor memory, or system memory as needed.In at least one embodiment, each graphics multiprocessor 4434 outputs a processed task to the data crossbar 4440 to provide a processed task to another processing cluster 4494 for further processing via the memory crossbar 4416, or to store a processed task in an L2 cache, local parallel processor memory, or system memory. In at least one embodiment, a pre-raster operations unit ("preROP") 4442 is configured to receive data from the graphics multiprocessor 4434 and direct data to the ROP units associated with the partition units described herein (e.g., partition units 4420A-4420N in FIG. Fig. 44). In at least one embodiment, PreROP 4442 may perform optimizations for color mixing, organizing pixel color data, and address translations.
[0359] Fig.44C shows a graphics multiprocessor 4496 according to at least one embodiment. In at least one embodiment, the graphics multiprocessor 4496 is the graphics multiprocessor 4434 of Fig. 44B. In at least one embodiment, graphics multiprocessor 4496 is coupled to pipeline manager 4432 of processing cluster 4494. In at least one embodiment, graphics multiprocessor 4496 has an execution pipeline including, among other things, an instruction cache 4452, an instruction unit 4454, an address mapping unit 4456, a register file 4458, one or more GPGPU cores 4462, and one or more LSUs 4466. GPGPU cores 4462 and LSUs 4466 are coupled to cache memory 4472 and shared memory 4470 via a memory and cache interconnect 4468.
[0360] In at least one embodiment, instruction cache 4452 receives a stream of instructions to be executed from pipeline manager 4432. In at least one embodiment, the instructions are cached in instruction cache 4452 and forwarded for execution by instruction unit 4454. In at least one embodiment, instruction unit 4454 may dispatch instructions as thread groups (e.g., warps), with each thread of a thread group associated with a different execution unit within GPGPU core 4462. In at least one embodiment, an instruction may access a local, shared, or global address space by specifying an address within a unified address space. In at least one embodiment, address mapping unit 4456 may be used to translate addresses in a unified address space into a unique memory address accessible by LSUs 4466.
[0361] In at least one embodiment, register file 4458 provides a set of registers for functional units of graphics multiprocessor 4496. In at least one embodiment, register file 4458 provides temporary storage for operands associated with data paths of functional units (e.g., GPGPU cores 4462, LSUs 4466) of graphics multiprocessor 4496. In at least one embodiment, register file 4458 is partitioned among individual functional units such that each functional unit is assigned its own section of register file 4458. In at least one embodiment, register file 4458 is partitioned among different thread groups executed by graphics multiprocessor 4496.
[0362] In at least one embodiment, GPGPU cores 4462 may each include FPUs and / or integer ALUs used to execute instructions of graphics multiprocessor 4496. GPGPU cores 4462 may be similar or different in architecture. In at least one embodiment, a first portion of GPGPU cores...
Claims
[1] A processor comprising: one or more circuits for scaling one or more clocks of one or more cores based at least in part on a proximity of the one or more cores to each other. [2] The processor of claim 1, wherein the one or more circuits scale the one or more clocks based at least in part on one or more core utilization patterns that include indications of one or more thermal conditions. [3] The processor of claim 1, wherein the one or more circuits are to avoid one or more patterns of core utilization that would result in one or more adverse thermal conditions. [4] The processor of claim 1, wherein the one or more circuits select one or more cores of the processor to execute one or more threads based at least in part on avoiding one or more patterns of core utilization indicative of one or more adverse thermal conditions. [5] The processor of claim 1, wherein the one or more further circuits scale the one or more clocks based on information indicative of one or more core utilization patterns stored in one or more configuration files. [6] The processor of claim 1, wherein the one or more circuits compare the core utilization of the processor to one or more core utilization patterns. [7] The processor of claim 1, wherein the proximity of the one or more cores corresponds to one or more positions of the one or more cores on a chip of the processor. [8] A system comprising: one or more processors for scaling one or more clocks of one or more cores based at least in part on the proximity of the one or more cores to each other. [9] The system of claim 8, wherein the one or more processors scale the one or more clocks based at least in part on one or more core utilization patterns indicative of one or more thermal conditions. [10] The system of claim 8, wherein the one or more processors avoid one or more patterns of core utilization that would result in one or more adverse thermal conditions. [11] The system of claim 8, wherein the one or more processors select one or more cores of the processor to execute one or more threads based at least in part on avoiding one or more patterns of core utilization that indicate one or more adverse thermal conditions. [12] The system of claim 8, wherein the one or more processors scale the one or more clocks based on information indicating one or more core utilization patterns stored in one or more configuration files. [13] The system of claim 8, wherein the one or more processors compare the core utilization of the processor to one or more core utilization patterns. [14] The system of claim 8, wherein the proximity of the one or more cores corresponds to one or more positions of the one or more cores on a chip of the processor. [15] A method comprising: scaling one or more clocks of one or more cores based at least in part on a proximity of the one or more cores to each other. [16] The method of claim 15, wherein the scaling of the one or more clocks is based at least in part on one or more core utilization patterns indicative of one or more thermal conditions. [17] The method of claim 15, wherein scaling one or more clocks avoids one or more core utilization patterns that would result in one or more adverse thermal conditions. [18] The method of claim 15, further comprising: Selecting one or more cores of the processor to execute one or more threads based at least in part on avoiding one or more patterns of core utilization that indicate one or more adverse thermal conditions. [19] The method of claim 15, wherein the scaling of one or more clocks is based on information indicating one or more core utilization patterns stored in one or more configuration files. [20] The method of claim 15, further comprising: Comparing the processor's core utilization with one or more core utilization patterns.