Intelligent Adaptive Heat Sink for Cooling Data Center Equipment

By using adaptive heat sinks in the data center cooling system, the problem of insufficient cooling capacity of the air cooling system is solved, and efficient and economical cooling effects are achieved to adapt to the different heat needs of the computing components.

CN116158201BActive Publication Date: 2025-08-05NVIDIA CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202080103971.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-08-24
Publication Date
2025-08-05
Estimated Expiration
2040-08-24

AI Technical Summary

Technical Problem

When the existing data center cooling system faces the sudden high heat demand of computing components, the air cooling system cannot effectively meet the cooling needs, and the liquid cooling system is uneconomical, resulting in insufficient cooling capacity.

Method used

Adaptive heat sinks are used to increase the heat dissipation area to meet different cooling needs by switching between shrinking and deploying configurations, and efficient cooling is achieved by combining the advantages of air and liquid cooling systems.

Benefits of technology

Improves the cooling capacity of the cooling system, economically meets different cooling needs, reduces operating costs, and improves the operational reliability of the data center.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116158201B_ABST
    Figure CN116158201B_ABST
Patent Text Reader

Abstract

A cooling system for data center equipment is disclosed. A heat sink (216) is disposed between a first plate (210) and a second plate (208) to dissipate a first amount of heat to an environment in a first configuration of the heat sink (216). The first plate (210) is movable relative to the second plate (208) to expose a surface area of the heat sink (216) to the environment in a second configuration of the heat sink (216).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] At least one embodiment relates to a cooling system for data center equipment. In at least one embodiment, a heat sink is disposed between a first plate and a second plate of a heat sink to dissipate a first amount of heat to an environment in a first configuration of the heat sink, and the first plate is movable relative to the second plate to expose a surface area of the heat sink to the environment in a second configuration of the heat sink. Background Art

[0002] Data center cooling systems typically use fans to circulate air through the server components. Some supercomputers or other high-capacity computers may use water or other cooling systems rather than air cooling systems to remove heat from the server components or racks in the data center to an area outside the data center. The cooling system may include chillers within the data center area, including an area outside the data center. The area outside the data center may be an area that includes a cooling tower or other external heat exchanger that receives heated coolant from the data center and disperses the heat to the environment (or external cooling medium) by forced ventilation or other means before the cooled coolant is recirculated back to the data center. In one example, the chiller and cooling tower together form a cooling facility with a pump that responds to temperatures measured by external equipment applied to the data center. An air cooling system alone may not absorb enough heat to support effective or efficient cooling in a data center, and a liquid cooling system may not be economical for the needs of the data center. BRIEF DESCRIPTION OF THE DRAWINGS

[0003] Various embodiments according to the present disclosure will be described with reference to the accompanying drawings, in which:

[0004] Figure 1 is a block diagram of an example data center having a cooling system that is subject to the improvements described in at least one embodiment;

[0005] Figure 2A is a block diagram illustrating mobile data center features of a cooling system incorporating adaptive fins in a first configuration of fins according to at least one embodiment;

[0006] Figure 2B is a block diagram illustrating a cooling system incorporating an adaptive heat sink in a second configuration of heat sinks according to at least one embodiment;

[0007] Figure 3A 、 3B and 3C illustrate aspects of an adaptive heat sink according to at least one embodiment;

[0008] Figure 4A 、 4Band 4C illustrate a subsystem supported by a processor to implement a cooling system incorporating an adaptive heat sink according to at least one embodiment;

[0009] Figure 4D and 4E shows a plan view of a board incorporating one or more processor-supported subsystems to implement a cooling system incorporating an adaptive heat sink in accordance with at least one embodiment;

[0010] Figure 5 is usable for use or manufacture according to at least one embodiment Figures 2A-4E and a process flow of the steps of the method of cooling the system 6A-17D;

[0011] Figure 6A An example data center is shown where data from Figure 2A-5 at least one embodiment of;

[0012] Figure 6B 、 6C Various embodiments are shown, such as Figure 6A and inference and / or training logic used in at least one embodiment of the present disclosure for implementing and / or supporting a cooling system incorporating an adaptive heat sink;

[0013] Figure 7A is a block diagram illustrating an exemplary computer system, which may be a system having interconnected devices and components, a system on a chip (SOC), or some combination thereof, formed together with a processor that may include an execution unit for executing instructions to support and / or implement a cooling system incorporating an adaptive heat sink as described herein, in accordance with at least one embodiment;

[0014] Figure 7B is a block diagram illustrating an electronic device for utilizing a processor to support and / or implement a cooling system incorporating an adaptive heat sink in accordance with at least one embodiment;

[0015] Figure 7C is a block diagram illustrating an electronic device for utilizing a processor to support and / or implement a cooling system incorporating an adaptive heat sink in accordance with at least one embodiment;

[0016] Figure 8 Another exemplary computer system is shown for implementing various processes and methods of a cooling system incorporating an adaptive heat sink as described throughout this disclosure, in accordance with at least one embodiment;

[0017] Figure 9AAn exemplary architecture is shown in which a GPU is communicatively coupled to a multi-core processor via a high-speed link for implementing and / or supporting a cooling system incorporating an adaptive heat sink, in accordance with at least one embodiment disclosed herein;

[0018] Figure 9B shows additional details of the interconnection between a multi-core processor and a graphics acceleration module according to an exemplary embodiment;

[0019] Figure 9C Another exemplary embodiment according to at least one embodiment disclosed herein is shown, wherein an accelerator integrated circuit is integrated within a processor for implementing and / or supporting a cooling system incorporating an adaptive heat sink;

[0020] Figure 9D An exemplary accelerator integration slice 990 is shown for implementing and / or supporting a cooling system incorporating an adaptive heat sink, in accordance with at least one embodiment disclosed herein;

[0021] Figure 9E shows additional details of an exemplary embodiment of a sharing model for implementing and / or supporting a cooling system incorporating an adaptive heat sink in accordance with at least one embodiment disclosed herein;

[0022] Figure 9F shows additional details of an exemplary embodiment of a unified memory addressable via a common virtual memory address space for accessing physical processor memory and GPU memory to implement and / or support a cooling system incorporating an adaptive heat sink in accordance with at least one embodiment disclosed herein;

[0023] Figure 10A An exemplary integrated circuit and associated graphics processor for use in a cooling system incorporating an adaptive heat sink according to the present disclosure are shown;

[0024] Figures 10B-10C An exemplary integrated circuit and associated graphics processor for supporting and / or implementing a cooling system incorporating an adaptive heat sink in accordance with at least one embodiment is shown;

[0025] Figures 10D-10E Additional exemplary graphics processor logic for supporting and / or implementing a cooling system incorporating an adaptive heat sink in accordance with at least one embodiment is shown;

[0026] Figure 11A is a block diagram illustrating a computing system for supporting and / or implementing a cooling system incorporating an adaptive heat sink in accordance with at least one embodiment;

[0027] Figure 11BA parallel processor for supporting and / or implementing a cooling system incorporating an adaptive heat sink in accordance with at least one embodiment is shown;

[0028] Figure 11C is a block diagram of a partitioning unit according to at least one embodiment;

[0029] Figure 11D A graphics multiprocessor for use with a cooling system incorporating an adaptive heat sink in accordance with at least one embodiment is shown;

[0030] Figure 11E A graphics multiprocessor is shown in accordance with at least one embodiment;

[0031] Figure 12A A multi-GPU computing system is shown in accordance with at least one embodiment;

[0032] Figure 12B is a block diagram of a graphics processor according to at least one embodiment;

[0033] Figure 13 is a block diagram illustrating a microarchitecture for a processor, which may include logic circuitry for executing instructions, according to at least one embodiment;

[0034] Figure 14 A deep learning application processor according to at least one embodiment is shown;

[0035] Figure 15 shows a block diagram of a neuromorphic processor according to at least one embodiment;

[0036] Figure 16A is a block diagram of a processing system according to at least one embodiment;

[0037] Figure 16B is a block diagram of a processor having one or more processor cores, an integrated memory controller, and an integrated graphics processor according to at least one embodiment;

[0038] Figure 16C is a block diagram of the hardware logic of a graphics processor core according to at least one embodiment;

[0039] Figures 16D-16E Thread execution logic including an array of processing elements of a graphics processor core is shown in accordance with at least one embodiment;

[0040] Figure 17A illustrates a parallel processing unit according to at least one embodiment;

[0041] Figure 17B illustrates a general processing cluster in accordance with at least one embodiment;

[0042] Figure 17C A memory partitioning unit of a parallel processing unit according to at least one embodiment is shown; and

[0043] Figure 17D A streaming multiprocessor in accordance with at least one embodiment is shown. DETAILED DESCRIPTION

[0044] Air cooling of high-density servers may not be efficient or may be ineffective given the sudden high heat demands caused by the varying computing loads in today's computing components. However, because demands can vary or tend to vary from minimum to maximum cooling requirements, these demands must be met in an economical manner. The varying cooling requirements also reflect the varying thermal characteristics of the data center. In at least one embodiment, the heat generated from components, servers, and racks is cumulatively referred to as a thermal signature or cooling demand because the cooling demand must fully address the thermal signature. In at least one embodiment, the thermal signature or cooling demand of a cooling system is the heat or cooling demand generated by the components, servers, or racks associated with the cooling system and can be a portion of the components, servers, and racks in the data center.

[0045] In at least one embodiment, the deployment of prefabricated (such as mobile) data centers enables the accommodation of data center information technology (IT) components within the prefabricated data center. In at least one aspect, the use of prefabricated data centers reduces costs, speeds up construction, enables relocation flexibility, and reduces operating costs of the deployed data center. Furthermore, these benefits also improve operational reliability. Furthermore, as power density in IT equipment increases and as the computing demands placed on these components increase, various components generate higher levels of heat. These components are data center equipment, which can include graphics processing units (GPUs), central processing units (CPUs), storage components, storage boxes, switches, network equipment, and auxiliary equipment. Due to the varying designs and structures of these components, unique challenges exist in the cooling design of such IT equipment in both prefabricated and permanent data centers. While air cooling through heat sinks can provide limited heat dissipation, at least one embodiment herein provides a method for passive fan heat sinks to react between a first, closed (retracted or overlapped) configuration and a second, extended (or exposed) configuration.

[0046] In at least one embodiment, for economic purposes, when the IT equipment indicates normal heating and sensors determine adequate (and normal) heat dissipation via air cooling of the heat sink, the heat sink remains retracted in a first configuration. In at least one embodiment, when the IT equipment indicates higher than normal heating due to a sudden computing demand, and when sensors determine an associated (and higher than normal) heat dissipation requirement via air cooling of the heat sink, the heat sink can be expanded in a second configuration to expose more surface area to the air cooling system's environment. This exposed area is able to exchange or dissipate more heat to the environment than when the heat sink is in the retracted configuration. In at least one embodiment, the exposed surface area provides proportional heat dissipation. Thus, there are intermediate configurations (variations in exposure) between the retracted and expanded configurations of the heat sink, and there are intermediate thermal characteristics (heat dissipation requirements) or cooling requirements that can be addressed in at least one embodiment. In at least one embodiment, each configuration of the heat sink achieves different cooling capabilities corresponding to different thermal characteristics (heat dissipation requirements) or cooling requirements of the cooling system incorporating the adaptive heat sink.

[0047] Figure 1 1 is a block diagram of an example data center 100 having a cooling system that is subject to the improvements described in at least one embodiment. Data center 100 may be one or more rooms 102 having racks 110 and auxiliary equipment for housing one or more servers on one or more server trays. Data center 100 is supported by a cooling tower 104 located outside data center 100. Cooling tower 104 removes heat from within data center 100 by acting on a primary cooling circuit 106. Furthermore, a cooling distribution unit (CDU) 112 is used between primary cooling circuit 106 and a secondary or secondary cooling circuit 108 to enable heat to be extracted from the secondary or secondary cooling circuit 108 to the primary cooling circuit 106. In one aspect, the secondary cooling circuit 108 can access different piping that feeds into the server trays as needed. Circuits 106, 108 are shown as line diagrams, but one of ordinary skill in the art will recognize that one or more piping features may be used. In one example, flexible polyvinyl chloride (PVC) tubing may be used with associated piping to move fluid along each of circuits 106, 108. In at least one embodiment, one or more coolant pumps can be used to maintain a pressure differential within the loops 106, 108 to enable coolant to move based on temperature sensors in various locations, including in the room, in one or more racks 110, and / or in server boxes or server trays within the racks 110.

[0048] In at least one embodiment, the coolant in the primary cooling loop 106 and the secondary cooling loop 108 can be at least water and an additive, such as ethylene glycol or propylene glycol. In operation, each of the primary cooling loop and the secondary cooling loop has its own coolant. In one aspect, the coolant in the secondary cooling loop can be dedicated to the needs of the components in the server tray or rack 110. The CDU 112 can independently or simultaneously control the coolant in the loops 106 and 108. For example, the CDU can be adapted to control the flow rate so that the coolant is appropriately distributed to absorb the heat generated within the rack 110. In addition, more flexible ducting 114 is provided from the secondary cooling loop 108 to enter each server tray and provide coolant to the electrical and / or computing components. In this disclosure, electrical and / or computing components are used interchangeably to refer to heat-generating components that benefit from the present data center cooling system. The tubing 118 that forms part of the secondary cooling loop 108 can be referred to as a room manifold. Separately, line 116, extending from line 118, may also be part of the auxiliary cooling circuit 108, but may be referred to as a row manifold. Line 114 enters the rack as part of the auxiliary cooling circuit 108, but may be referred to as a rack cooling manifold. Additionally, row manifold 116 extends to all racks along a row in the data center 100. The piping of the auxiliary cooling circuit 108, including manifolds 118, 116, and 114, may be improved by at least one embodiment of the present disclosure. Chiller 120 may be provided in the primary cooling circuit within the data center 102 to support cooling prior to the cooling tower. To the extent that additional circuits exist in the primary control circuit, a person of ordinary skill reading this disclosure will recognize that the additional circuits provide cooling external to the racks and external to the auxiliary cooling circuit; and may be combined with the primary cooling circuit of the present disclosure.

[0049] In at least one embodiment, during operation, heat generated within the server trays of rack 110 can be transferred to coolant exiting rack 110 via the flexible tubing of row manifold 114 of auxiliary cooling loop 108. Accordingly, secondary coolant from CDU 112 (in auxiliary cooling loop 108) used to cool rack 110 moves toward rack 110. Secondary coolant from CDU 112 is transferred from one side of the room manifold, which has tubing 118, to one side of rack 110 via row manifold 116 and passes through one side of the server trays via tubing 114. Spent secondary coolant (or exiting secondary coolant that has removed heat from the computing components) exits from the other side of the server trays (such as entering the left side of the rack for the server trays after circulating through the server trays or components on the server trays and exiting the right side of the rack). Spent secondary coolant exiting the server trays or rack 110 exits from a different side of tubing 114 (such as the outlet side) and moves to a parallel, also outlet side, of row manifold 116. The spent second coolant moves from the row manifolds 116 in a parallel portion of the chamber manifolds 118 traveling in an opposite direction to the incoming second coolant (which may also be refreshed second coolant) and toward the CDUs 112 .

[0050] In at least one embodiment, the spent secondary coolant exchanges its heat with the primary coolant in the primary cooling loop 106 via the CDU 112. The spent secondary coolant is refreshed (e.g., relatively cooler when compared to the temperature of the spent secondary coolant phase) and is ready to circulate back through the auxiliary cooling loop 108 to the computing components. Various flow and temperature control features in the CDU 112 enable control of the exchange of heat from the spent secondary coolant or the flow of secondary coolant into and out of the CDU 112. The CDU 112 can also control the flow of the primary coolant in the primary cooling loop 106.

[0051] In at least one embodiment, the economics of using an air cooling system are improved by the adaptive heat sink of the present invention. Thus, servers and some components within a rack may have cooling requirements that fall between those provided by an air cooling system and a liquid cooling system, and may not be able to obtain cooling capacity or capability beyond that provided by the static heat sink of the radiator in the air cooling system. Separately, such a problem may also exist in some liquid cooling systems that use heat sinks as heat exchange surfaces that contact a cooling medium (such as a coolant). In at least one embodiment of the present disclosure, the adaptive heat sink increases the cooling capacity or capability of a corresponding cooling system, such as at least an air cooling system. In at least one embodiment, the adaptive heat sink also increases the cooling capacity or capability of liquid cooling systems and hybrid cooling systems that combine air and liquid cooling systems.

[0052] Figure 2Ais a block diagram illustrating mobile data center features of a cooling system 200 incorporating an adaptive heat sink 216 having a first configuration of heat sinks in accordance with at least one embodiment. In at least one embodiment, the cooling system 200 is for use with data center equipment such as GPUs, CPUs, storage components, storage boxes, switches, network equipment, and auxiliary equipment. In at least one embodiment, the cooling system 200 includes a heat sink 202 having a heat sink 216 between a first plate 210 and a second plate 208. In at least one embodiment, the heat sink 216 can be two distinct heat sinks associated with the ribbon. In at least one embodiment, the heat sink 216 is a separate, unitary structure incorporating a flexible material throughout the heat sink. In at least one embodiment, only the second plate 208 has the heat sink 216 disposed thereon, wherein the heat sink 216 comprises a material capable of changing shape or structure upon application of heat, such as a bimorph material. The heat sink 216 is bent to include at least one overlapping portion, which in at least Figure 3A 、 3B and further discussed in related topics.

[0053] In at least one embodiment, Figure 2A As shown, the heat sink 216 can be in a first configuration with overlapping portions such that only a first surface area is exposed to the environment. In at least one embodiment, Figure 2A The bottom-most position 218 of the heat sink 216 is shown. The first surface area can be a default or primary surface area that is the outer portion of both sides of a single fin of the heat sink 216. In at least one embodiment, the surface area exposed in the second configuration is greater than the default or primary surface area of the heat sink in the first configuration. In at least one embodiment, the surface area refers to the total surface area of all of the heat sink 216 at a given time. In at least one embodiment, the individual fins of the heat sink 216 contribute to the first surface area, with the sides of all of the fins simultaneously in a single or unified first configuration (e.g., Figure 2A In at least one embodiment, each heat sink has at least two inner surfaces that can overlap with a band-like feature located therebetween, wherein the band-like feature forms an intermediate surface on either side thereof and between the two inner surfaces. In at least one embodiment, when the heat sink is a unitary structure, the middle portion of the heat sink can have an intermediate surface between the inner surfaces of the unitary structure.

[0054] In at least one embodiment, while the interior surface may receive some air in an air-cooled or air-cooled system, the interior surface in the first configuration may not adequately exchange heat with the environment. Thus, in at least one embodiment, in the first configuration of heat sink 216, heat sink 216 only dissipates a first amount of heat to the environment. In one example, the environment refers to air flowing through heat sink 216 from an air-cooled system. In at least one embodiment, the environment refers to the environment within a server box, rack, or data center. In at least one embodiment, the environment may be still air, which may be fan-driven. In at least one embodiment, the environment may be coolant-driven, coolant-driven, or refrigerant-driven.

[0055] In at least one embodiment, the first plate 210 is movable in a single relative direction relative to the second plate 208. In at least one embodiment, a horizontal heat sink configuration may be used wherein the first plate moves horizontally rather than vertically and the heat sinks are horizontal fins. In either case, the goal is to dissipate more heat to the environment when the heat sink is in the expanded configuration than when it is in the collapsed configuration. In at least one embodiment, the second plate 208 is secured to the data center equipment 206 in a suitable manner, either directly or via an intermediate cooling surface 204. In at least one embodiment, the interface between the heat sink 202 and the data center equipment 206 may be a thermal grease that facilitates heat transfer from the data center equipment 206 and the heat sink 202. In at least one embodiment, the thermal grease is a silver-based compound. In at least one embodiment, the intermediate cooling surface 204 may be an auxiliary cooling component in a hybrid cooling system.

[0056] In at least one embodiment, the intermediate cooling surface 204 can be a liquid cooling component adapted to allow coolant to flow through the tubes 220A, 220B. In at least one embodiment, movement of the first plate 210 relative to the second plate 208 causes the fins 216 to spread out and causes the overlapping surfaces (also referred to as inner surfaces) of the individual fins 216 to become non-overlapping with the intermediate portion or strip-like feature and, in at least one embodiment, non-overlapping with each other. In at least one embodiment, an adaptive fin can be constructed without an intermediate surface or strip-like feature, such that the fins are fully curved and overlap directly between the top and bottom of the fin, with a hinge therebetween.

[0057] In at least one embodiment, Figure 2AAs shown, the unfolding of heat sink 216 results in surface areas of the heat sink (previously in overlapping surfaces) being exposed to the environment in a manner similar to the outer surfaces on either side of heat sink 216. In at least one embodiment, the unfolding of heat sink 216 results in a second configuration of heat sink 216. The newly exposed and previously overlapping surface areas cause additional heat to be dissipated from the heat sink to the environment. The additional heat (or accumulated heat, together with the first amount of heat) is referred to as a second amount of heat, which is greater than the first amount of heat when the heat sink is in the first configuration.

[0058] In at least one embodiment, the heat sink 216 is assisted or unassisted in its change from the first configuration to the second configuration. In the auxiliary system, in at least one embodiment, a gear subsystem, an electromagnetic subsystem, a thermoelectric generator subsystem, a thermal reaction subsystem, or a pneumatic subsystem is provided via the motion feature 214 of the cooling system. In at least one embodiment, there can be more than one motion feature 214 throughout the second plate 208 to provide horizontal movement of the first surface or to provide equal exposure of all heat sinks 216 simultaneously. The motion feature, when more than one, is designed to address equal exposure of all overlapping surfaces of the heat sink 216. In at least one embodiment, the motion feature 214 is designed to address equal exposure of all overlapping surfaces of the heat sink 216. Figure 2A Reference to a single one of the heat sinks 216 shown in FIG. 1 or other figures herein is understood to apply equally to all of the heat sinks 216 provided for the heat sink 202 .

[0059] In at least one embodiment, when a motion feature 214 is provided, one or more support features 212 may be present to assist with the motion feature 214. In at least one embodiment, the one or more support features 212 provide stability when the motion feature 214 acts on an area of the first plate 210 to raise the first plate 210. In at least one embodiment, the support features provide at least tension-based support to maintain stability at the four corners of the first plate 210 as the first plate 210 is raised relative to the second plate 208. In at least one embodiment, the tension-based support is provided via an internal spring that is adapted to compress when the full load of the first plate 210 acts on the internal spring. However, in at least one embodiment, when the internal spring 210 is applied to the upward tension support of the first plate 210, the internal spring 210 can maintain its position to reduce the load from the first plate 210.

[0060] In at least one embodiment, in an unassisted system, cooling system 200 may not incorporate first plate 210. In at least one embodiment, cooling system 200 may not incorporate motion features 214 and may or may not incorporate support features 212. In at least one embodiment, in at least one embodiment of an unassisted system, if support features 212 are provided on second plate 210, first plate 208 serves as support features 212 to reduce the load of second plate 210 on heat sink 216. In at least one embodiment, an unassisted system is achieved by incorporating material within the integral structure of one or more heat sinks 216 or within the ribbon features of one or more heat sinks 216. In at least one embodiment, the material is a bimorph material, which is formed from at least two elements with different expansion coefficients, forming dual (or multiple) metal strips. This enables the ribbon features or intermediate portions of the integral structure of one or more heat sinks 216. In at least one embodiment, the bimorph material is caused to change shape or structure, for example, from a curved shape or structure to a relatively straightened shape or structure. In at least one embodiment, heat from associated data center components 206 causes the bimorph material to change shape. This change in shape or structure exposes at least more ambient air to the heat sink 216 than in a bent position. In at least one embodiment, assisted and unassisted systems can be used with the bimorph material to reduce the load on the moving features, or to activate the bimorph material to expand the heat sink 216 before the moving features begin to activate.

[0061] Figure 2B is a block diagram illustrating a cooling system 250 incorporating an adaptive heat sink in a second configuration of heat sinks 258 according to at least one embodiment. In at least one embodiment, a first plate 254 is shown relative to a heat sink such as Figure 2A The topmost position 264 of the second plate 252 of the heat sink 202). The heat sink 258 is relative to Figure 2A The heat sink 216 is in the second configuration. The heat sink 216 is shown in the first configuration. In at least one embodiment, the previously curved middle portion or band feature 262 of each heat sink is larger than when the heat sink is in the Figure 2A In at least one embodiment, the middle portion or strip feature 262 can be straight or curved to the extent that the topmost position 264 is relatively different from the bottommost position 218.

[0062] In at least one embodiment, when the intermediate portion 262 is a hinge or is in the form of a hinge, the intermediate portion 262 is not like Figure 2A 、 2BAs shown, however, the heat sink can have a surface area exposed to the environment that is more so in the second (expanded and exposed) configuration than in the first (contracted and overlapped) configuration. In at least one embodiment, the first and second configurations can be functionally enabled by the ability or capacity of a group of heat sinks to dissipate a first amount of heat that is different from a second amount of heat, respectively, over one or more cycles of data center equipment generating a certain amount of heat. The same functionality can be extrapolated to intermediate configurations and intermediate amounts of heat dissipated by the intermediate configurations.

[0063] In at least one embodiment, intermediate configurations of heat sinks between the first and second configurations can exist to provide intermediate surface areas capable of dissipating a third amount of heat relative to the respective intermediate configurations of heat sinks 258 (or 216). In at least one embodiment, in the auxiliary system, motion feature 260 can have at least one movable portion 260A and at least one fixed portion 260B. At least one support feature 256 can have similar fixed and movable portions. In at least one embodiment, at least one moving component is located within portions 260A, 260B of motion feature 260. In at least one embodiment, at least one moving component is a component that can be associated with a gear subsystem, an electromagnetic subsystem, a thermoelectric generator subsystem, a thermal reaction subsystem, or a pneumatic subsystem as discussed throughout this disclosure.

[0064] In at least one embodiment, cooling systems 200, 250 illustrate at least one band-like feature associated with the heat sink to enable the individual heat sinks to include an overlapping portion. The overlapping portion includes one or more surfaces of the individual heat sinks that are at least partially isolated from the environment in a first configuration and separated in a second configuration to expose the overlapping portion to the environment. Furthermore, in at least one embodiment, at least one band-like feature associated with the heat sink is partially formed from a bimorph material to enable the first plate to move relative to the second plate upon sensing heat acting on the bimorph material. This embodiment can represent the aforementioned unassisted system for moving the heat sink from a first configuration to a second configuration.

[0065] In at least one embodiment, for a pneumatic subsystem of an auxiliary system for a heat sink, a motion feature can utilize coolant from a liquid cooling system to provide motion of a first plate relative to a second plate. In at least one embodiment, a fluid or gas line (including a steam line carrying refrigerant) can receive a cooling fluid (or medium) from a cooling circuit of a data center hosting data center equipment. In at least one embodiment, the pneumatic subsystem uses the cooling fluid to extend a piston and move the first plate relative to the second plate, exposing a surface area of the heat sink in the second configuration of the heat sink.

[0066] Figure 3A and 3BAspects of adaptive heat sinks 300, 320, and 350 are shown in accordance with at least one embodiment. In at least one embodiment, heat sink 300 is either a unitary structure or a combination of multiple parts. In at least one embodiment, heat sink 300 comprises a first portion 302, a second portion 304, and a middle portion 306. In at least one embodiment, middle portion 306 may be a ribbon-like feature formed from a different material than the first and second portions. In at least one embodiment, middle portion 306 is formed from the same material as the first and second portions but may be a portion of different dimensions, allowing it to bend more than either the first or second portions. In at least one embodiment, the portion of different dimensions refers to a thinner portion of the same or similar material as the first and second portions, thereby allowing middle portion 306 to be a ribbon-like feature that is bendable relative to the first and second portions. In at least one embodiment, middle portion 306 is composed of a bimorph metal that changes shape or structure when heat is applied. This enables middle portion 306 to move relative to first portion 302, and also causes the attached second portion 304 to move. The end result is that at least the inner surface of one or more of the first and second portions of the heat sink is exposed to dissipate at least more heat than when the first portion is in its first configuration. In at least one embodiment, this represents an intelligent but processor-free adaptive heat sink for a heat sink.

[0067] In at least one embodiment, the middle portion has two surfaces that meet between the overlapping surfaces of the first and second portions 302, 304. In at least one embodiment, attachment material 308A, 308B is provided to associate the heat sink 300 to a plate of a heat sink, such as Figure 2A 、 2B In at least one embodiment, only the attachment material 308B at the bottom portion 302 is provided to attach the heat sink 300 to the base plate. This can be done without a secondary system heat sink, allowing the heat sink 300 to maintain a vertical configuration regardless of whether it is in the retracted or extended position. In at least one embodiment, the middle portion is sufficiently rigid to change shape or configuration from curved to nearly straight without causing the heat sink to sag.

[0068] In at least one embodiment, the heat sink 300 can be designed to maintain a vertical bottom configuration and a horizontal (or diagonal) top configuration. In at least one embodiment, when the heat sink 300 is in the second or extended configuration, the first portion 302 is vertical and the second portion 304 is horizontal or diagonal. This ensures that at least the inner surface of the first portion 302 is exposed to the environment to increase the heat dissipated from the heat sink 300. In at least one embodiment, the first portion 302 is vertical and the second portion 304 is also vertical. The topmost position in these different configurations can be the top position of the second portion 304 in a fully extended horizontal, diagonal, or vertical position.

[0069] Figure 3B Shown with Figure 3A The first position of the heat sink 300 is different from at least a second configuration of the heat sink 320 . Figure 3B Also shown is a mid-portion or band-like feature 326 having a modified shape or structure. In at least one embodiment, the modified shape is because the mid-portion 326 is not shaped in the same manner as the Figure 3A 306 of the heat sink 300. Even though the middle portion 326 still includes a bend or bend (indicating a shape or structure), the bend or bend is not as extensive as the bend or bend in the middle portion 306 of the heat sink 300. This allows at least the first portion 322 of the heat sink 320 to expose more of its inner surface 328B, and the second portion 324 to expose more of its inner surface 328A. While the inner surface 328A of the second portion 324 may be exposed to the environment due to the gap that may exist between the first portion 322 (see the gap between the first and second portions 302, 304 of the heat sink 300), the inner surface 328A is obstructed due to the first configuration and the gap therebetween. The exposed inner surface 328A allows the heat sink 320 to dissipate more heat in the second configuration than in the first configuration of the heat sink 300.

[0070] In at least one embodiment, Figure 3C It is shown that a plurality of portions 354A, 54B may be disposed on the first portion 352 of the heat sink 350, different from Figure 3A 、 3B 320. In at least one embodiment, the plurality of sections 354A, 354B may be separated by one or more intermediate sections or band features 356A, 356B. In at least one embodiment, the material strength of these sections may allow the second section 354A to extend (or expose surface area) before the third section 354B. In at least one embodiment, this adaptation may enable, for example, Figure 3CThe heat sink 350 can be configured to have multiple intermediate configurations. In at least one embodiment, the different portions 354A, 354B can have different angles in their second configuration. The angles can include vertical, horizontal, or diagonal extensions relative to the first portion 352 or relative to a second plate that is fixed and carries the heat sink 350.

[0071] Figure 4A 、 4B 4C shows a processor support subsystem 400; 430; 460 according to at least one embodiment for implementing a cooling system incorporating an adaptive heat sink, such as Figures 2A-3C The cooling system 200; 250 and the heat sink 300; 320; 350. The processor support subsystem 400; 430; 460 is at least Figure 2A 、 2B The processor-supported subsystems 400; 430; 460 support auxiliary systems for moving the first plate relative to the second plate, such as at least reference to Figure 2A 、 2B The processor-supported subsystem 400; 430; 460 is capable of sending at least a corresponding input to a corresponding controller to cause movement of the first plate relative to the second plate. Figure 4A Subsystem 400 is a processor-supported mechanical gear subsystem, which in turn is supported by controller inputs to and from an electromechanical or mechanical controller 406 and a servo motor 408. In at least one embodiment, when the processor senses a thermal signature or cooling need, the processor is capable of determining the cooling capacity or capability of the associated one or more heat sinks. The processor is adapted to react by at least activating the air cooling system (if not already activated) and by sending an input to the electromechanical or mechanical controller 406 to cause the heat sink to move to a second configuration. In at least one embodiment, the processor is capable of controlling multiple controllers for multiple heat sinks to simultaneously or independently move the respective heat sinks to the respective second configurations.

[0072] In use Figure 4AIn at least one embodiment of the processor-supported subsystem 400, a processor input is received in an electromechanical or mechanical controller 406, which sends another control input to a motor such as a servo motor 408. In at least one embodiment, the controller 406 is electromechanical and is adapted to provide signals to start, stop, and control the speed of the servo motor 408. In at least one embodiment, the controller 406 is mechanical and is adapted to maintain or remove a mechanical braking force associated with gears 410. In at least one embodiment, a set of gears 410 converts the mechanical output of the servo motor 408 into lateral motion of the piston 404. In at least one embodiment, the piston 404 has threads that are threadedly connected by at least one gear in the set of gears 410 to move the piston 404 in an upward or downward lateral motion. In at least one embodiment, the set of gears may include bevel gears for converting rotary motion into linear motion. In at least one embodiment, the top of the piston 404 may be associated with a first plate that is movable relative to a second plate. In at least one embodiment, the second plate houses a plurality of actuators for the piston 404. Figure 4A The processor associated with the electromechanical or mechanical controller 406 may reside within the housing 402 and may be part of a distributed control system. In at least one embodiment, the processor may be a Figure 9A In at least one embodiment, the processor is capable of Figures 4A-4C The detection of thermal requirements, the determination of configuration requirements, and the determination of a portion of a heat sink that needs to be activated to automatically change the configuration are performed within one of the subsystems 400; 430; 460 supported by the processor.

[0073] In at least one embodiment, Figure 4B The subsystem 430 is processor-supported based on an electromagnetic subsystem that is supported by processor inputs to and from an electromagnetic controller 436 and an electromagnet 440. In at least one embodiment, when the processor senses a thermal signature or cooling requirement, the processor is capable of determining the cooling capacity or capability of the associated one or more heat sinks. The processor is adapted to react by at least activating the air cooling system (if not already activated) and by sending an input to the electromagnetic controller 436 to cause the heat sink to move to the second configuration.

[0074] In use Figure 4BIn at least one embodiment of the processor-supported subsystem 430, a processor input is received in an electromagnetic controller 436, which sends another control input to one or more electromagnets 440. In at least one embodiment, the electromagnets 440 convert the electrical input of the electromagnetic controller 436 into a magnetic attraction or repulsion force, the sequence of which causes the piston 434 to move laterally from within the tube 438. In at least one embodiment, the top of the piston 434 can be associated with a first plate that is movable relative to a second plate. In at least one embodiment, the second plate houses a plurality of electrodes for the piston 434 to move laterally from within the tube 438. Figure 4B The housing 432 may include a processor supporting one or all of the components 436-440 of the subsystem 430. In at least one embodiment, the processor associated with the electromagnetic controller 436 may reside within the housing 432 and may be part of a distributed control system. In at least one embodiment, any of the controllers 406, 436, 466 may be adapted to return input or signals to the processor in response to a query regarding the position of the piston 404, 434, 436. In at least one embodiment, any of the controllers 406, 436, 466 may also be adapted to return input or signals to the processor in response to a query regarding the status of the controller 406, 436, 466.

[0075] In at least one embodiment, Figure 4C A processor-supported subsystem 460 is based on a pneumatic subsystem supported by processor inputs to and from a pneumatic controller 466 and a fluidics component 470A. In at least one embodiment, when the processor senses a thermal signature or cooling requirement, the processor is capable of determining the cooling capacity or capability of the associated one or more heat sinks. The processor is adapted to react by at least activating the air cooling system (if not already activated) and by sending an input to the pneumatic controller 436 to cause the heat sink to move to a second configuration. In at least one embodiment, when the processor senses a thermal signature or cooling requirement, the processor is capable of determining that the cooling capacity or capability of the associated one or more heat sinks is improved by a liquid cooling system. The processor is adapted to react by at least activating the air cooling system (if not already activated) and the liquid cooling system (if also not already activated) and by sending an input to the pneumatic controller 436 to cause the heat sink to move to the second configuration using fluid from the liquid cooling system as a secondary function of the liquid cooling system.

[0076] In use Figure 4CIn at least one embodiment of the processor supported subsystem 460, a processor input is received in a pneumatic controller 466 which sends another control input to one or more valves 470A via line 474. In at least one embodiment, the control input is electrical to an electro-pneumatic valve forming one or more valves 470A, or pneumatic to a fully pneumatic valve forming one or more valves 470A. In at least one embodiment, the one or more valves 470A convert the electrical input or pneumatic input of the pneumatic controller 436 into a fluid within tube 468 which causes lateral movement of the piston 464. In at least one embodiment, the fluid is from a liquid cooling system entering and exiting via lines or conduits 472A, 472B, or may be provided from a heat sink dedicated liquid storage tank for an adaptive heat sink via the same lines 472A, 472B. In at least one embodiment, the top of the piston 464 may be associated with a first plate that is movable relative to a second plate. In at least one embodiment, the second plate houses a liquid cooler for use in the heat sink. Figure 4B The housing 432 may contain one or all of the components 436-440 of the subsystem 430 supported by the processor. In at least one embodiment, the processor associated with the pneumatic controller 466 may reside within the housing 462 and may be part of a distributed control system. In at least one embodiment, the housing 462 may also be capable of detecting fluid leaks and may include a detector to notify the processor of fluid leaks as part of the status information queried by the processor. In at least one embodiment, each of the pistons 404, 434, 464 is returned to its initial position within the corresponding tube or housing by a spring acting on the bottom of the piston.

[0077] In at least one embodiment, for Figure 4C The pneumatic controller 466 can be used to provide a pneumatic control action via one or more bypass pumps or inline pumps on either outlet side of the pipes 472A and 472B. In addition, the pneumatic controller 466 can apply a pumping action on both the inlet and outlet sides of the pipes 472A and 472B. In this way, when the pneumatic controller 466 can issue a suction action instead of a pushing action. In at least one embodiment, a check valve can be used as one or more valves 470A to keep the fluid in the pipe 468 or keep the fluid outside the pipe 468 during the upward or downward pumping to move from the first configuration to the second configuration or from the second configuration to the first configuration. When the pneumatic controller 466 is causing the series action, there are both suction and pushing actions acting on the fluid in the pipes 472A and 472B. By causing the series flow, the flow rate achieved can be higher, so that the configuration switching can be achieved faster.

[0078] Figure 4D and 4EA plan view of a plate 480, 490 is shown, in accordance with at least one embodiment, incorporating one or more processor-supported subsystems to implement a cooling system incorporating an adaptive heat sink. In at least one embodiment, the plate is a second plate that remains fixedly or removably coupled to data center equipment or an intercooling surface. The plan view illustrates a first location or region 484 of one or more motion features and a second location or region 482 of one or more support features. In at least one embodiment, the motion features can include any of the processor-supported subsystems 400, 430, 460. In at least one embodiment, the plan view illustrates locations 484 within the plate 480, 490 available for supporting at least the dimensions of the corresponding housing 402, 432, 462. In at least one embodiment, the region 482 for support features can be significantly smaller than the region 484 for motion features. In at least one embodiment, as shown in the plan view of plate 490, there are only two motion features in region 494, and four support features are provided for these features in region 492. In at least one embodiment, there may be only one motion feature at the center of the plate 490 and four support features at the four corners of the plate. In at least one embodiment, the layout of the areas for motion features and support features is processed in part to maximize the space for the adaptive heat sink. In the case of a hybrid adaptive heat sink, the heat sink is able to support some of the weight of the first plate that moves relative to the base plate 480; 490. In this way, fewer support features or no support features may be required. In at least one embodiment, when the adaptive heat sink is not assisted, no area is required for motion or support features. In at least one embodiment, the entire plate 480; 490 is provided with an adaptive heat sink with dual piezoelectric chip capabilities.

[0079] Figure 5 is usable for use or manufacture according to at least one embodiment Figures 2A-4E 6A-17D. In at least one embodiment, step 502 is for providing a heat sink between a first plate and a second plate, wherein the first plate is movable relative to the second plate. In at least one embodiment, the first plate is movable with or without an assisted feature provided with the heat sink. In at least one embodiment, step 504 enables dissipation of a first amount of heat to the environment in a first configuration of the heat sink. In at least one embodiment, the first configuration is at least as described in Figure 2A As provided in the example of , where the heat sink has one or more overlapping portions.

[0080] In at least one embodiment, a determination is made, via step 506, that the associated data center equipment could benefit from additional heat dissipation (or additional cooling). In at least one embodiment, this determination is made based in part on the current amount of heat generated and the current amount of heat dissipated. When a positive difference is determined between the current amount of heat generated and the amount of heat dissipated, additional cooling may be required. In at least one embodiment, this determination is based in part on the active ongoing computations and the potential computation required of the data center equipment. When the additional load is expected to cause the potential computation to represent 80% to 100% utilization of the data center equipment, it can be expected that the amount of heat generated may increase beyond the current amount of heat dissipated. In at least one embodiment, a determination is also made, via step 508, whether the current configuration of the heat sink is unable to provide more than a first amount of heat. In at least one embodiment, this determination in step 508 can be made by allowing the generated heat to accumulate without dissipating it for a period of time (e.g., several seconds), or by noting the rate of increase in the amount of heat generated (and heat not dissipated). A positive determination in step 508 enables step 510 to move the first plate relative to the second plate to expose the surface area of the heat sink to the environment in a second configuration of the heat sink. In the second configuration, the heat sink dissipates a second amount of heat to the environment that is greater than the first amount of heat. When the determination in step 508 is negative, further determinations via step 506 are looped until a second configuration of heat sinks is desired.

[0081] In at least one embodiment, a learning subsystem having at least one processor may be used to determine when to move from a first configuration to a second configuration and back to the first configuration. In at least one embodiment, the learning subsystem is used to evaluate temperatures sensed from a component, server, or rack in a data center, where surface areas of heat sinks are associated with different amounts of heat dissipation. In at least one embodiment, the learning subsystem is adapted to operate with multiple heat sinks simultaneously or independently. In at least one embodiment, the learning subsystem may then use the collective heat sinks of the multiple heat sinks or the surface areas of all the heat sinks of the individual heat sinks. In at least one embodiment, the learning subsystem may then use the amounts of heat dissipation associated with the collective heat sinks of the multiple heat sinks or all the heat sinks of the individual heat sinks. In at least one embodiment, the learning subsystem may provide an output associated with at least one temperature to facilitate movement of the first plate relative to the second plate, thereby exposing the surface areas of the heat sinks.

[0082] In at least one embodiment, the learning subsystem may be implemented via a deep learning application processor, such as Figure 14 The processor 1400 in FIG. 1 may use a neuron 1502 and its components, which may be implemented using circuits or logic, including Figure 15Thus, the learning subsystem includes at least one processor for evaluating temperatures within servers of one or more racks, wherein surface areas of heat sinks are associated with different amounts of heat dissipation.

[0083] Once trained, the learning subsystem will be able to provide outputs having instructions associated with the surface area of the collective heat sinks of the plurality of heat sinks or the individual heat sinks. Figures 4A-4C As discussed, instructions are provided to the appropriate controller associated with the appropriate motion feature. In at least one embodiment, an output is provided from a learning subsystem executing a machine learning model. The machine learning model is adapted to process temperature using a plurality of neuron levels of the machine learning model, the plurality of neuron levels having temperatures and previously associated surface areas of the heat sink. The machine learning model is adapted to provide an output associated with at least one temperature to at least one controller. The output is provided after evaluating the previously associated surface area and the previously associated temperature of the heat sink.

[0084] Alternatively, in at least one embodiment, the differences in temperatures and subsequent temperatures from different heat sinks associated with different adaptive heat sinks, along with the necessary surface area available in different configurations for the different heat sinks, are used to train a learning system to recognize when to activate and deactivate associated controllers for different motion signatures of the different heat sinks. Once trained, the learning subsystem, through an appropriate controller (also referred to herein as a centralized control system or a distributed control system) described elsewhere in this disclosure, can provide one or more outputs to control the motion signatures of the different heat sinks in response to temperatures received from temperature sensors associated with different components, different servers, and different racks.

[0085] In at least one embodiment, aspects of the processing of the deep learning subsystem may be implemented using a method according to reference Figure 14 、 15 The collected information is processed using the discussed features. In one example, the processing of temperature uses multiple neuron levels of a machine learning model that are loaded with one or more of the above-mentioned collected temperature features and the corresponding surface areas of the different heat sinks (collectively or independently). The learning subsystem performs training, which can be represented as an evaluation of the temperature change associated with the previous surface area (or change in surface area) of each heat sink based on adjustments made to one or more controllers associated with the different heat sinks. The neuron level can store values associated with the evaluation process and can represent an association or correlation between the temperature change and the surface area of each of the different heat sinks.

[0086] In at least one embodiment, the learning subsystem, once trained, is able to determine the surface areas of different heat sinks (and appropriate positioning of corresponding pistons to achieve the surface areas) in an application to achieve cooling to temperatures (or changes, such as temperature reductions) associated with the cooling needs of, for example, different servers, different racks, or different components. Since the cooling needs must be within the cooling capabilities of the appropriate cooling subsystem, the learning subsystem is able to select which heat sinks (including types of air and / or liquid cooling systems) to activate to address the temperatures sensed from the different servers, different racks, or different components. The learning subsystem can use the collected temperatures and the previously associated surface areas of different heat sinks used to achieve the collected temperatures (e.g., or differences) to provide one or more outputs associated with the required surface areas (or piston movements) of different heat sinks to address the different cooling needs reflected by the temperatures (representing reductions) of different servers, different racks, or different components compared to the current temperatures.

[0087] In at least one embodiment, the result of the learning subsystem is one or more outputs to controllers associated with different heat sinks that modify the configuration of the heat sinks in the corresponding different heat sinks in response to sensed temperatures from different servers, different racks, or different components. The modification of the configuration of the heat sinks achieves a determined exposed surface area and a corresponding determined heat dissipation from different servers, different racks, or different components that require cooling. The modified surface area can be maintained until the temperature in the required area reaches a temperature associated with the surface area known to the learning subsystem. In at least one embodiment, the modified surface area can be maintained until the temperature in the area changes by a determined value. In at least one embodiment, the modified surface area can be maintained until the temperature in the area reaches a rated temperature for the different servers, different racks, different components, or different surface areas made possible by the configuration of the heat sinks.

[0088] In at least one embodiment, the processor and the controller can work in conjunction. A processor (also known as a centralized or distributed control system) is at least one processor having at least one logic unit to control a controller associated with one or more heat sinks. In at least one embodiment, at least one processor is located in a data center, such as Figure 7A The controller facilitates movement of respective heat sinks associated with respective heat sinks and facilitates cooling of an area in a data center in response to a temperature sensed in the area. In at least one embodiment, the at least one processor is such as Figure 9AIn at least one embodiment, the at least one logic unit may be adapted to receive temperature values from temperature sensors associated with the server or one or more racks and to facilitate movement of heat sinks associated with one or more heat sinks. In at least one embodiment, the controller includes a microprocessor to utilize its corresponding motion characteristics to perform its communication and control functions.

[0089] In at least one embodiment, such as Figure 9A The processors of the processor cores of the multi-core processors 905, 906 may include a learning subsystem for evaluating temperatures from sensors at various locations in the data center, such as various locations associated with a server or rack, or even components within a server, having a surface area associated with at least a heat sink. The learning subsystem may provide an output, such as instructions associated with at least the temperature or the surface area, for facilitating movement of a plate having a heat sink therein such that the heat sink expands from a first configuration to a second configuration to meet a different cooling requirement.

[0090] In at least one embodiment, the learning subsystem executes a machine learning model to process temperature using multiple neuron levels of the machine learning model that have temperatures and previously associated surface areas of heat sinks for different heat sinks. The machine learning model may use Figure 15 The neuronal structure described in Figure 14 The machine learning model provides outputs associated with the surface areas to one or more controllers based on a previously associated evaluation of the surface areas. In addition, the processor's instruction outputs, such as pins of a connector bus or balls of a ball grid array, enable the outputs to communicate with the one or more controllers to modify the surface areas of the first heat sink of the first heat sink while maintaining the second heat sink of the second heat sink in its existing configuration (existing surface areas exposed to the environment, unchanged). Thus, in at least one embodiment, the learning subsystem is capable of controlling the surface areas of the heat sinks of one or more heat sinks based on the cooling requirements of the one or more heat sinks by associated data center equipment.

[0091] In at least one embodiment, the present disclosure relates to at least one processor for a cooling system or a system having at least one processor. The at least one processor includes at least one logic unit for training one or more neural networks having hidden layers of neurons to evaluate temperatures sensed from components, servers, or racks in a data center, wherein the surface area of a heat sink is associated with different amounts of heat dissipated by the heat sink. The at least one logic unit is further adapted to provide an output associated with the at least one temperature for facilitating movement of a first plate relative to a second plate having a heat sink therebetween, thereby exposing the surface area of the heat sink. In at least one embodiment, the at least one logic unit is adapted to output at least one instruction associated with the at least one temperature for facilitating movement of the first plate relative to the second plate having a heat sink therebetween to expose the surface area of the heat sink.

[0092] In at least one embodiment, the present disclosure implements an unassisted system of heat sinks. In at least one embodiment, the unassisted system supports a heat sink having heat sinks that are movable to dissipate heat to the environment by exposing a surface area of the heat sink in a first configuration that is greater than the primary surface area of the heat sink in a second configuration. In at least one embodiment, the unassisted system is a processor-less subsystem that enables the heat sink to move. In at least one embodiment, the processor-less subsystem is implemented by a bimorph metal associated with each heat sink. In at least one embodiment, the bimorph metal unfolds a portion of each heat sink, thereby exposing the surface area of the heat sink. Because the bimorph material changes shape or structure in response to heat without the need for external instructions, the unassisted system of heat sinks has the intelligence to perform heat dissipation operations on its own. In at least one embodiment, the heat sink is adapted to expand in response to heat from an associated component above a threshold and to contract in response to heat below a threshold.

[0093] In at least one embodiment, the present disclosure is directed to a cooling system with an adaptable heat sink that can provide cooling from relatively low-density racks having approximately 10KW to higher density cooling of approximately 30KW using an air-based cooling subsystem, and additionally provide medium density cooling of between 30-50KW using deployable heat sinks; provide cooling from approximately 50KW to 60KW using a liquid-based cooling subsystem, but can also provide cooling from approximately 60KW to 80kW using deployable heat sinks with a liquid cooling system to dissipate heat via a combination of air and liquid cooling media.

[0094] In at least one embodiment, step 502 of method 500 includes associating at least one banding feature with the heat sink. In step 504, method 500 enables, via the at least one banding feature, each heat sink to include an overlapping portion that is at least partially shielded from the environment in a first configuration. Thus, in the first configuration, the heat sink dissipates a first amount of heat from the associated data center equipment. In at least one embodiment, step 510 includes enabling a configuration of the at least one banding feature to be varied to expose the overlapping portion to the environment in a second configuration. Furthermore, in at least one embodiment, step 510 is implemented using a fluid received from a cooling circuit of a data center that hosts the data center equipment using a fluid or gas line. Step 510 then includes extending a piston associated with a pneumatic subsystem using the cooling fluid to move the first plate relative to the second plate, thereby exposing a surface area of the heat sink in the second configuration of the heat sink.

[0095] Data Center

[0096] Figure 6A An example data center 600 is shown where data from Figure 2A-5 In at least one embodiment, the data center 600 includes a data center infrastructure layer 610, a framework layer 620, a software layer 630, and an application layer 640. In at least one embodiment, for example, Figure 2A-5 The features described in the components of the cooling system incorporating the adaptive heat sink can be implemented within or in conjunction with the exemplary data center 600. In at least one embodiment, the infrastructure layer 610, the framework layer 620, the software layer 630, and the application layer 640 can be provided in part or in whole by computing components located on server trays in the racks 210 of the data center 200. This enables the cooling system of the present disclosure to direct cooling to certain of the computing components in an efficient and effective manner. In addition, aspects of the data center including the data center infrastructure layer 610, the framework layer 620, the software layer 630, and the application layer 640 can be used to support at least the above-mentioned aspects of the present disclosure. Figure 2A-5 The intelligent control of the controller in the adaptive heat sink cooling system is discussed. For example, reference Figures 6A-17D The discussion can be understood as applying to the implementation or support of Figure 2A-5 The hardware and software features required for a data center cooling system combined with an adaptive heat sink.

[0097] In at least one embodiment, Figure 6AAs shown, the data center infrastructure layer 610 may include a resource coordinator 612, group computing resources 614, and node computing resources ("node CRs") 616(1)-616(N), where "N" represents any complete positive integer. In at least one embodiment, the node CRs 616(1)-616(N) may include, but are not limited to, any number of central processing units ("CPUs") or other processors (including accelerators, field programmable gate arrays (FPGAs), graphics processors, etc.), memory devices (such as dynamic read-only memories), storage devices (such as solid-state or disk drives), network input / output ("NW I / O") devices, network switches, virtual machines ("VMs"), power modules, and cooling modules. In at least one embodiment, one or more of the node CRs 616(1)-616(N) may be servers having one or more of the above-mentioned computing resources.

[0098] In at least one embodiment, the grouped computing resources 614 may include separate groups of node CRs housed in one or more racks (not shown), or may include many racks (also not shown) housed in data centers at various geographic locations. The separate groups of node CRs within the grouped computing resources 614 may include computing, networking, memory, or storage resources that may be configured or allocated to support groupings of one or more workloads. In at least one embodiment, several node CRs comprising a CPU or processor may be grouped in one or more racks to provide computing resources to support one or more workloads. In at least one embodiment, one or more racks may also include any number of power modules, cooling modules, and network switches, in any combination.

[0099] In at least one embodiment, resource coordinator 612 may configure or otherwise control one or more nodes CR 616(1)-616(N) and / or grouped computing resources 614. In at least one embodiment, resource coordinator 612 may comprise a software design infrastructure ("SDI") management entity for data center 600. In at least one embodiment, resource coordinator 612 may comprise hardware, software, or some combination thereof.

[0100] In at least one embodiment, Figure 6AAs shown, the framework layer 620 includes a job scheduler 622, a configuration manager 624, a resource manager 626, and a distributed file system 628. In at least one embodiment, the framework layer 620 may include a framework that supports software 632 of the software layer 630 and / or one or more applications 642 of the application layer 640. In at least one embodiment, the software 632 or the application 642 may include web-based service software or applications, such as those provided by Amazon Web Services, Google Cloud, and Microsoft Azure. In at least one embodiment, the framework layer 620 may include, but is not limited to, a free and open source software web application framework, such as Apache Spark™ (hereinafter referred to as "Spark"), which can utilize the distributed file system 628 for large-scale data processing (such as "big data"). In at least one embodiment, the job scheduler 622 may include a Spark driver to facilitate scheduling workloads supported by various layers of the data center 600. In at least one embodiment, the configuration manager 624 may be capable of configuring different layers, such as the software layer 630 and the framework layer 620, which includes Spark and a distributed file system 628 for supporting large-scale data processing. In at least one embodiment, the resource manager 626 can manage clustered or grouped computing resources that are mapped to or allocated to support the distributed file system 628 and the job scheduler 622. In at least one embodiment, the clustered or grouped computing resources can include grouped computing resources 614 on the data center infrastructure layer 610. In at least one embodiment, the resource manager 626 can coordinate with the resource coordinator 612 to manage these mapped or allocated computing resources.

[0101] In at least one embodiment, the software 632 included in the software layer 630 may include software used by at least a portion of the node CRs 616(1)-616(N), the grouped computing resources 614, and / or the distributed file system 628 of the framework layer 620. The one or more types of software may include, but are not limited to, Internet web page search software, email virus scanning software, database software, and streaming video content software.

[0102] In at least one embodiment, the one or more applications 642 included in the application layer 640 may include one or more types of applications used by at least a portion of the node CRs 616(1)-616(N), the grouped computing resources 614, and / or the distributed file system 628 of the framework layer 620. The one or more types of applications may include, but are not limited to, any number of genomics applications, cognitive computing, and machine learning applications, including training or inference software, machine learning framework software (such as PyTorch, TensorFlow, Caffe, etc.), or other machine learning applications used in conjunction with one or more embodiments.

[0103] In at least one embodiment, any of configuration manager 624, resource manager 626, and resource coordinator 612 can implement any number and type of self-modification actions based on any number and type of data obtained in any technically feasible manner. In at least one embodiment, the self-modification actions can relieve a data center operator of data center 600 from making potentially poor configuration decisions and can avoid underutilized and / or poorly performing portions of the data center.

[0104] In at least one embodiment, data center 600 may include tools, services, software, or other resources for training one or more machine learning models or using one or more machine learning models to predict or infer information according to one or more embodiments herein. In at least one embodiment, a machine learning model can be trained by calculating weight parameters according to a neural network architecture using the software and computing resources described above with respect to data center 600. In at least one embodiment, using the weight parameters calculated using one or more training techniques herein, a trained machine learning model corresponding to one or more neural networks can be used to infer or predict information using the resources described above with respect to data center 600. As previously discussed, deep learning techniques can be used to support intelligent control of controllers in the adaptive heat sink cooling system described herein by monitoring zone temperatures in the data center. Deep learning can be facilitated using any suitable learning network and computing capabilities of data center 600. Thus, hardware in the data center can simultaneously or concurrently support deep neural networks (DNNs), recurrent neural networks (RNNs), or convolutional neural networks (CNNs). For example, once a network is trained and successfully evaluated to identify data in a subset or slice, the trained network can provide similar representative data for use with the collected data.

[0105] In at least one embodiment, the data center 600 can use a CPU, an application-specific integrated circuit (ASIC), a GPU, an FPGA, or other hardware to use the above resources to perform training and / or reasoning. In addition, one or more of the above software and / or hardware resources can be configured as a service to allow users to train or perform information reasoning, such as pressure, flow rate, temperature, and location information or other artificial intelligence services.

[0106] Reasoning and training logic

[0107] Reasoning and / or training logic 615 may be used to perform reasoning and / or training operations associated with one or more embodiments. In at least one embodiment, reasoning and / or training logic 615 may be used to perform reasoning and / or training operations associated with one or more embodiments. Figure 6A Inference and / or training logic 615 is used to reason or predict operations based at least in part on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases herein. In at least one embodiment, reasoning and / or training logic 615 may include, but is not limited to, hardware logic in which computing resources are dedicated or otherwise uniquely used in conjunction with weight values or other information corresponding to one or more layers of neurons within a neural network. In at least one embodiment, reasoning and / or training logic 615 may be used in conjunction with an application specific integrated circuit (ASIC), such as an ASIC from Google. Processing unit, Inference Processing Unit (IPU) from GraphcoreTM or from Intel Corp (such as "LakeCrest") processor.

[0108] In at least one embodiment, the reasoning and / or training logic 615 can be used in conjunction with central processing unit (CPU) hardware, graphics processing unit (GPU) hardware, or other hardware (such as a field programmable gate array (FPGA)). In at least one embodiment, the reasoning and / or training logic 615 includes, but is not limited to, a code and / or data storage model that can be used to store code (e.g., graphics code), weight values, and / or other information, including bias values, gradient information, momentum values, and / or other parameter or hyperparameter information. In at least one embodiment, each code and / or data storage module is associated with a dedicated computing resource. In at least one embodiment, the dedicated computing resource includes computing hardware that also includes one or more ALUs that perform mathematical functions (e.g., linear algebra functions) only on the information stored in the code and / or data storage module, and stores the results stored therefrom in an activated storage module of the reasoning and / or training logic 615.

[0109] Figure 6B 、 Figure 6C Inference and / or training logic according to at least one embodiment is shown, such as in Figure 6A The reasoning and / or training logic 615 is used to perform reasoning and / or training operations associated with at least one embodiment of the present disclosure. Figure 6B and / or Figure 6C Provides details about the inference and / or training logic 615. Distinguished from the computational hardware 602, 606 by the use of an arithmetic logic unit (ALU) 610 Figure 6B and Figure 6C Inference and / or training logic 615. In at least one embodiment, each of computational hardware 602 and computational hardware 606 includes one or more ALUs that perform mathematical functions (e.g., linear algebraic functions) on information stored in code and / or data memory 601 and information stored in code and / or data memory 605, respectively, with the results stored in activation memory 620. Thus, unless otherwise specified, Figure 6B and Figure 6C may be substituted and used interchangeably.

[0110] In at least one embodiment, the inference and / or training logic 615 may include, but is not limited to, code and / or data storage 601 to store forward and / or output weights and / or input / output data and / or other parameters for neurons or layers of a neural network trained and / or used for inference in at least one embodiment. In at least one embodiment, the training logic 615 may include or be coupled to code and / or data storage 601 for storing graph code or other software to control timing and / or sequence, wherein weight and / or other parameter information is loaded to configure logic, including integer and / or floating point units (collectively referred to as arithmetic logic units (ALUs)). In at least one embodiment, the code (such as graph code) loads weight or other parameter information into the processor ALU based on the architecture of the neural network to which the code corresponds. In at least one embodiment, the code and / or data storage 601 stores input / output data and / or weight parameters during training and / or inference using aspects of at least one embodiment during forward propagation of the weight parameters for each layer of the neural network trained or used in conjunction with at least one embodiment. In at least one embodiment, any portion of code and / or data storage 601 may be included within other on-chip or off-chip data storage, including a processor's L1, L2, or L3 cache or system memory.

[0111] In at least one embodiment, any portion of code and / or data storage 601 may be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, code and / or code and / or data storage 601 may be cache memory, dynamically randomly addressable memory ("DRAM"), static randomly addressable memory ("SRAM"), non-volatile memory (such as flash memory), or other storage. In at least one embodiment, the choice of whether code and / or data storage 601 is internal or external to a processor, for example, or including DRAM, SRAM, flash memory, or some other storage type, may depend on the available storage space on-chip or off-chip, the latency requirements of the training and / or inference functions being performed, the data batch size used in inference and / or training of the neural network, or some combination of these factors.

[0112] In at least one embodiment, the inference and / or training logic 615 may include, but is not limited to, code and / or data storage 605 for storing backpropagation and / or output weights and / or input / output data corresponding to neurons or layers of a neural network trained and / or used for inference in at least one embodiment. In at least one embodiment, the code and / or data storage 605 stores weight parameters and / or input / output data for each layer of the neural network trained or used in conjunction with at least one embodiment during backpropagation of input / output data and / or weight parameters during training and / or inference using at least one embodiment. In at least one embodiment, the training logic 615 may include or be coupled to code and / or data storage 605 for storing graph code or other software to control the timing and / or sequence in which weight and / or other parameter information is loaded to configure logic, which includes integer and / or floating point units (collectively referred to as arithmetic logic units (ALUs)).

[0113] In at least one embodiment, code (such as graph code) loads weights or other parameter information into the processor ALU based on the architecture of the neural network to which the code corresponds. In at least one embodiment, any portion of code and / or data storage 605 can be included with other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory. In at least one embodiment, any portion of code and / or data storage 605 can be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, code and / or data storage 605 can be cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other storage. In at least one embodiment, the choice of whether code and / or data storage 605 is internal or external to the processor, for example, including DRAM, SRAM, flash memory, or some other storage type, depends on whether the available storage is on-chip or off-chip, the latency requirements of the training and / or inference functions being performed, the batch size of data used in inference and / or training of the neural network, or some combination of these factors.

[0114] In at least one embodiment, code and / or data store 601 and code and / or data store 605 may be separate storage structures. In at least one embodiment, code and / or data store 601 and code and / or data store 605 may be the same storage structure. In at least one embodiment, code and / or data store 601 and code and / or data store 605 may be partially the same storage structure and partially separate storage structures. In at least one embodiment, any portion of code and / or data store 601 and code and / or data store 605 may be included with other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory.

[0115] In at least one embodiment, inference and / or training logic 615 may include, but is not limited to, one or more arithmetic logic units (“ALUs”) 610 , including integer and / or floating point units, for performing logical and / or mathematical operations based at least in part on or directed by training and / or inference code (e.g., graph code), the results of which may produce activations (e.g., output values from a layer or neuron within a neural network) stored in activation storage 620 , which are functions of input / output and / or weight parameter data stored in code and / or data storage 601 and / or code and / or data storage 605 . In at least one embodiment, activations are performed in response to executing instructions or other code, linear algebra and / or matrix-based mathematics performed by ALU 610 to generate activations stored in activation storage 620, wherein weight values stored in code and / or data storage 605 and / or in code and / or data storage 601 are used as operands with other values, such as bias values, gradient information, momentum values, or other parameters or hyperparameters, any or all of which may be stored in code and / or data storage 605 and / or code and / or data storage 601 or other on-chip or off-chip storage.

[0116] In at least one embodiment, one or more ALUs 610 are included in one or more processors or other hardware logic devices or circuits, while in another embodiment, one or more ALUs 610 may be external to the processor or other hardware logic devices or circuits that use them (such as coprocessors). In at least one embodiment, one or more ALUs 610 may be included within an execution unit of a processor or otherwise included in a group of ALUs accessible by the execution units of a processor, which may be within the same processor or distributed between different processors of different types (such as a central processing unit, graphics processing unit, fixed function unit, etc.). In at least one embodiment, code and / or data storage 601, code and / or data storage 605, and activation storage 620 may be on the same processor or other hardware logic device or circuit, while in another embodiment, they may be on different processors or other hardware logic devices or circuits or some combination of the same and different processors or other hardware logic devices or circuits. In at least one embodiment, any portion of activation storage 620 may be included with other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory. Additionally, inference and / or training code may be stored with other code accessible to a processor or other hardware logic or circuitry and may be retrieved and / or processed using the processor's fetch, decode, schedule, execute, exit, and / or other logic circuitry.

[0117] In at least one embodiment, activation storage 620 may be cache memory, DRAM, SRAM, non-volatile memory (such as flash memory), or other storage. In at least one embodiment, activation storage 620 may be completely or partially internal or external to one or more processors or other logic circuits. In at least one embodiment, activation storage 620 may be selected to be internal or external to the processor, for example, or to comprise DRAM, SRAM, flash memory, or other storage types, depending on the storage available on or off chip, the latency requirements for performing training and / or inference functions, the batch size of data used in inferring and / or training neural networks, or some combination of these factors. In at least one embodiment, Figure 6B The inference and / or training logic 615 shown in FIG can be used in conjunction with an application specific integrated circuit (“ASIC”), such as the one from Google. Processing unit from Graphcore TM Inference Processing Unit (IPU) from Intel Corp. (e.g., "Lake Crest") processor. In at least one embodiment, Figure 6B The illustrated inference and / or training logic 615 may be used in conjunction with central processing unit (“CPU”) hardware, graphics processing unit (“GPU”) hardware, or other hardware such as a field programmable gate array (“FPGA”).

[0118] In at least one embodiment, Figure 6C Inference and / or training logic 615 is shown, which may include, but is not limited to, hardware logic, wherein computing resources are dedicated or otherwise uniquely used in conjunction with weight values or other information corresponding to one or more layers of neurons within a neural network, in accordance with at least one various embodiment. In at least one embodiment, Figure 6C The inference and / or training logic 615 shown in FIG can be used in conjunction with an application specific integrated circuit (ASIC), such as the ASIC from Google. Processing unit from Graphcore TM Inference Processing Unit (IPU) or from Intel Corp (such as "Lake Crest") processor. In at least one embodiment, Figure 6CThe inference and / or training logic 615 shown in can be used in conjunction with central processing unit (CPU) hardware, graphics processing unit (GPU) hardware, or other hardware such as a field programmable gate array (FPGA). In at least one embodiment, the inference and / or training logic 615 includes, but is not limited to, code and / or data storage 601 and code and / or data storage 605, which can be used to store code (such as graph code), weight values, and / or other information, including bias values, gradient information, momentum values, and / or other parameter or hyperparameter information. Figure 6C In at least one embodiment shown in FIG, code and / or data store 601 and code and / or data store 605 are each associated with dedicated computing resources (eg, computing hardware 602 and computing hardware 606), respectively.

[0119] In at least one embodiment, each of the code and / or data stores 601 and 605 and the corresponding computing hardware 602 and 606 corresponds to a different layer of a neural network, such that activations from one "storage / compute pair 601 / 602" of the code and / or data store 601 and computing hardware 602 are provided as input to the next "storage / compute pair 605 / 606" of the code and / or data store 605 and computing hardware 606, reflecting the conceptual organization of the neural network. In at least one embodiment, each storage / compute pair 601 / 602 and 605 / 606 can correspond to more than one neural network layer. In at least one embodiment, additional storage / compute pairs (not shown) can be included in the inference and / or training logic 615 after or in parallel with the storage / compute pairs 601 / 602 and 605 / 606.

[0120] Computer system

[0121] Figure 7A is a block diagram illustrating an exemplary computer system 700A, which may be a system of interconnected devices and components, a system on a chip (SOC), or some combination thereof, in accordance with at least one embodiment, and which may include an execution unit for executing instructions to support and / or implement intelligent control of a cooling system incorporating an adaptive heat sink as described herein. In at least one embodiment, the computer system 700A may include, but is not limited to, components such as the processor 702, to execute algorithms on process data using the execution unit including logic in accordance with the present disclosure, such as the embodiments herein. In at least one embodiment, the computer system 700A may include a processor, such as the Intel® processor 702, available from Intel Corporation of Santa Clara, California. Processor family, XeonTM, XScaleTM and / or StrongARMTM, CoreTM or Nervana TM microprocessor, but other systems may be used (including PCs with other microprocessors, engineering workstations, set-top boxes, etc.). In at least one embodiment, computer system 700B may execute a version of the WINDOWS operating system available from Microsoft Corporation of Redmond, Washington, although other operating systems (e.g., UNIX and Linux), embedded software, and / or graphical user interfaces may also be used.

[0122] In at least one embodiment, exemplary computer system 700A may incorporate components 110-116 (from Figure 1 ) to support processing aspects for intelligent control of cooling systems incorporating adaptive heat sinks. For at least this reason, in one embodiment, Figure 7A The system is shown as comprising interconnected hardware devices or "chips", while in other embodiments, Figure 7A An exemplary system-on-chip SoC may be shown. In at least one embodiment, Figure 7A The devices shown in FIG. 7 may be interconnected with a proprietary interconnect, a standardized interconnect (such as PCIe), or some combination thereof. In at least one embodiment, one or more components of computer system 700B are interconnected using a Compute Express Link (CXL) interconnect. Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments, for example, as previously described with respect to FIG. Figure 6A -C discussed. Figure 6A -C provides details about the reasoning and / or training logic 615. In at least one embodiment, the reasoning and / or training logic 615 may be implemented in the system Figure 7A for performing inference or prediction operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures, or neural network use cases herein.

[0123] Embodiments may be used in other devices, such as handheld devices and embedded applications. Some examples of handheld devices include cellular phones, Internet Protocol (IP) devices, digital cameras, personal digital assistants ("PDAs"), and handheld PCs. In at least one embodiment, the embedded application may include a microcontroller, a digital signal processor ("DSP"), a system on a chip, a network computer ("NetPC"), a set-top box, a network hub, a wide area network ("WAN") switch, or any other system that can execute one or more instructions according to at least one embodiment.

[0124] In at least one embodiment, the computer system 700A may include, but is not limited to, a processor 702, which may include, but is not limited to, one or more execution units 708 to perform machine learning model training and / or inference according to the techniques described herein. In at least one embodiment, the computer system 700A is a single-processor desktop or server system, but in another embodiment, the computer system 700A may be a multi-processor system. In at least one embodiment, the processor 702 may include, but is not limited to, a complex instruction set computer ("CISC") microprocessor, a reduced instruction set computing ("RISC") microprocessor, a very long instruction word ("VLIW") microprocessor, a processor that implements an instruction set combination, or any other processor device, such as a digital signal processor. In at least one embodiment, the processor 702 may be coupled to a processor bus 710, which may transmit data signals between the processor 702 and other components in the computer system 700A.

[0125] In at least one embodiment, processor 702 may include, but is not limited to, level 1 ("L1") internal cache memory ("cache") 704. In at least one embodiment, processor 702 may have a single internal cache or multiple levels of internal cache. In at least one embodiment, the cache memory may reside external to processor 702. Other embodiments may include a combination of internal and external caches, depending on the specific implementation and requirements. In at least one embodiment, register file 706 may store different types of data in various registers, including, but not limited to, integer registers, floating point registers, status registers, and an instruction pointer register.

[0126] In at least one embodiment, an execution unit 708, including but not limited to logic to perform integer and floating point operations, is also located in the processor 702. In at least one embodiment, the processor 702 may also include a microcode ("ucode") read-only memory ("ROM") for storing microcode for certain macroinstructions. In at least one embodiment, the execution unit 708 may include logic for processing a packed instruction set 709. In at least one embodiment, by including the packed instruction set 709 in the instruction set of the general-purpose processor, and the associated circuitry to execute the instructions, operations used by many multimedia applications may be performed using packed data in the general-purpose processor 702. In one or more embodiments, many multimedia applications may be accelerated and executed more efficiently by using the full width of the processor's data bus to perform operations on the packed data, which may not require transferring smaller units of data across the processor's data bus to perform one or more operations one data element at a time.

[0127] In at least one embodiment, execution unit 708 may also be used in a microcontroller, an embedded processor, a graphics device, a DSP, and other types of logic circuits. In at least one embodiment, computer system 700A may include, but is not limited to, memory 720. In at least one embodiment, memory 720 may be implemented as a dynamic random access memory ("DRAM") device, a static random access memory ("SRAM") device, a flash memory device, or other storage device. In at least one embodiment, memory 720 may store instructions 719 and / or data 721 represented by data signals that may be executed by processor 702.

[0128] In at least one embodiment, a system logic chip can be coupled to the processor bus 710 and the memory 720. In at least one embodiment, the system logic chip can include, but is not limited to, a memory controller hub ("MCH") 716, and the processor 702 can communicate with the MCH 716 via the processor bus 710. In at least one embodiment, the MCH 716 can provide a high-bandwidth memory path 718 to the memory 720 for instruction and data storage, as well as for storage of graphics commands, data, and textures. In at least one embodiment, the MCH 716 can initiate data signals between the processor 702, the memory 720, and other components in the computer system 700A, and bridge data signals between the processor bus 710, the memory 720, and the system I / O 722. In at least one embodiment, the system logic chip can provide a graphics port for coupling to a graphics controller. In at least one embodiment, the MCH 716 may be coupled to the memory 720 via a high-bandwidth memory path 718 , and the graphics / video card 712 may be coupled to the MCH 716 via an Accelerated Graphics Port (“AGP”) interconnect 714 .

[0129] In at least one embodiment, computer system 700A may use system I / O 722 as a proprietary hub interface bus to couple MCH 716 to I / O controller hub ("ICH") 730. In at least one embodiment, ICH 730 may provide direct connection to certain I / O devices via a local I / O bus. In at least one embodiment, the local I / O bus may include, but is not limited to, a high-speed I / O bus used to connect peripheral devices to memory 720, chipset, and processor 702. Examples may include, but are not limited to, an audio controller 729, a firmware hub ("Flash BIOS") 728, a wireless transceiver 726, a data store 724, a legacy I / O controller 723 including a user input and keyboard interface, a serial expansion port 727 (e.g., a Universal Serial Bus (USB) port), and a network controller 734. The data store 724 may include a hard drive, a floppy disk drive, a CD-ROM device, a flash memory device, or other mass storage device.

[0130] Figure 7B is a block diagram illustrating an electronic device 700B for utilizing a processor 710 to support and / or implement intelligent control of a cooling system incorporating an adaptive heat sink as described herein, in accordance with at least one embodiment. In at least one embodiment, the electronic device 700B may be, for example, but not limited to, a notebook computer, a tower server, a rack server, a blade server, a laptop computer, a desktop computer, a tablet computer, a mobile device, a phone, an embedded computer, or any other suitable electronic device. In at least one embodiment, the exemplary electronic device 700B may incorporate one or more components that support processing aspects of a cooling system incorporating an adaptive heat sink.

[0131] In at least one embodiment, system 700B may include, but is not limited to, a processor 710 communicatively coupled to any suitable number or kind of components, peripherals, modules, or devices. In at least one embodiment, processor 710 is coupled using a bus or interface, such as an I2C bus, a system management bus ("SMBus"), a low pin count (LPC) bus, a serial peripheral interface ("SPI"), a high-definition audio ("HDA") bus, a serial advanced technology attachment ("SATA") bus, a universal serial bus ("USB") (versions 1, 2, 3), or a universal asynchronous receiver / transmitter ("UART") bus. In at least one embodiment, Figure 7B shows a system comprising interconnected hardware devices or "chips", while in other embodiments, Figure 7B An exemplary system on a chip ("SoC") may be shown. In at least one embodiment, Figure 7BThe devices shown in can be interconnected with proprietary interconnects, standardized interconnects (such as PCIe), or some combination thereof. In at least one embodiment, Figure 7B One or more components of the system are interconnected using Compute Express Link (CXL) interconnect lines.

[0132] In at least one embodiment, Figure 7B The system may include a display 724, a touch screen 725, a touchpad 730, a near field communication unit ("NFC") 745, a sensor hub 740, a thermal sensor 746, a fast chipset ("EC") 735, a trusted platform module ("TPM") 738, a BIOS / firmware / flash memory ("BIOS, FW Flash") 722, a DSP 760, a drive 720 (e.g., a solid-state disk ("SSD") or a hard disk drive ("HDD")), a wireless local area network unit ("WLAN") 750, a Bluetooth unit 752, a wireless wide area network unit ("WWAN") 756, a global positioning system (GPS) unit 755, a camera ("USB 3.0 camera") 754 (e.g., a USB 3.0 camera), and / or a low-power double data rate ("LPDDR") memory unit ("LPDDR3") 715 implemented using, for example, the LPDDR3 standard. Each of these components may be implemented in any suitable manner.

[0133] In at least one embodiment, other components may be communicatively coupled to the processor 710 via the following components. In at least one embodiment, an accelerometer 741, an ambient light sensor ("ALS") 742, a compass 743, and a gyroscope 744 may be communicatively coupled to the sensor hub 740. In at least one embodiment, a thermal sensor 739, a fan 737, a keyboard 746, and a touchpad 730 may be communicatively coupled to the EC 735. In at least one embodiment, a speaker 763, an earpiece 764, and a microphone ("mic") 765 may be communicatively coupled to an audio unit ("audio codec and class-D amplifier") 762, which in turn may be communicatively coupled to the DSP 760. In at least one embodiment, the audio unit 764 may include, for example, but not limited to, an audio codec / decoder ("codec") and a class-D amplifier. In at least one embodiment, a SIM card ("SIM") 757 may be communicatively coupled to the WWAN unit 756. In at least one embodiment, components such as the WLAN unit 750 and the Bluetooth unit 752 and the WWAN unit 756 may be implemented as a next generation form factor (NGFF).

[0134] Reasoning and / or training logic 615 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 6B and / or Figure 6CProvides details about the reasoning and / or training logic 615. In at least one embodiment, the reasoning and / or training logic 615 may be implemented in the system Figure 7B for use in performing inference or prediction operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures, or neural network use cases herein.

[0135] Figure 7C A computer system 700C is shown for supporting and / or implementing intelligent control of a cooling system incorporating an adaptive heat sink as described herein, according to at least one embodiment. In at least one embodiment, computer system 700C includes, but is not limited to, a computer 771 and a USB drive 770. In at least one embodiment, computer 771 may include, but is not limited to, any number and type of processors (not shown) and memory (not shown). In at least one embodiment, computer 771 includes, but is not limited to, a server, a cloud instance, a laptop computer, and a desktop computer.

[0136] In at least one embodiment, the USB disk 770 includes, but is not limited to, a processing unit 772, a USB interface 774, and USB interface logic 773. In at least one embodiment, the processing unit 772 can be any instruction execution system, device, or device capable of executing instructions. In at least one embodiment, the processing unit 772 can include, but is not limited to, any number and type of processing cores (not shown). In at least one embodiment, the processing unit or core 772 includes an application specific integrated circuit ("ASIC") that is optimized to perform any number and type of operations associated with machine learning. For example, in at least one embodiment, the processing core 772 is a tensor processing unit ("TPC") that is optimized to perform machine learning reasoning operations. In at least one embodiment, the processing core 772 is a vision processing unit ("VPU") that is optimized to perform machine vision and machine learning reasoning operations.

[0137] In at least one embodiment, USB interface 774 can be any type of USB connector or USB receptacle. For example, in at least one embodiment, USB interface 774 is a USB 3.0 Type-C receptacle for data and power. In at least one embodiment, USB interface 774 is a USB 3.0 Type-A connector. In at least one embodiment, USB interface logic 773 can include any number and type of logic that enables processing unit 772 to connect to a device (such as computer 771) via USB connector 774.

[0138] Reasoning and / or training logic 615 (e.g., regarding Figure 6B and Figure 6C) is used to perform reasoning and / or training operations related to one or more embodiments. Figure 6B and Figure 6C Provides details about the reasoning and / or training logic 615. In at least one embodiment, the reasoning and / or training logic 615 may be used to Figure 7C In a system, an inference or prediction operation is performed based at least in part on weight parameters calculated using the neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.

[0139] Figure 8 A further exemplary computer system 800 is shown for implementing the various processes and methods of a cooling system incorporating an adaptive heat sink described throughout this disclosure, in accordance with at least one embodiment. In at least one embodiment, computer system 800 includes, but is not limited to, at least one central processing unit ("CPU") 802 connected to a communication bus 810 implemented using any suitable protocol, such as PCI ("Peripheral Component Interconnect"), Peripheral Component Interconnect Express ("PCI-Express"), AGP ("Accelerated Graphics Port"), HyperTransport, or any other bus or point-to-point communication protocol. In at least one embodiment, computer system 800 includes, but is not limited to, main memory 804 and control logic (e.g., implemented as hardware, software, or a combination thereof), and data may be stored in main memory 804 in the form of random access memory ("RAM"). In at least one embodiment, a network interface subsystem ("network interface") 822 provides an interface to other computing devices and networks for receiving data from computer system 800 and transmitting data to other systems.

[0140] In at least one embodiment, computer system 800 includes, but is not limited to, input device 808, parallel processing system 812, and display device 806, which can be implemented using cathode ray tubes ("CRTs"), liquid crystal displays ("LCDs"), light emitting diodes ("LEDs"), plasma displays, or other suitable display technologies. In at least one embodiment, user input is received from input device 808 (such as a keyboard, mouse, touchpad, microphone, and more). In at least one embodiment, each of the aforementioned modules can be located on a single semiconductor platform to form a processing system.

[0141] Reasoning and / or training logic 615 is used to perform reasoning and / or training operations associated with one or more embodiments, such as those previously described with respect to Figure 6A -C discussed. The following combination Figure 6A -C provides details about the reasoning and / or training logic 615. In at least one embodiment, the reasoning and / or training logic 615 may be implemented in the system Figure 8 In at least one embodiment, the inference and / or training logic 615 may be used in the system to perform inference or prediction operations based at least in part on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases herein. Figure 8 to perform inference or prediction operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures, or neural network use cases herein.

[0142] Figure 9A An exemplary architecture is shown in which multiple GPUs 910-913 are communicatively coupled to multiple multi-core processors 940-943 via high-speed links 905-906 (such as buses, point-to-point interconnects, etc.). In one embodiment, the high-speed links 940-943 support 4 GB / s, 30 GB / s, 80 GB / s, or higher communication throughput. Various interconnect protocols can be used, including but not limited to PCIe 4.0 or 5.0 and NVLink 2.0.

[0143] Furthermore, in one embodiment, two or more GPUs 910-913 are interconnected via high-speed links 929-930, which may be implemented using the same or different protocols / links as used for high-speed links 940-943. Similarly, two or more multi-core processors 905-906 may be connected via high-speed link 928, which may be a symmetric multiprocessor (SMP) bus running at 20 GB / s, 30 GB / s, 120 GB / s, or higher. Alternatively, the same protocol / links may be used (e.g., via a common interconnect fabric). Figure 9A All communications between the various system components shown in .

[0144] In one embodiment, each multi-core processor 905-906 is communicatively coupled to processor memory 901-902 via memory interconnects 926-927, respectively, and each GPU 910-913 is communicatively coupled to GPU memory 920-923 via GPU memory interconnects 950-953, respectively. Memory interconnects 926-927 and 950-953 can utilize the same or different memory access technologies. By way of example and not limitation, processor memory 901-902 and GPU memory 920-923 can be volatile memory, such as dynamic random access memory (DRAM) (including stacked DRAM), graphics DDR SDRAM (GDDR) (such as GDDR5, GDDR6), or high bandwidth memory (HBM), and / or can be non-volatile memory, such as 3D XPoint or Nano-Ram. In one embodiment, some portion of processor memory 901-902 can be volatile memory, while another portion can be non-volatile memory (such as using a two-level memory (2LM) hierarchy).

[0145] As described below, although the various processors 905-906 and GPUs 910-913 may each be physically coupled to a specific memory 901-902, 920-923, a unified memory architecture may be implemented in which a virtual system address space (also referred to as an "effective address" space) is distributed among the various physical memories. In at least one embodiment, the processor memories 901-902 may each include 64GB of system memory address space, and the GPU memories 920-923 may each include 32GB of system memory address space (resulting in a total addressable memory size of 256GB in this example).

[0146] As discussed elsewhere in this disclosure, at least flow rates and associated temperatures can be established for the first level of an intelligent learning system, such as a neural network system. Because the first level represents prior data, it also represents a smaller subset of data that can be used to improve the system by retraining the system. Testing and training can be performed in parallel using multiple processor units so that the intelligent learning system is robust. Figure 9A When the intelligent learning system achieves convergence, the number of data points used to cause convergence and the data in the data points are recorded. The data and data points can be used to fully control the cooling system incorporating the adaptive heat sink, exactly as in reference e.g. Figure 2A-5 discussed.

[0147] Figure 9B907 and a graphics acceleration module 946. The graphics acceleration module 946 may include one or more GPU chips integrated on a line card that is coupled to the processor 907 via a high-speed link 940. Alternatively, the graphics acceleration module 946 may be integrated with the processor 907 in the same package or chip.

[0148] In at least one embodiment, the illustrated processor 907 includes a plurality of cores 960A-960D, each core having a translation lookaside buffer 961A-961D and one or more caches 962A-962D. In at least one embodiment, the cores 960A-960D may include various other components not shown for executing instructions and processing data. The caches 962A-962D may include level 1 (L1) and level 2 (L2) caches. In addition, one or more shared caches 956 may be included in the caches 962A-962D and shared by each group of cores 960A-960D. In at least one embodiment, one embodiment of the processor 907 includes 24 cores, each core having its own L1 cache, twelve shared L2 caches, and twelve shared L3 caches. In this embodiment, two adjacent cores share one or more L2 and L3 caches. The processor 907 and the graphics acceleration module 946 are connected to a system memory 914, which may include Figure 9A Processor memory 901-902 in.

[0149] Coherence is maintained for data and instructions stored in the various caches 962A-962D, 956 and system memory 914 via inter-core communication over the coherence bus 964. In at least one embodiment, each cache may have cache coherence logic / circuitry associated therewith to communicate over the coherence bus 964 in response to detecting a read or write to a particular cache line. In one implementation, a cache snooping protocol is implemented over the coherence bus 964 to snoop cache accesses.

[0150] In at least one embodiment, the proxy circuitry 925 communicatively couples the graphics acceleration module 946 to the coherence bus 964, thereby allowing the graphics acceleration module 946 to participate in a cache coherence protocol as a peer of the cores 960A-960D. In particular, in at least one embodiment, the interface 935 provides a connection to the proxy circuitry 925 via a high-speed link 940 (such as a PCIe bus, NVLink, etc.), and the interface 937 connects the graphics acceleration module 946 to the link 940.

[0151] In one implementation, the accelerator integrated circuit 936 provides cache management, memory access, context management, and interrupt management services on behalf of the multiple graphics processing engines 931, 932, N of the graphics acceleration module. The graphics processing engines 931, 932, N may each include a separate graphics processing unit (GPU). In at least one embodiment, the graphics processing engines 931, 932, N may selectively include different types of graphics processing engines within the GPU, such as a graphics execution unit, a media processing engine (such as a video encoder / decoder), a sampler, and a blit engine. In at least one embodiment, the graphics acceleration module 946 may be a GPU having multiple graphics processing engines 931-932, N, or the graphics processing engines 931-932, N may be individual GPUs integrated on a common package, line card, or chip. As the case may be, Figure 9B The above determination of the reconstruction parameters and the reconstruction algorithm is performed in the GPU 931-N.

[0152] In one embodiment, the accelerator integrated circuit 936 includes a memory management unit (MMU) 939 for performing various memory management functions, such as virtual to physical memory translation (also known as effective to real memory translation), and a memory access protocol for accessing system memory 914. The MMU 939 may also include a translation lookaside buffer ("TLB") (not shown) for caching virtual / effective to physical / real address translations. In one implementation, a cache 938 may store commands and data for efficient access by the graphics processing engines 931-932, N. In at least one embodiment, data stored in the cache 938 and graphics memory 933-934, M may be kept consistent with the core caches 962A-962D, 956 and system memory 914. As before, this task may be accomplished via proxy circuitry 925 acting on behalf of cache 938 and graphics memory 933-934, M (such as sending updates related to modifications / accesses of cache lines on processor caches 962A-962D, 956 to cache 938 and receiving updates from cache 938).

[0153] A set of registers 945 stores context data for threads executed by graphics processing engines 931-932, N, and context management circuitry 948 manages thread contexts. In at least one embodiment, context management circuitry 948 can perform save and restore operations to save and restore the context of each thread during a context switch (such as where a first thread is saved and a second thread is stored so that the second thread can be executed by the graphics processing engine). In at least one embodiment, context management circuitry 948 can store current register values to a designated area in memory (such as identified by a context pointer) upon a context switch. The register values can then be restored upon returning to context. In one embodiment, interrupt management circuitry 947 receives and processes interrupts received from system devices.

[0154] In one implementation, the MMU 939 converts virtual / effective addresses from the graphics processing engine 931 into real / physical addresses in the system memory 914. One embodiment of the accelerator integrated circuit 936 supports multiple (e.g., 4, 8, 16) graphics accelerator modules 946 and / or other accelerator devices. The graphics accelerator module 946 can be dedicated to a single application executing on the processor 907, or can be shared between multiple applications. In one embodiment, a virtualized graphics execution environment is presented in which the resources of the graphics processing engines 931-932, N are shared with multiple applications or virtual machines (VMs). In at least one embodiment, resources can be subdivided into "slices" that are allocated to different VMs and / or applications based on processing requirements and priorities associated with the VMs and / or applications.

[0155] In at least one embodiment, the accelerator integrated circuit 936 acts as a bridge to the system for the graphics acceleration module 946 and provides address translation and system memory cache services. In addition, the accelerator integrated circuit 936 can provide virtualization facilities for the host processor to manage virtualization, interrupts, and memory management of the graphics processing engines 931-932, N.

[0156] Because the hardware resources of graphics processing engines 931-932, N are explicitly mapped into the real address space seen by host processor 907, any host processor can directly address these resources using effective address values. In at least one embodiment, one function of accelerator integrated circuit 936 is to physically separate graphics processing engines 931-932, N so that they appear as independent units to the system.

[0157] In at least one embodiment, one or more graphics memories 933-934, M are respectively coupled to each graphics processing engine 931-932, N. The graphics memories 933-934, M store instructions and data, which are processed by each graphics processing engine 931-932, N. The graphics memories 933-934, M can be volatile memory, such as DRAM (including stacked DRAM), GDDR memory (such as GDDR5, GDDR6), or HBM, and / or can be non-volatile memory, such as 3D XPoint or Nano-Ram.

[0158] In one embodiment, to reduce data traffic on link 940, a biasing technique is used to ensure that the data stored in graphics memories 933-934, M is the data most frequently used by graphics processing engines 931-932, N and is data that cores 960A-960D may not use (at least not frequently). Similarly, the biasing mechanism attempts to keep data needed by a core (and possibly not graphics processing engines 931-932, N) in caches 962A-962D, core 956, and system memory 914.

[0159] Figure 9C Another exemplary embodiment according to at least one embodiment disclosed herein is shown, in which an accelerator integrated circuit 936 is integrated within the processor 907 to implement and / or support intelligent control of a cooling system incorporating an adaptive heat sink. In at least this embodiment, the graphics processing engines 931-932, N communicate directly with the accelerator integrated circuit 936 via interfaces 937 and 935 (again, any form of bus or interface protocol may be utilized) over a high-speed link 940. The accelerator integrated circuit 936 may perform operations related to Figure 9B The operations described above are similar to those described above. However, due to its close proximity to the coherence bus 964 and caches 962A-962D, 956, higher throughput is possible. At least one embodiment supports different programming models, including a dedicated process programming model (without graphics acceleration module virtualization) and a shared programming model (with virtualization). The programming models can include a programming model controlled by the accelerator integrated circuit 936 and a programming model controlled by the graphics acceleration module 946.

[0160] In at least one embodiment, the graphics processing engines 931-932, N are dedicated to a single application or process under a single operating system. In at least one embodiment, a single application can funnel other application requests to the graphics processing engines 931-932, N, thereby providing virtualization within a VM / partition.

[0161] In at least one embodiment, graphics processing engines 931-932,N can be shared by multiple VM / application partitions. In at least one embodiment, the sharing model can use a hypervisor to virtualize graphics processing engines 931-932,N to allow each operating system to access them. For a single-partition system without a hypervisor, the operating system owns graphics processing engines 931-932,N. In at least one embodiment, the operating system can virtualize graphics processing engines 931-932,N to provide access to each process or application.

[0162] In at least one embodiment, the graphics acceleration module 946 or individual graphics processing engines 931-932, N use a process handle to select a process element. In at least one embodiment, the process element is stored in the system memory 914 and can be addressed using the effective address to real address translation techniques described herein. In at least one embodiment, the process handle can be an implementation-specific value that is provided to the host process when registering its context with the graphics processing engine 931-932, N (i.e., calling system software to add the process element to the process element linked list). In at least one embodiment, the lower 16 bits of the process handle can be the offset of the process element in the process element linked list.

[0163] Figure 9D An exemplary accelerator integrated slice 990 is shown for implementing and / or supporting intelligent control of a cooling system incorporating an adaptive heat sink in accordance with at least one embodiment disclosed herein. As used herein, a "slice" comprises a designated portion of the processing resources of an accelerator integrated circuit 936. An application is an effective address space 982 in system memory 914 that stores process elements 983. In at least one embodiment, process elements 983 are stored in response to a GPU call 981 from an application 980 executing on a processor 907. Process elements 983 contain the process state of the corresponding application 980. A work descriptor (WD) 984 contained in process element 983 may be a single job requested by an application, or may contain a pointer to a job queue. In at least one embodiment, WD 984 is a pointer to a job request queue in the application's address space 982.

[0164] The graphics acceleration module 946 and / or the individual graphics processing engines 931-932, N can be shared by all processes or a subset of processes in the system. In at least one embodiment, an infrastructure for setting process state and sending WD 984 to the graphics acceleration module 946 to start a job in a virtualized environment can be included.

[0165] In at least one embodiment, a dedicated process programming model is implementation-specific. In this model, a single process owns either a graphics acceleration module 946 or an individual graphics processing engine 931. When a graphics acceleration module 946 is owned by a single process, the hypervisor initializes the accelerator integrated circuit for the owned partition, and when a graphics acceleration module 946 is assigned, the operating system initializes the accelerator integrated circuit 936 for the owned process.

[0166] In operation, the WD fetch unit 991 in the accelerator integrated slice 990 fetches the next WD 984, which includes an indication of work to be completed by one or more graphics processing engines of the graphics acceleration module 946. Data from the WD 984 can be stored in registers 945 and used by the MMU 939, interrupt management circuitry 947, and / or context management circuitry 948, as shown. In at least one embodiment, one embodiment of the MMU 939 includes segment / page roaming circuitry for accessing segment / page tables 986 within the OS virtual address space 985. The interrupt management circuitry 947 can process interrupt events 992 received from the graphics acceleration module 946. In at least one embodiment, when performing graphics operations, effective addresses 993 generated by the graphics processing engines 931-932, N are converted into real addresses by the MMU 939.

[0167] In one embodiment, the same register set 945 is replicated for each graphics processing engine 931-932, N, and / or graphics acceleration module 946, and the same register set 945 can be initialized by the hypervisor or operating system. Each of these replicated registers can be included in the accelerator integration slice 990. Example registers that can be initialized by the hypervisor are shown in Table 1.

[0168] Table 1 – Hypervisor Initialization Registers

[0169]

[0170]

[0171] Example registers that may be initialized by the operating system are shown in Table 2.

[0172] Table 2 – Operating System Initialization Registers

[0173] 1 Process and thread identification 2 Effective Address (EA) context save / restore pointer 3 Virtual Address (VA) Accelerator Utilizes Record Pointers 4 Virtual Address (VA) Segment Table Pointer 5 Permission blocking 6 Job Descriptor

[0174] In at least one embodiment, each WD 984 is specific to a particular graphics acceleration module 946 and / or graphics processing engine 931-932, N. It contains all the information necessary for the graphics processing engine 931-932, N to complete the work, or it may be a pointer to a memory location where the application has set up a command queue for the work to be done.

[0175] Figure 9E 1 shows additional details of an exemplary embodiment of a sharing model. This embodiment includes a hypervisor real address space 998 in which a process element list 999 is stored. The hypervisor real address space 998 can be accessed via the hypervisor 996, which virtualizes the graphics acceleration module engine for the operating system 995.

[0176] In at least one embodiment, the shared programming model allows all processes or a subset of processes from all partitions or a subset of partitions in the system to use the graphics acceleration module 946. There are two programming models in which the graphics acceleration module 946 is shared by multiple processes and partitions, namely, time-sliced sharing and graphics-directed sharing.

[0177] In this model, the hypervisor 996 owns the graphics acceleration module 946 and makes its functionality available to all operating systems 995. For the graphics acceleration module 946 to support virtualization through the hypervisor 996, the graphics acceleration module 946 may adhere to the following requirements: (1) the application's job requests must be autonomous (i.e., no state needs to be maintained between jobs), or the graphics acceleration module 946 must provide a context save and restore mechanism, (2) the graphics acceleration module 946 guarantees that the application's job requests are completed within a specified amount of time, including any transition errors, or the graphics acceleration module 946 provides the ability to preempt job processing, and (3) fairness between graphics acceleration module 946 processes must be ensured when operating in a directed shared programming model.

[0178] In one embodiment, an application 980 is required to make an operating system 995 system call using a graphics acceleration module type, a work descriptor (WD), an authority mask register (AMR) value, and a context save / restore region pointer (CSRP). The graphics acceleration module type describes the target acceleration function for the system call. In at least one embodiment, the graphics acceleration module type can be a system-specific value. In at least one embodiment, the WD is formatted specifically for the graphics acceleration module 946 and can take the form of a graphics acceleration module 946 command, an effective address pointer to a user-defined structure, an effective address pointer to a command queue, or any other data structure describing work to be performed by the graphics acceleration module 946. In at least one embodiment, the AMR value is the AMR state for the current process. In at least one embodiment, the value passed to the operating system is similar to that of an application setting the AMR. If the implementation of the accelerator integrated circuit 936 and graphics acceleration module 946 does not support the User Authority Mask Override Register (UAMOR), the operating system can apply the current UAMOR value to the AMR value before passing the AMR in the hypervisor call. In at least one embodiment, the hypervisor 996 can apply the current privilege mask overwrite register (AMOR) value before placing the AMR into the process element 983. In at least one embodiment, the CSRP is one of the registers 945 that contains the effective address of an area in the application's effective address space 982 for the graphics acceleration module 946 to save and restore context state. This pointer is used in at least one embodiment but is optional if state does not need to be saved between jobs or when a job is preempted. In at least one embodiment, the context save / restore area can be fixed system memory.

[0179] Upon receiving the system call, the operating system 995 may verify that the application 980 has been registered and granted permission to use the graphics acceleration module 946. The operating system 995 then calls the hypervisor 996 using the information shown in Table 3.

[0180] Table 3 – OS to Hypervisor call parameters

[0181] 1 Work Descriptor (WD) 2 Access Mask Register (AMR) value (potentially masked) 3 Effective Address (EA) Context Save / Restore Region Pointer (CSRP) 4 Process ID (PID) and optional thread ID (TID) 5 Virtual Address (VA) Accelerator Usage Record Pointer (AURP) 6 Virtual address of the storage segment table pointer (SSTP) 7 Logical Interrupt Service Number (LISN)

[0182] Upon receiving the hypervisor call, the hypervisor 996 verifies that the operating system 995 has registered and been granted permission to use the graphics acceleration module 946. The hypervisor 996 then places the process element 983 into a linked list of process elements of the corresponding graphics acceleration module 946 type. The process element may include the information shown in Table 4.

[0183] Table 4 – Process element information

[0184]

[0185]

[0186] In at least one embodiment, the hypervisor initializes the plurality of accelerator integrated slice 990 registers 945 .

[0187] like Figure 9F As shown, in at least one embodiment, a unified memory is used that can be addressed via a common virtual memory address space for accessing physical processor memories 901-902 and GPU memories 920-923. In this implementation, operations executed on GPUs 910-913 utilize the same virtual / effective memory address space to access processor memories 901-902, and vice versa, thereby simplifying programmability. In one embodiment, a first portion of the virtual / effective address space is allocated to processor memory 901, a second portion is allocated to second processor memory 902, a third portion is allocated to GPU memory 920, and so on. In at least one embodiment, the entire virtual / effective memory space (sometimes referred to as the effective address space) is thus distributed across each of processor memories 901-902 and GPU memories 920-923, allowing any processor or GPU to access the memory using a virtual address mapped to any physical memory.

[0188] In one embodiment, bias / coherency management circuitry 994A-994E within one or more MMUs 939A-939E ensures cache coherency between the caches of one or more host processors (such as 905) and GPUs 910-913 and implements biasing techniques that indicate the physical memory where certain types of data should be stored. Figure 9F Multiple instances of bias / coherence management circuits 994A- 994E are shown in , but bias / coherence circuits may be implemented within an MMU of one or more host processors 905 and / or within an accelerator integrated circuit 936 .

[0189] One embodiment allows GPU-attached memory 920-923 to be mapped as part of system memory and accessed using shared virtual memory (SVM) technology, without suffering the performance drawbacks associated with full system cache coherence. In at least one embodiment, the ability to access GPU-attached memory 920-923 as system memory without heavy cache coherence overhead provides a favorable operating environment for GPU offloading. This arrangement allows host processor 905 software to set operands and access computation results without the overhead of traditional I / O DMA data copies. Such traditional copies include driver calls, interrupts, and memory-mapped I / O (MMIO) accesses, all of which are less efficient than simple memory accesses. In at least one embodiment, the ability to access GPU-attached memory 920-923 without cache coherence overhead can be critical to the execution time of offloaded computations. For example, in situations with heavy streaming write-to-memory traffic, cache coherence overhead can significantly reduce the effective write bandwidth seen by GPUs 910-913. In at least one embodiment, the efficiency of operand setup, result access, and GPU computation can play a role in determining the effectiveness of GPU offloading.

[0190] In at least one embodiment, the selection of GPU bias and host processor bias is driven by a bias tracker data structure. For example, a bias table can be used, which can be a page-granular structure (e.g., controlled at the granularity of a memory page) that includes 1 or 2 bits per GPU-attached memory page. In at least one embodiment, the bias table can be implemented in the stolen memory range of one or more GPU-attached memories 920-923, with or without a bias cache in the GPUs 910-913 (such as for caching frequently / recently used entries in the bias table). Alternatively, the entire bias table can be maintained within the GPU.

[0191] In at least one embodiment, before actually accessing the GPU memory, the bias table entry associated with each access to the GPU attached memory 920-923 is accessed, resulting in the following operations. Local requests from GPUs 910-913 whose pages are found in the GPU bias are forwarded directly to the corresponding GPU memory 920-923. Local requests from the GPU whose pages are found in the host bias are forwarded to the processor 905 (such as, via a high-speed link as above). In one embodiment, a request from the processor 905 to find the requested page in the host processor bias completes a request similar to a normal memory read. Alternatively, a request to a GPU biased page can be forwarded to the GPUs 910-913. In at least one embodiment, if the GPU is not currently using the page, the GPU can subsequently migrate the page to the host processor bias. In at least one embodiment, the bias state of a page can be changed by a software-based mechanism, a hardware-assisted software-based mechanism, or in limited cases by a purely hardware-based mechanism.

[0192] One mechanism for changing the bias state employs an API call (such as OpenCL), which in turn calls the GPU's device driver, which in turn sends a message (or queues a command descriptor) to the GPU, directing the GPU to change the bias state and, in some migrations, performs a cache flush operation in the host. In at least one embodiment, the cache flush operation is used for migrations from host processor 905 bias to GPU bias, but not for the reverse migration.

[0193] In one embodiment, cache coherence is maintained by temporarily rendering GPU-biased pages that cannot be cached by the host processor 905. To access these pages, the processor 905 may request access from the GPU 910, which may or may not immediately grant access. Therefore, to reduce communication between the processor 905 and the GPU 910, it is beneficial to ensure that the GPU-biased pages are the pages required by the GPU and not the host processor 905, and vice versa.

[0194] Reasoning and / or training logic 615 is used to execute one or more embodiments. Figure 6B and / or Figure 6C Details regarding the inference and / or training logic 615 are provided.

[0195] Figure 10AAn exemplary integrated circuit and associated graphics processor according to various embodiments described herein are shown, which can be manufactured using one or more IP cores to support and / or implement a cooling system incorporating an adaptive heat sink as described herein. In addition to the illustrations, other logic and circuitry can be included in at least one embodiment, including additional graphics processors / cores, peripheral interface controllers, or general-purpose processor cores.

[0196] Figure 10A is a block diagram illustrating an exemplary system on a chip integrated circuit 1000A that can be fabricated using one or more IP cores according to at least one embodiment. In at least one embodiment, integrated circuit 1000A includes one or more application processors 1005 (such as a CPU), at least one graphics processor 1010, and may additionally include an image processor 1015 and / or a video processor 1020, any of which may be modular IP cores. In at least one embodiment, integrated circuit 1000A includes peripheral or bus logic, including a USB controller 1025, a UART controller 1030, an SPI / SDIO controller 1035, and an I2S / I2C controller 1040. In at least one embodiment, integrated circuit 1000A may include a display device 1045 coupled to one or more of a High-Definition Multimedia Interface (HDMI) controller 1050 and a Mobile Industry Processor Interface (MIPI) display interface 1055. In at least one embodiment, storage may be provided by a flash memory subsystem 1060, including flash memory and a flash memory controller. In at least one embodiment, a memory interface may be provided via a memory controller 1065 for accessing SDRAM or SRAM memory devices. In at least one embodiment, some integrated circuits also include an embedded security engine 1070 .

[0197] The reasoning and / or training logic 615 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 6B and / or Figure 6C Details are provided regarding inference and / or training logic 615. In at least one embodiment, inference and / or training logic 615 may be used in integrated circuit 1000A to infer or predict operations based at least in part on weight parameters computed using neural network training operations, neural network functionality and / or architecture, or neural network use cases described herein.

[0198] Figures 10B-10CAn exemplary integrated circuit and associated graphics processor according to various embodiments described herein are shown, which can be manufactured using one or more IP cores to support and / or implement a cooling system incorporating an adaptive heat sink as described herein. In addition to the illustrations, other logic and circuitry can be included in at least one embodiment, including additional graphics processors / cores, peripheral interface controllers, or general-purpose processor cores.

[0199] Figures 10B-10C is a block diagram illustrating an exemplary graphics processor for use within a SoC according to embodiments described herein to support and / or implement a cooling system incorporating an adaptive heat sink as described herein. In one example, the graphics processor can be used in intelligent control of a cooling system incorporating an adaptive heat sink because existing math engines are able to process multi-level neural networks faster. Figure 10B An exemplary graphics processor 1010 of a system on a chip integrated circuit is shown, which may be manufactured using one or more IP cores, in accordance with at least one embodiment. Figure 10C Another exemplary graphics processor 1040 of a system on a chip integrated circuit is shown, which may be manufactured using one or more IP cores, in accordance with at least one embodiment. Figure 10B The graphics processor 1010 is a low power graphics processor core. In at least one embodiment, Figure 10C The graphics processor 1040 is a higher performance graphics processor core. In at least one embodiment, each graphics processor 1010, 1040 can be Figure 10A A variant of the graphics processor 1010.

[0200] In at least one embodiment, the graphics processor 1010 includes a vertex processor 1005 and one or more fragment processors 1015A-1015N (such as 1015A, 1015B, 1015C, 1015D through 1015N-1 and 1015N). In at least one embodiment, the graphics processor 1010 can execute different shader programs via separate logic, such that the vertex processor 1005 is optimized to perform operations for the vertex shader program, while the one or more fragment processors 1015A-1015N perform fragment (such as pixel) shading operations for the fragment or pixel or shader program. In at least one embodiment, the vertex processor 1005 performs the vertex processing stage of the 3D graphics pipeline and generates primitives and vertex data. In at least one embodiment, the one or more fragment processors 1015A-1015N use the primitives and vertex data generated by the vertex processor 1005 to generate a frame buffer for display on a display device. In at least one embodiment, one or more fragment processors 1015A-1015N are optimized to execute fragment shader programs as provided in the OpenGL API, which can be used to perform similar operations as pixel shader programs provided in the Direct 3D API.

[0201] In at least one embodiment, graphics processor 1010 additionally includes one or more memory management units (MMUs) 1020A-1020B, one or more caches 1025A-1025B, and one or more circuit interconnects 1030A-1030B. In at least one embodiment, one or more MMUs 1020A-1020B provide a mapping of virtual to physical addresses for graphics processor 1010, including for vertex processor 1005 and / or fragment processors 1015A-1015N, which may reference vertex or image / texture data stored in memory, in addition to vertex or image / texture data stored in one or more caches 1025A-1025B. In at least one embodiment, one or more MMUs 1020A-1020B may synchronize with other MMUs within the system, including with other MMUs. Figure 10A One or more MMUs associated with one or more application processors 1005, graphics processor 1015, and / or video processor 1020 enable each processor 1005-1020 to participate in a shared or unified virtual memory system. In at least one embodiment, one or more circuit interconnects 1030A-1030B enable graphics processor 1010 to connect to other IP cores within the SoC via an internal bus of the SoC or via a direct connection.

[0202] In at least one embodiment, graphics processor 1040 includes Figure 10AOne or more MMUs 1020A-1020B, caches 1025A-1025B, and circuit interconnects 1030A-1030B of graphics processor 1010. In at least one embodiment, graphics processor 1040 includes one or more shader cores 1055A-1055N (such as 1055A, 1055B, 1055C, 1055D, 1055E, 1055F through 1055N-1 and 1055N), such as Figure 10B As shown, it provides a unified shader core architecture in which a single core or type or core can execute all types of programmable shader code, including shader program code for implementing vertex shaders, fragment shaders, and / or compute shaders. In at least one embodiment, the number of shader cores can vary. In at least one embodiment, the graphics processor 1040 includes an inter-core task manager 1045 that acts as a thread dispatcher to dispatch execution threads to one or more shader cores 1055A-1055N and a tiling unit 1058 to accelerate tile-based rendering operations, in which rendering operations of a scene are subdivided in image space, for example, to exploit local spatial coherence within a scene or to optimize the use of internal caches.

[0203] Reasoning and / or training logic 615 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 6B and / or Figure 6C Provides details about the inference and / or training logic 615. In at least one embodiment, the inference and / or training logic 615 may be implemented in an integrated circuit. Figure 10A and / or Figure 10B for performing inference or prediction operations based at least in part on weight parameters computed using a neural network training operation, a neural network function or architecture, or a neural network use case herein.

[0204] Figures 10D-10E Additional exemplary graphics processor logic is shown according to embodiments described herein for supporting and / or implementing a cooling system incorporating an adaptive heat sink as described herein. In at least one embodiment, Figure 10D Shows that can be included in Figure 10A Graphics core 1000D within graphics processor 1010, and in at least one embodiment, may be such as Figure 10C Unified shader cores 1055A-1055N are shown. Figure 10B A highly parallel general purpose graphics processing unit ("GPGPU") 1030 suitable for deployment on a multi-chip module in at least one embodiment is shown.

[0205] In at least one embodiment, graphics core 1000D may include multiple slices 1001A-1001N, or partitions of each core, and a graphics processor may include multiple instances of graphics core 1000D. In at least one embodiment, slices 1001A-1001N may include support logic including local instruction caches 1004A-1004N, thread schedulers 1006A-1006N, thread dispatchers 1008A-1008N, and a set of registers 1010A-1010N. In at least one embodiment, slices 1001A-1001N may include a set of additional function units (AFUs 1012A-1012N), floating point units (FPUs 1014A-1014N), integer arithmetic logic units (ALUs 109A-109N), address calculation units (ACUs 1013A-1013N), double precision floating point units (DPFPUs 1015A-1015N), and matrix processing units (MPUs 1017A-1017N).

[0206] In at least one embodiment, the FPUs 1014A-1014N can perform single-precision (32-bit) and half-precision (16-bit) floating-point operations, while the DPFPUs 1015A-1015N perform double-precision (64-bit) floating-point operations. In at least one embodiment, the ALUs 1016A-1016N can perform variable-precision integer operations with 8-bit, 16-bit, and 32-bit precision, and can be configured for mixed-precision operations. In at least one embodiment, the MPUs 1017A-1017N can also be configured for mixed-precision matrix operations, including half-precision floating-point operations and 8-bit integer operations. In at least one embodiment, the MPUs 1017A-1010N can perform various matrix operations to accelerate machine learning application frameworks, including enabling support for accelerated general matrix-to-matrix multiplication (GEMM). In at least one embodiment, the AFUs 1012A-1012N can perform additional logical operations not supported by the floating-point or integer units, including trigonometric operations (such as sine, cosine, etc.).

[0207] As discussed elsewhere in this disclosure, the inference and / or training logic 615 (at least in Figure 6B 、 Figure 6C Reference) is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 6B and / or Figure 6C Details are provided regarding inference and / or training logic 615. In at least one embodiment, inference and / or training logic 615 may be used in graphics core 1000D to perform inference or prediction operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures, or neural network use cases herein.

[0208] Figure 11A A block diagram of a computer system 1100A according to at least one embodiment is shown. In at least one embodiment, computer system 1100A includes a processing subsystem 1101 having one or more processors 1102 and a system memory 1104, which communicates via an interconnect path that may include a memory hub 1105. In at least one embodiment, memory hub 1105 may be a separate component within a chipset assembly or may be integrated within one or more processors 1102. In at least one embodiment, memory hub 1105 is coupled to an I / O subsystem 1111 via a communication link 1106. In one embodiment, I / O subsystem 1111 includes an I / O hub 1107, which enables computer system 1100A to receive input from one or more input devices 1108. In at least one embodiment, I / O hub 1107 enables a display controller, which may be included in one or more processors 1102, to provide output to one or more display devices 1110A. In at least one embodiment, the one or more display devices 1110A coupled to the I / O hub 1107 may include local, internal, or embedded display devices.

[0209] In at least one embodiment, the processing subsystem 1101 includes one or more parallel processors 1112 coupled to the memory hub 1105 via a bus or other communication link 1113. In at least one embodiment, the communication link 1113 can use any of a number of standard-based communication link technologies or protocols, such as, but not limited to, PCI Express, or can be a vendor-specific communication interface or communication structure. In at least one embodiment, the one or more parallel processors 1112 form a parallel or vector processing system in a computational cluster, which can include a large number of processing cores and / or processing clusters, such as a multi-integrated core (MIC) processor. In at least one embodiment, the one or more parallel processors 1112 form a graphics processing subsystem that can output pixels to one of one or more display devices 1110A coupled via the I / O hub 1107. In at least one embodiment, the one or more parallel processors 1112 can also include a display controller and display interface (not shown) to enable direct connection to the one or more display devices 1110B.

[0210] In at least one embodiment, a system storage unit 1114 can be connected to the I / O hub 1107 to provide a storage mechanism for the computer system 1100A. In at least one embodiment, an I / O switch 1116 can be used to provide an interface mechanism to enable connections between the I / O hub 1107 and other components, such as a network adapter 1118 and / or a wireless network adapter 1119 that can be integrated into one or more platforms, as well as various other devices that can be added via one or more add-on devices 1120. In at least one embodiment, the network adapter 1118 can be an Ethernet adapter or another wired network adapter. In at least one embodiment, the wireless network adapter 1119 can include one or more of Wi-Fi, Bluetooth, near field communication (NFC), or other network devices including one or more radio devices.

[0211] In at least one embodiment, the computer system 1100A may include other components not explicitly shown, including USB or other port connections, optical storage drives, video capture devices, etc. Other components may also be connected to the I / O hub 1107. In at least one embodiment, the interconnection may be implemented using any suitable protocol, such as a PCI (Peripheral Component Interconnect)-based protocol (such as PCI-Express) or other bus or point-to-point communication interface and / or protocol. Figure 11A Communication paths for various components in a SoC, such as NV-Link high-speed interconnect or interconnect protocol.

[0212] In at least one embodiment, one or more parallel processors 1112 include circuits optimized for graphics and video processing, including, for example, video output circuitry, and constitute a graphics processing unit (GPU). In at least one embodiment, one or more parallel processors 1112 include circuits optimized for general-purpose processing. In at least one embodiment, the components of computer system 1100A can be integrated with one or more other system elements on a single integrated circuit. For example, in at least one embodiment, one or more parallel processors 1112, memory hub 1105, one or more processors 1102, and I / O hub 1107 can be integrated into a system-on-chip (SoC) integrated circuit. In at least one embodiment, the components of computer system 1100A can be integrated into a single package to form a system-in-package (SIP) configuration. In at least one embodiment, at least a portion of the components of computer system 1100A can be integrated into a multi-chip module (MCM), which can be interconnected with other multi-chip modules to form a modular computer system.

[0213] Reasoning and / or training logic 615 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 6B and / or Figure 6C Provides details about the reasoning and / or training logic 615. In at least one embodiment, the reasoning and / or training logic 615 may be Figure 11A for use in a system for performing inference or prediction operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures, or neural network use cases herein.

[0214] processor

[0215] Figure 11B FIG1 shows a parallel processor 1100B according to at least one embodiment. In at least one embodiment, the various components of the parallel processor 1100B may be implemented using one or more integrated circuit devices, such as a programmable processor, an application specific integrated circuit (ASIC), or a field programmable gate array (FPGA). In at least one embodiment, the parallel processor 1100B shown is a processor according to an exemplary embodiment. Figure 11B A variation of the one or more parallel processors 1112 is shown.

[0216] In at least one embodiment, parallel processor 1100B includes parallel processing unit 1102. In at least one embodiment, parallel processing unit 1102 includes an I / O unit 1104 that enables communication with other devices, including other instances of parallel processing unit 1102. In at least one embodiment, I / O unit 1104 can be directly connected to other devices. In at least one embodiment, I / O unit 1104 connects to other devices using a hub or switch interface (e.g., memory hub 1105). In at least one embodiment, the connection between memory hub 1105 and I / O unit 1104 forms a communication link 1113. In at least one embodiment, I / O unit 1104 is connected to a host interface 1106 and a memory crossbar switch 1116, where host interface 1106 receives commands for performing processing operations and memory crossbar switch 1116 receives commands for performing memory operations.

[0217] In at least one embodiment, when host interface 1106 receives command buffers via I / O unit 1104, host interface 1106 can direct work operations to execute those commands to front end 1108. In at least one embodiment, front end 1108 is coupled to scheduler 1110, which is configured to distribute commands or other work items to processing cluster array 1112. In at least one embodiment, scheduler 1110 ensures that processing cluster array 1112 is properly configured and in a valid state before distributing tasks to processing cluster array 1112. In at least one embodiment, scheduler 1110 is implemented by firmware logic executing on a microcontroller. In at least one embodiment, a microcontroller-implemented scheduler 1110 can be configured to perform complex scheduling and work distribution operations at both coarse and fine granularity, thereby enabling fast preemption and context switching of threads executing on processing array 1112. In at least one embodiment, host software can authenticate workloads for scheduling on processing array 1112 through one of multiple graphics processing doorbells. In at least one embodiment, the workload may then be automatically distributed across the processing array 1112 by scheduler 1110 logic within a microcontroller that includes scheduler 1110 .

[0218] In at least one embodiment, processing cluster array 1112 may include up to "N" processing clusters (e.g., cluster 1114A, cluster 1114B, through cluster 1114N). In at least one embodiment, each cluster 1114A-1114N of processing cluster array 1112 may execute a large number of concurrent threads. In at least one embodiment, scheduler 1110 may allocate work to clusters 1114A-1114N of processing cluster array 1112 using various scheduling and / or work distribution algorithms, which may vary depending on the workload generated by each program or computation type. In at least one embodiment, scheduling may be handled dynamically by scheduler 1110 or may be assisted in part by compiler logic during the compilation of program logic configured to be executed by processing cluster array 1112. In at least one embodiment, different clusters 1114A-1114N of processing cluster array 1112 may be assigned to process different types of programs or to perform different types of computations.

[0219] In at least one embodiment, processing cluster array 1112 can be configured to perform various types of parallel processing operations. In at least one embodiment, processing cluster array 1112 can be configured to perform general-purpose parallel computing operations. In at least one embodiment, processing cluster array 1112 can include logic to perform processing tasks including filtering video and / or audio data, performing modeling operations including physics operations, and performing data transformations.

[0220] In at least one embodiment, processing cluster array 1112 is configured to perform parallel graphics processing operations. In at least one embodiment, processing cluster array 1112 may include additional logic to support the execution of such graphics processing operations, including but not limited to texture sampling logic to perform texture operations, as well as tessellation logic and other vertex processing logic. In at least one embodiment, processing cluster array 1112 may be configured to execute shader programs related to graphics processing, such as, but not limited to, vertex shaders, tessellation shaders, geometry shaders, and pixel shaders. In at least one embodiment, parallel processing units 1102 may transfer data from system memory via I / O units 1104 for processing. In at least one embodiment, during processing, the transferred data may be stored in on-chip memory (such as parallel processor memory 1122) during processing and then written back to system memory.

[0221] In at least one embodiment, when parallel processing unit 1102 is used to perform graphics processing, scheduler 1110 can be configured to divide the processing workload into tasks of approximately equal size to better distribute graphics processing operations to multiple clusters 1114A-1114N of processing cluster array 1112. In at least one embodiment, portions of processing cluster array 1112 can be configured to perform different types of processing. In at least one embodiment, a first portion can be configured to perform vertex shading and topology generation, a second portion can be configured to perform tessellation and geometry shading, and a third portion can be configured to perform pixel shading or other screen-space operations to produce rendered images for display, as needed to simulate valve control of a cooling system incorporating an adaptive heat sink. In at least one embodiment, intermediate data generated by one or more of clusters 1114A-1114N can be stored in a buffer to allow the intermediate data to be transferred between clusters 1114A-1114N for further processing.

[0222] In at least one embodiment, the processing cluster array 1112 can receive processing tasks to be executed via the scheduler 1110, which receives commands defining the processing tasks from the front end 1108. In at least one embodiment, the processing tasks can include an index of data to be processed, such as surface (patch) data, primitive data, vertex data, and / or pixel data, as well as state parameters and commands defining how to process the data (such as what program to execute). In at least one embodiment, the scheduler 1110 can be configured to obtain an index corresponding to a task, or can receive the index from the front end 1108. In at least one embodiment, the front end 1108 can be configured to ensure that the processing cluster array 1112 is configured in a valid state before starting a workload specified by an incoming command buffer (such as a batch buffer, push buffer, etc.).

[0223] In at least one embodiment, each of one or more instances of parallel processing unit 1102 can be coupled to parallel processor memory 1122. In at least one embodiment, parallel processor memory 1122 can be accessed via memory crossbar 1116, which can receive memory requests from processing cluster array 1112 and I / O unit 1104. In at least one embodiment, memory crossbar 1116 can access parallel processor memory 1122 via memory interface 1118. In at least one embodiment, memory interface 1118 can include a plurality of partition units (such as partition unit 1120A, partition unit 1120B, and partition unit 1120N), each of which can be coupled to a portion of parallel processor memory 1122 (such as a memory unit). In at least one embodiment, the plurality of partition units 1120A-1120N are configured to be equal to the number of memory cells, such that the first partition unit 1120A has a corresponding first memory cell 1124A, the second partition unit 1120B has a corresponding memory cell 1124B, and the Nth partition unit 1120N has a corresponding Nth memory cell 1124N. In at least one embodiment, the number of partition units 1120A-1120N may not be equal to the number of memory devices.

[0224] In at least one embodiment, memory units 1124A-1124N may include various types of memory devices, including dynamic random access memory (DRAM) or graphics random access memory, such as synchronous graphics random access memory (SGRAM), including graphics double data rate (GDDR) memory. In at least one embodiment, memory units 1124A-1124N may also include 3D stacked memory, including but not limited to high bandwidth memory (HBM). In at least one embodiment, render targets such as frame buffers or texture maps may be stored across memory units 1124A-1124N, allowing partition units 1120A-1120N to write portions of each render target in parallel to efficiently use the available bandwidth of parallel processor memory 1122. In at least one embodiment, local instances of parallel processor memory 1122 may be eliminated in favor of a unified memory design utilizing system memory in combination with local cache memory.

[0225] In at least one embodiment, any of the clusters 1114A-1114N in the processing cluster array 1112 can process data to be written to any memory unit 1124A-1124N within the parallel processor memory 1122. In at least one embodiment, the memory crossbar 1116 can be configured to transmit the output of each cluster 1114A-1114N to any partition unit 1120A-1120N or another cluster 1114A-1114N, which can perform other processing operations on the output. In at least one embodiment, each cluster 1114A-1114N can communicate with a memory interface 1118 via the memory crossbar 1116 to read from or write to various external storage devices. In at least one embodiment, memory crossbar 1116 has connections to memory interface 1118 for communicating with I / O unit 1104, as well as connections to local instances of parallel processor memory 1102, thereby enabling processing units within different processing clusters 1114A-1114N to communicate with system memory or other memory that is not local to parallel processing unit 1102. In at least one embodiment, memory crossbar 1116 may use virtual channels to separate traffic flows between clusters 1114A-1114N and partition units 1120A-1120N.

[0226] In at least one embodiment, multiple instances of parallel processing unit 1102 can be provided on a single plug-in card, or multiple plug-in cards can be interconnected. In at least one embodiment, different instances of parallel processing unit 1102 can be configured to interoperate with each other, even if the different instances have different numbers of processing cores, different amounts of local parallel processor memory, and / or other configuration differences. In at least one embodiment, some instances of parallel processing unit 1102 can include higher precision floating point units relative to other instances. In at least one embodiment, a system incorporating one or more instances of parallel processing unit 1102 or parallel processor 1100B can be implemented in a variety of configurations and form factors, including but not limited to desktop, laptop, or handheld personal computers, servers, workstations, game consoles, and / or embedded systems.

[0227] Figure 11C is a block diagram of a partition unit 1120 according to at least one embodiment. In at least one embodiment, the partition unit 1120 is Figure 11B1120N。In at least one embodiment, the partition unit 1120 includes an L2 cache 1121, a frame buffer interface 1125 and an ROP 1126 (raster operation unit). The L2 cache 1121 is a read / write cache that is configured to perform load and store operations received from the memory crossbar switch 1116 and the ROP 1126. In at least one embodiment, the L2 cache 1121 outputs read misses and urgent write back requests to the frame buffer interface 1125 for processing. In at least one embodiment, updates can also be sent to the frame buffer via the frame buffer interface 1125 for processing. In at least one embodiment, the frame buffer interface 1125 communicates with memory units in the parallel processor memory (such as Figure 11B interacts with one of the memory units 1124A-1124N (such as within parallel processor memory 1122).

[0228] In at least one embodiment, ROP 1126 is a processing unit that performs raster operations such as stenciling, z-testing, blending, and the like. In at least one embodiment, ROP 1126 then outputs processed graphics data that is stored in graphics memory. In at least one embodiment, ROP 1126 includes compression logic to compress depth or color data written to memory and decompress depth or color data read from memory. In at least one embodiment, the compression logic can be lossless compression logic that utilizes one or more of a variety of compression algorithms. The compression logic implemented by ROP 1126 can vary based on the statistical characteristics of the data to be compressed. In at least one embodiment, incremental color compression is performed based on the depth and color data on a per-tile basis.

[0229] In at least one embodiment, ROP 1126 is included within each processing cluster (e.g., Figure 11B In at least one embodiment, read and write requests for pixel data are transmitted through the memory crossbar 1116 rather than through the pixel fragment data transfer. In at least one embodiment, the processed graphics data can be displayed on a display device such as a Figure 11A 110), is displayed on one of the one or more display devices 1110, is routed by the processor 1102 for further processing, or is displayed on one of the one or more display devices 1110 Figure 11B One of the processing entities within parallel processor 1100B is routed for further processing.

[0230] Figure 11D is a block diagram of a processing cluster 1114 within a parallel processing unit according to at least one embodiment. In at least one embodiment, a processing cluster is Figure 11BIn at least one embodiment, one or more processing clusters 1114 can be configured to execute many threads in parallel, where a "thread" refers to an instance of a particular program executed on a particular set of input data. In at least one embodiment, single instruction multiple data (SIMD) instruction issuance technology is used to support the parallel execution of a large number of threads without providing multiple independent instruction units. In at least one embodiment, single instruction multiple thread (SIMT) technology is used to support the parallel execution of a large number of synchronized threads, which uses a common instruction unit that is configured to issue instructions to a group of processing engines within each processing cluster.

[0231] In at least one embodiment, the operation of the processing cluster 1114 can be controlled by a pipeline manager 1132 that assigns processing tasks to SIMT parallel processors. In at least one embodiment, the pipeline manager 1132 Figure 11B The scheduler 1110 receives instructions and manages the execution of these instructions through the graphics multiprocessor 1134 and / or the texture unit 1136. In at least one embodiment, the graphics multiprocessor 1134 is an exemplary instance of a SIMT parallel processor. However, in at least one embodiment, various types of SIMT parallel processors with different architectures may be included within the processing cluster 1114. In at least one embodiment, one or more instances of the graphics multiprocessor 1134 may be included within the processing cluster 1114. In at least one embodiment, the graphics multiprocessor 1134 may process data, and the data crossbar 1140 may be used to distribute the processed data to one of multiple possible destinations (including other shader units). In at least one embodiment, the pipeline manager 1132 may facilitate the distribution of processed data by specifying the destination of the processed data to be distributed via the data crossbar 1140.

[0232] In at least one embodiment, each graphics multiprocessor 1134 within a processing cluster 1114 may include the same set of function execution logic (such as an arithmetic logic unit, a load-store unit, etc.). In at least one embodiment, the function execution logic may be configured in a pipelined manner, where a new instruction may be issued before a previous instruction has completed. In at least one embodiment, the function execution logic supports a variety of operations, including integer and floating-point arithmetic, comparison operations, Boolean operations, shifts, and calculations of various algebraic functions. In at least one embodiment, the same functional unit hardware may be utilized to perform different operations, and any combination of functional units may be present.

[0233] In at least one embodiment, instructions transmitted to processing cluster 1114 constitute threads. In at least one embodiment, a group of threads executed across a set of parallel processing engines is a thread group. In at least one embodiment, a thread group executes a program on different input data. In at least one embodiment, each thread within a thread group can be assigned to a different processing engine within graphics multiprocessor 1134. In at least one embodiment, a thread group can include fewer threads than the number of processing engines within graphics multiprocessor 1134. In at least one embodiment, when a thread group includes fewer threads than the number of processing engines, one or more processing engines may be idle during the processing of a loop by the thread group. In at least one embodiment, a thread group can also include more threads than the number of processing engines within graphics multiprocessor 1134. In at least one embodiment, when a thread group includes more threads than the number of processing engines within graphics multiprocessor 1134, processing can be performed within consecutive clock cycles. In at least one embodiment, multiple thread groups can be executed simultaneously on graphics multiprocessor 1134.

[0234] In at least one embodiment, the graphics multiprocessor 1134 includes internal cache memory to perform load and store operations. In at least one embodiment, the graphics multiprocessor 1134 can abandon the internal cache and use cache memory within the processing cluster 1114 (such as, L1 cache 1148). In at least one embodiment, each graphics multiprocessor 1134 can also access partition units (such as, Figure 11B 1120N) are shared across all processing clusters 1114 and can be used to transfer data between threads. In at least one embodiment, graphics multiprocessor 1134 can also access off-chip global memory, which can include one or more of local parallel processor memory and / or system memory. In at least one embodiment, any memory external to parallel processing unit 1102 can be used as global memory. In at least one embodiment, processing cluster 1114 includes multiple instances of graphics multiprocessor 1134, which can share common instructions and data, which can be stored in L1 cache 1148.

[0235] In at least one embodiment, each processing cluster 1114 may include a memory management unit ("MMU") 1145 configured to map virtual addresses to physical addresses. In at least one embodiment, one or more instances of MMU 1145 may reside in Figure 11B1148 . In at least one embodiment, the MMU 1145 includes a set of page table entries (PTEs) that are used to map virtual addresses to physical addresses of tiles and, in at least one embodiment, to cache line indices. In at least one embodiment, the MMU 1145 may include an address translation lookaside buffer (TLB) or a cache that may reside within the graphics multiprocessor 1134 or L1 cache 1148 or processing cluster 1114. In at least one embodiment, the physical address is processed to assign surface data access locality for efficient request interleaving between partition units. In at least one embodiment, the cache line index may be used to determine whether a request for a cache line is a hit or a miss.

[0236] In at least one embodiment, the processing clusters 1114 can be configured such that each graphics multiprocessor 1134 is coupled to a texture unit 1136 to perform texture mapping operations, such as determining texture sample locations, reading texture data, and filtering texture data. In at least one embodiment, texture data is read from an internal texture L1 cache (not shown) or from an L1 cache within the graphics multiprocessor 1134, and texture data is retrieved from an L2 cache, local parallel processor memory, or system memory as needed. In at least one embodiment, each graphics multiprocessor 1134 outputs one or more processed tasks to a data crossbar 1140 to provide the processed tasks to another processing cluster 1114 for further processing or to store the processed tasks in an L2 cache, local parallel processor memory, or system memory via the memory crossbar 1116. In at least one embodiment, a preROP 1142 (pre-raster operations unit) is configured to receive data from the graphics multiprocessor 1134 and direct the data to a ROP unit, which can communicate with a partitioning unit (such as a partitioning unit) as described herein. Figure 11B In at least one embodiment, the PreROP 1142 unit may perform optimizations for color blending, organize pixel color data, and perform address translation.

[0237] Reasoning and / or training logic 615 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 6B and / or Figure 6C Details are provided regarding inference and / or training logic 615. In at least one embodiment, inference and / or training logic 615 can be used in graphics processing cluster 1114 to perform inference or prediction operations based at least in part on weight parameters computed using the neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.

[0238] Figure 11E A graphics multiprocessor 1134 is shown in accordance with at least one embodiment. In at least one embodiment, the graphics multiprocessor 1134 is coupled to a pipeline manager 1132 of a processing cluster 1114. In at least one embodiment, the graphics multiprocessor 1134 has an execution pipeline that includes, but is not limited to, an instruction cache 1152, an instruction unit 1154, an address mapping unit 1156, a register file 1158, one or more general purpose graphics processing unit (GPGPU) cores 1162, and one or more load / store units 1166. The one or more GPGPU cores 1162 and the one or more load / store units 1166 are coupled to a cache memory 1172 and a shared memory 1170 via a memory and cache interconnect 1168.

[0239] In at least one embodiment, the instruction cache 1152 receives a stream of instructions to be executed from the pipeline manager 1132. In at least one embodiment, the instructions are cached in the instruction cache 1152 and dispatched for execution by the instruction unit 1154. In one embodiment, the instruction unit 1154 can dispatch instructions as thread groups (such as warps), assigning each thread group to a different execution unit within one or more GPGPU cores 1162. In at least one embodiment, instructions can access any local, shared, or global address space by specifying an address within the unified address space. In at least one embodiment, the address mapping unit 1156 can be used to convert addresses in the unified address space into different memory addresses that can be accessed by one or more load / store units 1166.

[0240] In at least one embodiment, register file 1158 provides a set of registers for the functional units of graphics multiprocessor 1134. In at least one embodiment, register file 1158 provides temporary storage for operands for the data paths of the functional units connected to graphics multiprocessor 1134, such as GPGPU core 1162 and load / store unit 1166. In at least one embodiment, register file 1158 is divided between each functional unit such that a dedicated portion of register file 1158 is allocated to each functional unit. In at least one embodiment, register file 1158 is divided between the different warps being executed by graphics multiprocessor 1134.

[0241] In at least one embodiment, the GPGPU cores 1162 may each include a floating-point unit (FPU) and / or an integer arithmetic logic unit (ALU) for executing instructions for the graphics multiprocessor 1134. The GPGPU cores 1162 may be architecturally similar or may differ in architecture. In at least one embodiment, a first portion of the GPGPU core 1162 includes a single-precision FPU and integer ALU, while a second portion of the GPGPU core includes a double-precision FPU. In at least one embodiment, the FPU may implement the IEEE 754-2008 standard for floating-point arithmetic or enable variable-precision floating-point arithmetic. In at least one embodiment, the graphics multiprocessor 1134 may additionally include one or more fixed-function or special-function units to perform specific functions, such as copying rectangles or pixel blending operations. In at least one embodiment, one or more of the GPGPU cores may also include fixed-function or special-function logic.

[0242] In at least one embodiment, the GPGPU core 1162 includes SIMD logic capable of executing a single instruction on multiple sets of data. In one embodiment, the GPGPU core 1162 can physically execute SIMD4, SIMD8, and SIMD16 instructions, and logically execute SIMD1, SIMD2, and SIMD32 instructions. In at least one embodiment, the SIMD instructions for the GPGPU core can be generated by a shader compiler at compile time, or automatically generated when executing a program written and compiled for a single program multiple data (SPMD) or SIMT architecture. In at least one embodiment, multiple threads of a program configured for a SIMT execution model can be executed by a single SIMD instruction. For example, in at least one embodiment, eight SIMT threads performing the same or similar operations can be executed in parallel by a single SIMD8 logic unit.

[0243] In at least one embodiment, the memory and cache interconnect 1168 is an interconnect network that connects each functional unit of the graphics multiprocessor 1134 to the register file 1158 and the shared memory 1170. In at least one embodiment, the memory and cache interconnect 1168 is a crossbar interconnect that allows the load / store unit 1166 to perform load and store operations between the shared memory 1170 and the register file 1158. In at least one embodiment, the register file 1158 can operate at the same frequency as the GPGPU core 1162, resulting in very low latency for data transfers between the GPGPU core 1162 and the register file 1158. In at least one embodiment, the shared memory 1170 can be used to enable communication between threads executing on the functional units within the graphics multiprocessor 1134. In at least one embodiment, the cache memory 1172 can be used, for example, as a data cache to cache texture data communicated between the functional units and the texture unit 1136. In at least one embodiment, the shared memory 1170 can also be used as a program-managed cache. In at least one embodiment, in addition to automatically cached data stored in cache memory 1172, threads executing on GPGPU core 1162 may programmatically store data in shared memory.

[0244] In at least one embodiment, a parallel processor or GPGPU as described herein is communicatively coupled to a host / processor core to accelerate graphics operations, machine learning operations, pattern analysis operations, and various general-purpose GPU (GPGPU) functions. In at least one embodiment, the GPU can be communicatively coupled to the host processor / core via a bus or other interconnect (e.g., a high-speed interconnect such as PCIe or NVLink). In at least one embodiment, the GPU can be integrated into the same package or chip as the core and communicatively coupled to the core via an internal processor bus / interconnect (inside the package or chip, in at least one embodiment). In at least one embodiment, regardless of the manner in which the GPU is connected, the processor core can assign work to the GPU in the form of a sequence of commands / instructions contained in a work descriptor. In at least one embodiment, the GPU then uses dedicated circuitry / logic to efficiently process these commands / instructions.

[0245] The reasoning and / or training logic 615 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 6B and / or Figure 6CDetails are provided regarding inference and / or training logic 615. In at least one embodiment, inference and / or training logic 615 may be used in graphics multiprocessor 1134 to perform inference or prediction operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.

[0246] FIG12A illustrates a multi-GPU computing system 1200A according to at least one embodiment. In at least one embodiment, multi-GPU computing system 1200A may include a processor 1202 coupled to a plurality of general-purpose graphics processing units (GPGPUs) 1206A-D via a host interface switch 1204. In at least one embodiment, host interface switch 1204 is a PCI Express switch device that couples processor 1202 to a PCI Express bus, through which processor 1202 can communicate with GPGPUs 1206A-D. GPGPUs 1206A-D may be interconnected via a set of high-speed P2P GPU-to-GPU links 1216. In at least one embodiment, GPU-to-GPU links 1216 connect to each of GPGPUs 1206A-D via a dedicated GPU link. In at least one embodiment, P2P GPU links 1216 enable direct communication between each GPGPU 1206A-D without requiring communication through host interface bus 1204 to which processor 1202 is connected. In at least one embodiment, host interface bus 1204 remains available for system memory access or communication with other instances of multi-GPU computing system 1200A, e.g., via one or more network devices, with GPU-to-GPU traffic directed to P2P GPU link 1216. While in at least one embodiment, GPGPUs 1206A-D are connected to processor 1202 via host interface switch 1204, in at least one embodiment, processor 1202 includes direct support for P2P GPU link 1216 and can connect directly to GPGPUs 1206A-D.

[0247] Reasoning and / or training logic 615 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 6B and / or Figure 6C Details are provided regarding inference and / or training logic 615. In at least one embodiment, inference and / or training logic 615 can be used in multi-GPU computing system 1200A to perform inference or prediction operations based at least in part on weight parameters computed using the neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.

[0248] Figure 12B FIG1 is a block diagram of a graphics processor 1200B according to at least one embodiment. In at least one embodiment, graphics processor 1200B includes a ring interconnect 1202, a pipeline front end 1204, a media engine 1237, and graphics cores 1280A-1280N. In at least one embodiment, ring interconnect 1202 couples graphics processor 1200B to other processing units, including other graphics processors or one or more general-purpose processor cores. In at least one embodiment, graphics processor 1200B is one of many processors integrated within a multi-core processing system.

[0249] In at least one embodiment, graphics processor 1200B receives batches of commands via ring interconnect 1202. In at least one embodiment, the incoming commands are interpreted by command streamer 1203 in pipeline front end 1204. In at least one embodiment, graphics processor 1200B includes scalable execution logic for performing 3D geometry processing and media processing via graphics cores 1280A-1280N. In at least one embodiment, for 3D geometry processing commands, command streamer 1203 provides commands to geometry pipeline 1236. In at least one embodiment, for at least some media processing commands, command streamer 1203 provides commands to video front end 1234, which is coupled to media engine 1237. In at least one embodiment, media engine 1237 includes a video quality engine (VQE) 1230 for video and image post-processing, and a multi-format encoding / decoding (MFX) 1233 engine for providing hardware-accelerated media data encoding and decoding. In at least one embodiment, the geometry pipeline 1236 and the media engine 1237 each generate execution threads for thread execution resources provided by at least one graphics core 1280A.

[0250] In at least one embodiment, graphics processor 1200B includes scalable thread execution resources featuring modular cores 1280A-1280N (sometimes referred to as core slices), each of which has multiple sub-cores 1250A-1250N, 1260A-1260N (sometimes referred to as core sub-slices). In at least one embodiment, graphics processor 1200B can have any number of graphics cores 1280A. In at least one embodiment, graphics processor 1200B includes graphics core 1280A having at least a first sub-core 1250A and a second sub-core 1260A. In at least one embodiment, graphics processor 1200B is a low-power processor having a single sub-core (such as 1250A). In at least one embodiment, graphics processor 1200B includes multiple graphics cores 1280A-1280N, each of which includes a set of first sub-cores 1250A-1250N and a set of second sub-cores 1260A-1260N. In at least one embodiment, each of the first sub-cores 1250A-1250N includes at least a first set of execution units 1252A-1252N and media / texture samplers 1254A-1254N. In at least one embodiment, each of the second sub-cores 1260A-1260N includes at least a second set of execution units 1262A-1262N and samplers 1264A-1264N. In at least one embodiment, each of the sub-cores 1250A-1250N, 1260A-1260N shares a set of shared resources 1270A-1270N. In at least one embodiment, the shared resources include shared cache memory and pixel operation logic.

[0251] Reasoning and / or training logic 615 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 6B and / or Figure 6C Details are provided regarding inference and / or training logic 615. In at least one embodiment, inference and / or training logic 615 may be used in graphics processor 1200B to perform inference or prediction operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.

[0252] Figure 13is a block diagram illustrating a microarchitecture for a processor 1300, which may include logic circuitry for executing instructions, according to at least one embodiment. In at least one embodiment, processor 1300 may execute instructions including x86 instructions, ARM instructions, specialized instructions for application-specific integrated circuits (ASICs), and the like. In at least one embodiment, processor 1300 may include registers for storing packed data, such as the 64-bit-wide MMX™ registers in microprocessors enabled with MMX technology from Intel Corporation of Santa Clara, California. In at least one embodiment, MMX registers, available in integer and floating-point form, may operate with packed data elements with Single Instruction Multiple Data ("SIMD") and Streaming SIMD Extensions ("SSE") instructions. In at least one embodiment, 128-bit-wide XMM registers associated with SSE2, SSE3, SSE4, AVX, or later (generally referred to as "SSEx") technology may store such packed data operands. In at least one embodiment, processor 1300 may execute instructions to accelerate machine learning or deep learning algorithms, training, or inference.

[0253] In at least one embodiment, the processor 1300 includes an in-order front end ("front end") 1301 to fetch instructions to be executed and prepare the instructions for later use in the processor pipeline. In at least one embodiment, the front end 1301 may include several units. In at least one embodiment, an instruction prefetcher 1326 retrieves instructions from memory and provides the instructions to an instruction decoder 1328, which in turn decodes or interprets the instructions. For example, in at least one embodiment, the instruction decoder 1328 decodes the received instructions into one or more operations called "microinstructions" or "micro-operations" (also referred to as "micro-ops" or "micro-instructions") that the machine can execute. In at least one embodiment, the instruction decoder 1328 parses the instructions into an opcode and corresponding data and control fields, which can be used by the microarchitecture to perform the operations according to at least one embodiment. In at least one embodiment, the trace cache 1330 can assemble the decoded microinstructions into a program-ordered sequence or trace in the microinstruction queue 1334 for execution. In at least one embodiment, when trace cache 1330 encounters a complex instruction, microcode ROM 1332 provides the microinstructions necessary to complete the operation.

[0254] In at least one embodiment, some instructions may be converted into a single micro-op, while other instructions may require several micro-ops to complete the entire operation. In at least one embodiment, if more than four micro-ops are required to complete an instruction, the instruction decoder 1328 may access the microcode ROM 1332 to execute the instruction. In at least one embodiment, an instruction may be decoded into a smaller number of micro-ops for processing at the instruction decoder 1328. In at least one embodiment, if multiple micro-ops are required to complete the operation, the instruction may be stored in the microcode ROM 1332. In at least one embodiment, the trace cache 1330 references the entry point programmable logic array ("PLA") to determine the correct micro-op pointer for reading the microcode sequence from the microcode ROM 1332 to complete one or more instructions according to at least one embodiment. In at least one embodiment, after the microcode ROM 1332 completes the micro-op sequencing for the instruction, the front end 1301 of the machine may resume fetching micro-ops from the trace cache 1330.

[0255] In at least one embodiment, an out-of-order execution engine ("OOO engine") 1303 can prepare instructions for execution. In at least one embodiment, the OOO logic has multiple buffers to smooth and reorder the instruction flow to optimize performance as instructions flow down the pipeline and are scheduled for execution. In at least one embodiment, the OOO engine 1303 includes, but is not limited to, an allocator / register renamer 1340, a memory microinstruction queue 1342, an integer / floating-point microinstruction queue 1344, a memory scheduler 1346, a fast scheduler 1302, a slow / general purpose floating-point scheduler ("slow / general purpose FP scheduler") 1304, and a simple floating-point scheduler ("simple FP scheduler") 1306. In at least one embodiment, the fast scheduler 1302, the slow / general purpose floating-point scheduler 1304, and the simple floating-point scheduler 1306 are also collectively referred to as "microinstruction schedulers 1302, 1304, 1306." In at least one embodiment, the allocator / register renamer 1340 allocates the machine buffers and resources required for each microinstruction to execute in sequence. In at least one embodiment, the allocator / register renamer 1340 renames logical registers into entries in the register file. In at least one embodiment, the allocator / register renamer 1340 also allocates an entry for each microinstruction in one of two microinstruction queues: a memory microinstruction queue 1342 for memory operations and an integer / floating-point microinstruction queue 1344 for non-memory operations, preceding the memory scheduler 1346 and the microinstruction schedulers 1302, 1304, 1306. In at least one embodiment, the microinstruction schedulers 1302, 1304, 1306 determine when a microinstruction is ready to execute based on the readiness of its dependent input register operand sources and the availability of the execution resource microinstructions that need to be completed. In at least one embodiment, the fast scheduler 1302 of at least one embodiment can schedule every half of the main clock cycle, while the slow / general floating-point scheduler 1304 and the simple floating-point scheduler 1306 can schedule once per main processor clock cycle. In at least one embodiment, microinstruction schedulers 1302, 1304, 1306 arbitrate on dispatch ports to schedule microinstructions for execution.

[0256] In at least one embodiment, execution block 1311 includes, but is not limited to, integer register file / bypass network 1308, floating point register file / bypass network ("FP register file / bypass network") 1310, address generation units ("AGUs") 1312 and 1314, fast arithmetic logic units ("fast ALUs") 1316 and 1318, slow arithmetic logic unit ("slow ALU") 1320, floating point ALU ("FP") 1322, and floating point move unit ("FP move") 1324. In at least one embodiment, integer register file / bypass network 1308 and floating point register file / bypass network 1310 are also referred to herein as "register files 1308, 1310." In at least one embodiment, AGUs 1312 and 1314, fast ALUs 1316 and 1318, slow ALU 1320, floating-point ALU 1322, and floating-point move unit 1324 are also referred to herein as "execution units 1312, 1314, 1316, 1318, 1320, 1322, and 1324." In at least one embodiment, execution block 1311 may include, but is not limited to, any number (including zero) and type of register files, bypass networks, address generation units, and execution units (in any combination).

[0257] In at least one embodiment, register networks 1308 and 1310 may be arranged between microinstruction schedulers 1302, 1304, and 1306 and execution units 1312, 1314, 1316, 1318, 1320, 1322, and 1324. In at least one embodiment, integer register file / bypass network 1308 performs integer operations. In at least one embodiment, floating-point register file / bypass network 1310 performs floating-point operations. In at least one embodiment, each of register networks 1308 and 1310 may include, but is not limited to, a bypass network that can bypass or forward recently completed results that have not yet been written to the register file to new dependent objects. In at least one embodiment, register networks 1308 and 1310 may communicate data with each other. In at least one embodiment, integer register file / bypass network 1308 may include, but is not limited to, two separate register files: one register file for low-order 32-bit data and a second register file for high-order 32-bit data. In at least one embodiment, floating point register file / bypass network 1310 may include, but is not limited to, 128-bit wide entries, as floating point instructions typically have operands that are 64 to 128 bits wide.

[0258] In at least one embodiment, execution units 1312, 1314, 1316, 1318, 1320, 1322, and 1324 may execute instructions. In at least one embodiment, register files 1308 and 1310 store integer and floating-point data operand values required for microinstructions to execute. In at least one embodiment, processor 1300 may include, but is not limited to, any number of execution units 1312, 1314, 1316, 1318, 1320, 1322, and 1324, and any combination thereof. In at least one embodiment, floating-point ALU 1322 and floating-point move unit 1324 may perform floating-point, MMX, SIMD, AVX, SSE, or other operations, including specialized machine learning instructions. In at least one embodiment, floating-point ALU 1322 may include, but is not limited to, a 64-bit by 64-bit floating-point divider to perform division, square root, and remainder micro-operations. In at least one embodiment, floating-point hardware may be used to process instructions involving floating-point values. In at least one embodiment, ALU operations can be passed to fast ALUs 1316 and 1318. In at least one embodiment, fast ALUs 1316 and 1318 can perform fast operations with an effective latency of half a clock cycle. In at least one embodiment, most complex integer operations go to slow ALU 1320, as slow ALU 1320 may include, but is not limited to, integer execution hardware for long-latency operations, such as multipliers, shifts, flag logic, and branch processing. In at least one embodiment, memory load / store operations can be performed by AGUs 1312 and 1314. In at least one embodiment, fast ALU 1316, fast ALU 1318, and slow ALU 1320 can perform integer operations on 64-bit data operands. In at least one embodiment, fast ALU 1316, fast ALU 1318, and slow ALU 1320 can be implemented to support various data bit sizes, including 16, 32, 128, 256, and the like. In at least one embodiment, the floating point ALU 1322 and floating point shift unit 1324 can be implemented to support a range of operands having bits of various widths. In at least one embodiment, the floating point ALU 1322 and floating point shift unit 1324 can operate on 128-bit wide packed data operands in conjunction with SIMD and multimedia instructions.

[0259] In at least one embodiment, the microinstruction schedulers 1302, 1304, 1306 schedule dependent operations before the parent load completes execution. In at least one embodiment, because microinstructions can be speculatively scheduled and executed in processor 1300, processor 1300 can also include logic for handling memory misses. In at least one embodiment, if a data load misses in the data cache, there may be dependent operations running in the pipeline that temporarily prevent the scheduler from having the correct data. In at least one embodiment, a replay mechanism tracks and re-executes instructions that use incorrect data. In at least one embodiment, it may be necessary to replay dependent operations and allow independent operations to complete. In at least one embodiment, the scheduler and replay mechanism of at least one embodiment of the processor can also be designed to capture instruction sequences for text string comparison operations.

[0260] In at least one embodiment, "register" may refer to an on-board processor storage location that can be used as part of an instruction to identify an operand. In at least one embodiment, registers may be those that can be used from outside the processor (from a programmer's perspective). In at least one embodiment, registers may not be limited to a particular type of circuit. Instead, in at least one embodiment, registers can store data, provide data, and perform the functions described herein. In at least one embodiment, the registers described herein can be implemented by circuitry within the processor using a variety of different techniques, such as dedicated physical registers, physical registers dynamically allocated using register renaming, a combination of dedicated and dynamically allocated physical registers, and the like. In at least one embodiment, integer registers store 32-bit integer data. The register file of at least one embodiment also includes eight multimedia SIMD registers for packing data.

[0261] Reasoning and / or training logic 615 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 6B and / or Figure 6C Details are provided regarding the inference and / or training logic 615. In at least one embodiment, some or all of the inference and / or training logic 615 may be incorporated into the execution block 1311 and other memories or registers, shown or not shown. For example, in at least one embodiment, the training and / or inference techniques described herein may utilize one or more ALUs shown in the execution block 1311. Additionally, weight parameters may be stored in on-chip or off-chip memory and / or registers (shown or not shown) that configure the ALUs of the execution block 1311 to execute one or more of the machine learning algorithms, neural network architectures, use cases, or training techniques described herein.

[0262] Figure 14A deep learning application processor 1400 is shown in accordance with at least one embodiment. In at least one embodiment, the deep learning application processor 1400 uses instructions that, if executed by the deep learning application processor 1400, cause the deep learning application processor 1400 to perform some or all of the processes and techniques described throughout this disclosure. In at least one embodiment, the deep learning application processor 1400 is an application-specific integrated circuit (ASIC). In at least one embodiment, the application processor 1400 performs matrix multiplication operations or "hardwires" them into hardware as a result of executing one or more instructions, or both. In at least one embodiment, the deep learning application processor 1400 includes, but is not limited to, processing clusters 1410(1)-1410(12), inter-chip links (“ICLs”) 1420(1)-1420(12), inter-chip controllers (“ICCs”) 1430(1)-1430(2), memory controllers (“Mem Ctrlr”) 1442(1)-1442(4), high bandwidth memory physical layer (“HBM PHY”) 1444(1)-1444(4), a management controller central processing unit (“management controller CPU”) 1450, serial peripheral interface, inter-integrated circuit, and general purpose input / output blocks (“SPI, I2C, GPIO”), a peripheral component interconnect express controller and direct memory access block (“PCIe controller and DMA”) 1470, and a sixteen-lane peripheral component interconnect express port (“PCI Express x 16”) 1480.

[0263] In at least one embodiment, processing cluster 1410 can perform deep learning operations, including inference or prediction operations based on weight parameters calculated by one or more training techniques, including those described herein. In at least one embodiment, each processing cluster 1410 can include, but is not limited to, any number and type of processors. In at least one embodiment, deep learning application processor 1400 can include any number and type of processing clusters 1400. In at least one embodiment, inter-chip link 1420 is bidirectional. In at least one embodiment, inter-chip link 1420 and inter-chip controller 1430 enable multiple deep learning application processors 1400 to exchange information, including activation information generated from executing one or more machine learning algorithms embodied in one or more neural networks. In at least one embodiment, deep learning application processor 1400 can include any number (including zero) and type of ICL 1420 and ICC 1430.

[0264] In at least one embodiment, HBM2 1440 provides a total of 32GB of memory. HBM2 1440(i) is associated with both a memory controller 1442(i) and an HBM PHY 1444(i). In at least one embodiment, any number of HBM2 1440 can provide any type and total amount of high-bandwidth memory and can be associated with any number (including zero) and type of memory controllers 1442 and HBM PHYs 1444. In at least one embodiment, SPI, I2C, GPIO 3360, PCIe controller 1460, DMA 1470, and / or PCIe 1480 can be replaced with any number and type of blocks to implement any number and type of communication standards in any technically feasible manner.

[0265] Reasoning and / or training logic 615 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 6B and / or Figure 6C Details are provided regarding the inference and / or training logic 615. In at least one embodiment, the deep learning application processor is used to train a machine learning model (e.g., a neural network) to predict or infer information provided to the deep learning application processor 1400. In at least one embodiment, the deep learning application processor 1400 is used to infer or predict information based on a trained machine learning model (such as a neural network) that has been trained by another processor or system or by the deep learning application processor 1400. In at least one embodiment, the processor 1400 can be used to perform one or more of the neural network use cases described herein.

[0266] Figure 15 is a block diagram of a neuromorphic processor 1500 according to at least one embodiment. In at least one embodiment, the neuromorphic processor 1500 can receive one or more inputs from a source external to the neuromorphic processor 1500. In at least one embodiment, these inputs can be transmitted to one or more neurons 1502 within the neuromorphic processor 1500. In at least one embodiment, the neurons 1502 and their components can be implemented using circuitry or logic including one or more arithmetic logic units (ALUs). In at least one embodiment, the neuromorphic processor 1500 can include, but is not limited to, thousands of instances of neurons 1502, although any suitable number of neurons 1502 can be used. In at least one embodiment, each instance of a neuron 1502 can include a neuron input 1504 and a neuron output 1506. In at least one embodiment, a neuron 1502 can generate an output that can be transmitted to the inputs of other instances of the neuron 1502. In at least one embodiment, the neuron input 1504 and the neuron output 1506 can be interconnected via a synapse 1508.

[0267] In at least one embodiment, the neurons 1502 and synapses 1508 can be interconnected so that the neuromorphic processor 1500 operates to process or analyze information received by the neuromorphic processor 1500. In at least one embodiment, the neuron 1502 can send an output pulse (or "trigger" or "spike") when the input received through the neuron input 1504 exceeds a threshold. In at least one embodiment, the neuron 1502 can sum or integrate the signals received at the neuron input 1504. For example, in at least one embodiment, the neuron 1502 can be implemented as a leaky integrate-and-trigger neuron, where if the sum (referred to as the "membrane potential") exceeds a threshold, the neuron 1502 can generate an output (or "trigger") using a transfer function such as a sigmoid or threshold function. In at least one embodiment, the leaky integrate-and-trigger neuron can sum the signals received at the neuron input 1504 into a membrane potential and can apply an application attenuation factor (or leakage) to reduce the membrane potential. In at least one embodiment, a leaky integrate-and-trigger neuron may trigger if multiple input signals are received at neuron input 1504 quickly enough to exceed a threshold (in at least one embodiment, before the membrane potential decays too low to trigger). In at least one embodiment, neuron 1502 may be implemented using circuitry or logic that receives input, integrates the input into a membrane potential, and decays the membrane potential. In at least one embodiment, the inputs may be averaged, or any other suitable transfer function may be used. Furthermore, in at least one embodiment, neuron 1502 may include, but is not limited to, comparator circuitry or logic that generates an output spike at neuron output 1506 when the result of applying the transfer function to neuron input 1504 exceeds a threshold. In at least one embodiment, once neuron 1502 triggers, it may ignore previously received input information by, for example, resetting the membrane potential to 0 or another suitable default value. In at least one embodiment, once the membrane potential is reset to 0, neuron 1502 may resume normal operation after a suitable period of time (or recovery period).

[0268] In at least one embodiment, neurons 1502 can be interconnected via synapses 1508. In at least one embodiment, synapses 1508 can be operable to transmit a signal from an output of a first neuron 1502 to an input of a second neuron 1502. In at least one embodiment, a neuron 1502 can transmit information across more than one instance of synapse 1508. In at least one embodiment, one or more instances of a neuron output 1506 can be connected to an instance of a neuron input 1504 in the same neuron 1502 via an instance of synapse 1508. In at least one embodiment, an instance of a neuron 1502 that generates an output to be transmitted across an instance of synapse 1508 can be referred to as a "presynaptic neuron" relative to that instance of synapse 1508. In at least one embodiment, an instance of a neuron 1502 that receives an input transmitted across an instance of synapse 1508 can be referred to as a "postsynaptic neuron" relative to an instance of synapse 1508. In at least one embodiment, with respect to the various instances of synapses 1508, because an instance of neuron 1502 can receive input from one or more instances of synapses 1508 and can also transmit output through one or more instances of synapses 1508, a single instance of neuron 1502 can be both a "pre-synaptic neuron" and a "post-synaptic neuron."

[0269] In at least one embodiment, neurons 1502 may be organized into one or more layers. Each instance of a neuron 1502 may have a neuron output 1506 that may fan out to one or more neuron inputs 1504 via one or more synapses 1508. In at least one embodiment, the neuron outputs 1506 of neurons 1502 in a first layer 1510 may be connected to the neuron inputs 1504 of neurons 1502 in a second layer 1512. In at least one embodiment, layers 1510 may be referred to as "feed-forward layers." In at least one embodiment, each instance of a neuron 1502 in an instance of the first layer 1510 may fan out to each instance of a neuron 1502 in a second layer 1512. In at least one embodiment, the first layer 1510 may be referred to as a "fully connected feed-forward layer." In at least one embodiment, each instance of a neuron 1502 in each instance of the second layer 1512 may fan out to fewer than all instances of neurons 1502 in a third layer 1514. In at least one embodiment, the second layer 1512 may be referred to as a "sparsely connected feed-forward layer." In at least one embodiment, neurons 1502 in the (same) second layer 1512 can fan out to neurons 1502 in multiple other layers, including also fanning out to neurons 1502 in the second layer 1512. In at least one embodiment, the second layer 1512 can be referred to as a "recurrent layer." In at least one embodiment, the neuromorphic processor 1500 can include, but is not limited to, any suitable combination of recurrent layers and feed-forward layers, including, but not limited to, sparsely connected feed-forward layers and fully connected feed-forward layers.

[0270] In at least one embodiment, the neuromorphic processor 1500 may include, but is not limited to, a reconfigurable interconnect architecture or a dedicated hardwired interconnect to connect synapses 1508 to neurons 1502. In at least one embodiment, the neuromorphic processor 1500 may include, but is not limited to, circuitry or logic that allows synapses to be assigned to different neurons 1502 as needed based on the neural network topology and neuron fan-in / fan-out. For example, in at least one embodiment, synapses 1508 may be connected to neurons 1502 using an interconnect architecture such as a network on a chip or through dedicated connections. In at least one embodiment, the synaptic interconnect and its components may be implemented using circuitry or logic.

[0271] Figure 16A16 shows a processing system according to at least one embodiment. In at least one embodiment, system 1600A includes one or more processors 1602 and one or more graphics processors 1608, and can be a single-processor desktop system, a multi-processor workstation system, or a server system with a large number of processors 1602 or processor cores 1607. In at least one embodiment, system 1600A is a processing platform incorporated within a system-on-chip (SoC) integrated circuit for use in mobile, handheld, or embedded devices.

[0272] In at least one embodiment, the system 1600A may include or be incorporated into a server-based gaming platform, including a gaming console, a mobile gaming console, a handheld gaming console, or an online gaming console. In at least one embodiment, the system 1600A is a mobile phone, a smartphone, a tablet computing device, or a mobile internet device. In at least one embodiment, the processing system 1600A may also include a device coupled to or integrated into a wearable device, such as a smartwatch wearable device, a smart glasses device, an augmented reality device, or a virtual reality device. In at least one embodiment, the processing system 1600A is a television or set-top box device having one or more processors 1602 and a graphical interface generated by one or more graphics processors 1608.

[0273] In at least one embodiment, one or more processors 1602 each include one or more processor cores 1607 to process instructions that, when executed, perform operations for system and user software. In at least one embodiment, each of the one or more processor cores 1607 is configured to process a specific instruction set 1609. In at least one embodiment, the instruction set 1609 can facilitate complex instruction set computing (CISC), reduced instruction set computing (RISC), or computing via very long instruction words (VLIW). In at least one embodiment, the processor cores 1607 can each process a different instruction set 1609, which can include instructions that facilitate emulating other instruction sets. In at least one embodiment, the processor cores 1607 can also include other processing devices, such as a digital signal processor (DSP).

[0274] In at least one embodiment, processor 1602 includes cache memory 1604. In at least one embodiment, processor 1602 can have a single internal cache or multiple levels of internal cache. In at least one embodiment, cache memory is shared among various components of processor 1602. In at least one embodiment, processor 1602 also uses an external cache (such as a level 3 (L3) cache or a last level cache (LLC)) (not shown), which can be shared among processor cores 1607 using known cache coherence techniques. In at least one embodiment, processor 1602 also includes a register file 1606. The processor can include different types of registers (such as integer registers, floating point registers, status registers, and instruction pointer registers) for storing different types of data. In at least one embodiment, register file 1606 can include general purpose registers or other registers.

[0275] In at least one embodiment, one or more processors 1602 are coupled to one or more interface buses 1610 to transmit communication signals, such as address, data, or control signals, between the processors 1602 and other components in the system 1600A. In at least one embodiment, the interface bus 1610 can be a processor bus, such as a version of a Direct Media Interface (DMI) bus. In at least one embodiment, the interface bus 1610 is not limited to a DMI bus and can include one or more peripheral component interconnect buses (such as PCI, PCI Express), a memory bus, or other types of interface buses. In at least one embodiment, the processor 1602 includes an integrated memory controller 1616 and a platform controller hub 1630. In at least one embodiment, the memory controller 1616 facilitates communication between memory devices and other components of the processing system 1600A, while the platform controller hub (PCH) 1630 provides connections to input / output (I / O) devices via a local I / O bus.

[0276] In at least one embodiment, memory device 1620 may be a dynamic random access memory (DRAM) device, a static random access memory (SRAM) device, a flash memory device, a phase change memory device, or other suitable memory device for use as processor memory. In at least one embodiment, memory device 1620 may be used as system memory for processing system 1600A to store data 1622 and instructions 1621 for use when one or more processors 1602 execute applications or processes. In at least one embodiment, memory controller 1616 is also coupled to an external graphics processor 1612 of at least one embodiment, which may communicate with one or more graphics processors 1608 in processor 1602 to perform graphics and media operations. In at least one embodiment, display device 1611 may be connected to processor 1602. In at least one embodiment, display device 1611 may include one or more internal display devices, such as in a mobile electronic device or laptop, or an external display device connected via a display interface (such as a DisplayPort). In at least one embodiment, the display device 1611 may include a head-mounted display (HMD), such as a stereoscopic display device used in virtual reality (VR) applications or augmented reality (AR) applications.

[0277] In at least one embodiment, the platform controller hub 1630 enables peripheral devices to connect to the storage device 1620 and the processor 1602 via a high-speed I / O bus. In at least one embodiment, the I / O peripherals include, but are not limited to, an audio controller 1646, a network controller 1634, a firmware interface 1628, a wireless transceiver 1626, a touch sensor 1625, and a data storage device 1624 (such as a hard drive, flash memory, etc.). In at least one embodiment, the data storage device 1624 can be connected via a storage interface (such as SATA) or via a peripheral bus, such as a peripheral component interconnect bus (e.g., PCI, PCIe). In at least one embodiment, the touch sensor 1625 can include a touch screen sensor, a pressure sensor, or a fingerprint sensor. In at least one embodiment, the wireless transceiver 1626 can be a Wi-Fi transceiver, a Bluetooth transceiver, or a mobile network transceiver, such as a 3G, 4G, or Long Term Evolution (LTE) transceiver. In at least one embodiment, the firmware interface 1628 enables communication with the system firmware and can be, for example, a Unified Extensible Firmware Interface (UEFI). In at least one embodiment, a network controller 1634 can enable network connectivity to a wired network. In at least one embodiment, a high-performance network controller (not shown) is coupled to the interface bus 1610. In at least one embodiment, the audio controller 1646 is a multi-channel high-definition audio controller. In at least one embodiment, the processing system 1600A includes a legacy I / O controller 1640 for coupling legacy (e.g., Personal System 2 (PS / 2)) devices to the system 1600A. In at least one embodiment, the platform controller hub 1630 can also be connected to one or more universal serial bus (USB) controllers 1642 that connect input devices such as a keyboard and mouse 1643 combination, a camera 1644, or other USB input devices.

[0278] In at least one embodiment, instances of memory controller 1616 and platform controller hub 1630 may be integrated into a discrete external graphics processor, such as external graphics processor 1612. In at least one embodiment, platform controller hub 1630 and / or memory controller 1616 may be external to one or more processors 1602. In at least one embodiment, system 1600A may include external memory controller 1616 and platform controller hub 1630, which may be configured as a memory controller hub and a peripheral controller hub in a system chipset that communicates with processor 1602.

[0279] Reasoning and / or training logic 615 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 6B and / or Figure 6C615. In at least one embodiment, some or all of the inference and / or training logic 615 may be incorporated into the graphics processor 1600A. In at least one embodiment, the training and / or inference techniques described herein may utilize one or more ALUs embodied in the graphics processor 1612. Additionally, in at least one embodiment, the inference and / or training operations described herein may utilize a number of ALUs other than the ALUs. Figure 6B and / or Figure 6C In at least one embodiment, the weight parameters may be stored in on-chip or off-chip memory and / or registers (shown or not shown), which configure the ALU of graphics processor 1600A to execute one or more of the machine learning algorithms, neural network architectures, use cases, or training techniques described herein.

[0280] Figure 16B is a block diagram of a processor 1600B having one or more processor cores 1602A-1602N, an integrated memory controller 1614, and an integrated graphics processor 1608, in accordance with at least one embodiment. In at least one embodiment, processor 1600B may include additional cores, up to and including additional core 1602N, which is represented by the dashed box. In at least one embodiment, each processor core 1602A-1602N includes one or more internal cache units 1604A-1604N. In at least one embodiment, each processor core may also have access to one or more shared cache units 1606.

[0281] In at least one embodiment, the internal cache units 1604A-1604N and the shared cache unit 1606 represent a cache memory hierarchy within the processor 1600B. In at least one embodiment, the cache memory units 1604A-1604N may include at least one level of instruction and data cache within each processor core and one or more levels of cache within a shared mid-level cache, such as level 2 (L2), level 3 (L3), level 4 (L4), or other levels of cache, with the highest level of cache before external memory being categorized as LLC. In at least one embodiment, cache coherence logic maintains coherence between the various cache units 1606 and 1604A-1604N.

[0282] In at least one embodiment, processor 1600B may also include a set of one or more bus controller units 1616 and a system agent core 1610. In at least one embodiment, the one or more bus controller units 1616 manage a set of peripheral buses, such as one or more PCI or PCIe buses. In at least one embodiment, the system agent core 1610 provides management functions for various processor components. In at least one embodiment, the system agent core 1610 includes one or more integrated memory controllers 1614 to manage access to various external memory devices (not shown).

[0283] In at least one embodiment, one or more processor cores 1602A-1602N include support for simultaneous multithreading. In at least one embodiment, system agent core 1610 includes components for coordinating and operating cores 1602A-1602N during multithreaded processing. In at least one embodiment, system agent core 1610 may additionally include a power control unit (PCU) that includes logic and components for regulating one or more power states of processor cores 1602A-1602N and graphics processor 1608.

[0284] In at least one embodiment, processor 1600B also includes a graphics processor 1608 for performing image processing operations. In at least one embodiment, graphics processor 1608 is coupled to a shared cache unit 1606 and a system agent core 1610 including one or more integrated memory controllers 1614. In at least one embodiment, system agent core 1610 also includes a display controller 1611 for driving the graphics processor output to one or more coupled displays. In at least one embodiment, display controller 1611 may also be a separate module coupled to graphics processor 1608 via at least one interconnect, or may be integrated within graphics processor 1608.

[0285] In at least one embodiment, a ring-based interconnect 1612 is used to couple the internal components of processor 1600B. In at least one embodiment, alternative interconnects may be used, such as point-to-point interconnects, switched interconnects, or other technologies. In at least one embodiment, graphics processor 1608 is coupled to ring interconnect 1612 via I / O link 1613.

[0286] In at least one embodiment, I / O link 1613 represents at least one of a variety of I / O interconnects, including an on-package I / O interconnect that facilitates communication between various processor components and a high-performance embedded memory module 1618 (e.g., an eDRAM module). In at least one embodiment, each of processor cores 1602A-1602N and graphics processor 1608 uses embedded memory module 1618 as a shared last-level cache.

[0287] In at least one embodiment, the processor cores 1602A-1602N are homogeneous cores that execute a common instruction set architecture. In at least one embodiment, the processor cores 1602A-1602N are heterogeneous in terms of instruction set architecture (ISA), wherein one or more processor cores 1602A-1602N execute a common instruction set, while one or more other processor cores 1602A-1602N execute a subset of the common instruction set or a different instruction set. In at least one embodiment, the processor cores 1602A-1602N are heterogeneous in terms of microarchitecture, wherein one or more cores with relatively high power consumption are coupled with one or more power cores with lower power consumption. In at least one embodiment, the processor 1600B can be implemented on one or more chips or as a SoC integrated circuit.

[0288] Reasoning and / or training logic 615 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 6B and / or Figure 6C 6. In at least one embodiment, some or all of the inference and / or training logic 615 may be incorporated into the processor 1600B. For example, in at least one embodiment, the training and / or inference techniques described herein may utilize one or more ALUs embodied in the processor 1600B. Figure 16A In addition, in at least one embodiment, the inference and / or training operations described herein may use a graphics core 1612, one or more processor cores 1602A-1602N, or other components. Figure 6B and / or Figure 6C In at least one embodiment, the weight parameters may be stored in on-chip or off-chip memory and / or registers (shown or not shown) that configure the ALU of the graphics processor 1600B to execute one or more machine learning algorithms, neural network architectures, use cases, or training techniques described herein.

[0289] Figure 16Cis a block diagram of the hardware logic of a graphics processor core 1600C according to at least one embodiment herein. In at least one embodiment, graphics processor core 1600C is included in a graphics core array. In at least one embodiment, graphics processor core 1600C (sometimes referred to as a core slice) can be one or more graphics cores within a modular graphics processor. In at least one embodiment, graphics processor core 1600C is an example of one graphics core slice, and the graphics processors herein can include multiple graphics core slices based on target power and performance envelopes. In at least one embodiment, each graphics core 1600C can include a fixed function block 1630 coupled to multiple sub-cores 1601A-1601F, also referred to as sub-slices, which include modular blocks of general-purpose and fixed-function logic.

[0290] In at least one embodiment, fixed function block 1630 includes a geometry fixed function pipeline 1636. For example, in lower performance and / or lower power graphics processor implementations, the geometry and fixed function pipeline 1636 can be shared by all sub-cores in graphics processor 1600C. In at least one embodiment, the geometry and fixed function pipeline 1636 includes a 3D fixed function pipeline, a video front end unit, a thread generator and thread dispatcher, and a unified return buffer manager that manages the unified return buffer.

[0291] In at least one fixed embodiment, fixed function block 1630 also includes a graphics SoC interface 1637, a graphics microcontroller 1638, and a media pipeline 1639. In at least one embodiment, fixed graphics SoC interface 1637 provides an interface between graphics core 1600C and other processor cores in the on-chip integrated circuit system. In at least one embodiment, graphics microcontroller 1638 is a programmable subprocessor that can be configured to manage various functions of graphics processor 1600C, including thread dispatching, scheduling, and preemption. In at least one embodiment, media pipeline 1639 includes logic that facilitates decoding, encoding, pre-processing, and / or post-processing of multimedia data, including image and video data. In at least one embodiment, media pipeline 1639 implements media operations via requests to computational or sampling logic within sub-cores 1601-1601F.

[0292] In at least one embodiment, SoC interface 1637 enables graphics core 1600C to communicate with a general-purpose application processor core (such as a CPU) and / or other components within the SoC, including memory hierarchy elements such as a shared last-level cache, system RAM, and / or embedded on-chip or packaged DRAM. In at least one embodiment, SoC interface 1637 may also enable communication with fixed-function devices within the SoC (e.g., a camera imaging pipeline) and enable the use and / or implementation of global memory atomics that can be shared between graphics core 1600C and the CPU within the SoC. In at least one embodiment, SoC interface 1637 may also implement power management controls for graphics core 1600C and enable interfaces between the clock domain of graphics core 1600C and other clock domains within the SoC. In at least one embodiment, SoC interface 1637 enables receiving command buffers from a command stream converter and a global thread dispatcher, which are configured to provide commands and instructions to each of one or more graphics cores within the graphics processor. In at least one embodiment, commands and instructions may be dispatched to the media pipeline 1639 when media operations are to be performed, or to the geometry and fixed function pipelines (such as the geometry and fixed function pipeline 1636, the geometry and fixed function pipeline 1614) when graphics processing operations are to be performed.

[0293] In at least one embodiment, graphics microcontroller 1638 can be configured to perform various scheduling and management tasks for graphics core 1600C. In at least one embodiment, graphics microcontroller 1638 can perform graphics and / or compute workload scheduling on various graphics parallel engines within execution unit (EU) arrays 1602A-1602F, 1604A-1604F in sub-cores 1601A-1601F. In at least one embodiment, host software executing on a CPU core of a SoC including graphics core 1600C can submit a workload to one of multiple graphics processor doorbells, which invokes scheduling operations on the appropriate graphics engine. In at least one embodiment, scheduling operations include determining which workload to run next, submitting the workload to the command stream converter, preempting existing workloads running on the engine, monitoring the progress of the workload, and notifying the host software when the workload is complete. In at least one embodiment, graphics microcontroller 1638 may also facilitate low power or idle states for graphics core 1600C, thereby providing graphics core 1600C with the ability to save and restore registers across low power state transitions within graphics core 1600C independent of the operating system and / or graphics driver software on the system.

[0294] In at least one embodiment, graphics core 1600C may have up to N modular sub-cores, more or less than the sub-cores 1601A-1601F shown. For each set of N sub-cores, in at least one embodiment, graphics core 1600C may also include shared function logic 1610, shared and / or cache memory 1612, geometry / fixed function pipelines 1614, and additional fixed function logic 1616 to accelerate various graphics and compute processing operations. In at least one embodiment, shared function logic 1610 may include logic units (such as samplers, math, and / or inter-thread communication logic) that may be shared by each of the N sub-cores within graphics core 1600C. In at least one embodiment, fixed, shared, and / or cache memory 1612 may serve as a last-level cache for the N sub-cores 1601A-1601F within graphics core 1600C and may also serve as shared memory accessible by multiple sub-cores. In at least one embodiment, geometry / fixed function pipeline 1614 may be included in place of geometry / fixed function pipeline 1636 within fixed function block 1630 and may include similar logic.

[0295] In at least one embodiment, graphics core 1600C includes additional fixed-function logic 1616, which may include various fixed-function acceleration logic for use with graphics core 1600C. In at least one embodiment, additional fixed-function logic 1616 includes an additional geometry pipeline for use in position-only shading. In position-only shading, there are at least two geometry pipelines, and within geometry and fixed-function pipelines 1614, 1636, there are at least two geometry pipelines, which are the full geometry pipeline and the culling pipeline, which may be included in additional fixed-function logic 1616. In at least one embodiment, the culling pipeline is a modified version of the full geometry pipeline. In at least one embodiment, the full pipeline and the culling pipeline can execute different instances of an application, each with a separate context. In at least one embodiment, position-only shading can hide long culling runs for discarded triangles, allowing shading to complete earlier in some cases. In at least one embodiment, the culling pipeline logic in the additional fixed function logic 1616 can execute position shaders in parallel with the main application and generate critical results faster than the full pipeline because the culling pipeline obtains and masks the position attributes of the vertices without having to perform rasterization and render the pixels to the frame buffer. In at least one embodiment, the culling pipeline can use the generated critical results to calculate visibility information for all triangles, regardless of whether they are culled. In at least one embodiment, the full pipeline (which in this case may be called a replay pipeline) can consume visibility information to skip culled triangles to mask only visible triangles that are ultimately passed to the rasterization stage.

[0296] In at least one embodiment, the additional fixed-function logic 1616 may also include machine learning acceleration logic, such as fixed-function matrix multiplication logic, to implement optimizations including for machine learning training or inference.

[0297] In at least one embodiment, a set of execution resources is included within each graphics sub-core 1601A-1601F that can be used to perform graphics, media, and compute operations in response to requests from the graphics pipeline, media pipeline, or shader programs. In at least one embodiment, the graphics sub-core 1601A-1601F includes multiple EU arrays 1602A-1602F, 1604A-1604F, thread dispatch and inter-thread communication (TD / IC) logic 1603A-1603F, 3D (such as texture) samplers 1605A-1605F, media samplers 1606A-1606F, shader processors 1607A-1607F, and shared local memory (SLM) 1608A-1608F. Each of EU arrays 1602A-1602F, 1604A-1604F includes multiple execution units, which are general-purpose graphics processing units capable of servicing graphics, media, or compute operations, executing floating-point and integer / fixed-point logic operations, including graphics, media, or compute shader programs. In at least one embodiment, TD / IC logic 1603A-1603F performs local thread dispatch and thread control operations for execution units within a sub-core and facilitates communication between threads executing on execution units within the sub-core. In at least one embodiment, 3D samplers 1605A-1605F can read texture or other 3D graphics-related data into memory. In at least one embodiment, 3D samplers can read texture data differently based on the configured sampling state and texture format associated with a given texture. In at least one embodiment, media samplers 1606A-1606F can perform similar read operations based on the type and format associated with the media data. In at least one embodiment, each graphics sub-core 1601A-1601F may alternatively include a unified 3D and media sampler. In at least one embodiment, threads executing on execution units within each sub-core 1601A-1601F may utilize shared local memory 1608A-1608F within each sub-core, enabling threads executing within a thread group to execute using a common pool of on-chip memory.

[0298] Reasoning and / or training logic 615 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 6B and / or Figure 6C615. In at least one embodiment, some or all of the inference and / or training logic 615 may be incorporated into the graphics processor 1610. In at least one embodiment, the training and / or inference techniques described herein may be used in Figure 16B In addition, in at least one embodiment, the inference and / or training operations described herein may use the addition of Figure 6B and / or Figure 6C In at least one embodiment, the weight parameters may be stored in on-chip or off-chip memory and / or registers (shown or not shown), which configure the ALU of the graphics processor 1600C to execute one or more of the machine learning algorithms, neural network architectures, use cases, or training techniques described herein.

[0299] Figures 16D-16E Thread execution logic 1600D is shown for an array of processing elements comprising a graphics processor core, in accordance with at least one embodiment. Figure 16D At least one embodiment is shown in which thread execution logic 1600D is used. Figure 16E Illustrative internal details of an execution unit are shown in accordance with at least one embodiment.

[0300] like Figure 16DAs shown in FIG, in at least one embodiment, thread execution logic 1600D includes a shader processor 1602, a thread dispatcher 1604, an instruction cache 1606, a scalable execution unit array including a plurality of execution units 1608A-1608N, one or more samplers 1610, a data cache 1612, and a data port 1614. In at least one embodiment, the scalable execution unit array can be dynamically scaled by enabling or disabling one or more execution units (such as any one of execution units 1608A, 1608B, 1608C, 1608D, 1608N-1, and 1608N), for example, based on the computational requirements of the workload. In at least one embodiment, the scalable execution units are interconnected via an interconnect structure that links to each execution unit. In at least one embodiment, thread execution logic 1600D includes one or more connections to memory (such as system memory or cache memory) through instruction cache 1606, data port 1614, sampler 1610, and one or more of execution units 1608A-1608N. In at least one embodiment, each execution unit (such as 1608A) is an independent programmable general-purpose computing unit that is capable of executing multiple simultaneous hardware threads while processing multiple data elements in parallel for each thread. In at least one embodiment, the array of execution units 1608A-1608N is scalable to include any number of individual execution units.

[0301] In at least one embodiment, execution units 1608A-1608N are primarily used to execute shader programs. In at least one embodiment, shader processor 1602 can process various shader programs and dispatch execution threads associated with the shader programs via thread dispatcher 1604. In at least one embodiment, thread dispatcher 1604 includes logic for arbitrating thread initialization requests from graphics and media pipelines and instantiating requested threads on one or more execution units 1608A-1608N. In at least one embodiment, in at least one embodiment, the geometry pipeline can dispatch vertex, tessellation, or geometry shaders to thread execution logic for processing. In at least one embodiment, thread dispatcher 1604 can also handle runtime thread generation requests from executing shader programs.

[0302] In at least one embodiment, execution units 1608A-1608N support an instruction set that includes native support for many standard 3D graphics shader instructions, enabling shader programs from graphics libraries such as Direct3D and OpenGL to execute with minimal translation. In at least one embodiment, the execution units support vertex and geometry processing (e.g., vertex programs, geometry programs, vertex shaders), pixel processing (e.g., pixel shaders, fragment shaders), and general-purpose processing (e.g., compute and media shaders). In at least one embodiment, each execution unit 1608A-1608N includes one or more arithmetic logic units (ALUs) capable of multi-issue single instruction, multiple data (SIMD) and multi-threaded operation, enabling an efficient execution environment despite higher latency memory accesses. In at least one embodiment, each hardware thread within each execution unit has a dedicated high-bandwidth register file and associated independent thread state. In at least one embodiment, execution is multiple issues per clock to the pipeline, which is capable of integer, single-precision, and double-precision floating-point operations, SIMD branching functions, logical operations, transcendental operations, and other operations. In at least one embodiment, while waiting for data from memory or one of the shared functions, dependency logic within execution units 1608A-1608N causes the waiting thread to sleep until the requested data is returned. In at least one embodiment, while the waiting thread is sleeping, hardware resources can be dedicated to processing other threads. In at least one embodiment, during the delay associated with vertex shader operations, the execution unit can perform operations on a pixel shader, a fragment shader, or another type of shader program (including a different vertex shader).

[0303] In at least one embodiment, each of execution units 1608A-1608N operates on an array of data elements. In at least one embodiment, the number of data elements is the "execution size" or number of lanes of an instruction. In at least one embodiment, an execution lane is the logic used to perform data element access, masking, and flow control within an instruction. In at least one embodiment, the number of lanes can be independent of the number of physical arithmetic logic units (ALUs) or floating point units (FPUs) used for a particular graphics processor. In at least one embodiment, execution units 1608A-1608N support integer and floating point data types.

[0304] In at least one embodiment, the execution unit instruction set includes SIMD instructions. In at least one embodiment, various data elements can be stored in registers as packed data types, and the execution unit will process various elements based on the data size of those elements. In at least one embodiment, in at least one embodiment, when operating on a 256-bit wide vector, the 256 bits of the vector are stored in registers, and the execution unit operates on the vector as four separate 64-bit packed data elements (quad word (QW) size data elements), eight separate 32-bit packed data elements (double word (DW) size data elements), sixteen separate 16-bit packed data elements (word (W) size data elements), or thirty-two separate 8-bit data elements (byte (B) size data elements). However, in at least one embodiment, different vector widths and register sizes are possible.

[0305] In at least one embodiment, one or more execution units can be combined into a fused execution unit 1609A-1609N having thread control logic (1607A-1607N) that executes for the fused EU. In at least one embodiment, multiple EUs can be merged into a EU group. In at least one embodiment, each EU in the fused EU group can be configured to execute a separate SIMD hardware thread. The number of EUs in the fused EU group can vary according to various embodiments. In at least one embodiment, each EU can execute various SIMD widths, including but not limited to SIMD8, SIMD16, and SIMD32. In at least one embodiment, each fused graphics execution unit 1609A-1609N includes at least two execution units. In at least one embodiment, in at least one embodiment, the fused execution unit 1609A includes a first EU 1608A, a second EU 1608B, and thread control logic 1607A shared by the first EU 1608A and the second EU 1608B. In at least one embodiment, thread control logic 1607A controls threads executing on fused graphics execution unit 1609A, allowing each EU within fused execution units 1609A-1609N to execute using a common instruction pointer register.

[0306] In at least one embodiment, one or more internal instruction caches (e.g., 1606) are included in thread execution logic 1600D to cache thread instructions for the execution units. In at least one embodiment, one or more data caches (e.g., 1612) are included to cache thread data during thread execution. In at least one embodiment, a sampler 1610 is included to provide texture sampling for 3D operations and media sampling for media operations. In at least one embodiment, sampler 1610 includes specialized texture or media sampling functionality to process texture or media data during the sampling process before providing the sampled data to the execution units.

[0307] During execution, in at least one embodiment, the graphics and media pipeline sends thread initiation requests to thread execution logic 1600D via thread generation and dispatch logic. In at least one embodiment, once a set of geometric objects has been processed and rasterized into pixel data, pixel processor logic (such as pixel shader logic, fragment shader logic, etc.) within shader processor 1602 is invoked to further calculate output information and cause the results to be written to output surfaces (such as a color buffer, depth buffer, stencil buffer, etc.). In at least one embodiment, the pixel shader or fragment shader calculates the values of various vertex attributes to be interpolated across the rasterized objects. In at least one embodiment, the pixel processor logic within shader processor 1602 then executes the pixel or fragment shader program provided by an application program interface (API). In at least one embodiment, to execute the shader program, shader processor 1602 dispatches threads to execution units (such as 1608A) via thread dispatcher 1604. In at least one embodiment, shader processor 1602 uses texture sampling logic in sampler 1610 to access texture data stored in texture maps in memory. In at least one embodiment, arithmetic operations on texture data and input geometry data calculate pixel color data for each geometry fragment, or discard one or more pixels for further processing.

[0308] In at least one embodiment, data port 1614 provides a memory access mechanism for thread execution logic 1600D to output processed data to memory for further processing on the graphics processor output pipeline. In at least one embodiment, data port 1614 includes or is coupled to one or more cache memories (e.g., data cache 1612) to cache data for memory access via the data port.

[0309] like Figure 16EAs shown, in at least one embodiment, graphics execution unit 1608 may include an instruction fetch unit 1637, a general register file array (GRF) 1624, an architectural register file array (ARF) 1626, a thread arbiter 1622, an issue unit 1630, a branch unit 1632, a set of SIMD floating point units (FPUs) 1634, and, in at least one embodiment, a set of dedicated integer SIMD ALUs 1635. In at least one embodiment, GRF 1624 and ARF 1626 include a set of general register files and architectural register files associated with each simultaneous hardware thread that can be active in graphics execution unit 1608. In at least one embodiment, per-thread architectural state is maintained in ARF 1626, while data used during thread execution is stored in GRF 1624. In at least one embodiment, the execution state of each thread, including the instruction pointer of each thread, may be maintained in thread-specific registers in ARF 1626.

[0310] In at least one embodiment, graphics execution unit 1608 has an architecture that is a combination of simultaneous multithreading (SMT) and fine-grained interleaved multithreading (IMT). In at least one embodiment, the architecture has a modular configuration that can be fine-tuned at design time based on the target number of simultaneous threads and the number of registers per execution unit, where execution unit resources are logically allocated for executing multiple simultaneous threads.

[0311] In at least one embodiment, graphics execution unit 1608 can collectively issue multiple instructions, each of which can be a different instruction. In at least one embodiment, the thread arbiter 1622 of a graphics execution unit thread 1608 can dispatch instructions to one of the issue unit 1630, branch unit 1632, or SIMD FPU 1632 for execution. In at least one embodiment, each execution thread can access 128 general purpose registers in GRF 1624, each of which can store 32 bytes and can be accessed as a SIMD 8-element vector of 32-bit data elements. In at least one embodiment, each execution unit thread can access 4KB of GRF 1624, although embodiments are not limited thereto and more or fewer register resources may be provided in other embodiments. In at least one embodiment, a maximum of seven threads can execute simultaneously, although the number of threads per execution unit may vary depending on the embodiment. In at least one embodiment where seven threads have access to 4KB, GRF 1624 can store a total of 28KB. In at least one embodiment, flexible addressing modes may allow registers to be addressed together to efficiently build wider registers or rectangular block data structures representing strides.

[0312] In at least one embodiment, memory operations, sampler operations, and other longer latency system communications are scheduled via "send" instructions executed by message passing send unit 1630. In at least one embodiment, dispatching branch instructions to a dedicated branch unit 1632 facilitates SIMD divergence and eventual convergence.

[0313] In at least one embodiment, the graphics execution unit 1608 includes one or more SIMD floating point units (FPUs) 1634 to perform floating point operations. In at least one embodiment, one or more FPUs 1634 also support integer computations. In at least one embodiment, one or more FPUs 1634 can SIMD perform up to M 32-bit floating point (or integer) operations, or SIMD perform up to 2M 16-bit integer or 16-bit floating point operations. In at least one embodiment, at least one of the one or more FPUs provides extended math capabilities to support high-throughput transcendental math functions and double-precision 64-bit floating point. In at least one embodiment, a set of 8-bit integer SIMD ALUs 1635 are also present and can be specifically optimized to perform operations related to machine learning computations.

[0314] In at least one embodiment, an array of multiple instances of graphics execution unit 1608 may be instantiated in graphics sub-core groupings (such as sub-slices). In at least one embodiment, execution unit 1608 may execute instructions across multiple execution lanes. In at least one embodiment, each thread executing on graphics execution unit 1608 executes on a different lane.

[0315] The reasoning and / or training logic 615 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 6B and / or Figure 6C Provides details about the reasoning and / or training logic 615. In at least one embodiment, some or all of the reasoning and / or training logic 615 may be incorporated into the execution logic 1600D. Additionally, in at least one embodiment, other than Figure 6B and / or Figure 6C In at least one embodiment, the weight parameters may be stored in on-chip or off-chip memory and / or registers (shown or not shown) that configure the ALUs of execution logic 1600D to execute one or more machine learning algorithms, neural network architectures, use cases, or training techniques described herein.

[0316] Figure 17AA parallel processing unit ("PPU") 1700A is shown in accordance with at least one embodiment. In at least one embodiment, PPU 1700A is configured with machine-readable code that, if executed by PPU 1700A, causes PPU 1700A to perform some or all of the processes and techniques described throughout this disclosure. In at least one embodiment, PPU 1700A is a multi-threaded processor implemented on one or more integrated circuit devices and utilizes multithreading as a latency hiding technique designed to process computer-readable instructions (also referred to as machine-readable instructions or simply instructions) executed in parallel on multiple threads. In at least one embodiment, a thread refers to an execution thread and is an instance of a group of instructions configured to be executed by PPU 1700A. In at least one embodiment, PPU 1700A is a graphics processing unit ("GPU") configured to implement a graphics rendering pipeline for processing three-dimensional ("3D") graphics data to generate two-dimensional ("2D") image data for display on a display device, such as a liquid crystal display ("LCD") device. In at least one embodiment, PPU 1700A is used to perform computations such as linear algebra operations and machine learning operations. Figure 17A The example parallel processor is shown for illustrative purposes only and should be construed as a non-limiting example of a processor architecture contemplated within the scope of the present disclosure, and any suitable processor may be employed in addition and / or in place thereof.

[0317] In at least one embodiment, one or more PPUs 1700A are configured to accelerate high-performance computing ("HPC"), data center, and machine learning applications. In at least one embodiment, the PPU 1700A is configured to accelerate deep learning systems and applications, including the following non-limiting examples: autonomous vehicle platforms, deep learning, high-precision speech, image, and text recognition systems, intelligent video analysis, molecular simulation, drug discovery, disease diagnosis, weather forecasting, big data analysis, astronomy, molecular dynamics simulation, financial modeling, robotics, factory automation, real-time language translation, online search optimization, and personalized user recommendations.

[0318] In at least one embodiment, PPU 1700A includes, but is not limited to, input / output ("I / O") units 1706, front-end units 1710, scheduler units 1712, work distribution units 1714, hubs 1716, crossbars ("Xbars") 1720, one or more general processing clusters ("GPCs") 1718, and one or more partitioning units ("memory partitioning units") 1722. In at least one embodiment, PPU 1700A is connected to a host processor or other PPUs 1700A via one or more high-speed GPU interconnects ("GPU interconnects") 1708. In at least one embodiment, PPU 1700A is connected to a host processor or other peripheral devices via interconnect 1702. In one embodiment, PPU 1700A is connected to local memory including one or more memory devices ("memory") 1704. In at least one embodiment, memory devices 1704 include, but are not limited to, one or more dynamic random access memory ("DRAM") devices. In at least one embodiment, one or more DRAM devices are configured and / or configurable as a high bandwidth memory ("HBM") subsystem with multiple DRAM dies stacked within each device.

[0319] In at least one embodiment, the high-speed GPU interconnect 1708 may refer to a wire-based, multi-lane communication link that a system uses to scale and includes one or more PPUs 1700A in conjunction with one or more central processing units ("CPUs"), supporting cache coherence between the PPUs 1700A and the CPUs and CPU mastering. In at least one embodiment, the high-speed GPU interconnect 1708 transmits data and / or commands to other units of the PPU 1700A, such as one or more copy engines, video encoders, video decoders, power management units, and / or other processors, via a hub 1716. Figure 17A Other components that may not be explicitly shown.

[0320] In at least one embodiment, I / O unit 1706 is configured to receive data from a host processor ( Figure 17A1700A). In at least one embodiment, I / O unit 1706 communicates with the host processor directly through interconnect 1702 or through one or more intermediate devices (e.g., a memory bridge). In at least one embodiment, I / O unit 1706 can communicate with one or more other processors (e.g., one or more PPUs 1700A) via interconnect 1702. In at least one embodiment, I / O unit 1706 implements a Peripheral Component Interconnect Express ("PCIe") interface for communicating over a PCIe bus. In at least one embodiment, I / O unit 1706 implements an interface for communicating with external devices.

[0321] In at least one embodiment, I / O unit 1706 decodes packets received via interconnect 1702. In at least one embodiment, at least some of the packets represent commands configured to cause PPU 1700A to perform various operations. In at least one embodiment, I / O unit 1706 sends the decoded commands to various other units of PPU 1700A as specified by the commands. In at least one embodiment, the commands are sent to front end unit 1710 and / or to hub 1716 or other units of PPU 1700A, such as one or more replication engines, video encoders, video decoders, power management units, etc. Figure 17A In at least one embodiment, I / O unit 1706 is configured to route communications between the various logical units of PPU 1700A.

[0322] In at least one embodiment, a program executed by a host processor encodes a command stream in a buffer that provides a workload to PPU 1700A for processing. In at least one embodiment, the workload includes instructions and data to be processed by those instructions. In at least one embodiment, the buffer is an area in memory that is accessible (e.g., read / write) by both the host processor and PPU 1700A—the host interface unit can be configured to access the buffer in system memory connected to interconnect 1702 via memory requests transmitted via I / O unit 1706 over interconnect 1702. In at least one embodiment, the host processor writes a command stream into the buffer and then sends a pointer indicating the beginning of the command stream to PPU 1700A, so that front end unit 1710 receives pointers to one or more command streams and manages the one or more command streams, reading commands from the command streams and forwarding the commands to various units of PPU 1700A.

[0323] In at least one embodiment, the front end unit 1710 is coupled to a scheduler unit 1712 that configures the various GPCs 1718 to process tasks defined by one or more command streams. In at least one embodiment, the scheduler unit 1712 is configured to track state information related to the various tasks managed by the scheduler unit 1712, where the state information may indicate which GPC 1718 the task is assigned to, whether the task is active or inactive, a priority associated with the task, and the like. In at least one embodiment, the scheduler unit 1712 manages multiple tasks that execute on one or more GPCs 1718.

[0324] In at least one embodiment, scheduler unit 1712 is coupled to work distribution unit 1714, which is configured to dispatch tasks for execution on GPCs 1718. In at least one embodiment, work distribution unit 1714 tracks a plurality of scheduled tasks received from scheduler unit 1712 and manages a pending task pool and an active task pool for each GPC 1718. In at least one embodiment, the pending task pool includes a plurality of time slots (such as 16 time slots) containing tasks assigned to be processed by a particular GPC 1718; the active task pool may include a plurality of time slots (such as 4 time slots) for tasks actively being processed by GPC 1718, such that as one of GPCs 1718 completes execution of a task, the task is evicted from the active task pool of GPC 1718 and one of the other tasks is selected from the pending task pool and scheduled for execution on GPC 1718. In at least one embodiment, if an active task is idle on GPC 1718, such as while waiting for data dependencies to be resolved, the active task is evicted from GPC 1718 and returned to the pending task pool, while another task in the pending task pool is selected and scheduled for execution on GPC 1718.

[0325] In at least one embodiment, work distribution unit 1714 communicates with one or more GPCs 1718 via XBar 1720. In at least one embodiment, XBar 1720 is an interconnect network that couples many units of PPU 1700A to other units of PPU 1700A and can be configured to couple work distribution unit 1714 to a specific GPC 1718. In at least one embodiment, one or more other units of PPU 1700A can also be connected to XBar 1716 via hub 1716.

[0326] In at least one embodiment, tasks are managed by a scheduler unit 1712 and assigned to one of the GPCs 1718 by a work distribution unit 1714. The GPCs 1718 are configured to process tasks and produce results. In at least one embodiment, the results can be consumed by other tasks in the GPC 1718, routed to a different GPC 1718 via the XBar 1716, or stored in memory 1704. In at least one embodiment, the results can be written to the memory 1704 via a partition unit 1722, which implements a memory interface for writing data to or reading data from the memory 1704. In at least one embodiment, the results can be transferred to another PPU 1704 or a CPU via a high-speed GPU interconnect 1708. In at least one embodiment, the PPU 1700A includes, but is not limited to, U partition units 1722, where the partition units 1722 are equal to the number of separate and distinct memory devices 1704 coupled to the PPU 1700A. In at least one embodiment, the following will be combined with Figure 17C The partition unit 1722 is described in more detail.

[0327] In at least one embodiment, the host processor executes a driver core that implements an application programming interface (API) that enables one or more applications executing on the host processor to schedule operations for execution on the PPU 1700A. In one embodiment, multiple computing applications are executed simultaneously by the PPU 1700A, and the PPU 1700A provides isolation, quality of service ("QoS"), and independent address spaces for the multiple computing applications. In at least one embodiment, the application generates instructions (such as, in the form of API calls) that cause the driver core to generate one or more tasks for execution by the PPU 1700A, and the driver core outputs the tasks to one or more streams processed by the PPU 1700A. In at least one embodiment, each task includes one or more related groups of threads, which may be referred to as warps. In at least one embodiment, a warp includes multiple related threads (such as 32 threads) that can be executed in parallel. In at least one embodiment, a cooperative thread may refer to multiple threads that include instructions for performing tasks and exchanging data through shared memory, combined with Figure 17C Threads and cooperating threads are described in greater detail according to at least one embodiment.

[0328] The reasoning and / or training logic 615 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 6B and / or Figure 6CProvides details regarding inference and / or training logic 615. In at least one embodiment, the deep learning application processor is used to train a machine learning model (such as a neural network) to predict or infer information provided to the PPU 1700A. In at least one embodiment, the PPU 1700A is used to infer or predict information based on a trained machine learning model (such as a neural network) that has been trained by another processor or system or the PPU 1700A. In at least one embodiment, the PPU 1700A can be used to perform one or more of the neural network use cases described herein.

[0329] Figure 17B A general processing cluster ("GPC") 1700B is shown in accordance with at least one embodiment. In at least one embodiment, GPC 1700B is Figure 17A 1718. In at least one embodiment, each GPC 1700B includes, but is not limited to, multiple hardware units for processing tasks, and each GPC 1700B includes, but is not limited to, a pipeline manager 1702, a pre-raster operations unit ("PROP") 1704, a raster engine 1708, a work distribution crossbar ("WDX") 1716, a memory management unit ("MMU") 1718, one or more data processing clusters ("DPCs") 1706, and any suitable combination of components.

[0330] In at least one embodiment, the operation of GPC 1700B is controlled by pipeline manager 1702. In at least one embodiment, pipeline manager 1702 manages the configuration of one or more DPCs 1706 to process tasks assigned to GPC 1700B. In at least one embodiment, pipeline manager 1702 configures at least one of one or more DPCs 1706 to implement at least a portion of a graphics rendering pipeline. In at least one embodiment, DPC 1706 is configured to execute vertex shader programs on a programmable streaming multiprocessor ("SM") 1714. In at least one embodiment, pipeline manager 1702 is configured to route packets received from a work distribution unit to appropriate logic within GPC 1700B, and in at least one embodiment, some packets may be routed to fixed-function hardware units in PROP 1704 and / or raster engine 1708, while other packets may be routed to DPC 1706 for processing by primitive engine 1712 or SM 1714. In at least one embodiment, pipeline manager 1702 configures at least one of DPCs 1706 to implement a neural network model and / or a computational pipeline.

[0331] In at least one embodiment, PROP unit 1704 is configured to route data generated by raster engine 1708 and DPC 1706 to a raster operations ("ROP") unit in partition unit 1722 in at least one embodiment, in conjunction with Figure 17A Described in more detail. In at least one embodiment, the PROP unit 1704 is configured to perform optimizations for color blending, organize pixel data, perform address translation, and the like. In at least one embodiment, the raster engine 1708 includes, but is not limited to, a plurality of fixed-function hardware units configured to perform various raster operations, and in at least one embodiment, the raster engine 1708 includes, but is not limited to, a setup engine, a coarse raster engine, a culling engine, a clipping engine, a fine raster engine, a tile aggregation engine, and any suitable combination thereof. In at least one embodiment, the setup engine receives transformed vertices and generates plane equations associated with the geometric primitives defined by the vertices; the plane equations are passed to the coarse raster engine to generate coverage information for the primitives (such as the x, y coverage mask for the tile); the output of the coarse raster engine is passed to the culling engine, where fragments associated with primitives that fail the z test are culled, and to the clipping engine, where fragments outside the viewing frustum are clipped. In at least one embodiment, the clipped and culled fragments are passed to the fine raster engine to generate properties for the pixel fragments based on the plane equations generated by the setup engine. In at least one embodiment, the output of raster engine 1708 includes fragments to be processed by any appropriate entity (e.g., by a fragment shader implemented within DPC 1706).

[0332] In at least one embodiment, each DPC 1706 included in a GPC 1700B includes, but is not limited to, an M-pipeline controller ("MPC") 1710; a primitive engine 1712; one or more SMs 1714; and any suitable combination thereof. In at least one embodiment, the MPC 1710 controls the operation of the DPC 1706, routing packets received from the pipeline manager 1702 to appropriate units within the DPC 1706. In at least one embodiment, packets associated with vertices are routed to the primitive engine 1712, which is configured to fetch vertex attributes associated with the vertices from memory; conversely, packets associated with shader programs may be sent to the SM 1714.

[0333] In at least one embodiment, SM 1714 includes, but is not limited to, a programmable streaming processor configured to process tasks represented by multiple threads. In at least one embodiment, SM 1714 is multithreaded and configured to simultaneously execute multiple threads (such as 32 threads) from a particular thread group and implements a single instruction, multiple data ("SIMD") architecture, in which each thread in a group of threads (such as a warp) is configured to process a different set of data based on the same instruction set. In at least one embodiment, all threads in a thread group execute the same instructions. In at least one embodiment, SM 1714 implements a single instruction, multiple thread ("SIMT") architecture, in which each thread in a group of threads is configured to process a different set of data based on the same instruction set, but in which individual threads in a thread group are allowed to diverge during execution. In at least one embodiment, a program counter, call stack, and execution state are maintained for each warp, thereby enabling concurrency between warps and serial execution within a warp when threads in the warp diverge. In another embodiment, a program counter, call stack, and execution state are maintained for each individual thread, thereby enabling equal concurrency between all threads within a warp and between warps. In at least one embodiment, execution state is maintained for each individual thread, and threads of the same instruction can be converged and executed in parallel to improve efficiency. At least one embodiment of SM 1714 is described in more detail below.

[0334] In at least one embodiment, the MMU 1718 is used between the GPC 1700B and the memory partition unit (such as Figure 17A The MMU 1718 provides an interface between the memory and the partition unit 1722, and provides virtual to physical address translation, memory protection, and arbitration of memory requests. In at least one embodiment, the MMU 1718 provides one or more translation lookaside buffers ("TLBs") for performing translation of virtual addresses to physical addresses in memory.

[0335] The reasoning and / or training logic 615 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 6B and / or Figure 6C Provides details regarding inference and / or training logic 615. In at least one embodiment, the deep learning application processor is used to train a machine learning model (such as a neural network) to predict or infer information provided to the GPC 1700B. In at least one embodiment, the GPC 1700B is used to infer or predict information based on a machine learning model (such as a neural network) that has been trained by another processor or system or the GPC 1700B. In at least one embodiment, the GPC 1700B can be used to perform one or more of the neural network use cases described herein.

[0336] Figure 17C A memory partition unit 1700C of a parallel processing unit ("PPU") is shown in accordance with at least one embodiment. In at least one embodiment, the memory partition unit 1700C includes, but is not limited to, a raster operations ("ROP") unit 1702; a level 2 ("L2") cache 1704; a memory interface 1706; and any suitable combination thereof. In at least one embodiment, the memory interface 1706 is coupled to a memory. In at least one embodiment, the memory interface 1706 can implement a 32-, 64-, 128-, 1024-bit data bus, or a similar implementation for high-speed data transfer. In at least one embodiment, the PPU includes U memory interfaces 1706, one memory interface 1706 for each pair of partition units 1700C, wherein each pair of partition units 1700C is connected to a corresponding memory device. In at least one embodiment, the PPU can be connected to up to Y memory devices, such as a high-bandwidth memory stack or graphics double data rate version 5 synchronous dynamic random access memory ("GDDR5 SDRAM").

[0337] In at least one embodiment, the memory interface 1706 implements a High Bandwidth Memory 2nd Generation ("HBM2") memory interface, and Y is equal to half of U. In at least one embodiment, the HBM2 memory stack is located on the same physical package as the PPU, providing significant power and area savings compared to GDDR5 SDRAM systems. In at least one embodiment, each HBM2 stack includes, but is not limited to, four memory dies, and Y=4, each HBM2 stack includes two 128-bit channels per die for a total of 8 channels and a data bus width of 1024 bits. In at least one embodiment, the memory supports single error correction double error detection ("SECDED") error correction code ("ECC") to protect data. In at least one embodiment, ECC provides higher reliability for computing applications that are sensitive to data corruption.

[0338] In at least one embodiment, the PPU implements a multi-level memory hierarchy. In at least one embodiment, the memory partition unit 1700C supports unified memory to provide a single unified virtual address space for the central processing unit ("CPU") and PPU memory, thereby enabling data sharing between virtual memory systems. In at least one embodiment, the frequency of PPU accesses to memory located on other processors is tracked to ensure that memory pages are moved to the physical memory of the PPU that accesses the pages more frequently. In at least one embodiment, the high-speed GPU interconnect 1708 supports address translation services that allow the PPU to directly access the CPU's page tables and provide full access to the CPU's memory through the PPU.

[0339] In at least one embodiment, the copy engine transfers data between multiple PPUs or between a PPU and a CPU. In at least one embodiment, the copy engine can generate a page fault for an address that is not mapped in a page table, and the memory partition unit 1700C then services the page fault, maps the address into a page table, and then the copy engine performs the transfer. In at least one embodiment, fixed (in at least one embodiment, non-pageable) memory is used for multiple copy engine operations between multiple processors, thereby substantially reducing the available memory. In at least one embodiment, in the event of a hardware page fault, the address can be passed to the copy engine without regard to whether it resides in the memory page, and the copy process is transparent.

[0340] According to at least one embodiment, Figure 17A Data from memory 1704 or other system memory is retrieved by memory partition unit 1700C and stored in L2 cache 1704, which is located on-chip and shared between the various GPCs. In at least one embodiment, each memory partition unit 1700C includes, but is not limited to, at least a portion of an L2 cache associated with a corresponding memory device. In at least one embodiment, lower-level caches are implemented in various units within a GPC. In at least one embodiment, each SM 1714 may implement a level 1 ("L1") cache, where the L1 cache is private memory dedicated to a particular SM 1714, and data is retrieved from L2 cache 1704 and stored in each L1 cache for processing within the functional units of SM 1714. In at least one embodiment, L2 cache 1704 is coupled to memory interface 1706 and XBar 1720.

[0341] In at least one embodiment, ROP unit 1702 performs graphics raster operations related to pixel color, such as color compression, pixel blending, and the like. In at least one embodiment, ROP unit 1702 performs depth testing in conjunction with raster engine 1708, receiving the depth of a sample location associated with a pixel fragment from the culling engine of raster engine 1708. In at least one embodiment, the depth is tested against the corresponding depth in the depth buffer associated with the sample location of the fragment. In at least one embodiment, if the fragment passes the depth test for the sample location, ROP unit 1702 updates the depth buffer and sends the result of the depth test to raster engine 1708. It will be appreciated that the number of partition units 1700C can differ from the number of GPCs, and therefore, in at least one embodiment, each ROP unit 1702 can be coupled to each GPC. In at least one embodiment, ROP unit 1702 tracks packets received from different GPCs and determines to which result generated by ROP unit 1702 via XBar 1720 to be routed.

[0342] Figure 17D Streaming multiprocessor ("SM") 1700D is shown in accordance with at least one embodiment. In at least one embodiment, SM 1700D is Figure 17BSM. In at least one embodiment, SM 1700D includes, but is not limited to, an instruction cache 1702; one or more scheduler units 1704; a register file 1708; one or more processing cores ("cores") 1710; one or more special function units ("SFUs") 1712; one or more load / store units ("LSUs") 1714; an interconnect network 1716; a shared memory / level 1 ("L1") cache 1718; and any suitable combination thereof. In at least one embodiment, a work distribution unit schedules tasks for execution on a general processing cluster ("GPC") of a parallel processing unit ("PPU"), with each task being assigned to a specific data processing cluster ("DPC") within the GPC, and if the task is associated with a shader program, the task is assigned to one of SMs 1700D. In at least one embodiment, scheduler unit 1704 receives tasks from the work distribution unit and manages instruction scheduling for one or more thread blocks assigned to SMs 1700D. In at least one embodiment, the scheduler unit 1704 schedules thread blocks for execution as warps of parallel threads, where each thread block is assigned at least one warp. In at least one embodiment, each warp executes a thread. In at least one embodiment, the scheduler unit 1704 manages a plurality of different thread blocks, assigns warps to different thread blocks, and then dispatches instructions from a plurality of different cooperating groups to various functional units (such as the processing core 1710, SFU 1712, and LSU 1714) during each clock cycle.

[0343] In at least one embodiment, cooperative groups can refer to a programming model for organizing groups of communicating threads that allows developers to express the granularity at which threads are communicating, thereby enabling the expression of richer, more efficient decompositions of parallelism. In at least one embodiment, a cooperative launch API supports synchronization between thread blocks to execute parallel algorithms. In at least one embodiment, applications of the programming model provide a single, simple construct for synchronizing cooperating threads: a barrier across all threads of a thread block (such as the syncthreads() function). However, in at least one embodiment, programmers can define thread groups at a granularity smaller than a thread block and synchronize within the defined group to achieve higher performance, design flexibility, and software reuse in the form of a collective group-wide function interface. In at least one embodiment, cooperative groups enable programmers to explicitly define thread groups at sub-block (in at least one embodiment, as small as a single thread) and multi-block granularity and perform collective operations, such as synchronizing threads within a cooperative group. In at least one embodiment, the programming model supports clean composition across software boundaries, so that libraries and utility functions can safely synchronize in their local environment without making assumptions about convergence. In at least one embodiment, the cooperation group primitive enables new patterns of cooperative parallelism, including but not limited to producer-consumer parallelism, opportunistic parallelism, and global synchronization across an entire grid of thread blocks.

[0344] In at least one embodiment, the scheduling unit 1706 is configured to send instructions to one or more of the functional units, and the scheduler unit 1704 includes, but is not limited to, two scheduling units 1706 that enable two different instructions from the same warp to be scheduled per clock cycle. In at least one embodiment, each scheduler unit 1704 includes a single scheduling unit 1706 or additional scheduling units 1706.

[0345] In at least one embodiment, each SM 1700D includes, but is not limited to, a register file 1708 that provides a set of registers for the functional units of SM 1700D. In at least one embodiment, register file 1708 is partitioned between each functional unit, allocating a dedicated portion of register file 1708 to each functional unit. In at least one embodiment, register file 1708 is partitioned between the different warps executed by SM 1700D, and register file 1708 provides temporary storage for operands connected to the data paths of the functional units. In at least one embodiment, each SM 1700D includes, but is not limited to, a plurality of L processing cores 1710. In at least one embodiment, SM 1700D includes, but is not limited to, a large number (such as 128 or more) of different processing cores 1710. In at least one embodiment, each processing core 1710 includes, but is not limited to, a fully pipelined, single-precision, double-precision, and / or mixed-precision processing unit, including, but not limited to, a floating-point arithmetic logic unit and an integer arithmetic logic unit. In at least one embodiment, the floating point arithmetic logic unit implements the IEEE 754-2008 standard for floating point arithmetic. In at least one embodiment, the processing core 1710 includes, but is not limited to, 64 single precision (32-bit) floating point cores, 64 integer cores, 32 double precision (64-bit) floating point cores, and 8 tensor cores.

[0346] According to at least one embodiment, the tensor core is configured to perform matrix operations. In at least one embodiment, one or more tensor cores are included in processing core 1710. In at least one embodiment, the tensor core is configured to perform deep learning matrix arithmetic, such as convolution operations for neural network training and inference. In at least one embodiment, each tensor core operates on a 4×4 matrix and performs a matrix multiplication and accumulation operation D=A×B+C, where A, B, C, and D are 4×4 matrices.

[0347] In at least one embodiment, the matrix multiplication inputs A and B are 16-bit floating-point matrices, and the accumulation matrices C and D are 16-bit floating-point or 32-bit floating-point matrices. In at least one embodiment, the tensor core performs a 32-bit floating-point accumulation operation on the 16-bit floating-point input data. In at least one embodiment, the 16-bit floating-point multiplication uses 64 operations and results in a full-precision product, which is then accumulated with other intermediate products using 32-bit floating-point addition to perform a 4x4x4 matrix multiplication. In at least one embodiment, the tensor core is used to perform larger two-dimensional or higher-dimensional matrix operations composed of these smaller elements. In at least one embodiment, an API (such as the CUDA 9 C++ API) exposes specialized matrix load, matrix multiplication and accumulation, and matrix store operations to efficiently use the tensor cores from a CUDA-C++ program. In at least one embodiment, at the CUDA level, the warp-level interface assumes a 16×16 matrix size that spans all 32 warp threads.

[0348] In at least one embodiment, each SM 1700D includes, but is not limited to, M SFUs 1712 that perform specialized functions (such as attribute evaluation, reciprocal square root, etc.). In at least one embodiment, SFUs 1712 include, but are not limited to, tree traversal units configured to traverse a hierarchical tree data structure. In at least one embodiment, SFUs 1712 include, but are not limited to, texture units configured to perform texture map filtering operations. In at least one embodiment, the texture units are configured to load texture maps (such as a 2D array of texels) from memory and sample the texture maps to generate sampled texture values for use by shader programs executed by the SM 1700D. In at least one embodiment, the texture maps are stored in shared memory / L1 cache 1718. In at least one embodiment, the texture units implement texture operations (such as filtering operations) using mip-maps (such as texture maps with different levels of detail), according to at least one embodiment. In at least one embodiment, each SM 1700D includes, but is not limited to, two texture units.

[0349] In at least one embodiment, each SM 1700D includes, but is not limited to, N LSUs 1714 that implement load and store operations between the shared memory / L1 cache 1718 and the register file 1708. In at least one embodiment, an interconnection network 1716 connects each functional unit to the register file 1708, and the LSUs 1714 connect to both the register file 1708 and the shared memory / L1 cache 1718. In at least one embodiment, the interconnection network 1716 is a crossbar switch that can be configured to connect any functional unit to any register in the register file 1708, and to connect the LSUs 1714 to memory locations in the register file 1708 and the shared memory / L1 cache 1718.

[0350] In at least one embodiment, shared memory / L1 cache 1718 is an array of on-chip memory that, in at least one embodiment, allows for data storage and communication between SM 1700D and primitive engines, as well as between threads within SM 1700D. In at least one embodiment, shared memory / L1 cache 1718 includes, but is not limited to, 128KB of storage capacity and is located in the path from SM 1700D to the partition unit. In at least one embodiment, shared memory / L1 cache 1718 is used to cache reads and writes in at least one embodiment. In at least one embodiment, one or more of shared memory / L1 cache 1718, L2 cache, and memory is a backing store.

[0351] In at least one embodiment, combining data cache and shared memory functionality into a single memory block provides improved performance for both types of memory accesses. In at least one embodiment, capacity is used by programs that do not use the shared memory or use it as a cache, for example, if the shared memory is configured to use half of its capacity, and texture and load / store operations can use the remaining capacity. According to at least one embodiment, integration within shared memory / L1 cache 1718 enables shared memory / L1 cache 1718 to function as a high-throughput pipeline for streaming data, while providing high-bandwidth and low-latency access to frequently reused data. In at least one embodiment, when configured for general-purpose parallel computing, a simpler configuration can be used compared to graphics processing. In at least one embodiment, the fixed-function graphics processing unit is bypassed, creating a simpler programming model. In at least one embodiment, in a general-purpose parallel computing configuration, the work distribution unit directly allocates and distributes blocks of threads to DPCs. In at least one embodiment, the threads in a block execute a common program, use unique thread IDs in computations to ensure that each thread generates unique results, use SM 1700D to execute the program and perform computations, use shared memory / L1 cache 1718 to communicate between threads, and use LSU 1714 to read and write global memory through shared memory / L1 cache 1718 and a memory partitioning unit. In at least one embodiment, when configured for general parallel computation, SM 1700D writes commands to scheduler unit 1704 that can be used to start new work on a DPC.

[0352] In at least one embodiment, the PPU is included in or coupled to a desktop computer, a laptop computer, a tablet computer, a server, a supercomputer, a smartphone (e.g., a wireless, handheld device), a personal digital assistant ("PDA"), a digital camera, a vehicle, a head-mounted display, a handheld electronic device, etc. In at least one embodiment, the PPU is implemented on a single semiconductor substrate. In at least one embodiment, the PPU is included in a system-on-chip ("SoC") along with one or more other devices (e.g., an additional PPU, memory, a reduced instruction set computer ("RISC") CPU, one or more memory management units ("MMUs"), a digital-to-analog converter ("DAC"), etc.

[0353] In at least one embodiment, the PPU can be included on a graphics card that includes one or more storage devices. The graphics card can be configured to connect to a PCIe slot on a desktop computer motherboard. In at least one embodiment, the PPU can be an integrated graphics processing unit ("iGPU") included in a chipset on the motherboard.

[0354] Reasoning and / or training logic 615 is used to perform reasoning and / or training operations related to one or more embodiments. Figure 6B and / or Figure 6C Provides details regarding inference and / or training logic 615. In at least one embodiment, the deep learning application processor is used to train a machine learning model (such as a neural network) to predict or infer information provided to the SM 1700D. In at least one embodiment, the SM 1700D is used to infer or predict information based on a machine learning model (such as a neural network) that has been trained by another processor or system or by the SM 1700D. In at least one embodiment, the SM 1700D can be used to perform one or more of the neural network use cases described herein.

[0355] In at least one embodiment, a single semiconductor platform can refer to a single semiconductor-based integrated circuit or chip. In at least one embodiment, a multi-chip module with increased connectivity can be used, which emulates on-chip operations and provides substantial improvements over implementations utilizing a central processing unit ("CPU") and a bus. In at least one embodiment, the various modules can also be placed separately or in various combinations of semiconductor platforms, depending on user needs.

[0356] In at least one embodiment, a computer program in the form of machine-readable executable code or computer control logic algorithms is stored in main memory 804 and / or secondary storage. According to at least one embodiment, if executed by one or more processors, the computer program enables system 800 to perform various functions. In at least one embodiment, memory 804, storage, and / or any other storage are possible examples of computer-readable media. In at least one embodiment, secondary storage can refer to any suitable storage device or system, such as a hard drive and / or removable storage drive, which represents a floppy disk drive, a tape drive, an optical drive, a digital versatile disk ("DVD") drive, a recording device, a universal serial bus ("USB") flash memory, etc. In at least one embodiment, the architecture and / or functionality of each of the previous figures is implemented in the context of a CPU 802; a parallel processing system 812; an integrated circuit capable of having at least some of the capabilities of two CPUs 802; a parallel processing system 812; a chipset (e.g., a group of integrated circuits designed to operate and sell as a unit to perform related functions, etc.); and any suitable combination of integrated circuits.

[0357] In at least one embodiment, the architecture and / or functionality of the various previous figures are implemented in the context of a general-purpose computer system, a circuit board system, a game console system dedicated for entertainment purposes, a dedicated system, etc. In at least one embodiment, the computer system 800 can take the form of a desktop computer, a laptop computer, a tablet computer, a server, a supercomputer, a smartphone (such as a wireless, handheld device), a personal digital assistant ("PDA"), a digital camera, a vehicle, a head-mounted display, a handheld electronic device, a mobile telephone device, a television, a workstation, a game console, an embedded system, and / or any other type of logic.

[0358] In at least one embodiment, the parallel processing system 812 includes, but is not limited to, a plurality of parallel processing units ("PPUs") 814 and associated memory 816. In at least one embodiment, the PPUs 814 are connected to a host processor or other peripheral devices via an interconnect 818 and a switch 820 or multiplexer. In at least one embodiment, the parallel processing system 812 distributes computational tasks across parallelizable PPUs 814, for example, as part of a distribution of computational tasks across multiple graphics processing unit ("GPU") thread blocks. In at least one embodiment, memory is shared and accessed (e.g., for read and / or write access) between some or all of the PPUs 814, although such shared memory may incur a performance penalty relative to using local memory and registers resident on the PPUs 814. In at least one embodiment, the operation of the PPUs 814 is synchronized using a command (such as __syncthreads()), where all threads in a block (e.g., executing across multiple PPUs 814) reach a certain code execution point before proceeding.

[0359] Other variations are within the spirit of the present disclosure. Thus, while the disclosed technology is susceptible to various modifications and alternative constructions, certain illustrated embodiments thereof...

Claims

1. A cooling system for data center equipment, comprising: a plurality of heat sinks between a first plate and a second plate for dissipating a first amount of heat to an environment in a first configuration of the plurality of heat sinks, the first plate being movable relative to the second plate to expose a surface area of the plurality of heat sinks to the environment in a second configuration of the plurality of heat sinks and dissipate a second amount of heat greater than the first amount of heat; wherein, in the first configuration, each heat sink is bent to include an overlapping portion, the overlapping portion being at least partially isolated from the environment; and In the second configuration, the individual fins are unfolded and the overlapping portions are separated to expose the overlapping portions to the environment.

2. The cooling system according to claim 1, further comprising: The surface area exposed in the second configuration is greater than the surface area of the plurality of fins exposed in the first configuration.

3. The cooling system according to claim 1, further comprising: Intermediate configurations of the plurality of fins between the first configuration and the second configuration are configured to provide an intermediate surface area to enable dissipation of a third amount of heat associated with each intermediate configuration of the plurality of fins.

4. The cooling system according to claim 1, further comprising: A gear subsystem, an electromagnetic subsystem, a thermoelectric generator subsystem, a thermal reaction subsystem, or a pneumatic subsystem for moving the first plate relative to the second plate in response to sensed heat from an associated computing component or from the environment.

5. The cooling system according to claim 1, further comprising: At least one strip portion associated with the plurality of fins is configured to enable each fin to include the overlapping portion.

6. The cooling system according to claim 1, further comprising: At least one strip portion associated with the plurality of fins and formed in part of a bimorph material is configured to enable movement of the first plate relative to the second plate upon sensing heat acting on the bimorph material.

7. The cooling system according to claim 1, further comprising: a fluid or gas line for receiving cooling fluid from a cooling circuit of a data center housing said data center equipment; A pneumatic subsystem is configured to use a cooling fluid to extend a piston and move the first plate relative to the second plate to expose a surface area of the plurality of fins in the second configuration of the plurality of fins.

8. The cooling system according to claim 7, further comprising: A learning subsystem comprising at least one processor configured to: evaluate temperatures sensed from components, servers, or racks in the data center, wherein surface areas of the plurality of heat sinks are associated with different amounts of heat dissipation; and providing an output correlated to at least one temperature to facilitate movement of the first plate relative to the second plate to expose surface areas of the plurality of fins.

9. The cooling system according to claim 8, further comprising: at least one controller that facilitates movement of the first plate relative to the second plate; and the learning subsystem executing a machine learning model to: Processing the temperature using a multi-level neuron of the machine learning model, wherein the multi-level neuron is loaded with the temperature and the surface areas of the plurality of heat sinks associated with the temperature; as well as The output associated with at least one temperature is provided to the at least one controller, the output being provided after evaluating the surface areas of the plurality of heat sinks associated with previously collected temperatures and the previously collected temperatures.

10. At least one processor for a cooling system, comprising: at least one logic unit for controlling movement associated with a plurality of heat sinks between a first plate and a second plate, the plurality of heat sinks dissipating a first amount of heat to an environment in a first configuration of the plurality of heat sinks, the first plate being movable relative to the second plate to expose surface areas of the plurality of heat sinks to the environment in a second configuration of the plurality of heat sinks; wherein, in the first configuration, each heat sink is bent to include an overlapping portion, the overlapping portion being at least partially isolated from the environment; and In the second configuration, the individual fins are unfolded and the overlapping portions are separated to expose the overlapping portions to the environment.

11. The at least one processor of claim 10, further comprising: A learning subsystem for: evaluating temperatures sensed from a component, server, or rack in a data center, wherein surface areas of the plurality of heat sinks are associated with different amounts of heat dissipation; and providing an output associated with at least one temperature to facilitate movement of the first plate relative to the second plate such that surface areas of the plurality of heat sinks exposed to the environment are increased.

12. The at least one processor of claim 11, further comprising: The learning subsystem executes machine learning models to: Processing the temperature using a multi-level neuron of the machine learning model, wherein the multi-level neuron is loaded with the temperature and the surface areas of the plurality of heat sinks associated with the temperature; as well as The output associated with at least one temperature is provided to at least one controller, the output provided after evaluating the surface areas of the plurality of heat sinks associated with previously collected temperatures and the previously collected temperatures.

13. The at least one processor of claim 10, further comprising: A command output is configured to communicate an output associated with a controller to facilitate movement of the first plate relative to the second plate.

14. The at least one processor of claim 10, further comprising: The at least one logic unit is adapted to receive a temperature value from a temperature sensor associated with data center equipment, the temperature sensor being adapted to facilitate movement of the first plate movable relative to the second plate.

15. At least one processor for a cooling system, comprising: at least one logic unit for training one or more neural networks having layers of hidden neurons to: evaluate temperatures sensed from components, servers, or racks in a data center, wherein surface areas of a plurality of heat sinks are associated with different amounts of heat dissipation from the plurality of heat sinks; and provide an output associated with at least one temperature to facilitate movement of a first plate relative to a second plate to expose the surface areas of the plurality of heat sinks, the first plate and the second plate having the plurality of heat sinks therebetween; wherein, before the first plate moves relative to the second plate, each heat sink is bent to include an overlapping portion, the overlapping portion being at least partially isolated from the environment; and After the first plate is moved relative to the second plate, the respective fins are unfolded and the overlapping portions are separated to expose the overlapping portions to the environment.

16. The at least one processor of claim 15, further comprising: The at least one logic unit is configured to output at least one instruction associated with at least one temperature to facilitate movement of the first plate relative to the second plate to expose the surface area of the plurality of heat sinks, with the plurality of heat sinks between the first plate and the second plate.

17. The at least one processor of claim 15, further comprising: At least one command output for communicating said output associated with a controller to facilitate movement of said first plate relative to said second plate.

18. The at least one processor of claim 15, further comprising: The at least one logic unit is adapted to receive a temperature value from a temperature sensor associated with data center equipment, the temperature sensor being adapted to facilitate movement of the first plate movable relative to the second plate.

19. A data center cooling system comprising: at least one processor configured to: train one or more neural networks having layers of hidden neurons to: evaluate temperatures sensed from components, servers, or racks in a data center, wherein surface areas of a plurality of heat sinks are associated with different amounts of heat dissipation from the plurality of heat sinks; and provide an output associated with at least one temperature to facilitate movement of a first plate relative to a second plate to expose the surface areas of the plurality of heat sinks, the first plate and the second plate having the plurality of heat sinks therebetween; wherein, before the first plate moves relative to the second plate, each heat sink is bent to include an overlapping portion, the overlapping portion being at least partially isolated from the environment; and After the first plate is moved relative to the second plate, the respective fins are unfolded and the overlapping portions are separated to expose the overlapping portions to the environment.

20. The data center cooling system of claim 19, further comprising: The at least one processor is configured to output at least one instruction associated with at least one temperature to facilitate movement of the first plate relative to the second plate to expose the surface area of the plurality of heat sinks, with the plurality of heat sinks between the first plate and the second plate.

21. The data center cooling system of claim 19, further comprising: At least one instruction output of the at least one processor is configured to communicate the output associated with a controller to facilitate movement of the first plate relative to the second plate.

22. The data center cooling system of claim 19, further comprising: The at least one processor is adapted to receive a temperature value from a temperature sensor associated with data center equipment, the temperature sensor being adapted to facilitate movement of the first plate movable relative to the second plate.

23. A heat sink comprising a plurality of heat sinks between a first plate and a second plate, wherein: the first plate being movable relative to the second plate to dissipate heat to an environment by exposing a surface area of the plurality of fins in a second configuration, the surface area of the plurality of fins exposed in the second configuration being greater than the surface area of the plurality of fins exposed in the first configuration; In the first configuration, each heat sink is bent to include an overlapping portion, the overlapping portion being at least partially isolated from the environment; as well as In the second configuration, the individual fins are unfolded and the overlapping portions are separated to expose the overlapping portions to the environment.

24. The heat sink according to claim 23, further comprising: A processor-less subsystem is provided for moving the plurality of heat sinks.

25. The heat sink of claim 24, further comprising: The processor-less subsystem is implemented by bimorph metal associated with each heat sink, the bimorph metal being used to unfold a portion of each heat sink to expose a surface area of the plurality of heat sinks.

26. The heat sink of claim 23, further comprising: The plurality of fins are adapted to expand in response to heat from an associated component exceeding a threshold value, and to contract in response to the heat falling below the threshold value.

27. A method for cooling data center equipment, comprising: providing a plurality of heat sinks between a first plate and a second plate to dissipate a first amount of heat to an environment in a first configuration of the plurality of heat sinks, the first plate being movable relative to the second plate to expose a surface area of the plurality of heat sinks to the environment in a second configuration of the plurality of heat sinks and dissipate a second amount of heat greater than the first amount of heat; wherein, in the first configuration, each heat sink is bent to include an overlapping portion, the overlapping portion being at least partially isolated from the environment; and In the second configuration, the individual fins are unfolded and the overlapping portions are separated to expose the overlapping portions to the environment.

28. The method of claim 27, further comprising: A gear subsystem, an electromagnetic subsystem, a thermoelectric generator subsystem, a thermal reaction subsystem, or a pneumatic subsystem is provided to move the first plate relative to the second plate in response to sensed heat from an associated computing component or from the environment.

29. The method of claim 27, further comprising: associating at least one strip portion with the plurality of fins; By means of the at least one strip portion, each heat sink is enabled to include the overlapping portion; as well as The structure of the at least one strip portion is enabled to change to expose the overlapping portion to the environment in the second configuration.

30. The method of claim 27, further comprising: receiving cooling fluid from a cooling circuit of a data center housing the data center equipment using a fluid or gas line; A piston associated with a pneumatic subsystem is extended using the cooling fluid to move the first plate relative to the second plate to expose a surface area of the plurality of fins in the second configuration of the plurality of fins.

Citation Information

Patent Citations

  • Ribbed radiator with changeable dimension

    CN103841808A