Intelligent Threshold Leakage Remedy for Data Center Cooling Systems

By using AI/ML algorithms to monitor and adjust cooling system parameters in the data center liquid cooling system, the threshold leakage problem is solved, ensuring that the equipment operates at normal temperature, reducing coolant loss, protecting equipment safety, and improving system stability and efficiency.

CN114126341BActive Publication Date: 2025-07-18NVIDIA CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110996106.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-08-27
Filing Date
2021-08-27
Publication Date
2025-07-18
Estimated Expiration
2041-08-27

AI Technical Summary

Technical Problem

There is a threshold leakage problem in the liquid cooling system in the data center, which causes the computing device to operate within the normal temperature threshold but improperly lose coolant, which may affect the normal operation and safety of the device.

Method used

Using algorithms based on artificial intelligence and machine learning, the cooling system parameters are monitored through sensors, potential leakage is predicted and power state and coolant flow is adjusted, reducing dependence on coolant and achieving intelligent remediation.

Benefits of technology

Effectively detect and deal with threshold leakage, reduce dependence on coolant, protect computing equipment from damage, maintain computing status, and improve system robustness and reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114126341B_ABST
    Figure CN114126341B_ABST
Patent Text Reader

Abstract

Intelligent Threshold Leak Remedy for Data Center Cooling Systems is disclosed, specifically a remedy system for threshold leaks in a liquid cooling system for a data center. The system includes a flow controller and a power controller, which are adapted to receive inputs from a learning subsystem that can determine that a threshold leak has occurred even while the computing components are operating normally, thereby changing the power state to reduce reliance on the coolant, which can cause a change in the flow rate of the coolant.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] At least one embodiment relates to a remediation system for threshold leakage in a data center cooling system. In at least one embodiment, a flow controller and a power controller are adapted to receive inputs from a learning subsystem that can determine that a threshold leakage has occurred even while computing components are operating normally, causing a change in power state to reduce reliance on coolant and a change in the flow rate of the coolant. Background Art

[0002] Data center cooling systems typically use fans to circulate air through server components. Certain supercomputers or other high-capacity computers may use water or other cooling systems instead of air-cooling systems to absorb heat from server components or racks in the data center to an area outside the data center. The cooling system may include a chiller within the data center area (including areas outside the data center). The area outside the data center may be an area including a cooling tower or other external heat exchanger that receives heated coolant from the data center and dissipates heat to the environment (or an external cooling medium) by air-cooling or other means before the cooled coolant is recirculated back to the data center. In an example, the chiller and the cooling tower together form a cooling facility, where pumps respond to temperatures measured by external devices applied to the data center. A separate air-cooling system may not be able to absorb enough heat to support effective or efficient cooling of the data center, while liquid-cooling systems, while meeting the needs of the data center, are prone to leakage problems that can cause equipment short circuits and damage. Brief Description of the Drawings

[0003] Embodiments in accordance with the present disclosure will be described with reference to the drawings, where:

[0004] Figure 1 is a block diagram of an example data center having a cooling system subject to improvements described in at least one embodiment;

[0005] Figure 2 is a block diagram showing server-level features of a remediation system for threshold leakage in a data center liquid cooling system according to at least one embodiment;

[0006] Figure 3 is a block diagram showing rack-level features of a remediation system for threshold leakage in a data center liquid cooling system according to at least one embodiment;

[0007] Figure 4 is a block diagram showing data center-level features of a remediation system for threshold leakage in a data center liquid cooling system according to at least one embodiment;

[0008] Figure 5is a processing flow of steps of a method for a cooling system usable for use or manufacture according to at least one embodiment Figure 2-4 and Figures 6A-17D ;

[0009] Figure 6A illustrates an example data center in which at least one embodiment from Figure 2-5 can be used;

[0010] Figure 6B 、 Figure 6C illustrates inference and / or training logic according to various embodiments for implementing and / or supporting a remediation system for threshold leakage in a data center liquid cooling system, such as the inference and / or training logic used in Figure 6A and at least one embodiment of the present disclosure;

[0011] Figure 7A is a block diagram illustrating an exemplary computer system according to at least one embodiment, the exemplary computer system can be a system with interconnected devices and components, a system on a chip (SOC), or some combination thereof formed with a processor, the processor can include execution units for executing instructions to support and / or implement the remediation system for threshold leakage in a data center liquid cooling system described herein;

[0012] Figure 7B is a block diagram illustrating an electronic device for using a processor to support and / or implement a remediation system for threshold leakage in a data center liquid cooling system;

[0013] Figure 7C illustrates a block diagram of an electronic device for using a processor to support and / or implement a remediation system for threshold leakage in a data center liquid cooling system;

[0014] Figure 8 illustrates another exemplary computer system according to at least one embodiment, which is used to implement various processes and methods of the remediation system for threshold leakage in a data center liquid cooling system described throughout the present disclosure;

[0015] Figure 9A illustrates an exemplary architecture according to at least one embodiment of the present disclosure, in which a GPU is communicatively coupled to a multi-core processor through a high-speed link to implement and / or support a remediation system for threshold leakage in a data center liquid cooling system;

[0016] Figure 9B illustrates additional details of the interconnection between a multi-core processor and a graphics acceleration module according to one exemplary embodiment;

[0017] Figure 9C Shows another exemplary embodiment according to at least one embodiment of the present disclosure, in which an accelerator integrated circuit is integrated within a processor to implement and / or support a remediation system for threshold leakage in a data center liquid cooling system;

[0018] Figure 9D Shows an exemplary accelerator integration slice for implementing and / or supporting a remediation system for threshold leakage in a data center liquid cooling system according to at least one embodiment of the present disclosure;

[0019] Figure 9E Shows additional details of an exemplary embodiment of a shared model for implementing and / or supporting a remediation system for threshold leakage in a data center liquid cooling system according to at least one embodiment of the present disclosure;

[0020] Figure 9F Shows additional details of an exemplary embodiment of a unified memory, which can be addressed via a common virtual memory address space for accessing physical processor memory and GPU memory, to implement and / or support a remediation system for threshold leakage in a data center liquid cooling system;

[0021] Figure 10A Shows an exemplary integrated circuit and an associated graphics processor for a remediation system for threshold leakage in a data center liquid cooling system according to the embodiments described herein;

[0022] Figures 10B-10C Shows an exemplary integrated circuit and an associated graphics processor for supporting and / or implementing a remediation system for threshold leakage in a data center liquid cooling system according to at least one embodiment;

[0023] Figures 10D-10E Shows additional exemplary graphics processor logic for supporting and / or implementing a remediation system for threshold leakage in a data center liquid cooling system according to at least one embodiment;

[0024] Figure 11A Is a block diagram of a computing system for supporting and / or implementing a remediation system for threshold leakage in a data center liquid cooling system according to at least one embodiment;

[0025] Figure 11B Shows a parallel processor for supporting and / or implementing a remediation system for threshold leakage in a data center liquid cooling system according to at least one embodiment;

[0026] Figure 11C Is a block diagram of a partitioning unit according to at least one embodiment;

[0027] Figure 11D Shows a graphics multiprocessor for a remedial system for threshold leakage in a data center liquid cooling system according to at least one embodiment;

[0028] Figure 11E Shows a graphics multiprocessor according to at least one embodiment;

[0029] Figure 12A Shows a multi-GPU computing system according to at least one embodiment;

[0030] Figure 12B Is a block diagram of a graphics processor according to at least one embodiment;

[0031] Figure 13 Is a block diagram showing a microarchitecture for a processor according to at least one embodiment, the processor may include logic circuitry for executing instructions;

[0032] Figure 14 Shows a deep learning application processor according to at least one embodiment;

[0033] Figure 15 Shows a block diagram of a neuromorphic processor according to at least one embodiment;

[0034] Figure 16A Is a block diagram of a processing system according to at least one embodiment;

[0035] Figure 16B Is a block diagram of a processor having one or more processor cores, an integrated memory controller, and an integrated graphics processor according to at least one embodiment;

[0036] Figure 16C Is a block diagram of the hardware logic of a graphics processor core according to at least one embodiment;

[0037] Figures 16D-16E Shows thread execution logic according to at least one embodiment, which includes an array of processing elements of a graphics processor core.

[0038] Figure 17A Shows a parallel processing unit according to at least one embodiment;

[0039] Figure 17B Shows a general processing cluster according to at least one embodiment;

[0040] Figure 17C Shows a memory partition unit of a parallel processing unit according to at least one embodiment; and

[0041] Figure 17D Shows a streaming multiprocessor according to at least one embodiment. Detailed implementation mode

[0042] Air cooling of high-density servers may be inefficient or ineffective because sudden high heat demands are caused by changing computational loads in today's computing components. However, since the requirements can change or tend to a range of different cooling demands from minimum to maximum, a suitable cooling system must be used to meet these requirements in an economical way. For medium to high cooling demands, a liquid cooling system can be used. Different cooling demands also reflect different thermal characteristics of the data center. In at least one embodiment, the heat generated from components, servers, and racks is cumulatively referred to as the thermal characteristic or cooling demand because the cooling demand must fully address the thermal characteristic. In at least one embodiment, the thermal characteristic or cooling demand of a cooling system is the heat or cooling demand generated by components, servers, or racks associated with the cooling system and can be part of the components, servers, and racks of a data center.

[0043] In at least one embodiment, a remedial system for threshold leakage in a data center liquid cooling system is disclosed. The remedial system is adapted to intelligently detect potential liquid leakage in different liquid cooling components (such as liquid-cooled server racks), herein referred to as threshold leakage. In at least one embodiment, the threshold leakage is a potential leakage where no physical leakage has occurred. In at least one embodiment, the threshold leakage is a limited leakage that does not affect the normal operation of at least one computing device (also referred to herein as a data center device). In at least one embodiment, normal operation is when at least one computing component operates within a normal temperature threshold and receives coolant from the data center liquid cooling system (or associated therewith) that is experiencing threshold leakage.

[0044] In at least one embodiment, the remedial system includes a rack-mounted power distribution unit (PDU) having a plurality of wired or wireless sensors, such as temperature, humidity, fluid flow rate, leakage sensors, and other relevant sensors with artificial intelligence or machine learning (AI / ML)-based algorithms, which are adapted to determine threshold leakage by at least predicting parameter changes within a range (which can be a moving range). These changes may not affect the normal operation of at least one computing device, but provide information on a greater leakage that may occur and potentially damage at least one computing device. The parameters can be associated with characteristics or aspects of a data center liquid cooling system installed to cool multiple data center racks. In response to the information characteristics of the threshold leakage, the AI / ML-based algorithm notifies a power controller to cause at least one computing component to change its power state to reduce reliance on coolant and notifies a flow controller to cause a change in the flow of coolant to at least one computing component.

[0045] In at least one embodiment, a remediation system includes a central or distributed control system integrated within or associated with a PDU of a single or multiple racks in a data center. The PDU is enhanced with one or more of the above sensors, which endow it with the ability to monitor and measure various physical system characteristics or parameters associated with a liquid cooling system of the data center. The characteristics or parameters may include the temperature of critical components (e.g., pipe temperature, component temperature, fluid flow rate, humidity, relative humidity, and one or more leak signals). In at least one embodiment, at least one parameter associated with the liquid cooling system of the data center is a cooling response, which may be a proportional parameter that correlates a cooling demand with the power drawn by the system or component being cooled. In at least one embodiment, the more power is drawn, the more heat is generated and the more cooling is required.

[0046] In at least one embodiment, the effects of the present disclosure enable the use of at least one or more of the above parameters to predict the likelihood of small-scale or large-scale fluid leaks from threshold leaks, which may occur anywhere in the fluid flow path of a liquid-cooled component (e.g., a server, network component, storage component, or rack component). The disclosed AI / ML-based algorithms are adapted to be trained using at least one or more of the above parameters to achieve an inference of a threshold leak and subsequent rapid response. In at least one embodiment, the rapid response may be an appropriate signal sent from the PDU to at least one flow controller and at least one power controller to adjust or shut off the coolant or other fluid flow to the affected computing device (or other liquid-cooled information technology (I.T. or IT) device) of the data center liquid cooling system, while acting to suddenly or normally shut down the computing device. In at least one embodiment, these actions minimize system damage while maintaining the status of the computing state of the IT servers and network devices. In at least one embodiment, the data center or computing device may be a graphics processing unit (GPU), a central processing unit (CPU), and a switch.

[0047] In at least one embodiment, the present disclosure implements a plug-and-play remediation system for leaks, in part, based on intelligently determining threshold leaks and preventing degradation of a data center liquid cooling system. The AI / ML aspects are provided via one or more processors, as described throughout the present disclosure. In at least one embodiment, a remediation system for threshold leaks in a data center liquid cooling system includes a flow controller and a power controller within a power distribution unit (PDU). The flow controller and the power controller are adapted to receive inputs from a learning subsystem that is adapted to determine that at least one parameter associated with the data center liquid cooling system is outside a determined range, indicating that a threshold leak of coolant has occurred. The threshold leak is associated with at least one computing component that operates within normal temperature thresholds and receives coolant. The power controller causes at least one computing component to change power states to reduce reliance on coolant, while the flow controller causes a change in the flow of coolant to at least one computing component. Thus, in at least one embodiment, the present disclosure is a simple and readily available system for protecting expensive IT infrastructure from simple water damage caused by liquids within a data center liquid cooling system.

[0048] Figure 1 FIG. 4 is a block diagram of an example data center 100 having a cooling system that undergoes the improvements described in at least one embodiment. The data center 100 can be one or more rooms 102 that house racks 110 and ancillary equipment to accommodate one or more servers on one or more server trays. The data center 100 is supported by a cooling tower 104 located outside the data center 100. The cooling tower 104 dissipates heat from within the data center 100 by acting on a main cooling loop 106. Additionally, a cooling distribution unit (CDU) 112 is used between the main cooling loop 106 and a second or auxiliary cooling loop 108 to enable extraction of heat from the second or auxiliary cooling loop 108 into the main cooling loop 106. In one aspect, the auxiliary cooling loop 108 can connect various conduits directly into server trays as needed. Loops 106, 108 are shown as line diagrams, but one of ordinary skill in the art will recognize that one or more conduit features can be used. In one case, flexible polyvinyl chloride (PVC) tubing can be used with associated piping to move fluid along each of the loops 106, 108. In at least one embodiment, one or more coolant pumps can be used to maintain a pressure differential within the loops 106, 108 to enable coolant to move based on temperature sensors at different locations, including within the room, within one or more racks 110, and / or within server enclosures or server trays within the racks 110.

[0049] In at least one embodiment, the coolant in the primary cooling loop 106 and the secondary cooling loop 108 can be at least water and an additive, such as ethylene glycol or propylene glycol. In operation, each of the primary cooling loop and the secondary cooling loop has its own coolant. In one aspect, the coolant in the secondary cooling loop can be dedicated to the needs of the components in the server tray or rack 110. The CDU 112 is capable of precisely controlling the coolant in the loops 106, 108 independently or simultaneously. For example, the CDU can be adapted to control the flow rate so as to appropriately distribute the coolant to extract the heat generated within the rack 110. Additionally, more flexible tubes 114 are provided from the secondary cooling loop 108 to enter each server tray and supply coolant to the electrical and / or computing components. In the present disclosure, the electrical and / or computing components can be used interchangeably to refer to the heat-generating components that benefit from the data center cooling system. The pipe 118 forming part of the secondary cooling loop 108 can be referred to as a room manifold. Separately, the pipe 116 extending from the pipe 118 can also be part of the secondary cooling loop 108, but can be referred to as a row manifold. The tube 114 enters the rack as part of the secondary cooling loop 108, but can be called a rack cooling manifold. Additionally, the row manifold 116 extends along a row in the data center 100 to all the racks. The pipe of the secondary cooling loop 108 including the manifolds 118, 116, and 114 can be improved by at least one embodiment of the present disclosure. An optional cooler 120 can be provided in the primary cooling loop within the data center 102 to support cooling before the cooling tower. In terms of the presence of additional loops in the primary control loop, those of ordinary skill in the art reading the present disclosure will recognize that the additional loops provide cooling outside the rack and outside the secondary cooling loop; and can be used with the primary cooling loop for the present disclosure.

[0050] In at least one embodiment, in operation, heat generated within the server trays of rack 110 can be transferred via the flexible tubes of manifold 114 of the second cooling loop 108 to the coolant exiting rack 110. Correlatively, a second coolant (in the auxiliary cooling loop 108) for cooling rack 110 from CDU 112 moves towards rack 110. The second coolant from CDU 112 is transferred from one side of the room manifold having pipes 118 to one side of rack 110 via manifold 116 and through one side of the server trays via pipe 114. The used second coolant (or the heat-carrying second coolant exiting the computing components) exits from the other side of the server tray (e.g., enters the left side of the rack of the server tray and exits from the right side of the rack after circulating through the server tray or components on the server tray). The used second coolant exiting the server tray or rack 110 exits from a different side (e.g., the outlet side) of pipe 114 and moves to the parallel but also the outlet side of manifold 116. From manifold 116, the used second coolant moves in a direction opposite to that of the incoming second coolant (which can also be fresher second coolant) in the parallel portion of the room manifold 118 and towards CDU 112.

[0051] In at least one embodiment, the used second coolant exchanges its heat with the primary coolant in the primary cooling loop 106 via CDU 112. The used second coolant is refreshed (e.g., relatively cooled when compared to the temperature at the used coolant) and is ready to cycle back to the computing components through the second cooling loop 108. The various flow and temperature control features in CDU 112 enable control of the heat exchanged from the used second coolant or the flow of the second coolant into and out of CDU 112. CDU 112 is also capable of controlling the flow of the primary coolant in the primary cooling loop 106. Since there are many flow lines or pipes and there can be ports, couplers, and adapters between the flow lines or pipes and other components (such as CDUs and cold plates), there are many opportunities for leaks that can cause damage to the IT infrastructure.

[0052] Figure 2FIG. 0 is a block diagram showing server-level features 200 of a remedial system for threshold leakage in a data center liquid cooling system. In at least one embodiment, the remedial system is partially or fully within a server chassis or server tray 202. In at least one embodiment, the remedial system is adapted to identify and address threshold leakage in the data center liquid cooling system. The server chassis or server tray 202 may include a coolant distribution unit (CDU) 204 having associated flow loops 212A, 212B. Each flow loop 212A, 212B has an inlet flow line 214 and an outlet flow line 218. One or more intermediate flow lines 216 effect continuous cooling of one or more cold plates 210A, 210B. Similar features described with respect to loop 212A may be used for loop 212B, including providing cooling to one or more cold plates 210C, 210D. Thus, at least one cooling assembly 220A, B, C, D associated with one or more of the cold plates 210A-D may be subject to the cooling of the data center liquid cooling system.

[0053] In at least one embodiment, flow controllers 222A; 222B; 226 are provided for controlling different levels of flow in the data center. In at least one embodiment, one or more flow controllers 222A, 222B may be partially or fully disposed within the PDU 204 for controlling the flow of coolant or fluid from outside the server chassis or tray 202 into the server chassis or tray 202. In at least one embodiment, one or more flow controllers 222A, 222B may be disposed within the PDU 204 for controlling the flow of coolant or fluid from a rack housing the server chassis or tray 202 to the cold plates 210A-D within the server chassis or tray 202. In at least one embodiment, the cold plates 210A-D may be independently associated with a flow controller, such as flow controller 226 provided at a fluid port 224A associated with cold plate 210A. The fluid ports 224A-D couple the inlet and outlet (or intermediate) flow lines into and out of the cold plates.

[0054] In at least one embodiment, at least one flow controller 222A;B can be a micropump or valve that can be remotely controlled to provide control or regulation of coolant flow, as described throughout this disclosure. In at least one embodiment, at least one flow controller 222A;B can be located at the input (or inlet, such as flow controllers 222A, 222B) or output (or outlet, such as flow controllers 222C, 222D) of a pipe or flow line. In at least one embodiment, when located at the outlet of a pipe or flow line, the action permitted by at least one flow controller 222C;D is a suction or pulling action applied to the coolant by a micropump or a coolant outflow action of opening the outflow of the coolant using a valve. In at least one embodiment, when located at the inlet of one or more loops 212A, 212B, the action permitted by at least one flow controller 222A;B is a pumping action applied to the coolant by a micropump or a coolant inflow action of opening the inflow of the coolant using a valve.

[0055] In at least one embodiment, the micropump or valve includes wired and wireless control sub-assemblies located with the PDU. In at least one embodiment, the pump or valve control is within the PDU, while the pump and valve can be within the CDU 204. In at least one embodiment, the pump or valve operates only from the PDU. In at least one embodiment, the reference to a flow controller thus refers to a feature capable of causing a change in the flow of coolant to at least one computing component. In at least one embodiment, one or more flow controllers 222A, B are partially within the CDU 204 and partially within the PDU of the rack housing the server tray or enclosure 202.

[0056] In at least one embodiment, the wired and wireless control sub-assemblies are the electronic controller part of the flow controller. The electronic controller of the flow controller can be a pump speed controller with a shut-off configuration. In at least one embodiment, at least partially based on a determined threshold leakage, the pump speed controller can be adjusted in a fuzzy (non-binary) manner through an input in a flow regulation circuit according to the requirements of the learning subsystem to achieve different and desired flow rates from the pump (including a fully closed flow). In at least one embodiment, the wired and wireless control sub-assemblies are a valve controller with a shut-off configuration, forming the electronic controller part of the flow controller. In at least one embodiment, partially based on the determined threshold leakage, the valve controller can be adjusted in a fuzzy manner according to the requirements of the learning subsystem to achieve different and desired flow rates through the valve (including a fully closed flow). In at least one embodiment, the pump or valve forms the mechanical controller part of the flow controller. In at least one embodiment, the flow controller is just a feature in the PDU that controls the pump or valve, because the pump or valve in the CDU cannot operate without the action of the flow controller in the PDU.

[0057] Figure 3 is a block diagram showing rack-level features 300 of a remedial system for threshold leakage in a data center liquid cooling system. The rack-level features 300 include one or more brackets 304, 306, which are associated with one or more PDUs 308, 310 and have one or more inlet ducts 312A, B. The one or more brackets 304, 306 also support one or more CDUs 316, 318 and have one or more inlet pipes or flow lines 314A, B. A flow controller having one or more parts 320A, B is located wholly or partially in a respective one of the one or more PDUs 308, 310. In at least one embodiment, an electronic controller part (e.g., at least one processor) is in the PDU, while a mechanical controller part (pump or valve) is in the CDU. In at least one embodiment, the location of the parts of the flow controller depends on whether the flow controller is on the outlet side, inlet side, or both sides of the server tray or enclosure 328. In at least one embodiment, although the electronic and mechanical controller parts are coupled via a link 324, the electronic controller part 320A can be coupled to multiple mechanical controller parts 320B via multiple such links, enabling a single electronic controller part to control multiple mechanical controller parts of different server enclosures or trays 328 simultaneously or individually.

[0058] In at least one embodiment, the link is a high-speed communication or electronic link, and although illustrated as physically coupled, it can be a wireless coupling between the electronic controller part and the mechanical controller part. There may be a distributed control system between the electronic controller part and the mechanical controller part. In at least one embodiment, the electronic controller part performs substantial operations to control the mechanical controller part, which may be designed to have limited circuitry or processing capabilities. In at least one embodiment, the electronic controller assembly can be a circuit board having at least one processor and having an independent or shared part of an AI / ML algorithm for determining threshold leakage and for acting on the determination of threshold leakage. In at least one embodiment, the action characteristics of the electronic controller part can cause a change in the flow of coolant to an associated at least one computing component via the mechanical controller part acting on the coolant through the port 318 of the rack-level features 300.

[0059] In at least one embodiment, one or more power controllers 322 are provided within the PDUs 308, 310, which also house the electronic controller portion 320A of the (host) flow controller. In at least one embodiment, the flow controller can be referred to by separately referring to the electronic controller portion 320A. This enables the PDUs 308; 310 to house components that make the power supply a controller and control the flow of the coolant. The electronic controller portion 320A is associated with one or more mechanical controller portions in a rack that houses the server trays or enclosures 328. In at least one embodiment, one or more power controllers 322 include at least one processor and are self - sufficient or part of a distributed control system, either together with or independent of the at least one processor of the electronic controller portion 320A. In at least one embodiment, one or more power controllers 322 can communicate directly with at least one computing device in the server tray or enclosure 328 via a separate link 326 that serves as a high - speed communication or electrical link. In the manner of link 324, link 326 can be a physical coupling or a wireless coupling between one or more power controllers 322 and at least one computing device.

[0060] In at least one embodiment, the server tray or enclosure 328 receives power and communication via the cables of one or more PDUs 308, 310. The power from the PDU can be used to power the mechanical controller portion of the flow controller and the computing devices within the server trays or enclosures of the rack. Thus, one or more power controllers 322 within one or more PDUs 308, 310 can control the power supply to at least one computing device via signals to the power supply or transmitter within the PDUs 308, 310. At the same time, one or more power controllers 322 also supply power to one or more flow controllers, such as both the electronic controller portion and the mechanical controller portion. One or more power controllers 322 can ensure that power continues to be supplied to the flow controller before shutting down the flow controller and then the computing device and finally itself, so as to enable the change of the coolant flow. In at least one embodiment, the shutdown of one or more power controllers occurs in sequence to cause the shutdown of the coolant flow, the shutdown of the computing device, and then its own shutdown. Thus, in at least one embodiment that includes such a sequence, the flow controller can be indirectly controlled by the power controller. In at least one embodiment, the link 326 can be from one or more power controllers 322 to the power supply or transmitter within one or more PDUs 308, 310 and is not visible outside the PDU.

[0061] In at least one embodiment, the flow controller and the power controller within the PDU are adapted to receive inputs from the learning subsystem. In at least one embodiment, the electronic controller portion 320A and one or more power controllers 322 are part of a distributed control system that is capable of performing some or all of the AI / ML algorithms of the learning subsystem to remediate, at least in part, based on a threshold leak in the data center cooling system. However, thus, the learning subsystem includes at least one processor, even if the at least one processor is a shared processor. In at least one embodiment, the electronic controller portion of the flow controller, one or more power controllers, and the learning subsystem are different modules of the remediation system, since each module performs at least different operations.

[0062] In at least one embodiment, the learning subsystem is adapted to determine that at least one parameter associated with the data center liquid cooling system is outside a determined range, such that a threshold leak of coolant has occurred. In at least one embodiment, the threshold leak is associated with at least one computing component of a respective server tray or enclosure 328 that operates within a normal temperature threshold and receives coolant from an inlet line 330. The power controller 322 causes at least one computing component of the respective server tray or enclosure to change its power state to reduce its reliance on coolant, at least in part based on an input from the learning subsystem. Separately, the electronic controller portion 320A, as well as the mechanical controller portion 320B, of the flow controller are caused to vary the flow of coolant to at least one computing component, at least in part based on another input (or the same input) from the learning subsystem.

[0063] In at least one embodiment, the threshold leak indicates that a first amount of coolant is incorrectly leaving the data center liquid cooling system while at least one computing component operates within a normal temperature threshold and receives coolant. In other words, there is at least one parameter that is abnormal, but most of the remaining parameters monitored by the learning subsystem are likely normal. However, correlatively, at least one computing device does not exhibit signs of being cooling- and operation-limited. In at least one embodiment, the threshold leak occurs before a normal leak occurs, where the normal leak is different from the threshold leak in that it indicates that a second amount of coolant greater than the first amount is incorrectly leaving the data center liquid cooling system. In such a scenario, the leak is also associated with at least one computing component not being able to operate within a normal temperature threshold and not receiving coolant to maintain the normal temperature threshold.

[0064] Figure 4 is a block diagram showing a data center-level feature 400 of a remediation system for a threshold leak in a data center liquid cooling system. In at least one embodiment, a plurality of racks 402 are equipped with Figure 2 and Figure 3The features described in. In at least one embodiment, the drain manifold 404 has ports coupled to the rack manifolds 406, 408. In at least one embodiment, each of the manifolds 404 - 408 is separately designated for inlet cooling fluid and for outlet cooling fluid. Among the manifolds 404 - 408, a first set of manifolds performs the function of one or more drain manifolds, which are coupled to a second set of manifolds that perform the function of one or more rack manifolds.

[0065] In at least one embodiment, for a remediation system that supports data center - level feature 400, one or more flow controllers 410 have their electronic controller portions within the PDU of the respective rack 402, but also implement the mechanical controller portion associated with the drain manifold. The flow controllers 410 act at susceptible points, such as at port 412 between the drain manifold and the rack manifold. In fact, one or more flow controllers 410 can address threshold leakage at the drain manifold level using the control form provided within the rack's PDU. In at least one embodiment, one or more flow controllers 410 can include the electronic and mechanical controller portions within a single enclosure and can be part of a distributed control system via link 414. In at least one embodiment, the distributed control system is implemented by at least one processor within the single enclosure of each flow controller, such as flow controller 410; and by at least one processor within one or more power controllers 418. At least one processor of the distributed control system is associated with a learning subsystem to control the flow controllers 410 and the power controllers 418 using the inputs to each controller. In at least one embodiment, a first input among the inputs will cause the shutdown of at least one computing component and a second input among the inputs will cause the shutdown of the coolant.

[0066] In at least one embodiment, the remediation system can be implemented using at least one processor of a computing device within the server box. In at least one embodiment, the at least one processor can be Figure 14 the processor 1400 in, and can use the neuron 1502 and its components, which can be implemented using circuitry or logic, including one or more arithmetic - logic units (ALUs) as described in Figure 15 to implement the AI / ML algorithms described throughout this disclosure.

[0067] In at least one embodiment, the remediation system may rely on at least one processor to implement a load transfer subsystem that transfers the load associated with at least one computing component experiencing threshold leakage before at least one computing component shuts down. In at least one embodiment, the load may be transferred to at least one second computing component that receives the coolant or a second coolant not affected by the threshold leakage. The remediation system may transfer the load from a server chassis of a computing component in one rack 402 to a different rack 402, or to a different server chassis in the same rack.

[0068] In at least one embodiment, the remediation system has at least one processor associated with a learning subsystem such that the learning subsystem can determine that a threshold leakage of the coolant has occurred by determining that the pressure, flow rate, or temperature of the coolant going to or from at least one computing component is outside of a normal threshold within a determined range. In at least one embodiment, the coolant may be associated with different pressures, flow rates, or temperatures. In at least one embodiment, the manufacturer may define these different pressures, flow rates, or temperatures for the effective operation of the coolant. The range may be taken from a previously specified range for the effective operation of the coolant. One range may be defined as a tolerance of 10% to 2% of the lower limit of the maximum specified pressure, temperature, or flow rate for effective operation. This range is the warning range. A smaller tolerance, such as 2% below the warning range, may be taken from the range for effective operation. The two percentage points of tolerance (2% and below of the effective operation range) are defined by at least an alarm threshold (e.g., at the 1% mark of the maximum point of the effective operation range) at which at least one computing component can no longer operate within the normal temperature threshold and no longer receives coolant to maintain the normal temperature threshold. In effect, the coolant no longer operates effectively as a coolant at the alarm threshold.

[0069] In at least one embodiment, the present disclosure enables at least a remediation system to operate at tolerances different from those specified by the manufacturer. In at least one embodiment, the warning range is the movement range. In at least one embodiment, since the coolant can be designated to change its flow rate to achieve different cooling levels, the coolant can have a wide range of operations, but the wide range of operations may not track a series of operations of at least one computing device. The present disclosure can relate the effective operating parameters of the coolant (extrapolated to a data center liquid cooling system) to the functions of at least one computing device (which can also be extrapolated to enter server trays or racks or more). In at least one embodiment, if the coolant is designated to reduce the heat generation of 250 KW, the warning range can be the heat generated by at least one computing device receiving the coolant ranging from 225 KW to 245 KW. The alarm threshold can range from 245 KW to 250 KW. In at least one embodiment, at least one computing component can operate with a heat generation between 225 KW and 240 KW. Thus, the movement range is the range identified by the learning subsystem, ignoring the range from 225 KW to 245 KW, supporting the range from 225 KW to 240 KW of at least one computing component, and being able to act on this range to cause at least one computing component to change its power state, so as to reduce the dependence on the coolant and cause a change in the flow rate of the coolant to at least one computing component through the corresponding flow controller and power controller.

[0070] Thus, in at least one embodiment, the remediation system works with the computing devices in the data center and the coolant used to determine the best actions to prevent damage to the computing devices. Even if the coolant is designated for a higher cooling capacity but fails to achieve this to address the high heat generated by at least one computing component (which may be shown as a parameter change rather than a physical leak in at least one embodiment due to an undetected leak), the remediation system can react before at least one computing device is damaged or before a physical leak occurs. However, in at least one embodiment, the threshold leak is a physical leak outside the coolant pipe or joint (such as a fluid adapter) that cannot be detected by the leak detector, but is also a leak that does not affect the normal operation of at least one computing component. In at least one embodiment, normal operation means that the heat generated by at least one computing component is within the normal operation threshold of the at least one computing component and does not cause any damage to the at least one computing component. In at least one embodiment, the threshold leak may be the result of inadvertently diverting the coolant flow to a computing component that does not require additional coolant, and this may result in a reduction in the coolant flow rate or pressure at a different computing component that requires additional coolant.

[0071] In one embodiment, the remediation system includes a distributed control system having one or more of a primary flow controller, a primary power controller, and at least a first portion of a learning subsystem within a first PDU; and having an auxiliary flow controller, an auxiliary power controller, and one or more of a second portion of a learning subsystem located within an auxiliary PDU. Thus, one or more processors located at different locations within the data center can provide intelligence to the learning subsystem to make it robust.

[0072] In at least one embodiment, the distributed control system communicates inputs from the learning subsystem with one or more of the auxiliary flow controller and the auxiliary power controller. The auxiliary flow controller and the auxiliary power controller can be associated with a server tray or enclosure different from the primary flow controller and the primary power controller, or can be associated with a rack completely different from the rack housing the primary flow controller and the primary power controller. In at least one embodiment, this feature enables at least one of the auxiliary flow controller and the auxiliary power controller to cause at least one auxiliary computing component to change to a second power state to reduce reliance on coolant and to cause a change in the flow of coolant to at least one auxiliary computing component. Thus, at least one computing component directly associated with the primary flow controller and the primary power controller is not affected, but the processing capabilities within these controllers are used to affect power and coolant changes for at least one auxiliary computing component in a different server tray or enclosure or a completely different rack.

[0073] In at least one embodiment, the learning subsystem includes a processor associated with one or more of a flow controller of a PDU, a power controller of a PDU, an auxiliary flow controller of an auxiliary PDU, and an auxiliary power controller of an auxiliary PDU. In at least one embodiment, this configuration illustrates the distributed nature of the control system. In at least one embodiment, the learning subsystem is associated with at least one processor to evaluate at least one parameter associated with a liquid cooling system of the data center using the foregoing moving range including a warning range (referred to as a determination range). The moving range represents a parameter value of at least one parameter, such as a parameter for the ability to address heat generated by at least one computing component. The parameter value is within a threshold of at least one parameter (e.g., 8% to 10% of the maximum heat generation capacity rating by the coolant as specified by the manufacturer).

[0074] In at least one embodiment, the learning subsystem is trained to recognize the flow rate of the coolant and the pressure of the coolant such that a higher cooling than indicated by a measured warning value must be achieved. At parameter values within the warning range, at least one computing component is still operating within the normal temperature threshold and is still receiving coolant, but there is an indication of ineffective coolant. The normal temperature threshold may be referred to as a functional parameter associated with at least one computing device. The indication of ineffectiveness using the warning threshold is also the case when not otherwise identified using only the functional parameters of at least one computing device indicated as normal. In at least one embodiment, this symbolizes a threshold leak of the coolant.

[0075] In at least one embodiment, a remedial system for threshold leaks in a data center liquid cooling system is capable of effecting a display and an intervention associated with the threshold leak. In at least one embodiment, the intervention system receives an input from the learning subsystem that is adapted to determine that at least one parameter associated with the data center liquid cooling system is outside a determined range and thus a threshold leak of the coolant has occurred. The intervention system provides a warning indication to effect an intervention to address the threshold leak associated with at least one computing component operating within the normal temperature threshold and receiving coolant. In at least one embodiment, if at least one parameter is determined to be a false alarm, the intervention system enables feedback to be provided to the learning subsystem. The intervention system can be used to enable a power controller to continue its operation and enable a flow controller to continue its operation.

[0076] In at least one embodiment, the learning subsystem can compare the flow rate, pressure, and cooling achieved for heat previously generated by at least one computing component with the current heat generated at the same flow rate and the same pressure. When the learning subsystem determines that the cooling is not as previously achieved while everything else is similar or within a reasonable range from a previous measurement (including the power drawn, indicating no change in workload; or the pipe temperature, indicating similar cooling within the pipe with coolant but no cooling plate), then the learning subsystem is able to initiate a power reduction or shutdown and is able to initiate a coolant shutdown. The shutdown can occur before transferring the processing load of the problematic computing device to a different computing device not affected by the problems experienced by the computing device being shut down.

[0077] In at least one embodiment, the learning subsystem provides a first input in the inputs to a power controller to cause at least one computing component to change its power state to reduce dependence on coolant, and provides a second input in the inputs to a flow controller to change the flow rate of coolant to at least one computing component. In at least one embodiment, at least one parameter includes one or more of the following: the temperature of the coolant; the temperature of at least one computing component or a first area having at least one computing component; the temperature of a pipe carrying the coolant; the humidity or relative humidity of the first area or a second area including the flow controller; the flow rate of coolant into the first area; the flow rate of coolant out of the first area; a proportional cooling response to the power drawn by at least one computing component; and the fluid leakage rate of the coolant. In at least one embodiment, when the learning subsystem tracks the resulting coolant pressure to address cooling requirements or reach a cooling level rated by its cooling capacity, if the pressure fails to rise to meet the same cooling requirements or reach the same cooling level, the learning subsystem can indicate that there may be a fluid leak. Over time, during a period of monitoring this issue, the pressure deficit may translate into a fluid leakage rate.

[0078] In at least one embodiment, the learning subsystem executes a machine learning model. The machine learning model enables processing of parameter values associated with at least one parameter using multiple neuron layers of the machine learning model. The machine learning model has parameter values and has a previously associated flow rate of coolant and a previously associated power state of at least one computing component. After evaluating the parameter values using the previously associated flow rate of coolant and the previously associated power state, the machine learning model provides inputs to the flow controller and the power controller.

[0079] In at least one embodiment, a PDU or a data center management system (DMS) includes a learning subsystem having at least one processor to execute the machine learning model. In at least one embodiment, the volume of coolant or the flow rate of coolant to initiate cooling or cause the temperature to be addressed or changed can be referenced against the expected cooling for at least one input temperature. In at least one embodiment, the learning subsystem can be implemented via a deep learning application processor (e.g., Figure 14 the processor 1400 in Figure 15 and can use neurons 1502 and their components, which can be implemented using circuitry or logic, including one or more arithmetic logic units (ALUs) as described in

[0080] In at least one embodiment, at least one processor can be used for a remediation system. In at least one embodiment, at least one processor can use a deep learning application processor (e.g., Figure 14in the processor 1400) and can be implemented using the neuron 1502 and its components, which can be implemented using circuitry or logic, including one or more arithmetic logic units (ALUs) as described in Figure 15 As described. In at least one embodiment, the at least one processor thus has at least one logic unit to control at least one flow controller and at least one power controller.

[0081] In at least one embodiment, the present disclosure relates to at least one processor for remedying a threshold leak in a liquid cooling system, which can be implemented in an existing data center liquid cooling system with minor physical modifications. In at least one embodiment, the at least one processor includes at least one logic unit for controlling a flow controller and a power controller within a power distribution unit (PDU), as discussed with reference to Figure 2-4 As discussed. The flow controller and the power controller receive inputs from a learning subsystem adapted to determine that at least one parameter associated with the liquid cooling system is outside a normal threshold within a determined range, indicating that a threshold leak of coolant has occurred.

[0082] In at least one embodiment, the threshold leak is associated with at least one computing component that operates within a normal temperature threshold and receives coolant. In at least one embodiment, as noted in the foregoing features, the functional parameters of the at least one computing component may not indicate a problem brewed with at least one liquid cooling system. The indication of the threshold leak is at least partially based on a warning threshold that may be a moving range. The moving range includes values of flow rate, flow, and power status learned by the learning subsystem. Flow rate, flow, temperature, humidity, and power status may have wide tolerances, depending on the operational aspects of the data center. Thus, a change in temperature (e.g., an increase) that may be considered a threshold leak and is not resolved by an expected increase in the flow rate of coolant to the affected computing device may not indicate a threshold leak. This is because there may be other coolant lines connected to other computing devices that draw from the same coolant source. The moving range takes into account these changes in expected flow rate or pressure that do not represent a threshold leak.

[0083] In at least one embodiment, a moving range is achieved by correlating temperature changes of multiple computing devices associated with the same coolant line with at least the flow rate, flow pressure, and flow volume of the coolant. As additional computing devices are loaded, the moving range is improved so that they perform at the highest allowed performance and require the maximum available cooling. In at least one example, a support vector machine (SVM) can be used to fit different data sets within a margin, and the fitted data can be used to train a neural network to track a parameter based on the moving range of the parameter. In at least one example, the unfixed moving range also inherently includes a proportion that may be betrayed by at least one parameter representing a threshold leakage situation.

[0084] In at least one embodiment, using an expected increased flow rate from the above temperature change (e.g., an increase) that is not resolved by the expected increased coolant flow rate flowing to the affected computing device, a machine learning model supporting SVM or one or more neural networks can correlate the temperature change with different coolant flow rates that can resolve the temperature change. Even if the coolant diverts in different branches of the cooling loop, the learning subsystem can consider variances (by at least using the margin of the SVM in at least one embodiment) before increasing the indication of threshold leakage if the associated computing device otherwise operates within normal temperature thresholds and receives coolant. In at least one embodiment, each parameter including its margin represents the moving range of the parameter. In at least one embodiment, other parameters described throughout this disclosure can be fitted or correlated with the margin in a similar manner so that different parameters can have a moving range.

[0085] In at least one embodiment, at least one logic unit associated with the learning subsystem enables the load transfer subsystem to transfer the load associated with at least one computing component before at least one computing component shuts down. The load is transferred to at least one second computing component that receives coolant or receives a second coolant that is not affected by threshold leakage. In at least one embodiment, at least one logic unit is associated with the learning subsystem to determine that a threshold leakage of coolant has occurred by determining that the pressure, flow rate, or temperature of the coolant entering and leaving at least one computing component is outside the normal threshold of a determined range. Instead, the values of these parameters may be within a warning range. The warning range is a range that is entirely within the alarm threshold and is thus defined by at least one alarm threshold at which at least one computing component no longer operates within the normal temperature threshold and no longer receives coolant to maintain the normal temperature threshold.

[0086] In at least one embodiment, at least one processor is adapted to include an instruction output to transfer an input from a learning subsystem to a flow controller and a power controller. In at least one embodiment, at least one logic unit of at least one processor is adapted to receive parameter values from a parameter sensor associated with a data center liquid cooling system. The at least one logic unit is also adapted to receive functional information from a functional sensor of at least one computing device. The functional information is also referred to as functional parameters throughout the present disclosure. The at least one logic unit is adapted to facilitate a change in power state and a change in coolant flow rate.

[0087] In at least one embodiment, the present disclosure relates to at least one processor for a liquid cooling system, the liquid cooling system for training or trainable. The at least one processor has at least one logic unit for training one or more neural networks having a hidden layer of neurons to evaluate parameter values associated with at least one parameter of the liquid cooling system. In at least one embodiment, the at least one logic unit must load these parameter values under a training scenario in order to be able to evaluate them. In at least one embodiment, the one or more neural networks use the parameter values and the previously associated flow rate of the coolant and the previously associated power state of at least one computing component to perform the evaluation. In at least one embodiment, the evaluation is for determining a range of movement of at least one parameter outside a normal threshold and indicating that a threshold leak of the coolant has occurred.

[0088] In at least one embodiment, as described throughout the present disclosure, the threshold leak is associated with at least one computing component operating within a normal temperature threshold and receiving coolant. In at least one embodiment, one or more neural networks are used to provide an output associated with a first change in the power state of at least one computing component (to reduce reliance on coolant) and a second change in the flow rate of coolant to at least one computing component.

[0089] In at least one embodiment, at least one processor capable of performing the above training has a learning subsystem within at least one logic unit for executing a machine learning model. The machine learning model can process parameter values associated with at least one parameter using the previously associated flow rate of the coolant and the previously associated power state of at least one computing component. The machine learning model can provide an input after evaluating the parameter values using the previously associated flow rate of the coolant and the previously associated power state. The input is to a flow controller and a power controller associated with the learning subsystem.

[0090] In at least one embodiment, at least one processor further includes at least one logic unit for outputting at least one instruction respectively associated with a first change in power state and a second change in coolant flow rate. In at least one embodiment, the at least one processor further includes an instruction output for transmitting an output from a learning subsystem that executes one or more neural networks to cause a first change in the power state of at least one computing component and to cause a second change in the flow rate of coolant to at least one computing component.

[0091] In at least one embodiment, at least one processor includes at least one logic unit adapted to receive parameter values from a parameter sensor associated with a data center liquid cooling system. The at least one logic unit is further adapted to receive functional information from a functional sensor of at least one computing device. This information can be whether the computing device is operating within the normal temperature value of its core or outside the normal temperature value of its core. When the parameter values are inconsistent and the computing device is operating normally, there may still be undetected problems, and the learning subsystem of the present disclosure is adapted to respectively facilitate a first change in power state and a second change in coolant flow rate.

[0092] At least one embodiment of the present disclosure relates to a remedial system for threshold leakage in a liquid cooling system having at least one processor adapted to be trained and provide an input for an output associated with a first change in the power state of at least one computing component (to reduce dependence on coolant) and associated with a second change in the flow rate of coolant to at least one computing component. In at least one embodiment, the remedial system uses at least one processor of a learning subsystem that executes a machine learning model as described in the description with reference to Figure 2-4 as described therein.

[0093] Figure 5 is a process flow for steps of a method 500 for a cooling system that can be used to use or manufacture Figure 2-4 and Figures 6A-17D Step 502 provides a flow controller and a power controller within a power distribution unit (PDU). Step 504 enables a learning subsystem to determine that at least one parameter associated with a data center liquid cooling system is outside a determined range. Step 506 can be part of the learning subsystem and performs a determination that a threshold leakage of coolant has occurred. Step 508 is performed based in part on the result of step 506. When at least one computing component associated with the cooling system is operating within a normal temperature threshold and receiving coolant, step 508 determines that a threshold leakage has occurred. Otherwise, step 504 continues to monitor the parameter values via the learning subsystem.

[0094] Step 510 enables the flow controller and the power controller to receive inputs from the learning subsystem. Step 512 causes at least one computing component to change its power state by the power controller to reduce reliance on the coolant. Additionally, step 514 causes a change in the coolant flow rate to at least one computing component by the flow controller. In at least one embodiment, remediation method 500 enables the learning subsystem to determine a threshold leak through at least two monitored actions to implement the substeps of steps 504 - 508. In a first monitored action, the learning subsystem has an indication that a first amount of coolant with previously recorded parameters is leaving the data center liquid cooling system incorrectly.

[0095] In at least one embodiment, the learning subsystem tracks one or more of the following: the temperature of the coolant; the temperature of at least one computing component or a first area having at least one computing component; the temperature of the pipes carrying the coolant; the humidity or relative humidity of the first area or a second area including the flow controller; the flow rate of the coolant flowing into the first area; the flow rate of the coolant from the first area; the proportional cooling response to the power drawn by at least one computing component; and the fluid leak rate of the coolant. In at least one embodiment, as the power drawn by at least one computing component increases, the heat generated from at least one computing component also increases, and the cooling demand increases. Knowing the power drawn informs the AI / ML learning subsystem of the change in cooling demand to be addressed by the data center liquid cooling system.

[0096] In at least one embodiment using parameters, the learning subsystem can determine that the temperature of the coolant is too high even when the cooling system is attempting to cool computing equipment that is typically cooled by the same coolant. In at least one embodiment, the learning subsystem can determine that the temperature of at least one computing component or a first area having at least one computing component is too high when previously controlled by the same coolant. In at least one embodiment, the learning subsystem can determine that the temperature of the pipes carrying the coolant is high when this is not the case for a previously established thermal signature or cooling demand.

[0097] In at least one embodiment, even if the coolant is flowing normally, the learning subsystem can determine that the humidity or relative humidity in the first area or the second area with a flow controller is outside the normal range. In at least one embodiment, the learning subsystem can determine that the flow rate of the coolant flowing to the first area is outside the normal range, even if the pressure is maintained or has a tolerance change. In at least one embodiment, the learning subsystem can determine that the flow rate of the coolant from the first area is outside the range, even when the applied pressure is maintained or has a tolerance change. In at least one embodiment, even if other parameters are within tolerance or are maintained, the learning subsystem can determine that a fluid leakage rate of the coolant is detected. In each of these examples, for the threshold leakage to be determined, only one parameter may be outside the tolerance.

[0098] In at least one embodiment, the threshold leakage is thus notified by at least a second monitored action, where the learning subsystem has at least one computing component operating within a normal temperature threshold and receiving an indication of the coolant. This is different from a normal leakage of a second quantity of coolant that exceeds a first quantity, where the second quantity of coolant incorrectly leaves the data center liquid cooling system such that at least one computing component cannot operate within the normal temperature threshold and cannot receive the coolant to maintain the normal temperature threshold.

[0099] In at least one embodiment, the remediation method 500 further includes features within step 510 for using at least one processor associated with the learning subsystem to control the flow controller and the power controller using the input. In at least one embodiment, a first input in the input causes the shutdown of at least one computing component and a second input in the input causes the shutdown of the coolant.

[0100] In at least one embodiment, the remediation method 500 further includes features within step 512 for using at least one processor to implement a load transfer subsystem to transfer the load associated with at least one computing component before the at least one computing component shuts down. In at least one embodiment, the load is transferred to at least one second computing component that receives the coolant or receives a second coolant that is not affected by the threshold leakage.

[0101] In at least one embodiment, the learning subsystem can be implemented via a deep learning application processor (e.g., Figure 14 the processor 1400 in Figure 15 ), and can use neurons 1502 and their components, which can be implemented using circuits or logic, including one or more arithmetic logic units (ALUs) as described inFigure 14 , Figure 15 The information collected for the feature processing being discussed. In at least one embodiment, the processing of the parameters uses multiple neuron levels of a machine learning model, which are loaded with one or more of the above - collected parameter values of the cooling system and the corresponding power states of at least one computing component. In at least one embodiment, different coolants can be used to perform testing or training. The learning subsystem performs training, which can be represented as an evaluation of changes in parameter values associated with a previous flow rate or flow associated with an adjustment to one or more flow controllers and a previous power state associated with an adjustment to a power controller. The neuron levels can store values associated with the evaluation process and can represent the association or correlation between parameter changes (even within a moving range) and the power state.

[0102] In at least one embodiment, the processor and the flow controller can work together. The processor (also referred to as a centralized or distributed control system) is at least one processor that has at least one logic unit for controlling the flow controller and the power controller associated with one or more cooling loops. In at least one embodiment, at least one processor is within a data center, such as Figure 7A processor 702. The controller facilitates the movement of the corresponding coolant and the power adjustment of the corresponding computing component. In at least one embodiment, at least one processor is a processor core of a multi - core processor, such as Figure 9A the multi - core processors 905, 906 in

[0103] In at least one embodiment, the processor (such as Figure 9A the processor cores of the multi - core processors 905, 906 in

[0104] Data center

[0105] Figure 6A illustrates an example data center 600, where from Figure 2-5At least one embodiment may be used. In at least one embodiment, data center 600 includes a data center infrastructure layer 610, a framework layer 620, a software layer 630, and an application layer 640. In at least one embodiment, for example, as referenced Figure 2-5 As described, features in components of a remediation system for threshold leakage in a data center liquid cooling system may be performed within or in cooperation with example data center 600. In at least one embodiment, infrastructure layer 610, framework layer 620, software layer 630, and application layer 640 may be provided, in part or in whole, by computing components located on server trays within racks 210 of data center 200. This enables the cooling system of the present disclosure to directly cool certain of the computing components in an efficient and effective manner. Additionally, various aspects of the data center, including data center infrastructure layer 610, framework layer 620, software layer 630, and application layer 640, may be used to support intelligent control of a controller in the threshold leakage remediation system referenced at least above Figure 2-5 discussed. Thus, the discussion referenced Figures 6A-17D may be understood to apply to the hardware and software features required to implement or support a remediation system for threshold leakage in a data center liquid cooling system for Figure 2-5 a data center.

[0106] In at least one embodiment, as Figure 6A shown, data center infrastructure layer 610 may include a resource coordinator 612, grouped computing resources 614, and node computing resources (“node C.R.”) 616(1)-616(N), where “N” represents any whole positive integer. In at least one embodiment, node C.R. 616(1)-616(N) may include, but is not limited to, any number of central processing units (“CPU”) or other processors (including accelerators, field programmable gate arrays (FPGA), graphics processors, etc.), memory devices (such as dynamic read-only memory), storage devices (such as solid state or disk drives), network input / output (“NW I / O”) devices, network switches, virtual machines (“VM”), power modules, and cooling modules, etc. In at least one embodiment, one or more of node C.R. 616(1)-616(N) may be a server having one or more of the above computing resources.

[0107] In at least one embodiment, the grouped computing resources 614 may include separate groupings (not shown) of node C.R.s housed within one or more racks, or many racks (also not shown) within data centers at various geographical locations. The separate groupings of node C.R.s within the grouped computing resources 614 may include grouped computing, network, memory, or storage resources that may be configured or allocated to support one or more workloads. In at least one embodiment, several node C.R.s including CPUs or processors may be grouped within one or more racks to provide computing resources to support one or more workloads. In at least one embodiment, one or more racks may also include any number of power modules, cooling modules, and network switches, in any combination.

[0108] In at least one embodiment, the resource coordinator 612 may configure or otherwise control one or more node C.R.s 616(1)-616(N) and / or the grouped computing resources 614. In at least one embodiment, the resource coordinator 612 may include a software design infrastructure (“SDI”) management entity for the data center 600. In at least one embodiment, the resource coordinator may include hardware, software, or some combination thereof.

[0109] In at least one embodiment, as Figure 6AAs shown, the framework layer 620 includes a job scheduler 622, a configuration manager 624, a resource manager 626, and a distributed file system 628. In at least one embodiment, the framework layer 620 may include a framework that supports software 632 of the software layer 630 and / or one or more applications 642 of the application layer 640. In at least one embodiment, the software 632 or the application 642 may respectively include web-based service software or applications, such as services or applications provided by Amazon Web Services, Google Cloud, and Microsoft Azure. In at least one embodiment, the framework layer 620 may be, but is not limited to, a free and open-source software web application framework, such as Apache SparkTM (hereinafter referred to as "Spark") that can utilize the distributed file system 628 for large-scale data processing (e.g., "big data"). In at least one embodiment, the job scheduler 622 may include a Spark driver to facilitate scheduling of workloads supported by the various layers of the data center 600. In at least one embodiment, the configuration manager 624 may be able to configure different layers, such as the software layer 630 and the framework layer 620 including Spark and the distributed file system 628 for supporting large-scale data processing. In at least one embodiment, the resource manager 626 is capable of managing cluster or grouped computing resources mapped to or allocated for supporting the distributed file system 628 and the job scheduler 622. In at least one embodiment, the cluster or grouped computing resources may include grouped computing resources 614 on the data center infrastructure layer 610. In at least one embodiment, the resource manager 626 may coordinate with the resource coordinator 612 to manage these mapped or allocated computing resources.

[0110] In at least one embodiment, the software 632 included in the software layer 630 may include software used by at least a portion of the nodes C.R. 616(1)-616(N), the grouped computing resources 614, and / or the distributed file system 628 of the framework layer 620. One or more types of software may include, but are not limited to, Internet web search software, email virus scanning software, database software, and streaming video content software.

[0111] In at least one embodiment, one or more applications 642 included in the application layer 640 may include one or more types of applications used by at least a portion of nodes C.R. 616(1)-616(N), the packet computing resources 614, and / or the distributed file system 628 of the framework layer 620. The one or more types of applications may include, but are not limited to, any number of genomics applications, cognitive computing, and machine learning applications, including training or inference software, machine learning framework software (such as PyTorch, TensorFlow, Caffe, etc.) or other machine learning applications used in conjunction with one or more embodiments.

[0112] In at least one embodiment, any one of the configuration manager 624, the resource manager 626, and the resource coordinator 612 may implement any number and type of self-modifying actions based on any amount and type of data obtained in any technically feasible manner. In at least one embodiment, the self-modifying actions may relieve the data center operator of the data center 600 from making potentially bad configuration decisions and may avoid underutilization and / or poorly performing parts of the data center.

[0113] In at least one embodiment, the data center 600 may include tools, services, software, or other resources to train one or more machine learning models or use one or more machine learning models to predict or infer information according to one or more embodiments herein. In at least one embodiment, a machine learning model may be trained by computing weight parameters according to a neural network architecture by using the software and computing resources described above with respect to the data center 600. In at least one embodiment, by using the weight parameters calculated by one or more training techniques herein, the resources described above with respect to the data center 600 may be used to infer or predict information using the trained machine learning model corresponding to one or more neural networks. As previously mentioned, deep learning techniques may be used to support intelligent control of the controller in a remediation system for threshold leakage in a data center liquid cooling system by monitoring the regional temperature of the data center. Any suitable learning network and the computing power of the data center 600 may be used to advance deep learning. Thus, in this way, the hardware in the data center may support deep neural networks (DNNs), recurrent neural networks (RNNs), or convolutional neural networks (CNNs) simultaneously or concurrently. For example, once the network has been trained and successfully evaluated to identify data in a subset or slice, the trained network may provide similar representative data for use with the collected data.

[0114] In at least one embodiment, data center 600 may use a CPU, an application specific integrated circuit (ASIC), a GPU, an FPGA, or other hardware to perform training and / or inference using the above resources. Additionally, one or more of the above software and / or hardware resources may be configured as a service to allow users to train or perform information inference, such as pressure, flow rate, temperature, and location information, or other artificial intelligence services.

[0115] Inference and training logic

[0116] Inference and / or training logic 615 may be used to perform inference and / or training operations associated with one or more embodiments. In at least one embodiment, the inference and / or training logic 615 may be in a system Figure 6A for inferring or predicting operations based at least in part on weight parameters calculated using neural network training operations \ neural network functions and / or architectures or neural network use cases herein. In at least one embodiment, the inference and / or training logic 615 may include, but is not limited to, hardware logic in which computing resources are dedicated or otherwise uniquely used in conjunction with weight values or other information corresponding to one or more layers of neurons within a neural network. In at least one embodiment, the inference and / or training logic 615 may be used in conjunction with an application specific integrated circuit (ASIC), such as the processing unit from Google, the TM inference processing unit (IPU) from Graphcore or the (e.g., “LakeCrest”) processor from Intel Corp.

[0117] In at least one embodiment, the inference and / or training logic 615 may be used in conjunction with central processing unit (CPU) hardware, graphics processing unit (GPU) hardware, or other hardware (such as a field programmable gate array (FPGA)). In at least one embodiment, the inference and / or training logic 615 includes, but is not limited to, code and / or data storage models that may be used to store code (such as graph code), weight values, and / or other information, including bias values, gradient information, momentum values, and / or other parameter or hyperparameter information. In at least one embodiment, each code and / or data storage module is associated with dedicated computing resources. In at least one embodiment, the dedicated computing resources include computing hardware that further includes one or more ALUs that perform only mathematical functions (such as linear algebra functions) on information stored in the code and / or data storage module and store the results therefrom in the activation storage module of the inference and / or training logic 615.

[0118] Figure 6B 、 Figure 6Cillustrates inference and / or training logic according to at least one embodiment, such as the inference and / or training logic used in Figure 6A and at least one embodiment of the present disclosure. The inference and / or training logic 615 is used to perform inference and / or training operations associated with at least one embodiment. Details regarding the inference and / or training logic 615 are provided below in conjunction with Figure 6B and / or Figure 6C The inference and / or training logic 615 is distinguished by using an arithmetic logic unit (ALU) 610 with computing hardware 602, 606. Figure 6B and Figure 6C In at least one embodiment, each of the computing hardware 602 and the computing hardware 606 includes one or more ALUs, and the ALUs only perform mathematical functions (such as linear algebra functions) on the information stored in the code and / or data memory 601 and the information stored in the code and / or data memory 605 respectively, and the results are stored in the activation memory 620. Thus, unless otherwise specified, Figure 6B and Figure 6C may be alternative and may be used interchangeably.

[0119] In at least one embodiment, the inference and / or training logic 615 may include, but is not limited to, the code and / or data storage 601 to store the forward and / or output weights and / or input / output data and / or other parameters of the neurons or layers of the neural network trained and / or used for inference in at least one embodiment. In at least one embodiment, the training logic 615 may include or be coupled to the code and / or data storage 601 for storing graphical code or other software to control timing and / or sequence, where the weight and / or other parameter information is loaded to configure the logic, including integer and / or floating-point units (collectively referred to as arithmetic logic unit (ALU)). In at least one embodiment, the code (such as graph code) loads the weight or other parameter information into the processor ALU based on the architecture of the neural network corresponding to the code. In at least one embodiment, the code and / or data storage 601 stores the weight parameters and / or input / output data of each layer of the neural network combined with at least one embodiment during the forward propagation of the input / output data and / or weight parameters during training and / or inference using aspects of at least one embodiment. In at least one embodiment, any part of the code and / or data storage 601 may be included in other on-chip or off-chip data storage, including the L1, L2, or L3 cache of the processor or the system memory.

[0120] In at least one embodiment, any portion of the code and / or data storage 601 can be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, the code and / or data storage 601 can be cache memory, dynamic random-access memory (“DRAM”), static random-access memory (“SRAM”), non-volatile memory (such as flash memory), or other storage. In at least one embodiment, the choice of whether the code and / or data storage 601 is internal or external to the processor, for example, or includes DRAM, SRAM, flash memory, or some other storage type, can depend on the available storage space on or off the chip, the latency requirements for performing training and / or inference functions, the data batch size used in the inference and / or training of the neural network, or some combination of these factors.

[0121] In at least one embodiment, the inference and / or training logic 615 can include, but is not limited to, the code and / or data storage 605 to store the backward and / or output weights and / or input / output data corresponding to the neurons or layers of the neural network that are trained as and / or used for inference in at least one embodiment. In at least one embodiment, during the training and / or inference using at least one embodiment, the code and / or data storage 605 stores the weight parameters and / or input / output data of each layer of the neural network that are combined with at least one embodiment during the backpropagation of the input / output data and / or weight parameters. In at least one embodiment, the training logic 615 can include or be coupled to the code and / or data storage 605 for storing the graph code or other software to control the timing and / or sequence, where the weights and / or other parameter information are loaded to configure the logic, which includes integer and / or floating-point units (collectively referred to as arithmetic logic units (ALUs)).

[0122] In at least one embodiment, code (such as graph code) loads weight or other parameter information into the processor ALU based on the architecture of the neural network corresponding to the code. In at least one embodiment, any portion of the code and / or data storage 605 may be included with other on-chip or off-chip data storage, including the L1, L2, or L3 cache of the processor or system memory. In at least one embodiment, any portion of the code and / or data storage 605 may be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, the code and / or data storage 605 may be cache memory, DRAM, SRAM, non-volatile memory (such as flash memory), or other storage. In at least one embodiment, the choice of whether the code and / or data storage 605 is internal or external to the processor, e.g., including DRAM, SRAM, flash memory, or some other storage type, depends on whether the available storage is on-chip or off-chip, the latency requirements of the training and / or inference functions being performed, the data batch size used in the inference and / or training of the neural network, or some combination of these factors.

[0123] In at least one embodiment, the code and / or data storage 601 and the code and / or data storage 605 may be separate storage structures. In at least one embodiment, the code and / or data storage 601 and the code and / or data storage 605 may be the same storage structure. In at least one embodiment, the code and / or data storage 601 and the code and / or data storage 605 may be partially the same storage structure and partially separate storage structures. In at least one embodiment, any portion of the code and / or data storage 601 and the code and / or data storage 605 may be included with other on-chip or off-chip data storage, including the L1, L2, or L3 cache of the processor or system memory.

[0124] In at least one embodiment, the inference and / or training logic 615 can include, but is not limited to, one or more arithmetic logic units (“ALUs”) 610, which include integer and / or floating point units for performing logical and / or mathematical operations, at least in part, based on and / or as instructed by training and / or inference code (e.g., graph code), the results of which may generate activations (e.g., output values from layers or neurons within a neural network) stored in the activation storage 620, which are a function of input / output and / or weight parameter data stored in the code and / or data storage 601 and / or the code and / or data storage 605. In at least one embodiment, the activations are generated by linear algebra and / or matrix-based mathematics performed by the ALU 610 in response to executing instructions or other code, where the weight values stored in the code and / or data storage 605 and / or the code and / or data storage 601 are used as operands with other values, such as bias values, gradient information, momentum values, or other parameters or hyperparameters, any or all of which may be stored in the code and / or data storage 605 and / or the code and / or data storage 601 or other on-chip or off-chip storage.

[0125] In at least one embodiment, one or more ALUs 610 are included in one or more processors or other hardware logic devices or circuits, while in another embodiment, one or more ALUs 610 can be outside of the processor or other hardware logic device or circuit using them (e.g., a coprocessor). In at least one embodiment, one or more ALUs 610 can be included within the execution units of a processor or otherwise included in a group of ALUs accessible by the execution units of a processor, which can be within the same processor or distributed among different types of different processors (e.g., central processing units, graphics processing units, fixed function units, etc.). In at least one embodiment, the code and / or data storage 601, the code and / or data storage 605, and the activation storage 620 can be on the same processor or other hardware logic device or circuit, while in another embodiment, they can be on different processors or other hardware logic devices or circuits or some combination of the same and different processors or other hardware logic devices or circuits. In at least one embodiment, any portion of the activation storage 620 can be included with other on-chip or off-chip data storage, including the L1, L2, or L3 cache of the processor or system memory. Additionally, the inference and / or training code can be stored with other code accessible to the processor or other hardware logic or circuit and can be fetched and / or processed using the fetch, decode, schedule, execute, retire, and / or other logic circuits of the processor.

[0126] In at least one embodiment, the activation store 620 can be cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other storage. In at least one embodiment, the activation store 620 can be wholly or partially internal or external to one or more processors or other logic circuits. In at least one embodiment, whether the activation store 620 is internal or external to the processor can be selected depending on on-chip or off-chip available storage, the latency requirements for training and / or inference functions, the batch size of the data used in inferring and / or training a neural network, or some combination of these factors, e.g., or including DRAM, SRAM, flash memory, or other storage types. In at least one embodiment, Figure 6B the inference and / or training logic 615 shown in can be used in conjunction with an application specific integrated circuit (“ASIC”), such as the TM processing unit from Google, the inference processing unit (IPU) from Graphcore, or the Figure 6B processor (e.g., “Lake Crest”) from Intel Corp. In at least one embodiment,

[0127] In at least one embodiment, Figure 6C the inference and / or training logic 615 shown in accordance with at least one respective embodiment can include, but is not limited to, hardware logic where computing resources are dedicated or otherwise uniquely used in conjunction with weight values or other information corresponding to one or more layers of neurons within a neural network. In at least one embodiment, Figure 6C the inference and / or training logic 615 shown in can be used in conjunction with an application specific integrated circuit (ASIC), such as the TM processing unit from Google, the inference processing unit (IPU) from Graphcore, or the Figure 6CThe inference and / or training logic 615 shown in FIG. may be used in conjunction with central processing unit (CPU) hardware, graphics processing unit (GPU) hardware, or other hardware (such as field programmable gate arrays (FPGAs)). In at least one embodiment, the inference and / or training logic 615 includes, but is not limited to, code and / or data storage 601 and code and / or data storage 605, which may be used to store code (such as graph code), weight values, and / or other information, including bias values, gradient information, momentum values, and / or other parameter or hyperparameter information. In Figure 6C In at least one embodiment shown in FIG., each of code and / or data storage 601 and code and / or data storage 605 is respectively associated with dedicated computing resources (such as computing hardware 602 and computing hardware 606).

[0128] In at least one embodiment, each of code and / or data storage 601 and 605 and corresponding computing hardware 602 and 606 respectively corresponds to a different layer of a neural network, such that the activations obtained from one "storage / compute pair 601 / 602" of code and / or data storage 601 and computing hardware 602 are provided as inputs to the next "storage / compute pair 605 / 606" of code and / or data storage 605 and computing hardware 606, so as to reflect the conceptual organization of the neural network. In at least one embodiment, each storage / compute pair 601 / 602 and 605 / 606 may correspond to more than one neural network layer. In at least one embodiment, additional storage / compute pairs (not shown) may be included in the inference and / or training logic 615 after or in parallel with the storage compute pairs 601 / 602 and 605 / 606.

[0129] Computer system

[0130] Figure 7A FIG. shows a block diagram of an exemplary computer system 700A according to at least one embodiment. The exemplary computer system may be a system with interconnected devices and components, a system on a chip (SOC), or some combination formed with a processor, which may include execution units to execute instructions to support and / or implement intelligent control of a remedial system for threshold leakage in a data center liquid cooling system herein. In at least one embodiment, the computer system 700A according to the present disclosure (such as the embodiments herein) may include, but is not limited to, components (such as a processor 702), to execute algorithms on process data using execution units including logic. In at least one embodiment, the computer system 700A may include a processor, such as those available from Intel Corporation of Santa Clara, California processor family, XeonTM, XScaleTM, and / or StrongARMTM, Core TM or Nervana TM microprocessors, but other systems can also be used (including PCs with other microprocessors, engineering workstations, set-top boxes, etc.). In at least one embodiment, although other operating systems (such as UNIX and Linux), embedded software, and / or graphical user interfaces can also be used, the computer system 700B can execute a version of the WINDOWS operating system available from Microsoft Corporation, Redmond, Washington.

[0131] In at least one embodiment, the exemplary computer system 700A can incorporate one or more of the components 110 - 116 (from Figure 1 ) to support the processing aspects of an intelligent control for a remediation system for threshold leakage in a data center liquid cooling system. At least for this reason, in one embodiment, Figure 7A a system including interconnected hardware devices or "chips" is shown, while in other embodiments, Figure 7A an exemplary system-on-chip SoC can be shown. In at least one embodiment, Figure 7A the devices shown in Figure 6A -C can be interconnected with a proprietary interconnect, a standardized interconnect (e.g., PCIe), or some combination thereof. In at least one embodiment, one or more components of the computer system 700B use a Compute Express Link (CXL) interconnect to interconnect. Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments, e.g., as previously discussed with respect to Figure 6A -C. Details regarding the inference and / or training logic 615 are provided below in connection with Figure 7A -C. In at least one embodiment, the inference and / or training logic 615 can be used in a system

[0132] to infer or predict operations at least in part based on weight parameters calculated using neural network training operations \ neural network functions and / or architectures or neural network use cases herein.

[0133] In at least one embodiment, computer system 700A may include, but is not limited to, a processor 702, which may include, but is not limited to, one or more execution units 708 to perform machine learning model training and / or inference according to the techniques described herein. In at least one embodiment, computer system 700A is a single-processor desktop or server system, but in another embodiment, computer system 700A may be a multi-processor system. In at least one embodiment, processor 702 may include, but is not limited to, a complex instruction set computer (“CISC”) microprocessor, a reduced instruction set computing (“RISC”) microprocessor, a very long instruction word (“VLIW”) microprocessor, a processor implementing an instruction set combination, or any other processor device, such as a digital signal processor. In at least one embodiment, processor 702 may be coupled to a processor bus 710, which may transfer data signals between processor 702 and other components in computer system 700A.

[0134] In at least one embodiment, processor 702 may include, but is not limited to, a level 1 (“L1”) internal cache memory (“cache”) 704. In at least one embodiment, processor 702 may have a single internal cache or multiple levels of internal caches. In at least one embodiment, the cache memory may reside external to processor 702. Other embodiments may also include a combination of internal and external caches, depending on the particular implementation and requirements. In at least one embodiment, register file 706 may store different types of data in various registers, including, but not limited to, integer registers, floating-point registers, status registers, and instruction pointer registers.

[0135] In at least one embodiment, execution units 708, including, but not limited to, logic for performing integer and floating-point operations, are also located in processor 702. In at least one embodiment, processor 702 may also include a microcode (“ucode”) read-only memory (“ROM”) for storing microcode for certain macro instructions. In at least one embodiment, execution units 708 may include logic for processing a packed instruction set 709. In at least one embodiment, by including the packed instruction set 709 in the instruction set of a general-purpose processor, and the associated circuitry for the instructions to be executed, operations used by many multimedia applications may be performed using packed data in general-purpose processor 702. In one or more embodiments, operations may be performed on packed data by using the full width of the processor's data bus to accelerate and more efficiently execute many multimedia applications, which may not require transferring smaller data units on the processor's data bus to perform one or more operations on one data element at a time.

[0136] In at least one embodiment, execution unit 708 may also be used in a microcontroller, an embedded processor, a graphics device, a DSP, and other types of logic circuits. In at least one embodiment, computer system 700A may include, but is not limited to, memory 720. In at least one embodiment, memory 720 may be implemented as a dynamic random access memory (“DRAM”) device, a static random access memory (“SRAM”) device, a flash memory device, or other storage devices. In at least one embodiment, memory 720 may store instructions 719 and / or data 721 represented by data signals that may be executed by processor 702.

[0137] In at least one embodiment, the system logic chip may be coupled to processor bus 710 and memory 720. In at least one embodiment, the system logic chip may include, but is not limited to, a memory controller hub (“MCH”) 716, and processor 702 may communicate with MCH 716 via processor bus 710. In at least one embodiment, MCH 716 may provide a high-bandwidth memory path 718 to memory 720 for instruction and data storage, as well as for storage of graphics commands, data, and textures. In at least one embodiment, MCH 716 may initiate data signals among processor 702, memory 720, and other components in computer system 700A, and bridge data signals among processor bus 710, memory 720, and system I / O 722. In at least one embodiment, the system logic chip may provide a graphics port for coupling to a graphics controller. In at least one embodiment, MCH 716 may be coupled to memory 720 via high-bandwidth memory path 718, and graphics / video card 712 may be coupled to MCH 716 via an Accelerated Graphics Port (“AGP”) interconnect 714.

[0138] In at least one embodiment, the computer system 700A may use the system I / O 722 as a proprietary hub interface bus to couple the MCH 716 to an I / O controller hub (“ICH”) 730. In at least one embodiment, the ICH 730 may provide direct connections to certain I / O devices via a local I / O bus. In at least one embodiment, the local I / O bus may include, but is not limited to, high-speed I / O buses for connecting peripheral devices to the memory 720, chipset, and processor 702. Examples may include, but are not limited to, an audio controller 729, a firmware hub (“Flash BIOS”) 728, a wireless transceiver 726, a data storage 724, a legacy I / O controller 723 that includes user input and a keyboard interface, a serial expansion port 727 (such as a Universal Serial Bus (USB) port), and a network controller 734. The data storage 724 may include a hard disk drive, a floppy disk drive, a CD-ROM device, a flash memory device, or other mass storage devices.

[0139] Figure 7B is a block diagram showing an electronic device 700B for utilizing a processor 710 in accordance with at least one embodiment to support and / or implement intelligent control of a remedial system for threshold leakage in the data center liquid cooling system described herein. In at least one embodiment, the electronic device 700B may be, for example but not limited to, a laptop computer, a tower server, a rack server, a blade server, a notebook computer, a desktop computer, a tablet computer, a mobile device, a phone, an embedded computer, or any other suitable electronic device. In at least one embodiment, the exemplary electronic device 700B may incorporate one or more components that support processing aspects for a remedial system for threshold leakage in a data center liquid cooling system.

[0140] In at least one embodiment, the system 700B may include, but is not limited to, a processor 710 communicatively coupled to any suitable number or type of components, peripherals, modules, or devices. In at least one embodiment, the processor 710 is coupled using a bus or interface, such as an I℃ bus, a system management bus (“SMBus”), a low pin count (LPC) bus, a serial peripheral interface (“SPI”), a high definition audio (“HDA”) bus, a serial advanced technology attachment (“SATA”) bus, a Universal Serial Bus (“USB”) (versions 1, 2, 3), or a Universal Asynchronous Receiver / Transmitter (“UART”) bus. In at least one embodiment, Figure 7B shows a system that includes interconnected hardware devices or “chips,” while in other embodiments, Figure 7B an exemplary system-on-a-chip (“SoC”) may be shown. In at least one embodiment, Figure 7BThe devices shown can be interconnected with proprietary interconnects, standardized interconnects (e.g., PCIe), or some combination thereof. In at least one embodiment, Figure 7B one or more components of Figure 7B are interconnected using Compute Express Link (CXL) interconnects.

[0141] In at least one embodiment, Figure 7B it may include a display 724, a touch screen 725, a touchpad 730, a Near Field Communication unit (“NFC”) 745, a sensor hub 740, a thermal sensor 746, an Embedded Controller (“EC”) 735, a Trusted Platform Module (“TPM”) 738, BIOS / Firmware / Flash (“BIOS, FW Flash”) 722, a DSP 760, a drive 720 (e.g., a Solid State Drive (“SSD”) or a Hard Disk Drive (“HDD”)), a Wireless Local Area Network unit (“WLAN”) 750, a Bluetooth unit 752, a Wireless Wide Area Network unit (“WWAN”) 756, a Global Positioning System (GPS) unit 755, a camera (“USB 3.0 camera”) 754 (e.g., a USB 3.0 camera), and / or a Low Power Double Data Rate (“LPDDR”) memory unit (“LPDDR3”) 715 implemented in accordance with, for example, the LPDDR3 standard. These components can each be implemented in any suitable manner.

[0142] In at least one embodiment, other components can be communicatively coupled to the processor 710 through the following components. In at least one embodiment, an accelerometer 741, an Ambient Light Sensor (“ALS”) 742, a compass 743, and a gyroscope 744 can be communicatively coupled to the sensor hub 740. In at least one embodiment, a thermal sensor 739, a fan 737, a keyboard 746, and a touchpad 730 can be communicatively coupled to the EC 735. In at least one embodiment, a speaker 763, headphones 764, and a microphone (“mic”) 765 can be communicatively coupled to an audio unit (“audio codec and class-D amplifier”) 762, which in turn can be communicatively coupled to the DSP 760. In at least one embodiment, the audio unit 764 can include, for example but not limited to, an audio encoder / decoder (“codec”) and a class-D amplifier. In at least one embodiment, a Subscriber Identity Module (“SIM”) 757 can be communicatively coupled to the WWAN unit 756. In at least one embodiment, components (such as the WLAN unit 750, the Bluetooth unit 752, and the WWAN unit 756) can be implemented in a Next Generation Form Factor (NGFF).

[0143] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. As described herein in connection with Figure 6B and / or Figure 6CProvide details regarding inference and / or training logic 615. In at least one embodiment, the inference and / or training logic 615 may be used in a system Figure 7B to infer or predict operations based at least in part on weight parameters calculated using neural network training operations, neural network functions, and / or architectures or neural network use cases herein.

[0144] Figure 7C FIG. shows a computer system 700C according to at least one embodiment, which is used to support and / or implement intelligent control of a remedial system for threshold leakage in a data center liquid cooling system described herein. In at least one embodiment, the computer system 700C includes, but is not limited to, a computer 771 and a USB drive 770. In at least one embodiment, the computer 771 may include, but is not limited to, any number and type of processors (not shown) and memories (not shown). In at least one embodiment, the computer 771 includes, but is not limited to, a server, a cloud instance, a laptop computer, and a desktop computer.

[0145] In at least one embodiment, the USB drive 770 includes, but is not limited to, a processing unit 772, a USB interface 774, and USB interface logic 773. In at least one embodiment, the processing unit 772 may be any instruction execution system, apparatus, or device capable of executing instructions. In at least one embodiment, the processing unit 772 may include, but is not limited to, any number and type of processing cores (not shown). In at least one embodiment, the processing unit or core 772 includes an application specific integrated circuit (“ASIC”) that is optimized to perform any number and type of operations associated with machine learning. For example, in at least one embodiment, the processing core 772 is a tensor processing unit (“TPC”) that is optimized to perform machine learning inference operations. In at least one embodiment, the processing core 772 is a vision processing unit (“VPU”) that is optimized to perform machine vision and machine learning inference operations.

[0146] In at least one embodiment, the USB interface 774 may be any type of USB connector or USB socket. For example, in at least one embodiment, the USB interface 774 is a USB 3.0 Type-C socket for data and power. In at least one embodiment, the USB interface 774 is a USB 3.0 Type-A connector. In at least one embodiment, the USB interface logic 773 may include any number and type of logic that enables the processing unit 772 to connect to a device (such as the computer 771) via the USB connector 774.

[0147] The inference and / or training logic 615 (as described with respect to Figure 6B and Figure 6Cdescribed) for performing inference and / or training operations associated with one or more embodiments. The following is combined with Figure 6B and Figure 6C Provide details about the inference and / or training logic 615. In at least one embodiment, the inference and / or training logic 615 can be used to Figure 7C In a system of, to infer or predict operations based at least in part on weight parameters calculated using the neural network training operations, neural network functions, and / or architectures, or neural network use cases herein.

[0148] Figure 8 FIG. shows a further exemplary computer system 800 for implementing the various processes and methods of a remedial system for threshold leakage in a data center liquid cooling system described throughout this disclosure. In at least one embodiment, the computer system 800 includes, but is not limited to, at least one central processing unit (“CPU”) 802, which is connected to a communication bus 810 implemented using any suitable protocol, such as PCI (“Peripheral Component Interconnect”), Peripheral Component Interconnect Express (“PCI-Express”), AGP (“Accelerated Graphics Port”), HyperTransport, or any other bus or point-to-point communication protocol. In at least one embodiment, the computer system 800 includes, but is not limited to, a main memory 804 and control logic (e.g., implemented as hardware, software, or a combination thereof), and data can be stored in the main memory 804 in the form of random access memory (“RAM”). In at least one embodiment, a network interface subsystem (“network interface”) 822 provides an interface to other computing devices and networks for receiving data from the computer system 800 and transmitting data to other systems.

[0149] In at least one embodiment, the computer system 800 includes, but is not limited to, an input device 808, a parallel processing system 812, and a display device 806 in at least one embodiment, which can be implemented using a cathode ray tube (“CRT”), a liquid crystal display (“LCD”), a light emitting diode (“LED”), a plasma display, or other suitable display technology. In at least one embodiment, user input is received from the input device 808 (such as a keyboard, mouse, touchpad, microphone, and more). In at least one embodiment, each of the above modules can be located on a single semiconductor platform to form a processing system.

[0150] The inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments, such as previously described with respect to Figure 6A -C. The following is combined with Figure 6A-C provides details regarding inference and / or training logic 615. In at least one embodiment, the inference and / or training logic 615 can be used in a system Figure 8 to perform inference or prediction operations at least in part based on weight parameters calculated using neural network training operations, neural network functions, and / or architectures or neural network use cases herein. In at least one embodiment, the inference and / or training logic 615 can be used in a system Figure 8 to perform inference or prediction operations at least in part based on weight parameters calculated using neural network training operations, neural network functions, and / or architectures or neural network use cases herein.

[0151] Figure 9A An exemplary architecture is shown in which multiple GPUs 910 - 913 are communicatively coupled to multiple multi - core processors 940 - 943 via high - speed links 905 - 906 (e.g., bus / peer - to - peer interconnects, etc.). In one embodiment, the high - speed links 940 - 943 support a communication throughput of 4GB / s, 30GB / s, 80GB / s, or higher. Various interconnect protocols can be used, including but not limited to PCIe 4.0 or 5.0 and NVLink 2.0.

[0152] Furthermore, in one embodiment, two or more GPUs 910 - 913 are interconnected via high - speed links 929 - 930, which can be implemented using the same or different protocols / links as those used for high - speed links 940 - 943. Similarly, two or more multi - core processors 905 - 906 can be connected via high - speed link 928, which can be a symmetric multi - processor (SMP) bus operating at a speed of 20GB / s, 30GB / s, 120GB / s, or higher. Alternatively, the same protocol / link (e.g., via a common interconnect structure) can be used to accomplish Figure 9A all communications between the various system components shown.

[0153] In one embodiment, each multi-core processor 905-906 is communicatively coupled to processor memories 901-902 via memory interconnects 926-927 respectively, and each GPU 910-913 is communicatively coupled to GPU memories 920-923 via GPU memory interconnects 950-953 respectively. The memory interconnects 926-927 and 950-953 may utilize the same or different memory access technologies. By way of example and not limitation, the processor memories 901-902 and GPU memories 920-923 may be volatile memories such as dynamic random access memory (DRAM) (including stacked DRAM), graphics DDR SDRAM (GDDR) (e.g., GDDR5, GDDR6), or high bandwidth memory (HBM), and / or may be non-volatile memories such as 3D XPoint or Nano-Ram. In one embodiment, some portions of the processor memories 901-902 may be volatile memories while other portions may be non-volatile memories (e.g., using a two-level memory (2LM) hierarchy).

[0154] As follows, although the various processors 905-906 and GPUs 910-913 may be physically coupled to specific memories 901-902, 920-923 respectively, a unified memory architecture may be implemented where the virtual system address space (also referred to as the "effective address" space) is distributed among the various physical memories. In at least one embodiment, each of the processor memories 901-902 may include 64GB of system memory address space, and each of the GPU memories 920-923 may include 32GB of system memory address space (resulting in a total addressable memory size of 256GB in this example).

[0155] As discussed elsewhere in this disclosure, a flow rate and an associated temperature may be established at least for a first level of an intelligent learning system (e.g., a neural network system). Since the first level represents prior data, it also represents a smaller subset of data, which can be used to improve the system by retraining the system. Multiple processor units may be used to perform testing and training in parallel such that the intelligent learning system is robust. An architecture such as Figure 9A may be used. When convergence is achieved for the intelligent learning system, the number of data points recorded and the data in the data points used to cause the convergence are recorded. The data and data points may be used entirely to control a remedial system for threshold leakage in a data center liquid cooling system as referred to, for example, in Figure 2-5 as discussed.

[0156] Figure 9BAdditional details of the interconnection between the multi-core processor 907 and the graphics acceleration module 946 according to an exemplary embodiment are shown. The graphics acceleration module 946 may include one or more GPU chips integrated on a line card that is coupled to the processor 907 via a high-speed link 940. Optionally, the graphics acceleration module 946 may be integrated with the processor 907 in the same package or chip.

[0157] In at least one embodiment, the illustrated processor 907 includes a plurality of cores 960A - 960D, each core having a translation lookaside buffer 961A - 961D and one or more caches 962A - 962D. In at least one embodiment, the cores 960A - 960D may include various other components (not shown) for executing instructions and processing data. The caches 962A - 962D may include level 1 (L1) and level 2 (L2) caches. Additionally, one or more shared caches 956 may be included within the caches 962A - 962D and shared by groups of cores 960A - 960D. In at least one embodiment, one embodiment of the processor 907 includes 24 cores, each core having its own L1 cache, twelve shared L2 caches, and twelve shared L3 caches. In this embodiment, two adjacent cores share one or more L2 and L3 caches. The processor 907 and the graphics acceleration module 946 are connected to the system memory 914, which may include Figure 9A processor memories 901 - 902 therein.

[0158] Coherence for data and instructions stored in the respective caches 962A - 962D, 956, and the system memory 914 is maintained via the coherence bus 964 for inter-core communication. In at least one embodiment, each cache may have cache coherence logic / circuit associated therewith to communicate via the coherence bus 964 in response to detecting a read or write to a particular cache line. In one implementation, a cache snooping protocol is implemented via the coherence bus 964 to snoop cache accesses.

[0159] In at least one embodiment, the proxy circuit 925 communicatively couples the graphics acceleration module 946 to the coherence bus 964, thereby allowing the graphics acceleration module 946 to participate in the cache coherence protocol as a peer of the cores 960A - 960D. In particular, in at least one embodiment, the interface 935 provides a connection to the proxy circuit 925 via the high-speed link 940 (e.g., PCIe bus, NVLink, etc.), and the interface 937 connects the graphics acceleration module 946 to the link 940.

[0160] In one implementation, accelerator integrated circuit 936 represents multiple graphics processing engines 931, 932, N of a graphics acceleration module that provide cache management, memory access, context management, and interrupt management services. Graphics processing engines 931, 932, N may each include a separate graphics processing unit (GPU). In at least one embodiment, graphics processing engines 931, 932, N may selectively include different types of graphics processing engines within a GPU, such as graphics execution units, media processing engines (e.g., video encoder / decoder), samplers, and blit engines. In at least one embodiment, graphics acceleration module 946 may be a GPU having multiple graphics processing engines 931-932, N, or graphics processing engines 931-932, N may be individual GPUs integrated on a common package, line card, or chip. Depending on the circumstances, the determination of the reconstruction parameters and reconstruction algorithm described above may be performed in Figure 9B the GPUs 931-N.

[0161] In one embodiment, accelerator integrated circuit 936 includes a memory management unit (MMU) 939 for performing various memory management functions, such as virtual-to-physical memory translation (also known as effective-to-real memory translation), and also includes a memory access protocol for accessing system memory 914. MMU 939 may also include a translation lookaside buffer (“TLB”) (not shown) for caching virtual / effective-to-physical / real address translations. In one implementation, cache 938 may store commands and data for efficient access by graphics processing engines 931-932, N. In at least one embodiment, data stored in cache 938 and graphics memories 933-934, M may be kept consistent with core caches 962A-962D, 956, and system memory 914. As before, this task may be accomplished via proxy circuitry 925 representing cache 938 and graphics memories 933-934, M (e.g., sending updates related to modifications / accesses of cache lines on processor caches 962A-962D, 956 to cache 938 and receiving updates from cache 938).

[0162] A set of registers 945 stores context data of threads executed by the graphics processing engines 931-932,N, and the context management circuitry 948 manages thread contexts. In at least one embodiment, the context management circuitry 948 may perform save and restore operations to save and restore the contexts of individual threads during a context switch (e.g., where the first thread is saved and the second thread is stored so that the second thread can be executed by the graphics processing engine). In at least one embodiment, the context management circuitry 948, upon a context switch, may store the current register values into a specified area in memory (e.g., identified by a context pointer). The register values can then be restored when the context is returned. In one embodiment, the interrupt management circuitry 947 receives and processes interrupts received from system devices.

[0163] In one implementation, the MMU 939 translates virtual / valid addresses from the graphics processing engine 931 into real / physical addresses in the system memory 914. One embodiment of the accelerator integrated circuit 936 supports multiple (e.g., 4, 8, 16) graphics accelerator modules 946 and / or other accelerator devices. The graphics accelerator modules 946 may be dedicated to a single application executing on the processor 907 or may be shared among multiple applications. In one embodiment, a virtualized graphics execution environment is presented where the resources of the graphics processing engines 931-932,N are shared among multiple applications or virtual machines (VMs). In at least one embodiment, the resources may be subdivided into "slices" based on processing requirements and priorities associated with the VMs and / or applications, and these slices are allocated to different VMs and / or applications.

[0164] In at least one embodiment, the accelerator integrated circuit 936 acts as a bridge for the system of graphics accelerator modules 946 and provides address translation and system memory cache services. Additionally, the accelerator integrated circuit 936 may provide virtualization facilities for the host processor to manage the virtualization, interrupts, and memory management of the graphics processing engines 931-932,N.

[0165] Since the hardware resources of the graphics processing engines 931-932,N are explicitly mapped to the real address space seen by the host processor 907, any host processor can directly address these resources using valid address values. In at least one embodiment, one function of the accelerator integrated circuit 936 is to physically isolate the graphics processing engines 931-932,N such that they appear as independent units to the system.

[0166] In at least one embodiment, one or more graphics memories 933 - 934, M are respectively coupled to each graphics processing engine 931 - 932, N. The graphics memories 933 - 934, M store instructions and data that are processed by each graphics processing engine 931 - 932, N. The graphics memories 933 - 934, M can be volatile memories such as DRAM (including stacked DRAM), GDDR memories (e.g., GDDR5, GDDR6) or HBM, and / or can be non - volatile memories such as 3DXPoint or Nano - Ram.

[0167] In one embodiment, to reduce data traffic on link 940, a biasing technique is used to ensure that the data stored in the graphics memories 933 - 934, M is the data most frequently used by the graphics processing engines 931 - 932, N, and data that the cores 960A - 960D may not use (at least not frequently). Similarly, the biasing mechanism attempts to keep the data required by the cores (and which may not be the graphics processing engines 931 - 932, N) in caches 962A - 962D, core 956, and system memory 914.

[0168] Figure 9C Another exemplary embodiment is shown, where accelerator integrated circuit 936 integrates in processor 907 a processor 907 for implementing and / or supporting intelligent control of a remedial system for threshold leakage in a data center liquid cooling system according to at least one embodiment disclosed herein. In at least this embodiment, the graphics processing engines 931 - 932, N communicate directly with the accelerator integrated circuit 936 via interface 937 and interface 935 (again, any form of bus or interface protocol can be utilized) over high - speed link 940. The accelerator integrated circuit 936 can perform operations similar to those Figure 9B described. However, due to its close proximity to coherence bus 964 and caches 962A - 962D, 956, it may have higher throughput. At least one embodiment supports different programming models, including a dedicated process programming model (without graphics acceleration module virtualization) and a shared programming model (with virtualization), and the programming models can include a programming model controlled by the accelerator integrated circuit 936 and a programming model controlled by the graphics acceleration module 946.

[0169] In at least one embodiment, the graphics processing engines 931 - 932, N are dedicated to a single application or process under a single operating system. In at least one embodiment, a single application can funnel requests from other applications to the graphics processing engines 931 - 932, N, thereby providing virtualization within a VM / partition.

[0170] In at least one embodiment, the graphics processing engines 931-932,N can be shared by multiple VM / application partitions. In at least one embodiment, the sharing model can use a hypervisor to virtualize the graphics processing engines 931-932,N to allow each operating system to access them. For a single-partition system without a hypervisor, the operating system owns the graphics processing engines 931-932,N. In at least one embodiment, the operating system can virtualize the graphics processing engines 931-932,N to provide access to each process or application.

[0171] In at least one embodiment, the graphics acceleration module 946 or individual graphics processing engines 931-932,N use a process handle to select a process element. In at least one embodiment, the process element is stored in the system memory 914 and can be addressed using the effective address to real address translation techniques described herein. In at least one embodiment, the process handle can be an implementation-specific value that is provided to the host process when registering its context with the graphics processing engines 931-932,N (i.e., calling the system software to add the process element to the process element linked list). In at least one embodiment, the lower 16 bits of the process handle can be the offset of the process element in the process element linked list.

[0172] Figure 9D An exemplary accelerator integrated die 990 is shown for implementing and / or supporting intelligent control of a remedial system for threshold leakage in a data center liquid cooling system in accordance with at least one embodiment disclosed herein. As used herein, a "slice" includes a designated portion of the processing resources of the accelerator integrated circuit 936. An application is an effective address space 982 in the system memory 914 that stores process elements 983. In at least one embodiment, the process elements 983 are stored in response to a GPU call 981 from an application 980 executing on the processor 907. The process element 983 contains the process state of the corresponding application 980. The work descriptor (WD) 984 contained in the process element 983 can be a single job requested by the application or can contain a pointer to a job queue. In at least one embodiment, the WD 984 is a pointer to a job request queue in the address space 982 of the application.

[0173] The graphics acceleration module 946 and / or individual graphics processing engines 931-932,N can be shared by all processes or a subset of processes in the system. In at least one embodiment, an infrastructure can be included for setting the process state and sending the WD 984 to the graphics acceleration module 946 to start a job in a virtualized environment.

[0174] In at least one embodiment, the dedicated process programming model is implementation - specific. In this model, a single process owns the graphics acceleration module 946 or an individual graphics processing engine 931. When the graphics acceleration module 946 is owned by a single process, the hypervisor initializes the accelerator integrated circuit for the owned partition, and when the graphics acceleration module 946 is assigned, the operating system initializes the accelerator integrated circuit 936 for the owned process.

[0175] In operation, the WD fetch unit 991 in the accelerator integrated slice 990 fetches the next WD 984, which includes an indication of work to be completed by one or more graphics processing engines of the graphics acceleration module 946. Data from the WD 984 can be stored in the register 945 and used by the MMU 939, the interrupt management circuit 947, and / or the context management circuit 948, as shown. In at least one embodiment, one embodiment of the MMU 939 includes a segment / page walk circuit for accessing the segment / page table 986 within the OS virtual address space 985. The interrupt management circuit 947 can process the interrupt event 992 received from the graphics acceleration module 946. In at least one embodiment, when a graphics operation is executed, the effective address 993 generated by the graphics processing engines 931 - 932, N is translated to a real address by the MMU 939.

[0176] In one embodiment, the same register set 945 is replicated for each graphics processing engine 931 - 932, N and / or the graphics acceleration module 946, and the same register set 945 can be initialized by the hypervisor or the operating system. Each of these replicated registers can be included in the accelerator integrated slice 990. Exemplary registers that can be initialized by the hypervisor are shown in Table 1.

[0177] Table 1 – Hypervisor - Initialized Registers

[0178] 1 Slice Control Register 2 Real Address (RA) Scheduler Process Area Pointer 3 Privilege Mask Override Register 4 Interrupt Vector Table Entry Offset 5 Interrupt Vector Table Entry Limit 6 Status Register 7 Logical Partition Identifier 8 Real Address (RA) Hypervisor Accelerator Utilization Record Pointer 9 Storage Descriptor Register

[0179] Exemplary registers that can be initialized by the operating system are shown in Table 2.

[0180] Table 2 – Operating - System - Initialized Registers

[0181] 1 Process and Thread Identification 2 Effective Address (EA) Context Save / restore Pointer 3 Virtual Address (VA) Accelerator Utilization Record Pointer 4 Virtual Address (VA) Storage Segment Table Pointer 5 Privilege Mask 6 Work Descriptor

[0182] In at least one embodiment, each WD 984 is specific to a particular graphics acceleration module 946 and / or graphics processing engines 931 - 932, N. It contains all the information required for the graphics processing engines 931 - 932, N to complete the work, or it can be a pointer to a memory location where the application has set up a command queue for the work to be completed.

[0183] Figure 9E Additional details of an exemplary embodiment of the shared model are shown. This embodiment includes a hypervisor real address space 998 in which a list of process elements 999 is stored. The hypervisor real address space 998 can be accessed via a hypervisor 996 that virtualizes a graphics acceleration module engine for an operating system 995.

[0184] In at least one embodiment, the shared programming model allows all processes or a subset of processes from all partitions or a subset of partitions in the system to use the graphics acceleration module 946. There are two programming models in which the graphics acceleration module 946 is shared by multiple processes and partitions, namely, time slice sharing and graphics directed sharing.

[0185] In this model, the system hypervisor 996 owns the graphics acceleration module 946 and makes its functionality available to all operating systems 995. For the graphics acceleration module 946 to support virtualization through the system hypervisor 996, the graphics acceleration module 946 may comply with the following requirements: (1) the job requests of the application must be autonomous (i.e., do not need to maintain state between jobs), or the graphics acceleration module 946 must provide a context save and restore mechanism, (2) the graphics acceleration module 946 guarantees that the job requests of the application are completed within a specified amount of time, including any translation errors, or the graphics acceleration module 946 provides the ability to preempt job processing, and (3) fairness between processes of the graphics acceleration module 946 must be ensured when operating in the directed sharing programming model.

[0186] In one embodiment, application 980 needs to use a graphics acceleration module type, a work descriptor (WD), a permission mask register (AMR) value, and a context save / restore area pointer (CSRP) to make an operating system 995 system call. The graphics acceleration module type describes the target acceleration function for the system call. In at least one embodiment, the graphics acceleration module type can be a system-specific value. In at least one embodiment, the WD is formatted specifically for the graphics acceleration module 946 and can take the form of a graphics acceleration module 946 command, a valid address pointer to a user-defined structure, a valid address pointer to a command queue, or any other data structure that describes the work to be done by the graphics acceleration module 946. In at least one embodiment, the AMR value is the AMR state for the current process. In at least one embodiment, the value passed to the operating system is similar to the application that sets the AMR. If the implementation of the accelerator integrated circuit 936 and the graphics acceleration module 946 does not support the user authority mask override register (UAMOR), the operating system can apply the current UAMOR value to the AMR value before passing the AMR in the hypervisor call. In at least one embodiment, the hypervisor 996 can apply the current authority mask override register (AMOR) value before putting the AMR into the process element 983. In at least one embodiment, the CSRP is one of the registers 945 that contains the valid address of a region in the valid address space 982 of the application for the graphics acceleration module 946 to save and restore the context state. This pointer is used in at least one embodiment and is optional if there is no need to save state between jobs or when a job is preempted. In at least one embodiment, the context save / restore area can be a fixed system memory.

[0187] Upon receiving the system call, the operating system 995 can verify that the application 980 is registered and has been granted permission to use the graphics acceleration module 946. Then, the operating system 995 uses the information shown in Table 3 to call the hypervisor 996.

[0188] Table 3 – OS to Hypervisor Call Parameters

[0189] 1 Work Descriptor (WD) 2 Privilege Mask Register (AMR) Value (Potentially Masked) 3 Effective Address (EA) Context Save / restore Area Pointer (CSRP) 4 Process ID (PID) and Optional Thread ID (TID) 5 Virtual Address (VA) Accelerator Usage Record Pointer (AURP) 6 Virtual Address of Storage Segment Table Pointer (SSTP) 7 Logical Interrupt Service Number (LISN)

[0190] Upon receiving the hypervisor call, the hypervisor 996 verifies that the operating system 995 is registered and has been granted permission to use the graphics acceleration module 946. Then, the hypervisor 996 places the process element 983 into the process element linked list of the corresponding graphics acceleration module 946 type. The process element can include the information shown in Table 4.

[0191] Table 4 – Process Element Information

[0192] 1 Work Descriptor (WD) 2 Privilege Mask Register (AMR) Value (Potentially Masked). 3 Effective Address (EA) Context Save / restore Area Pointer (CSRP) 4 Process ID (PID) and Optional Thread ID (TID) 5 Virtual Address (VA) Accelerator Usage Record Pointer (AURP) 6 Virtual Address of Storage Segment Table Pointer (SSTP) 7 Logical Interrupt Service Number (LISN) 8 Interrupt Vector Table, Derived from Hypervisor Call Parameters 9 Status Register (SR) Value 10 Logical Partition ID (LPID) 11 Real Address (RA) Hypervisor Accelerator Usage Record Pointer 12 Storage Descriptor Register (SDR)

[0193] In at least one embodiment, the hypervisor initializes a plurality of accelerator integrated slice 990 registers 945.

[0194] As Figure 9F shown, in at least one embodiment, a unified memory is used, and the unified memory can be addressed via a common virtual memory address space for accessing the physical processor memories 901 - 902 and the GPU memories 920 - 923. In this implementation, operations executed on the GPUs 910 - 913 utilize the same virtual / effective memory address space to access the processor memories 901 - 902, and vice versa, thus simplifying programmability. In one embodiment, a first portion of the virtual / effective address space is allocated to the processor memory 901, a second portion is allocated to the second processor memory 902, a third portion is allocated to the GPU memory 920, and so on. In at least one embodiment, the entire virtual / effective memory space (sometimes referred to as the effective address space) is thus distributed among each of the processor memories 901 - 902 and the GPU memories 920 - 923, allowing any processor or GPU to access the memory using a virtual address mapped to any physical memory.

[0195] In one embodiment, the bias / coherency management circuits 994A - 994E within one or more MMUs 939A - 939E ensure cache coherency between one or more host processors (e.g., host processor 905) and the caches of the GPUs 910 - 913, and implement a biasing technique for physical memory indicating where certain types of data should be stored. Although Figure 9F multiple instances of the bias / coherency management circuits 994A - 994E are shown, the bias / coherency circuits can be implemented within the MMU of one or more host processors 905 and / or within the accelerator integrated circuit 936.

[0196] One embodiment allows the GPU attached memories 920 - 923 to be mapped as part of the system memory and accessed using shared virtual memory (SVM) techniques without suffering the performance penalties associated with full system cache coherence. In at least one embodiment, the ability to access the GPU attached memories 920 - 923 as system memory without heavy cache coherence overhead provides a favorable operating environment for GPU offloading. This arrangement allows the host processor 905 software to set operands and access computation results without the overhead of traditional I / O DMA data copies. Such traditional copies include driver calls, interrupts, and memory mapped I / O (MMIO) accesses, which are all less efficient than simple memory accesses. In at least one embodiment, the ability to access the GPU attached memories 920 - 923 without cache coherence overhead can be critical to the execution time of offloaded computations. For example, in the presence of a large amount of streaming write memory traffic, the cache coherence overhead can significantly reduce the effective write bandwidth seen by the GPUs 910 - 913. In at least one embodiment, the efficiency of operand setting, result access, and GPU computation can play a role in determining the effectiveness of GPU offloading.

[0197] In at least one embodiment, the selection of GPU bias and host processor bias is driven by a bias tracker data structure. For example, a bias table can be used, which can be a page granularity structure (e.g., controlled at the granularity of memory pages), and the page granularity structure includes 1 or 2 bits per GPU attached memory page. In at least one embodiment, the bias table can be implemented in the stolen memory ranges of one or more GPU attached memories 920 - 923 with or without a bias cache (e.g., for caching frequently / most recently used entries of the bias table) in the GPUs 910 - 913. Alternatively, the entire bias table can be maintained within the GPU.

[0198] In at least one embodiment, before actually accessing the GPU memory, the bias table entry associated with each access to the GPU attached memory 920-923 is accessed, thereby causing the following operations. Local requests from the GPUs 910-913 that find their pages in the GPU bias are directly forwarded to the corresponding GPU memories 920-923. Local requests from the GPUs that find their pages in the host bias are forwarded to the processor 905 (e.g., via the high-speed link as above). In one embodiment, a request from the processor 905 that finds the requested page in the host processor bias completes a request similar to a normal memory read. Alternatively, a request pointing to a GPU bias page can be forwarded to the GPUs 910-913. In at least one embodiment, if the GPU is not currently using a page, the GPU can subsequently migrate the page to the host processor bias. In at least one embodiment, the bias state of a page can be changed by a software-based mechanism, a software mechanism assisted by hardware, or, in limited cases, by a purely hardware-based mechanism.

[0199] A mechanism for changing the bias state employs an API call (e.g., OpenCL), which then calls the device driver of the GPU. The device driver then sends a message (or enqueues a command descriptor) to the GPU, instructing the GPU to change the bias state and perform a cache flush operation in the host in some migrations. In at least one embodiment, the cache flush operation is used for migrations from the host processor 905 bias to the GPU bias, but not for the reverse migration.

[0200] In one embodiment, cache coherence is maintained by temporarily rendering GPU bias pages that cannot be cached by the host processor 905. To access these pages, the processor 905 can request access from the GPU 910, and the GPU 910 may or may not immediately grant the access. Therefore, to reduce the communication between the processor 905 and the GPU 910, it is beneficial to ensure that the GPU bias pages are the pages required by the GPU rather than the host processor 905, and vice versa.

[0201] The inference and / or training logic 615 is used to execute one or more embodiments. Details regarding the inference and / or training logic 615 may be provided hereinafter in conjunction with Figure 6B and / or Figure 6C provide details about the inference and / or training logic 615.

[0202] Figure 10AIllustrated is an exemplary integrated circuit and associated graphics processor in accordance with various embodiments herein, which may be fabricated using one or more IP cores to support and / or implement a remedy system for threshold leakage in a data center liquid cooling system as described herein. In addition to the illustration, other logic and circuitry may be included in at least one embodiment, including additional graphics processors / cores, peripheral interface controllers, or general-purpose processor cores.

[0203] Figure 10A FIG. 4 is a block diagram illustrating an exemplary system on a chip integrated circuit 1000A that may be fabricated using one or more IP cores. In at least one embodiment, integrated circuit 1000A includes one or more application processors 1005 (e.g., CPUs), at least one graphics processor 1010, and may additionally include an image processor 1015 and / or a video processor 1020, any of which may be a modular IP core. In at least one embodiment, integrated circuit 1000A includes peripheral or bus logic, which includes a USB controller 1025, a UART controller 1030, an SPI / SDIO controller 1035, and an I 2 S / I 2 2 C controller 1040. In at least one embodiment, integrated circuit 1000A may include a display device 1045 coupled to one or more of a high-definition multimedia interface (HDMI) controller 1050 and a mobile industry processor interface (MIPI) display interface 1055. In at least one embodiment, storage may be provided by a flash memory subsystem 1060, including flash memory and a flash memory controller. In at least one embodiment, a memory interface may be provided via a memory controller 1065 for accessing SDRAM or SRAM memory devices. In at least one embodiment, some integrated circuits also include an embedded security engine 1070.

[0204] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 615 are provided below in conjunction with Figure 6B and / or Figure 6C FIGS. 6A-6D. In at least one embodiment, inference and / or training logic 615 may be in integrated circuit 1000A to infer or predict operations based at least in part on weight parameters calculated using neural network training operations, neural network functions, and / or architectures or neural network use cases described herein.

[0205] Figures 10B-10CAn exemplary integrated circuit and associated graphics processor in accordance with various embodiments of the present disclosure are shown, which may be fabricated using one or more IP cores to support and / or implement a system for remedying threshold leakage in a data center liquid cooling system. In addition to the illustration, other logic and circuitry may be included in at least one embodiment, including additional graphics processor(s) / core(s), peripheral interface controllers, or general purpose processor cores.

[0206] Figures 10B-10C FIG. is a block diagram showing an exemplary graphics processor used within a SoC in accordance with embodiments described herein to support and / or implement a system for remedying threshold leakage in a data center liquid cooling system as described herein. In an example, the graphics processor may be used for intelligent control of a system for remedying threshold leakage in a data center liquid cooling system because existing math engines are capable of processing multi-level neural networks more quickly. Figure 10B FIG. shows an exemplary graphics processor 1010 of a system-on-chip integrated circuit in accordance with at least one embodiment, which may be fabricated using one or more IP cores. Figure 10C FIG. shows another exemplary graphics processor 1040 of a system-on-chip integrated circuit in accordance with at least one embodiment, which may be fabricated using one or more IP cores. In at least one embodiment, Figure 10B the graphics processor 1010 is a low-power graphics processor core. In at least one embodiment, Figure 10C the graphics processor 1040 is a higher-performance graphics processor core. In at least one embodiment, each of the graphics processors 1010, 1040 may be Figure 10A a variant of the graphics processor 1010.

[0207] In at least one embodiment, the graphics processor 1010 includes a vertex processor 1005 and one or more fragment processors 1015A - 1015N (e.g., 1015A, 1015B, 1015C, 1015D to 1015N - 1, and 1015N). In at least one embodiment, the graphics processor 1010 can execute different shader programs via separate logic such that the vertex processor 1005 is optimized to perform operations for vertex shader programs, while the one or more fragment processors 1015A - 1015N perform fragment (e.g., pixel) shading operations for fragment or pixel or shader programs. In at least one embodiment, the vertex processor 1005 executes the vertex processing stage of the 3D graphics pipeline and generates primitives and vertex data. In at least one embodiment, the one or more fragment processors 1015A - 1015N use the primitives and vertex data generated by the vertex processor 1005 to generate a frame buffer to be displayed on a display device. In at least one embodiment, the one or more fragment processors 1015A - 1015N are optimized to execute fragment shader programs as provided in the OpenGL API, which can be used to perform operations similar to pixel shader programs provided in the Direct 3D API.

[0208] In at least one embodiment, the graphics processor 1010 additionally includes one or more memory management units (MMUs) 1020A - 1020B, one or more caches 1025A - 1025B, and one or more circuit interconnects 1030A - 1030B. In at least one embodiment, the one or more MMUs 1020A - 1020B provide a virtual - to - physical address mapping for the graphics processor 1010, including for the vertex processor 1005 and / or the fragment processors 1015A - 1015N, which may reference vertex or image / texture data stored in memory, in addition to vertex or image / texture data stored in the one or more caches 1025A - 1025B. In at least one embodiment, the one or more MMUs 1020A - 1020B can be synchronized with other MMUs within the system, including one or more MMUs associated with Figure 10A one or more application processors 1005, image processors 1015, and / or video processors 1020 such that each processor 1005 - 1020 can participate in a shared or unified virtual memory system. In at least one embodiment, the one or more circuit interconnects 1030A - 1030B enable the graphics processor 1010 to connect to other IP cores within the SoC via the internal bus of the SoC or via a direct connection.

[0209] In at least one embodiment, the graphics processor 1040 includes Figure 10AOne or more MMUs 1020A - 1020B, caches 1025A - 1025B, and circuit interconnections 1030A - 1030B of the graphics processor 1010. In at least one embodiment, the graphics processor 1040 includes one or more shader cores 1055A - 1055N (e.g., 1055A, 1055B, 1055C, 1055D, 1055E, 1055F to 1055N - 1, and 1055N), as Figure 10B shown, which provides a unified shader core architecture where a single core or type of core can execute all types of programmable shader code, including shader program code for implementing vertex shaders, fragment shaders, and / or compute shaders. In at least one embodiment, the multiple shader cores can vary. In at least one embodiment, the graphics processor 1040 includes an inter - core task manager 1045, which acts as a thread dispatcher to dispatch execution threads to one or more shader cores 1055A - 1055N and a tiling unit 1058 to accelerate tiling operations for tile - based rendering, where the rendering operation of the scene is subdivided in the image space, e.g., to utilize local spatial coherence within the scene or optimize the use of internal caches.

[0210] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. Details regarding the inference and / or training logic 615 are provided herein in connection with Figure 6B and / or Figure 6C In at least one embodiment, the inference and / or training logic 615 can be in an integrated circuit Figure 10A and / or Figure 10B for performing inference or prediction operations based at least in part on weight parameters calculated using neural network training operations, neural network functions or architectures, or neural network use cases herein.

[0211] Figures 10D - 10E Shows additional exemplary graphics processor logic according to embodiments described herein to support and / or implement a remediation system for threshold leakage in a data center liquid cooling system as described herein. In at least one embodiment, Figure 10D Shows a graphics core 1000D that can be included within Figure 10A the graphics processor 1010, and in at least one embodiment, it can be a unified shader core 1055A - 1055N as Figure 10C shown. Figure 10B Shows a highly parallel general - purpose graphics processing unit (“GPGPU”) 1030 suitable for deployment on a multi - chip module in at least one embodiment.

[0212] In at least one embodiment, the graphics core 1000D may include a plurality of slices 1001A - 1001N or partitions per core, and the graphics processor may include multiple instances of the graphics core 1000D. In at least one embodiment, the slices 1001A - 1001N may include support logic, which includes local instruction caches 1004A - 1004N, thread schedulers 1006A - 1006N, thread dispatchers 1008A - 1008N, and a set of registers 1010A - 1010N. In at least one embodiment, the slices 1001A - 1001N may include a set of additional functional units (AFU 1012A - 1012N), floating - point units (FPU 1014A - 1014N), integer arithmetic logic units (ALU 109A - 109N), address calculation units (ACU 1013A - 1013N), double - precision floating - point units (DPFPU 1015A - 1015N), and matrix processing units (MPU 1017A - 1017N).

[0213] In at least one embodiment, the FPU 1014A - 1014N may perform single - precision (32 - bit) and half - precision (16 - bit) floating - point operations, while the DPFPU 1015A - 1015N performs double - precision (64 - bit) floating - point operations. In at least one embodiment, the ALU1016A - 1016N may perform variable - precision integer operations in 8 - bit, 16 - bit, and 32 - bit precision and may be configured for mixed - precision operations. In at least one embodiment, the MPU 1017A - 1017N may also be configured for mixed - precision matrix operations, including half - precision floating - point operations and 8 - bit integer operations. In at least one embodiment, the MPU 1017A - 1010N may perform various matrix operations to accelerate machine - learning application frameworks, including enabling accelerated general matrix - to - matrix multiplication (GEMM). In at least one embodiment, the AFU 1012A - 1012N may perform additional logic operations not supported by the floating - point or integer units, including trigonometric operations (e.g., sine, cosine, etc.).

[0214] As discussed elsewhere in this disclosure, the inference and / or training logic 615 (referred to at least in Figure 6B 、 Figure 6C is used to perform inference and / or training operations associated with one or more embodiments. Details regarding the inference and / or training logic 615 are provided below in conjunction with Figure 6B and / or Figure 6C In at least one embodiment, the inference and / or training logic 615 may be used in the graphics core 1000D to infer or predict operations based at least in part on weight parameters calculated using neural - network training operations, neural - network functions, and / or architectures or neural - network use cases herein.

[0215] Figure 11A FIG. 3 shows a block diagram of a computer system 1100A according to at least one embodiment. In at least one embodiment, computer system 1100A includes a processing subsystem 1101 having one or more processors 1102 and a system memory 1104, and the system memory 1104 communicates via an interconnect path that may include a memory hub 1105. In at least one embodiment, the memory hub 1105 may be a separate component within a chipset component or may be integrated within one or more processors 1102. In at least one embodiment, the memory hub 1105 is coupled to an I / O subsystem 1111 via a communication link 1106. In one embodiment, the I / O subsystem 1111 includes an I / O hub 1107, and the I / O hub may enable computer system 1100A to receive input from one or more input devices 1108. In at least one embodiment, the I / O hub 1107 may enable a display controller to provide output to one or more display devices 1110A, and the display controller may be included within one or more processors 1102. In at least one embodiment, one or more display devices 1110A coupled to the I / O hub 1107 may include local, internal, or embedded display devices.

[0216] In at least one embodiment, the processing subsystem 1101 includes one or more parallel processors 1112 coupled to the memory hub 1105 via a bus or other communication link 1113. In at least one embodiment, the communication link 1113 may use any one of a number of standard-based communication link technologies or protocols, such as but not limited to PCI Express, or may be a vendor-specific communication interface or communication fabric. In at least one embodiment, one or more parallel processors 1112 form a parallel or vector processing system in a computing center, and the system may include a large number of processing cores and / or processing clusters, such as a many integrated core (MIC) processor. In at least one embodiment, one or more parallel processors 1112 form a graphics processing subsystem, and the graphics processing subsystem may output pixels to one of the one or more display devices 1110A coupled via the I / O hub 1107. In at least one embodiment, one or more parallel processors 1112 may also include a display controller and a display interface (not shown) to enable direct connection to one or more display devices 1110B.

[0217] In at least one embodiment, the system storage unit 1114 may be connected to the I / O hub 1107 to provide a storage mechanism for the computer system 1100A. In at least one embodiment, the I / O switch 1116 may be used to provide an interface mechanism to enable connections between the I / O hub 1107 and other components, such as network adapter 1118 and / or wireless network adapter 1119 that may be integrated into one or more platforms, and various other devices that may be added via one or more additional devices 1120. In at least one embodiment, the network adapter 1118 may be an Ethernet adapter or another wired network adapter. In at least one embodiment, the wireless network adapter 1119 may include one or more of Wi-Fi, Bluetooth, Near Field Communication (NFC), or other network devices including one or more radio devices.

[0218] In at least one embodiment, the computer system 1100A may include other components not explicitly shown, such as USB or other port connections, optical storage drives, video capture devices, etc., and these other components may also be connected to the I / O hub 1107. In at least one embodiment, any suitable protocol (such as a PCI (Peripheral Component Interconnect)-based protocol (such as PCI-Express) or other bus or point-to-point communication interface and / or protocol) may be used to implement the communication paths between the various components, such as NV-Link high-speed interconnect or interconnect protocol. Figure 11A For the communication paths of the various components, such as NV-Link high-speed interconnect or interconnect protocol.

[0219] In at least one embodiment, one or more parallel processors 1112 include circuitry optimized for graphics and video processing, the circuitry including, for example, video output circuitry, and constitute a Graphics Processing Unit (GPU). In at least one embodiment, one or more parallel processors 1112 include circuitry optimized for general-purpose processing. In at least one embodiment, the components of the computer system 1100A may be integrated with one or more other system elements on a single integrated circuit. For example, in at least one embodiment, one or more parallel processors 1112, memory hub 1105, one or more processors 1102, and I / O hub 1107 may be integrated into a System-on-Chip (SoC) integrated circuit. In at least one embodiment, the components of the computer system 1100A may be integrated into a single package to form a System-in-Package (SIP) configuration. In at least one embodiment, at least a portion of the components of the computer system 1100A may be integrated into a Multi-Chip Module (MCM), and the multi-chip module may be interconnected with other multi-chip modules into a modular computer system.

[0220] The inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. This is described in conjunction with Figure 6B and / or Figure 6C which provides details regarding the inference and / or training logic 615. In at least one embodiment, the inference and / or training logic 615 may be used in a Figure 11A system for inferring or predicting operations based at least in part on weight parameters calculated using neural network training operations, neural network functions, and / or architectures or neural network use cases herein.

[0221] Processor

[0222] Figure 11B illustrates a parallel processor 1100B according to at least one embodiment. In at least one embodiment, the various components of the parallel processor 1100B may be implemented using one or more integrated circuit devices, such as programmable processors, application specific integrated circuits (ASICs), or field programmable gate arrays (FPGAs). In at least one embodiment, the illustrated parallel processor 1100B is a variant of the Figure 11B one or more parallel processors 1112 shown in the exemplary embodiment.

[0223] In at least one embodiment, the parallel processor 1100B includes a parallel processing unit 1102. In at least one embodiment, the parallel processing unit 1102 includes an I / O unit 1104 that enables communication with other devices, including other instances of the parallel processing unit 1102. In at least one embodiment, the I / O unit 1104 may be directly connected to other devices. In at least one embodiment, the I / O unit 1104 is connected to other devices using a hub or switch interface (e.g., a memory hub 1105). In at least one embodiment, the connection between the memory hub 1105 and the I / O unit 1104 forms a communication link 1113. In at least one embodiment, the I / O unit 1104 is connected to a host interface 1106 and a memory crossbar 1116, where the host interface 1106 receives commands for performing processing operations and the memory crossbar 1116 receives commands for performing memory operations.

[0224] In at least one embodiment, when the host interface 1106 receives a command buffer via the I / O unit 1104, the host interface 1106 may initiate work operations to execute those commands to the front end 1108. In at least one embodiment, the front end 1108 is coupled to a scheduler 1110 configured to allocate commands or other work items to an array of processing clusters 1112. In at least one embodiment, the scheduler 1110 ensures that the array of processing clusters 1112 is properly configured and in an active state before tasks are allocated to the array of processing clusters 1112. In at least one embodiment, the scheduler 1110 is implemented by firmware logic executed on a microcontroller. In at least one embodiment, the microcontroller-implemented scheduler 1110 may be configured to perform complex scheduling and work allocation operations at both a coarse-grained and a fine-grained level, enabling fast preemption and context switching of threads executing on the processing array 1112. In at least one embodiment, host software may attest to a workload for scheduling on the processing array 1112 via one of the multi-graphics processing doorbells. In at least one embodiment, the workload may then be automatically allocated on the processing array 1112 by the scheduler 1110 logic within a microcontroller that includes the scheduler 1110.

[0225] In at least one embodiment, the array of processing clusters 1112 may include up to “N” processing clusters (e.g., clusters 1114A, 1114B through 1114N). In at least one embodiment, each of the clusters 1114A - 1114N of the array of processing clusters 1112 may execute a large number of concurrent threads. In at least one embodiment, the scheduler 1110 may use various scheduling and / or work allocation algorithms to allocate work to the clusters 1114A - 1114N of the array of processing clusters 1112, which may vary according to the workload generated by each type of program or computation. In at least one embodiment, scheduling may be handled dynamically by the scheduler 1110, or may be assisted in part by compiler logic during the compilation of program logic configured to be executed by the array of processing clusters 1112. In at least one embodiment, different ones of the clusters 1114A - 1114N of the array of processing clusters 1112 may be allocated for processing different types of programs or for performing different types of computations.

[0226] In at least one embodiment, the array of processing clusters 1112 may be configured to perform various types of parallel processing operations. In at least one embodiment, the array of processing clusters 1112 is configured to perform general-purpose parallel computing operations. In at least one embodiment, the array of processing clusters 1112 may include logic for performing processing tasks that include filtering of video and / or audio data, performing modeling operations, including physical operations, and performing data conversions.

[0227] In at least one embodiment, the processing cluster array 1112 is configured to perform parallel graphics processing operations. In at least one embodiment, the processing cluster array 1112 may include additional logic to support the execution of such graphics processing operations, including but not limited to texture sampling logic for performing texture operations, and tessellation logic and other vertex processing logic. In at least one embodiment, the processing cluster array 1112 may be configured to execute shader programs related to graphics processing, such as but not limited to vertex shaders, tessellation shaders, geometry shaders, and pixel shaders. In at least one embodiment, the parallel processing unit 1102 may transfer data from the system memory via the I / O unit 1104 for processing. In at least one embodiment, during processing, the transferred data may be stored in on-chip memory (e.g., parallel processor memory 1122) during processing and then written back to the system memory.

[0228] In at least one embodiment, when the parallel processing unit 1102 is used to perform graphics processing, the scheduler 1110 may be configured to divide the processing workload into tasks of approximately equal size to better distribute the graphics processing operations to the multiple clusters 1114A - 1114N of the processing cluster array 1112. In at least one embodiment, portions of the processing cluster array 1112 may be configured to perform different types of processing. In at least one embodiment, if a valve control for a remediation system that simulates a threshold leak in a data center liquid cooling system is needed, a first portion may be configured to perform vertex shading and topology generation, a second portion may be configured to perform tessellation and geometry shading, and a third portion may be configured to perform pixel shading or other screen space operations to generate a rendered image for display. In at least one embodiment, intermediate data generated by one or more of the clusters 1114A - 1114N may be stored in a buffer to allow transfer of the intermediate data between the clusters 1114A - 1114N for further processing.

[0229] In at least one embodiment, the processing cluster array 1112 may receive processing tasks to be executed via a scheduler 1110, which receives commands defining the processing tasks from a front end 1108. In at least one embodiment, the processing tasks may include an index of data to be processed, e.g., surface (patch) data, raw data, vertex data, and / or pixel data, as well as status parameters and commands defining how to process the data (e.g., what program to execute). In at least one embodiment, the scheduler 1110 may be configured to obtain an index corresponding to the task, or may receive the index from the front end 1108. In at least one embodiment, the front end 1108 may be configured to ensure that the processing cluster array 1112 is configured in an effective state before starting a workload specified by an incoming command buffer (e.g., a batch-buffer, a push buffer, etc.).

[0230] In at least one embodiment, each of one or more instances of the parallel processing units 1102 may be coupled to a parallel processor memory 1122. In at least one embodiment, the parallel processor memory 1122 may be accessed via a memory crossbar 1116, which may receive memory requests from the processing cluster array 1112 as well as the I / O unit 1104. In at least one embodiment, the memory crossbar 1116 may access the parallel processor memory 1122 via a memory interface 1118. In at least one embodiment, the memory interface 1118 may include a plurality of partitioning units (e.g., partitioning unit 1120A, partitioning unit 1120B through partitioning unit 1120N), each of which may be coupled to a portion (e.g., a memory unit) of the parallel processor memory 1122. In at least one embodiment, the plurality of partitioning units 1120A-1120N are configured to be equal to the number of memory units such that the first partitioning unit 1120A has a corresponding first memory unit 1124A, the second partitioning unit 1120B has a corresponding memory unit 1124B, and the Nth partitioning unit 1120N has a corresponding Nth memory unit 1124N. In at least one embodiment, the number of partitioning units 1120A-1120N may not be equal to the number of memory devices.

[0231] In at least one embodiment, the memory units 1124A - 1124N may include various types of memory devices, including dynamic random access memory (DRAM) or graphics random access memory, such as synchronous graphics random access memory (SGRAM), including graphics double data rate (GDDR) memory. In at least one embodiment, the memory units 1124A - 1124N may further include 3D stacked memory, including but not limited to high bandwidth memory (HBM). In at least one embodiment, render targets such as frame buffers or texture maps may be stored across the memory units 1124A - 1124N, allowing the partitioning units 1120A - 1120N to write portions of each render target in parallel, to effectively utilize the available bandwidth of the parallel processor memory 1122. In at least one embodiment, a local instance of the parallel processor memory 1122 may be excluded in favor of a unified memory design that utilizes system memory in combination with local cache memory.

[0232] In at least one embodiment, any one of the clusters 1114A - 1114N in the cluster array 1112 of processing clusters may process data to be written into any of the memory units 1124A - 1124N within the parallel processor memory 1122. In at least one embodiment, the memory crossbar 1116 may be configured to transfer the output of each cluster 1114A - 1114N to any of the partitioning units 1120A - 1120N or to another cluster 1114A - 1114N, where the cluster 1114A - 1114N may perform additional processing operations on the output. In at least one embodiment, each cluster 1114A - 1114N may communicate with the memory interface 1118 through the memory crossbar 1116 to read from or write to various external storage devices. In at least one embodiment, the memory crossbar 1116 has a connection to the memory interface 1118 to communicate with the I / O unit 1104, and a connection to a local instance of the parallel processor memory 1102, enabling processing units within different processing clusters 1114A - 1114N to communicate with system memory or other memory that is not local to the parallel processing unit 1102. In at least one embodiment, the memory crossbar 1116 may use virtual channels to separate the traffic flow between the clusters 1114A - 1114N and the partitioning units 1120A - 1120N.

[0233] In at least one embodiment, multiple instances of the parallel processing unit 1102 may be provided on a single insertion card, or multiple insertion cards may be interconnected. In at least one embodiment, different instances of the parallel processing unit 1102 may be configured to operate with each other even if the different instances have different numbers of processing cores, different numbers of local parallel processor memories, and / or other configuration differences. In at least one embodiment, some instances of the parallel processing unit 1102 may include floating-point units with higher precision relative to other instances. In at least one embodiment, systems incorporating one or more instances of the parallel processing unit 1102 or the parallel processor 1100B may be implemented in a variety of configurations and form factors, including but not limited to desktop computers, laptop computers, or handheld personal computers, servers, workstations, gaming consoles, and / or embedded systems.

[0234] Figure 11C is a block diagram of a partitioning unit 1120 according to at least one embodiment. In at least one embodiment, the partitioning unit 1120 is Figure 11B an instance of one of the partitioning units 1120A - 1120N. In at least one embodiment, the partitioning unit 1120 includes an L2 cache 1121, a frame buffer interface 1125, and a ROP 1126 (raster operation unit). The L2 cache 1121 is a read / write cache configured to perform load and store operations received from the memory crossbar 1116 and the ROP 1126. In at least one embodiment, the L2 cache 1121 outputs read misses and urgent write-back requests to the frame buffer interface 1125 for processing. In at least one embodiment, updates may also be sent to the frame buffer via the frame buffer interface 1125 for processing. In at least one embodiment, the frame buffer interface 1125 interacts with one of the memory units (such as Figure 11B the memory units 1124A - 1124N (e.g., within the parallel processor memory 1122)) of the parallel processor memory.

[0235] In at least one embodiment, the ROP 1126 is a processing unit that performs raster operations such as stencil, z-test, blending, and so on. In at least one embodiment, the ROP 1126 then outputs the processed graphics data stored in the graphics memory. In at least one embodiment, the ROP 1126 includes compression logic to compress depth or color data written to the memory and decompress depth or color data read from the memory. In at least one embodiment, the compression logic may be lossless compression logic that utilizes one or more of a variety of compression algorithms. The compression logic performed by the ROP 1126 may vary based on the statistical characteristics of the data to be compressed. In at least one embodiment, incremental color compression is performed based on depth and color data on a per-tile basis.

[0236] In at least one embodiment, the ROP 1126 is included within each processing cluster (e.g., Figure 11B clusters 1114A - 1114N), rather than within the partitioning unit 1120. In at least one embodiment, read and write requests for pixel data are made via the memory crossbar 1116 rather than pixel fragment data transfer. In at least one embodiment, the processed graphics data can be displayed on a display device (such as Figure 11A one of one or more display devices 1110), routed by the processor 1102 for further processing, or routed by Figure 11B one of the processing entities within the parallel processor 1100B for further processing.

[0237] Figure 11D is a block diagram of a processing cluster 1114 within a parallel processing unit according to at least one embodiment. In at least one embodiment, the processing cluster is Figure 11B an instance of one of the processing clusters 1114A - 1114N. In at least one embodiment, one or more processing clusters 1114 can be configured to execute a number of threads in parallel, where a "thread" refers to an instance of a particular program executed on a particular set of input data. In at least one embodiment, single instruction multiple data (SIMD) instruction issue techniques are used to support the parallel execution of a large number of threads without providing multiple independent instruction units. In at least one embodiment, single instruction multiple thread (SIMT) techniques are used to support the parallel execution of a large number of synchronized threads, which uses a common instruction unit configured to issue instructions to a set of processing engines within each processing cluster.

[0238] In at least one embodiment, the operation of the processing cluster 1114 can be controlled by assigning processing tasks to the pipeline manager 1132 of the SIMT parallel processor. In at least one embodiment, the pipeline manager 1132 receives from Figure 11BThe scheduler 1110 receives instructions and manages the execution of these instructions via the graphics multiprocessor 1134 and / or the texture unit 1136. In at least one embodiment, the graphics multiprocessor 1134 is an exemplary instance of a SIMT parallel processor. However, in at least one embodiment, various types of SIMT parallel processors with different architectures may be included within the processing cluster 1114. In at least one embodiment, one or more instances of the graphics multiprocessor 1134 may be included within the processing cluster 1114. In at least one embodiment, the graphics multiprocessor 1134 may process data, and the data crossbar 1140 may be used to distribute the processed data to one of multiple possible destinations (including other shader units). In at least one embodiment, the pipeline manager 1132 may facilitate the distribution of the processed data by specifying the destination of the processed data to be allocated via the data crossbar 1140.

[0239] In at least one embodiment, each graphics multiprocessor 1134 within the processing cluster 1114 may include the same set of functional execution logic (e.g., arithmetic logic units, load store units, etc.). In at least one embodiment, the functional execution logic may be configured in a pipeline manner, where new instructions may be issued before the completion of previous instructions. In at least one embodiment, the functional execution logic supports a variety of operations, including integer and floating-point arithmetic, comparison operations, boolean operations, shifts, and the calculation of various algebraic functions. In at least one embodiment, the same functional unit hardware may be utilized to perform different operations, and any combination of functional units may exist.

[0240] In at least one embodiment, the instructions transmitted to the processing cluster 1114 constitute threads. In at least one embodiment, a set of threads executed across a set of parallel processing engines is a thread group. In at least one embodiment, the thread group executes a program on different input data. In at least one embodiment, each thread within the thread group may be assigned to a different processing engine within the graphics multiprocessor 1134. In at least one embodiment, the thread group may include fewer threads than the number of processing engines within the graphics multiprocessor 1134. In at least one embodiment, when the number of threads included in the thread group is less than the number of processing engines, one or more processing engines may be idle during the loop in which the thread group is being processed. In at least one embodiment, the thread group may also include more threads than the number of processing engines within the graphics multiprocessor 1134. In at least one embodiment, when the thread group includes more threads than the processing engines within the graphics multiprocessor 1134, processing may be performed in consecutive clock cycles. In at least one embodiment, multiple thread groups may be executed simultaneously on the graphics multiprocessor 1134.

[0241] In at least one embodiment, the graphics multiprocessor 1134 includes an internal cache memory to perform load and store operations. In at least one embodiment, the graphics multiprocessor 1134 may forego the internal cache and use the cache memory (e.g., L1 cache 1148) within the processing cluster 1114. In at least one embodiment, each graphics multiprocessor 1134 may also access the L2 cache within the partition units (e.g., Figure 11B partition units 1120A - 1120N) of Figure 11B , which are shared among all processing clusters 1114 and can be used to transfer data between threads. In at least one embodiment, the graphics multiprocessor 1134 may also access off-chip global memory, which may include one or more of local parallel processor memory and / or system memory. In at least one embodiment, any memory external to the parallel processing unit 1102 can be used as global memory. In at least one embodiment, the processing cluster 1114 includes multiple instances of the graphics multiprocessor 1134, which may share common instructions and data that can be stored in the L1 cache 1148.

[0242] In at least one embodiment, each processing cluster 1114 may include a memory management unit (“MMU”) 1145 configured to map virtual addresses to physical addresses. In at least one embodiment, one or more instances of the MMU 1145 may reside within Figure 11B the memory interface 1118 of Figure 11B . In at least one embodiment, the MMU 1145 includes a set of page table entries (PTEs) that are used to map virtual addresses to the physical addresses of tiles and, in at least one embodiment, to cache line indices. In at least one embodiment, the MMU 1145 may include a translation lookaside buffer (TLB) or a cache that may reside within the graphics multiprocessor 1134 or the L1 cache 1148 or the processing cluster 1114. In at least one embodiment, physical addresses are processed to allocate surface data access locality for efficient request interleaving between partition units. In at least one embodiment, the cache line index can be used to determine whether a request to a cache line is a hit or a miss.

[0243] In at least one embodiment, the processing cluster 1114 can be configured such that each graphics multiprocessor 1134 is coupled to a texture unit 1136 to perform texture mapping operations that determine texture sample locations, read texture data, and filter texture data. In at least one embodiment, texture data is read as needed from an internal texture L1 cache (not shown) or from an L1 cache within the graphics multiprocessor 1134, and texture data is fetched from an L2 cache, local parallel processor memory, or system memory. In at least one embodiment, each graphics multiprocessor 1134 outputs one or more processed tasks to a data crossbar 1140 to provide the processed tasks to another processing cluster 1114 for further processing or to store one or more processed tasks in an L2 cache, local parallel processor memory, or system memory via a memory crossbar 1116. In at least one embodiment, the preROP 1142 (pre-raster operation unit) is configured to receive data from the graphics multiprocessor 1134 and direct the data to a ROP unit that can be located with the partitioning units (e.g., Figure 11B partitioning units 1120A - 1120N) of this document. In at least one embodiment, the PreROP 1142 unit can perform optimizations for color blending, organize pixel color data, and perform address translation.

[0244] The inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. Details regarding the inference and / or training logic 615 are provided below in conjunction with Figure 6B and / or Figure 6C In at least one embodiment, the inference and / or training logic 615 can be in the graphics processing cluster 1114 to perform inference or prediction operations at least in part based on weight parameters calculated using the neural network training operations, neural network functions, and / or architectures or neural network use cases described herein.

[0245] Figure 11EIllustrates a graphics multiprocessor 1134 according to at least one embodiment. In at least one embodiment, the graphics multiprocessor 1134 is coupled to the pipeline manager 1132 of the processing cluster 1114. In at least one embodiment, the graphics multiprocessor 1134 has an execution pipeline that includes, but is not limited to, an instruction cache 1152, an instruction unit 1154, an address mapping unit 1156, a register file 1158, one or more general-purpose graphics processing unit (GPGPU) cores 1162, and one or more load / store units 1166. The one or more GPGPU cores 1162 and the one or more load / store units 1166 are coupled to a cache memory 1172 and a shared memory 1170 via a memory and cache interconnect 1168.

[0246] In at least one embodiment, the instruction cache 1152 receives a stream of instructions to be executed from the pipeline manager 1132. In at least one embodiment, the instructions are cached in the instruction cache 1152 and dispatched for execution by the instruction unit 1154. In one embodiment, the instruction unit 1154 may dispatch instructions as a thread group (e.g., a warp), with each thread group assigned to a different execution unit within one or more GPGPU cores 1162. In at least one embodiment, instructions can access any local, shared, or global address space by specifying an address within a unified address space. In at least one embodiment, the address mapping unit 1156 can be used to convert an address in the unified address space into a different memory address that can be accessed by one or more load / store units 1166.

[0247] In at least one embodiment, the register file 1158 provides a set of registers for the functional units of the graphics multiprocessor 1134. In at least one embodiment, the register file 1158 provides temporary storage for the operands of the data paths of the functional units (e.g., GPGPU cores 1162, load / store units 1166) connected to the graphics multiprocessor 1134. In at least one embodiment, the register file 1158 is divided among each functional unit such that a dedicated portion of the register file 1158 is allocated to each functional unit. In at least one embodiment, the register file 1158 is divided among different warps being executed by the graphics multiprocessor 1134.

[0248] In at least one embodiment, the GPGPU cores 1162 may each include a floating-point unit (FPU) and / or an integer arithmetic logic unit (ALU) for executing instructions of the graphics multiprocessors 1134. The GPGPU cores 1162 may be architecturally similar or may have different architectures. In at least one embodiment, a first portion of the GPGPU cores 1162 includes a single-precision FPU and an integer ALU, while a second portion of the GPGPU cores includes a double-precision FPU. In at least one embodiment, the FPU may implement the IEEE 754-2008 standard for floating-point algorithms or enable variable-precision floating-point algorithms. In at least one embodiment, the graphics multiprocessors 1134 may additionally include one or more fixed-function or special-function units to perform specific functions such as copy rectangle or pixel blend operations. In at least one embodiment, one or more of the GPGPU cores may also include fixed or special-function logic.

[0249] In at least one embodiment, the GPGPU cores 1162 include SIMD logic capable of executing a single instruction on multiple sets of data. In one embodiment, the GPGPU cores 1162 may physically execute SIMD4, SIMD8, and SIMD16 instructions and logically execute SIMD1, SIMD2, and SIMD32 instructions. In at least one embodiment, the SIMD instructions for the GPGPU cores may be generated by a shader compiler at compile time or automatically generated when executing a program written and compiled for a single-program multiple-data (SPMD) or SIMT architecture. In at least one embodiment, multiple threads of a program configured for the SIMT execution model may be executed by a single SIMD instruction. For example, in at least one embodiment, eight SIMT threads performing the same or similar operations may be executed in parallel by a single SIMD8 logic unit.

[0250] In at least one embodiment, the memory and cache interconnect 1168 is an interconnect network that connects each functional unit of the graphics multiprocessor 1134 to the register file 1158 and the shared memory 1170. In at least one embodiment, the memory and cache interconnect 1168 is a crossbar interconnect that allows the load / store unit 1166 to perform load and store operations between the shared memory 1170 and the register file 1158. In at least one embodiment, the register file 1158 can operate at the same frequency as the GPGPU core 1162, resulting in very low latency for data transfer between the GPGPU core 1162 and the register file 1158. In at least one embodiment, the shared memory 1170 can be used to enable communication between threads executing on functional units within the graphics multiprocessor 1134. In at least one embodiment, the cache memory 1172 can be used as, for example, a data cache to cache texture data communicated between the functional units and the texture unit 1136. In at least one embodiment, the shared memory 1170 can also be used as a program-managed cache. In at least one embodiment, in addition to the automatically cached data stored in the cache memory 1172, threads executing on the GPGPU core 1162 can also programmatically store data in the shared memory.

[0251] In at least one embodiment, a parallel processor or GPGPU as described herein is communicatively coupled to a host / processor core to accelerate graphics operations, machine learning operations, pattern analysis operations, and various general-purpose GPU (GPGPU) functions. In at least one embodiment, the GPU can be communicatively coupled to the host processor / core via a bus or other interconnect (e.g., a high-speed interconnect such as PCIe or NVLink). In at least one embodiment, the GPU can be integrated with the core on the same package or chip and communicatively coupled to the core via an internal processor bus / interconnect (internal to the package or chip in at least one embodiment). In at least one embodiment, regardless of how the GPU is connected, the processor core can allocate work to the GPU in the form of a sequence of commands / instructions included in a work descriptor. In at least one embodiment, the GPU then uses dedicated circuitry / logic to efficiently process these commands / instructions.

[0252] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. As described below in conjunction with Figure 6B and / or Figure 6CProvide details regarding inference and / or training logic 615. In at least one embodiment, the inference and / or training logic 615 may be used in the graphics multiprocessor 1134 to perform inference or prediction operations based at least in part on weight parameters computed using neural network training operations, neural network functions, and / or architectures or neural network use cases described herein.

[0253] Figure 12A FIG. 1200A illustrates a multi-GPU computing system 1200A according to at least one embodiment. In at least one embodiment, the multi-GPU computing system 1200A may include a processor 1202 coupled to a plurality of general-purpose graphics processing units (GPGPUs) 1206A-D via a host interface switch 1204. In at least one embodiment, the host interface switch 1204 is a PCI Express switch device that couples the processor 1202 to a PCI Express bus, and the processor 1202 may communicate with the GPGPUs 1206A-D via the PCI Express bus. The GPGPUs 1206A-D may be interconnected via a set of high-speed P2P GPU-to-GPU links 1216. In at least one embodiment, the GPU-to-GPU link 1216 is connected to each of the GPGPUs 1206A-D via a dedicated GPU link. In at least one embodiment, the P2P GPU link 1216 enables direct communication between each of the GPGPUs 1206A-D without communicating through the host interface bus 1204 to which the processor 1202 is connected. In at least one embodiment, in the case where GPU-to-GPU traffic is directed to the P2P GPU link 1216, the host interface bus 1204 remains available for system memory access or communication with other instances of the multi-GPU computing system 1200A via, for example, one or more network devices. Although in at least one embodiment, the GPGPUs 1206A-D are connected to the processor 1202 via the host interface switch 1204, in at least one embodiment, the processor 1202 includes direct support for the P2P GPU link 1216 and may be directly connected to the GPGPUs 1206A-D.

[0254] The inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. As described herein in connection with Figure 6B and / or Figure 6C Provide details regarding inference and / or training logic 615. In at least one embodiment, the inference and / or training logic 615 may be used in the multi-GPU computing system 1200A to perform inference or prediction operations based at least in part on weight parameters computed using neural network training operations, neural network functions, and / or architectures or neural network use cases described herein.

[0255] Figure 12B It is a block diagram of a graphics processor 1200B according to at least one embodiment. In at least one embodiment, the graphics processor 1200B includes a ring interconnect 1202, a pipeline front end 1204, a media engine 1237, and graphics cores 1280A - 1280N. In at least one embodiment, the ring interconnect 1202 couples the graphics processor 1200B to other processing units, and the processing units include other graphics processors or one or more general - purpose processor cores. In at least one embodiment, the graphics processor 1200B is one of many processors integrated within a multi - core processing system.

[0256] In at least one embodiment, the graphics processor 1200B receives multiple batches of commands via the ring interconnect 1202. In at least one embodiment, the input commands are interpreted by a command streamer 1203 in the pipeline front end 1204. In at least one embodiment, the graphics processor 1200B includes scalable execution logic for performing 3D geometry processing and media processing via the graphics cores 1280A - 1280N. In at least one embodiment, for 3D geometry processing commands, the command streamer 1203 provides the commands to the geometry pipeline 1236. In at least one embodiment, for at least some media processing commands, the command streamer 1203 provides the commands to the video front end 1234, which is coupled to the media engine 1237. In at least one embodiment, the media engine 1237 includes a video quality engine (VQE) 1230 for video and image post - processing, and a multi - format encoding / decoding (MFX) 1233 engine for providing hardware - accelerated encoding and decoding of media data. In at least one embodiment, the geometry pipeline 1236 and the media engine 1237 each generate execution threads for thread execution resources provided by at least one graphics core 1280A.

[0257] In at least one embodiment, the graphics processor 1200B includes scalable thread execution resources having module cores 1280A - 1280N (sometimes referred to as core slices), each graphics core having multiple sub - cores 1250A - 1250N, 1260A - 1260N (sometimes referred to as core sub - slices). In at least one embodiment, the graphics processor 1200B can have any number of graphics cores 1280A. In at least one embodiment, the graphics processor 1200B includes a graphics core 1280A having at least a first sub - core 1250A and a second sub - core 1260A. In at least one embodiment, the graphics processor 1200B is a low - power processor having a single sub - core (e.g., the first sub - core 1250A). In at least one embodiment, the graphics processor 1200B includes multiple graphics cores 1280A - 1280N, each graphics core including a set of first sub - cores 1250A - 1250N and a set of second sub - cores 1260A - 1260N. In at least one embodiment, each of the first sub - cores 1250A - 1250N includes at least a first set of execution units 1252A - 1252N and media / texture samplers 1254A - 1254N. In at least one embodiment, each of the second sub - cores 1260A - 1260N includes at least a second set of execution units 1262A - 1262N and samplers 1264A - 1264N. In at least one embodiment, each of the sub - cores 1250A - 1250N, 1260A - 1260N shares a set of shared resources 1270A - 1270N. In at least one embodiment, the shared resources include shared cache memory and pixel operation logic.

[0258] The inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. Details regarding the inference and / or training logic 615 are provided herein in connection with Figure 6B and / or Figure 6C In at least one embodiment, the inference and / or training logic 615 can be in the graphics processor 1200B to perform inference or prediction operations at least in part based on weight parameters calculated using the neural network training operations, neural network functions, and / or architectures or neural network use cases described herein.

[0259] Figure 13is a block diagram of a microarchitecture for a processor 1300 according to at least one embodiment. The processor 1300 may include logic circuitry for executing instructions. In at least one embodiment, the processor 1300 may execute instructions including x86 instructions, ARM instructions, special instructions for application specific integrated circuits (ASICs), etc. In at least one embodiment, the processor 1300 may include registers for storing packed data, such as the 64-bit wide MMXTM registers in a microprocessor enabled with MMX technology by Intel Corporation in Santa Clara, California. In at least one embodiment, the MMX registers available in integer and floating-point forms may operate with packed data elements, and the packed data elements accompany single instruction multiple data (“SIMD”) and streaming SIMD extensions (“SSE”) instructions. In at least one embodiment, 128-bit wide XMM registers related to SSE2, SSE3, SSE4, AVX or later versions (generally referred to as “SSEx” technologies) may hold such packed data operands. In at least one embodiment, the processor 1300 may execute instructions to accelerate machine learning or deep learning algorithms, training or inference.

[0260] In at least one embodiment, the processor 1300 includes an in-order front end (“front end”) 1301 to fetch instructions to be executed and prepare the instructions for later use in the processor pipeline. In at least one embodiment, the front end 1301 may include several units. In at least one embodiment, the instruction prefetcher 1326 fetches instructions from memory and provides the instructions to the instruction decoder 1328, which in turn decodes or interprets the instructions. For example, in at least one embodiment, the instruction decoder 1328 decodes the received instructions into one or more operations of so-called “microinstructions” or “micro-operations” (also referred to as “micro-ops” or “microinstructions”) that are machine-executable. In at least one embodiment, the instruction decoder 1328 parses the instructions into an opcode and corresponding data and control fields, which may be used by the microarchitecture to perform operations according to at least one embodiment. In at least one embodiment, the trace cache 1330 may assemble the decoded microinstructions into a program-ordered sequence or trace in the microinstruction queue 1334 for execution. In at least one embodiment, when the trace cache 1330 encounters a complex instruction, the microcode ROM 1332 provides the microinstructions required to complete the operation.

[0261] In at least one embodiment, some instructions can be converted into a single micro-operation, while other instructions require several micro-operations to complete the entire operation. In at least one embodiment, if more than four micro-instructions are required to complete an instruction, the instruction decoder 1328 can access the microcode ROM 1332 to execute the instruction. In at least one embodiment, instructions can be decoded into a small number of micro-instructions for processing at the instruction decoder 1328. In at least one embodiment, if multiple micro-instructions are required to complete an operation, the instruction can be stored in the microcode ROM 1332. In at least one embodiment, the trace cache 1330 refers to the entry point programmable logic array (“PLA”) to determine the correct micro-instruction pointer for reading the microcode sequence from the microcode ROM 1332 to complete one or more instructions according to at least one embodiment. In at least one embodiment, after the microcode ROM 1332 completes the micro-operation sequencing of the instruction, the front end 1301 of the machine can resume fetching micro-operations from the trace cache 1330.

[0262] In at least one embodiment, the out-of-order execution engine ("out-of-order engine") 1303 may prepare instructions for execution. In at least one embodiment, the out-of-order execution logic has multiple buffers to smooth and reorder the instruction stream to optimize performance as the instructions descend along the pipeline and are scheduled for execution. In at least one embodiment, the out-of-order execution engine 1303 includes, but is not limited to, an allocator / register renamer 1340, a memory micro-instruction queue 1342, an integer / floating-point micro-instruction queue 1344, a memory scheduler 1346, a fast scheduler 1302, a slow / general floating-point scheduler ("slow / general FP scheduler") 1304, and a simple floating-point scheduler ("simple FP scheduler") 1306. In at least one embodiment, the fast scheduler 1302, the slow / general floating-point scheduler 1304, and the simple floating-point scheduler 1306 are also collectively referred to as "micro-instruction schedulers 1302, 1304, 1306". In at least one embodiment, the allocator / register renamer 1340 allocates the machine buffers and resources required for each micro-instruction to be executed in sequence. In at least one embodiment, the allocator / register renamer 1340 renames logical registers to entries in the register file. In at least one embodiment, the allocator / register renamer 1340 also allocates entries for each micro-instruction in one of the two micro-instruction queues, the memory micro-instruction queue 1342 for memory operations and the integer / floating-point micro-instruction queue 1344 for non-memory operations, ahead of the memory scheduler 1346 and the micro-instruction schedulers 1302, 1304, 1306. In at least one embodiment, the micro-instruction schedulers 1302, 1304, 1306 determine when a micro-instruction is ready for execution based on the readiness of their dependent input register operand sources and the availability of the execution resources micro-instructions that need to be completed. In at least one embodiment, the fast scheduler 1302 of at least one embodiment may be scheduled on each half of the main clock cycle, while the slow / general floating-point scheduler 1304 and the simple floating-point scheduler 1306 may be scheduled once per main processor clock cycle. In at least one embodiment, the micro-instruction schedulers 1302, 1304, 1306 arbitrate the scheduling ports to schedule micro-instructions for execution.

[0263] In at least one embodiment, execution block 1311 includes, but is not limited to, integer register file / bypass network 1308, floating-point register file / bypass network (“FP register file / bypass network”) 1310, address generation units (“AGU”) 1312 and 1314, fast arithmetic logic units (“fast ALU”) 1316 and 1318, slow arithmetic logic unit (“slow ALU”) 1320, floating-point ALU (“FP”) 1322, and floating-point move unit (“FP move”) 1324. In at least one embodiment, integer register file / bypass network 1308 and floating-point register file / bypass network 1310 are also referred to herein as “register files 1308, 1310”. In at least one embodiment, AGU 1312 and 1314, fast ALU 1316 and 1318, slow ALU 1320, floating-point ALU 1322, and floating-point move unit 1324 are also referred to herein as “execution units 1312, 1314, 1316, 1318, 1320, 1322, and 1324”. In at least one embodiment, execution block 1311 may include, but is not limited to, any number (including zero) and type of register files, bypass networks, address generation units, and execution units (in any combination).

[0264] In at least one embodiment, register networks 1308, 1310 may be arranged between microinstruction schedulers 1302, 1304, 1306 and execution units 1312, 1314, 1316, 1318, 1320, 1322, and 1324. In at least one embodiment, integer register file / bypass network 1308 performs integer operations. In at least one embodiment, floating-point register file / bypass network 1310 performs floating-point operations. In at least one embodiment, each of register networks 1308, 1310 may include, but is not limited to, a bypass network that may bypass or forward a just-completed result that has not yet been written to the register file to a new dependent. In at least one embodiment, register networks 1308, 1310 may communicate data with each other. In at least one embodiment, integer register file / bypass network 1308 may include, but is not limited to, two separate register files, one register file for low-order 32-bit data and a second register file for high-order 32-bit data. In at least one embodiment, floating-point register file / bypass network 1310 may include, but is not limited to, 128-bit-wide entries, since floating-point instructions typically have operands with widths of 64 to 128 bits.

[0265] In at least one embodiment, execution units 1312, 1314, 1316, 1318, 1320, 1322, 1324 may execute instructions. In at least one embodiment, register files 1308, 1310 store integer and floating-point data operand values that the microinstructions need to execute. In at least one embodiment, the processor 1300 may include, but is not limited to, any number of execution units 1312, 1314, 1316, 1318, 1320, 1322, 1324 and their combinations. In at least one embodiment, the floating-point ALU 1322 and the floating-point move unit 1324 may execute floating-point, MMX, SIMD, AVX, and SSE or other operations, including specialized machine learning instructions. In at least one embodiment, the floating-point ALU 1322 may include, but is not limited to, a 64-bit by 64-bit floating-point divider to perform division, square root, and remainder micro-operations. In at least one embodiment, instructions involving floating-point values may be processed with floating-point hardware. In at least one embodiment, ALU operations may be passed to the fast ALUs 1316, 1318. In at least one embodiment, the fast ALUs 1316, 1318 may execute fast operations with an effective latency of half a clock cycle. In at least one embodiment, most complex integer operations enter the slow ALU 1320 because the slow ALU 1320 may include, but is not limited to, integer execution hardware for long-latency type operations such as multipliers, shifts, flag logic, and branch processing. In at least one embodiment, memory load / store operations may be performed by the AGUs 1312, 1314. In at least one embodiment, the fast ALU 1316, the fast ALU 1318, and the slow ALU 1320 may perform integer operations on 64-bit data operands. In at least one embodiment, the fast ALU 1316, the fast ALU 1318, and the slow ALU 1320 may be implemented to support various data bit sizes including sixteen, thirty-two, 128, 256, etc. In at least one embodiment, the floating-point ALU 1322 and the floating-point move unit 1324 may be implemented to support a certain range of operands with bits of various widths. In at least one embodiment, the floating-point ALU 1322 and the floating-point move unit 1324 may operate on 128-bit wide packed data operands in combination with SIMD and multimedia instructions.

[0266] In at least one embodiment, the microinstruction schedulers 1302, 1304, 1306 schedule dependent operations before the completion of the execution of the parent load. In at least one embodiment, since microinstructions can be speculatively scheduled and executed in the processor 1300, the processor 1300 may also include logic for handling memory misses. In at least one embodiment, if a data load miss occurs in the data cache, there may be dependent operations running in the pipeline that leave the scheduler temporarily without the correct data. In at least one embodiment, a replay mechanism tracks and re-executes instructions that used incorrect data. In at least one embodiment, it may be necessary to replay dependent operations and independent operations may be allowed to complete. In at least one embodiment, the scheduler and replay mechanism of at least one embodiment of the processor may also be designed to capture sequences of instructions for text string comparison operations.

[0267] In at least one embodiment, a "register" may refer to an on-chip processor storage location that can be part of an instruction used to identify an operand. In at least one embodiment, registers may be those that can be used from outside the processor (from the programmer's perspective). In at least one embodiment, registers may not be limited to a particular type of circuit. Instead, in at least one embodiment, registers can store data, provide data, and perform the functions described herein. In at least one embodiment, the registers described herein may be implemented using a variety of different techniques by circuits within the processor, such as dedicated physical registers, physical registers dynamically allocated using register renaming, combinations of dedicated and dynamically allocated physical registers, etc. In at least one embodiment, integer registers store 32-bit integer data. The register file of at least one embodiment also contains eight multimedia SIMD registers for packaging data.

[0268] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 615 are provided in conjunction with Figure 6B and / or Figure 6C In at least one embodiment, some or all of the inference and / or training logic 615 may be incorporated into execution block 1311 and other memories or registers shown or not shown. For example, in at least one embodiment, the training and / or inference techniques described herein may use one or more ALUs shown in execution block 1311. Additionally, weight parameters may be stored in on-chip or off-chip memories and / or registers (shown or not shown) that configure the ALUs of execution block 1311 to perform one or more of the machine learning algorithms, neural network architectures, use cases, or training techniques described herein.

[0269] Figure 14FIG. 0 shows a deep learning application processor 1400 according to at least one embodiment. In at least one embodiment, the deep learning application processor 1400 uses instructions that, if executed by the deep learning application processor 1400, cause the deep learning application processor 1400 to perform some or all of the processes and techniques described throughout this disclosure. In at least one embodiment, the deep learning application processor 1400 is an application specific integrated circuit (ASIC). In at least one embodiment, the application processor 1400 performs matrix multiplication operations or is “hardwired” into the hardware as a result of executing one or more instructions or both. In at least one embodiment, the deep learning application processor 1400 includes, but is not limited to, processing clusters 1410(1)-1410(12), inter-chip links (“ICL”) 1420(1)-1420(12), inter-chip controllers (“ICC”) 1430(1)-1430(2), memory controllers (“Mem Ctrlr”) 1442(1)-1442(4), high bandwidth memory physical layers (“HBM PHY”) 1444(1)-1444(4), management controller central processing units (“management controller CPU”) 1450, serial peripheral interfaces, internal integrated circuits, and general purpose input / output blocks (“SPI, I2C, GPIO”), peripheral component interconnect express controllers and direct memory access blocks (“PCIe controllers and DMA”) 1470, and sixteen-channel peripheral component interconnect express ports (“PCI Express x16”) 1480.

[0270] In at least one embodiment, the processing clusters 1410 can perform deep learning operations, including inference or prediction operations based on weight parameters calculated using one or more training techniques, including those techniques described herein. In at least one embodiment, each processing cluster 1410 can include, but is not limited to, any number and type of processors. In at least one embodiment, the deep learning application processor 1400 can include any number and type of processing clusters 1400. In at least one embodiment, the inter-chip link 1420 is bi-directional. In at least one embodiment, the inter-chip link 1420 and the inter-chip controller 1430 enable multiple deep learning application processors 1400 to exchange information, including activation information resulting from executing one or more machine learning algorithms embodied in one or more neural networks. In at least one embodiment, the deep learning application processor 1400 can include any number (including zero) and type of ICL 1420 and ICC 1430.

[0271] In at least one embodiment, the HBM2 1440 provides a total of 32GB of memory. The HBM2 1440 is associated with both a memory controller 1442(i) and an HBM PHY 1444(i). In at least one embodiment, any number of HBM2 1440s can provide any type and total amount of high-bandwidth memory and can be associated with any number (including zero) and type of memory controllers 1442 and HBM PHYs 1444. In at least one embodiment, any number and type of blocks can replace the SPI, I2C, GPIO 3360, PCIe controller 1460, and DMA 1470 and / or PCIe 1480 to implement any number and type of communication standards in any technically feasible manner.

[0272] The inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. Details regarding the inference and / or training logic 615 are provided herein in connection with Figure 6B and / or Figure 6C In at least one embodiment, a deep learning application processor is used to train a machine learning model (e.g., a neural network) to predict or infer information provided to the deep learning application processor 1400. In at least one embodiment, the deep learning application processor 1400 is used to infer or predict information based on a trained machine learning model (e.g., a neural network) that has been trained by another processor or system or by the deep learning application processor 1400. In at least one embodiment, the processor 1400 can be used to execute one or more of the neural network use cases herein.

[0273] Figure 15 is a block diagram of a neuromorphic processor 1500 according to at least one embodiment. In at least one embodiment, the neuromorphic processor 1500 can receive one or more inputs from sources external to the neuromorphic processor 1500. In at least one embodiment, these inputs can be transmitted to one or more neurons 1502 within the neuromorphic processor 1500. In at least one embodiment, the neurons 1502 and their components can be implemented using circuitry or logic including one or more arithmetic logic units (ALUs). In at least one embodiment, the neuromorphic processor 1500 can include, but is not limited to, thousands of instances of neurons 1502, but any suitable number of neurons 1502 can be used. In at least one embodiment, each instance of the neuron 1502 can include a neuron input 1504 and a neuron output 1506. In at least one embodiment, the neuron 1502 can generate an output that can be an input to other instances of the neuron 1502. In at least one embodiment, the neuron input 1504 and the neuron output 1506 can be interconnected via synapses 1508.

[0274] In at least one embodiment, neurons 1502 and synapses 1508 may be interconnected such that the neuromorphic processor 1500 operates to process or analyze information received by the neuromorphic processor 1500. In at least one embodiment, when the input received through neuron input 1504 exceeds a threshold, neuron 1502 may send an output pulse (or "fire" or "spike"). In at least one embodiment, neuron 1502 may sum or integrate the signals received at neuron input 1504. For example, in at least one embodiment, neuron 1502 may be implemented as a leaky integrate-and-fire neuron, where if the sum (referred to as the "membrane potential") exceeds a threshold, neuron 1502 may use a transfer function such as a sigmoid or threshold function to produce an output (or "fire"). In at least one embodiment, the leaky integrate-and-fire neuron may sum the signals received at neuron input 1504 into a membrane potential and may apply a decay factor (or leak) to reduce the membrane potential. In at least one embodiment, if multiple input signals are received at neuron input 1504 fast enough to exceed the threshold (in at least one embodiment, before the membrane potential decays too low to fire), the leaky integrate-and-fire neuron may fire. In at least one embodiment, neuron 1502 may be implemented using circuitry or logic that receives inputs, integrates the inputs into a membrane potential, and decays the membrane potential. In at least one embodiment, the inputs may be averaged, or any other suitable transfer function may be used. Additionally, in at least one embodiment, neuron 1502 may include, but is not limited to, comparator circuitry or logic that produces an output spike at neuron output 1506 when the result of applying the transfer function to neuron input 1504 exceeds a threshold. In at least one embodiment, once neuron 1502 fires, it may ignore previously received input information, for example, by resetting the membrane potential to 0 or another suitable default value. In at least one embodiment, once the membrane potential is reset to 0, neuron 1502 may resume normal operation after a suitable period of time (or refractory period).

[0275] In at least one embodiment, neurons 1502 may be interconnected by synapses 1508. In at least one embodiment, synapses 1508 may operate to transmit a signal from the output of a first neuron 1502 to the input of a second neuron 1502. In at least one embodiment, a neuron 1502 may transmit information over more than one instance of a synapse 1508. In at least one embodiment, one or more instances of a neuron output 1506 may be connected to an instance of a neuron input 1504 in the same neuron 1502 by an instance of a synapse 1508. In at least one embodiment, an instance of a neuron 1502 that produces an output to be transmitted over an instance of a synapse 1508 may be referred to as a "presynaptic neuron" with respect to that instance of the synapse 1508. In at least one embodiment, an instance of a neuron 1502 that receives an input transmitted over an instance of a synapse 1508 may be referred to as a "postsynaptic neuron" with respect to the instance of the synapse 1508. In at least one embodiment, with respect to various instances of synapses 1508, since an instance of a neuron 1502 may receive inputs from one or more instances of synapses 1508 and may also transmit outputs over one or more instances of synapses 1508, a single instance of a neuron 1502 may be both a "presynaptic neuron" and a "postsynaptic neuron".

[0276] In at least one embodiment, neurons 1502 may be organized into one or more layers. Each instance of a neuron 1502 may have a neuron output 1506 that may fan out through one or more synapses 1508 to one or more neuron inputs 1504. In at least one embodiment, the neuron outputs 1506 of neurons 1502 in a first layer 1510 may be connected to the neuron inputs 1504 of neurons 1502 in a second layer 1512. In at least one embodiment, the layer 1510 may be referred to as a "feedforward layer". In at least one embodiment, each instance of a neuron 1502 in an instance of the first layer 1510 may fan out to each instance of a neuron 1502 in the second layer 1512. In at least one embodiment, the first layer 1510 may be referred to as a "fully connected feedforward layer". In at least one embodiment, each instance of a neuron 1502 in each instance of the second layer 1512 fans out to fewer than all instances of neurons 1502 in a third layer 1514. In at least one embodiment, the second layer 1512 may be referred to as a "sparsely connected feedforward layer". In at least one embodiment, neurons 1502 in the (same) second layer 1512 may fan out to neurons 1502 in multiple other layers, including also fanning out to neurons 1502 in the second layer 1512. In at least one embodiment, the second layer 1512 may be referred to as a "recurrent layer". In at least one embodiment, the neuromorphic processor 1500 may include any suitable combination of, but not limited to, recurrent layers and feedforward layers, including but not limited to sparsely connected feedforward layers and fully connected feedforward layers.

[0277] In at least one embodiment, the neuromorphic processor 1500 may include, but is not limited to, a reconfigurable interconnect architecture or dedicated hardwired interconnections to connect synapses 1508 to neurons 1502. In at least one embodiment, the neuromorphic processor 1500 may include, but is not limited to, circuitry or logic that allows synapses to be assigned to different neurons 1502 as needed, based on the neural network topology and neuron fan-in / fan-out. For example, in at least one embodiment, synapses 1508 may be connected to neurons 1502 using an interconnect structure such as a network-on-chip or through dedicated connections. In at least one embodiment, circuitry or logic may be used to implement the synapse interconnections and their components.

[0278] Figure 16AA processing system according to at least one embodiment is shown. In at least one embodiment, system 1600A includes one or more processors 1602 and one or more graphics processors 1608, and can be a single-processor desktop system, a multi-processor workstation system, or a server system with a large number of processors 1602 or processor cores 1607. In at least one embodiment, system 1600A is a processing platform integrated within a system-on-chip (SoC) integrated circuit for use in mobile, handheld, or embedded devices.

[0279] In at least one embodiment, system 1600A can be included in or incorporated into a server-based gaming platform, including a game console such as a game and media console, a mobile game console, a handheld game console, or an online game console. In at least one embodiment, system 1600A is a mobile phone, a smartphone, a tablet computing device, or a mobile Internet device. In at least one embodiment, the processing system 1600A can also be coupled to or integrated within a wearable device, such as a smartwatch wearable device, a smart glasses device, an augmented reality device, or a virtual reality device. In at least one embodiment, the processing system 1600A is a television or a set-top box device having one or more processors 1602 and a graphical interface generated by one or more graphics processors 1608.

[0280] In at least one embodiment, each of the one or more processors 1602 includes one or more processor cores 1607 to process instructions that, when executed, perform operations for system and user software. In at least one embodiment, each of the one or more processor cores 1607 is configured to process a specific instruction set 1609. In at least one embodiment, the instruction set 1609 can facilitate complex instruction set computing (CISC), reduced instruction set computing (RISC), or computing via very long instruction words (VLIW). In at least one embodiment, the processor cores 1607 can each process different instruction sets 1609, which can include instructions that help to emulate other instruction sets. In at least one embodiment, the processor cores 1607 can also include other processing devices, such as a digital signal processor (DSP).

[0281] In at least one embodiment, the processor 1602 includes a cache memory 1604. In at least one embodiment, the processor 1602 may have a single internal cache or multiple levels of internal caches. In at least one embodiment, the cache memory is shared among the various components of the processor 1602. In at least one embodiment, the processor 1602 also uses an external cache (e.g., a level 3 (L3) cache or a last level cache (LLC)) (not shown), and the external cache can be shared among the processor cores 1607 using known cache coherence techniques. In at least one embodiment, the processor 1602 further includes a register file 1606, and the processor may include different types of registers for storing different types of data (e.g., integer registers, floating-point registers, status registers, and instruction pointer registers). In at least one embodiment, the register file 1606 may include general-purpose registers or other registers.

[0282] In at least one embodiment, one or more processors 1602 are coupled to one or more interface buses 1610 to transfer communication signals, such as address, data, or control signals, between the processor 1602 and other components in the system 1600A. In at least one embodiment, the interface bus 1610 may be a processor bus, such as a version of the direct media interface (DMI) bus, in one embodiment. In at least one embodiment, the interface bus 1610 is not limited to the DMI bus and may include one or more peripheral component interconnect buses (e.g., PCI, PCI Express), a memory bus, or other types of interface buses. In at least one embodiment, the processor 1602 includes an integrated memory controller 1616 and a platform controller hub 1630. In at least one embodiment, the memory controller 1616 facilitates communication between the memory device and other components of the processing system 1600A, while the platform controller hub (PCH) 1630 provides connections to input / output (I / O) devices via a local I / O bus.

[0283] In at least one embodiment, the memory device 1620 may be a dynamic random access memory (DRAM) device, a static random access memory (SRAM) device, a flash memory device, a phase change memory device, or have suitable performance to be used as a processor memory. In at least one embodiment, the storage device 1620 may be used as the system memory of the processing system 1600A to store data 1622 and instructions 1621 for use when one or more processors 1602 execute an application or process. In at least one embodiment, the memory controller 1616 is also coupled to an external graphics processor 1612 of at least one embodiment, which may communicate with one or more graphics processors 1608 in the processor 1602 to perform graphics and media operations. In at least one embodiment, the display device 1611 may be connected to the processor 1602. In at least one embodiment, the display device 1611 may include one or more of internal display devices, such as in a mobile electronic device or a laptop device, or an external display device connected through a display interface (such as DisplayPort, etc.). In at least one embodiment, the display device 1611 may include a head-mounted display (HMD), such as a stereoscopic display device for virtual reality (VR) applications or augmented reality (AR) applications.

[0284] In at least one embodiment, the platform controller hub 1630 enables peripheral devices to be connected to the storage device 1620 and the processor 1602 via a high-speed I / O bus. In at least one embodiment, the I / O peripheral devices include, but are not limited to, an audio controller 1646, a network controller 1634, a firmware interface 1628, a wireless transceiver 1626, a touch sensor 1625, and a data storage device 1624 (e.g., a hard disk drive, a flash memory, etc.). In at least one embodiment, the data storage device 1624 can be connected via a storage interface (e.g., SATA) or via a peripheral bus, such as a Peripheral Component Interconnect bus (e.g., PCI, PCIe). In at least one embodiment, the touch sensor 1625 can include a touch screen sensor, a pressure sensor, or a fingerprint sensor. In at least one embodiment, the wireless transceiver 1626 can be a Wi-Fi transceiver, a Bluetooth transceiver, or a mobile network transceiver, such as a 3G, 4G, or Long Term Evolution (LTE) transceiver. In at least one embodiment, the firmware interface 1628 enables communication with the system firmware and can be, for example, a Unified Extensible Firmware Interface (UEFI). In at least one embodiment, the network controller 1634 can enable a network connection to a wired network. In at least one embodiment, a high-performance network controller (not shown) is coupled to the interface bus 1610. In at least one embodiment, the audio controller 1646 is a multi-channel high-definition audio controller. In at least one embodiment, the processing system 1600A includes a legacy I / O controller 1640 for coupling legacy (e.g., Personal System 2 (PS / 2)) devices to the system 1600A. In at least one embodiment, the platform controller hub 1630 can also be connected to one or more Universal Serial Bus (USB) controllers 1642, which connect input devices, such as a keyboard and mouse 1643 combination, a camera 1644, or other USB input devices.

[0285] In at least one embodiment, instances of the memory controller 1616 and the platform controller hub 1630 can be integrated into a discrete external graphics processor, such as the external graphics processor 1612. In at least one embodiment, the platform controller hub 1630 and / or the memory controller 1616 can be external to one or more of the processors 1602. In at least one embodiment, the system 1600A can include an external memory controller 1616 and a platform controller hub 1630, which can be configured as a memory controller hub and a peripheral controller hub in a system chipset that communicates with the processor 1602.

[0286] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. This is described in conjunction with Figure 6B and / or Figure 6CProvide details regarding inference and / or training logic 615. In at least one embodiment, some or all of the inference and / or training logic 615 may be incorporated into a graphics processor 1600A. In at least one embodiment, the training and / or inference techniques described herein may use one or more ALUs embodied in a graphics processor 1612. Additionally, in at least one embodiment, the inference and / or training operations described herein may be completed using logic other than the logic shown in Figure 6B and / or Figure 6C The weight parameters may be stored in on-chip or off-chip memory and / or registers (shown or not shown) that configure the ALUs of the graphics processor 1600A to perform one or more of the machine learning algorithms, neural network architectures, use cases, or training techniques described herein.

[0287] Figure 16B is a block diagram of a processor 1600B having one or more processor cores 1602A - 1602N, an integrated memory controller 1614, and an integrated graphics processor 1608, according to at least one embodiment. In at least one embodiment, the processor 1600B may include additional cores, up to and including the additional core 1602N shown in dashed boxes. In at least one embodiment, each processor core 1602A - 1602N includes one or more internal cache units 1604A - 1604N. In at least one embodiment, each processor core may also access one or more shared cache units 1606.

[0288] In at least one embodiment, the internal cache units 1604A - 1604N and the shared cache units 1606 represent the cache memory hierarchy within the processor 1600B. In at least one embodiment, the cache memory units 1604A - 1604N may include at least one level of instruction and data cache within each processor core and one or more levels of cache in a shared mid - level cache, such as level 2 (L2), level 3 (L3), level 4 (L4), or other levels of cache, where the highest level of cache before external memory is classified as the LLC. In at least one embodiment, cache coherence logic maintains coherence between the various cache units 1606 and 1604A - 1604N.

[0289] In at least one embodiment, the processor 1600B may further include a set of one or more bus controller units 1616 and a system agent core 1610. In at least one embodiment, the one or more bus controller units 1616 manage a set of peripheral buses, such as one or more PCI or PCIe buses. In at least one embodiment, the system agent core 1610 provides management functions for various processor components. In at least one embodiment, the system agent core 1610 includes one or more integrated memory controllers 1614 to manage access to various external memory devices (not shown).

[0290] In at least one embodiment, the one or more processor cores 1602A - 1602N include support for simultaneous multi-threading. In at least one embodiment, the system agent core 1610 includes components for coordinating and operating the cores 1602A - 1602N during multi-threaded processing. In at least one embodiment, the system agent core 1610 may additionally include a power control unit (PCU) that includes logic and components for regulating one or more power states of the processor cores 1602A - 1602N and the graphics processor 1608.

[0291] In at least one embodiment, the processor 1600B further includes a graphics processor 1608 for performing image processing operations. In at least one embodiment, the graphics processor 1608 is coupled to the shared cache unit 1606 and the system agent core 1610 that includes one or more integrated memory controllers 1614. In at least one embodiment, the system agent core 1610 further includes a display controller 1611 for driving the graphics processor output to one or more coupled displays. In at least one embodiment, the display controller 1611 may also be an independent module coupled to the graphics processor 1608 via at least one interconnect, or may be integrated within the graphics processor 1608.

[0292] In at least one embodiment, a ring-based interconnect unit 1612 is used to couple the internal components of the processor 1600B. In at least one embodiment, alternative interconnect units may be used, such as point-to-point interconnects, switched interconnects, or other technologies. In at least one embodiment, the graphics processor 1608 is coupled to the ring interconnect 1612 via an I / O link 1613.

[0293] In at least one embodiment, I / O link 1613 represents at least one of a variety of I / O interconnections, including a package I / O interconnection that facilitates communication between various processor components and a high-performance embedded memory module 1618 (e.g., an eDRAM module). In at least one embodiment, each of processor cores 1602A - 1602N and graphics processor 1608 uses embedded memory module 1618 as a shared last-level cache.

[0294] In at least one embodiment, processor cores 1602A - 1602N are homogeneous cores that execute a common instruction set architecture. In at least one embodiment, processor cores 1602A - 1602N are heterogeneous in terms of instruction set architecture (ISA), where one or more of processor cores 1602A - 1602N execute a common instruction set, while one or more other processor cores 1602A - 1602N execute a subset of the common instruction set or a different instruction set. In at least one embodiment, processor cores 1602A - 1602N are heterogeneous in terms of microarchitecture, where one or more cores with relatively high power consumption are coupled with one or more power cores with lower power consumption. In at least one embodiment, processor 1600B can be implemented on one or more chips or as a SoC integrated circuit.

[0295] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 615 are provided herein in connection with Figure 6B and / or Figure 6C In at least one embodiment, some or all of inference and / or training logic 615 can be incorporated into processor 1600B. For example, in at least one embodiment, the training and / or inference techniques described herein can use one or more ALUs, where the ALUs are embodied in Figure 16A graphics core 1612, one or more processor cores 1602A - 1602N, or other components shown. Additionally, in at least one embodiment, the inference and / or training operations described herein can be completed using logic other than Figure 6B and / or Figure 6C the logic shown. In at least one embodiment, weight parameters can be stored in on-chip or off-chip memory and / or registers (shown or not shown), which configure the ALUs of graphics processor 1600B to execute one or more machine learning algorithms, neural network architectures, use cases, or training techniques introduced herein.

[0296] Figure 16CIt is a block diagram of the hardware logic of the graphics processing unit core 1600C according to at least one embodiment of the present disclosure. In at least one embodiment, the graphics processing unit core 1600C is included in a graphics core array. In at least one embodiment, the graphics processing unit core 1600C (sometimes referred to as a core slice) can be one or more graphics cores within a modular graphics processing unit. In at least one embodiment, the graphics processing unit core 1600C is an example of a graphics core slice, and the graphics processing unit herein can include multiple graphics core slices based on a target power and performance envelope. In at least one embodiment, each graphics core 1600C can include a fixed function block 1630 coupled to a plurality of sub-cores 1601A - 1601F, also referred to as sub-slices, which include modular blocks of general and fixed function logic.

[0297] In at least one embodiment, the fixed function block 1630 includes a geometry fixed function pipeline 1636. For example, in a lower performance and / or lower power graphics processing unit implementation, this geometry and fixed function pipeline 1636 can be shared by all sub-cores in the graphics processing unit 1600C. In at least one embodiment, the geometry and fixed function pipeline 1636 includes a 3D fixed function pipeline, a video front-end unit, a thread generator and a thread dispatcher, and a unified return buffer manager that manages a unified return buffer.

[0298] In at least one fixed embodiment, the fixed function block 1630 further includes a graphics SoC interface 1637, a graphics microcontroller 1638, and a media pipeline 1639. In at least one embodiment, the fixed graphics SoC interface 1637 provides an interface between the graphics core 1600C and other processor cores in the on-chip integrated circuit system. In at least one embodiment, the graphics microcontroller 1638 is a programmable sub-processor that can be configured to manage various functions of the graphics processing unit 1600C, including thread dispatching, scheduling, and preemption. In at least one embodiment, the media pipeline 1639 includes logic that helps decode, encode, preprocess, and / or postprocess multimedia data including image and video data. In at least one embodiment, the media pipeline 1639 implements media operations via requests to the computing or sampling logic within the sub-cores 1601 - 1601F.

[0299] In at least one embodiment, the SoC interface 1637 enables the graphics core 1600C to communicate with a general-purpose application processor core (e.g., a CPU) and / or other components within the SoC, including memory hierarchy elements such as a shared last-level cache, system RAM, and / or embedded on-chip or package DRAM. In at least one embodiment, the SoC interface 1637 may also enable communication with fixed-function devices within the SoC (e.g., a camera imaging pipeline), and enable the use and / or implementation of global memory atoms that may be shared between the graphics core 1600C and the CPU within the SoC. In at least one embodiment, the SoC interface 1637 may also implement power management control for the graphics core 1600C and enable an interface between the clock domain of the graphics core 1600C and other clock domains within the SoC. In at least one embodiment, the SoC interface 1637 enables receipt of command buffers from a command stream converter and a global thread dispatcher, which are configured to provide commands and instructions to each of one or more graphics cores within the graphics processor. In at least one embodiment, when a media operation is to be performed, the commands and instructions may be dispatched to the media pipeline 1639, or when a graphics processing operation is to be performed, they may be assigned to the geometry and fixed-function pipelines (e.g., geometry and fixed-function pipeline 1636, geometry and fixed-function pipeline 1614).

[0300] In at least one embodiment, the graphics microcontroller 1638 may be configured to perform various scheduling and management tasks for the graphics core 1600C. In at least one embodiment, the graphics microcontroller 1638 may perform graphics and / or compute workload scheduling on various graphics parallel engines within the execution unit (EU) arrays 1602A-1602F, 1604A-1604F in the sub-cores 1601A-1601F. In at least one embodiment, host software executing on a CPU core of an SoC that includes the graphics core 1600C may submit a workload to one of multiple graphics processor doorbells, which invoke scheduling operations on an appropriate graphics engine. In at least one embodiment, the scheduling operations include determining which workload to run next, submitting the workload to a command stream converter, preempting an existing workload running on the engine, monitoring the progress of the workload, and notifying the host software when the workload is complete. In at least one embodiment, the graphics microcontroller 1638 may also facilitate a low-power or idle state for the graphics core 1600C, thereby providing the graphics core 1600C with the ability to save and restore registers across low-power state transitions independently of the operating system and / or graphics driver software on the system.

[0301] In at least one embodiment, the graphics core 1600C can have up to N more or fewer modular sub-cores than the shown sub-cores 1601A - 1601F. For each group of N sub-cores, in at least one embodiment, the graphics core 1600C can further include shared functional logic 1610, shared and / or cache memory 1612, a geometry / fixed-function pipeline 1614, and additional fixed-function logic 1616 to accelerate various graphics and compute processing operations. In at least one embodiment, the shared functional logic 1610 can include logic units (e.g., samplers, math, and / or inter-thread communication logic) that can be shared by each of the N sub-cores within the graphics core 1600C. In at least one embodiment, the fixed, shared, and / or cache memory 1612 can be the last-level cache for the N sub-cores 1601A - 1601F within the graphics core 1600C and can also be used as shared memory accessible by multiple sub-cores. In at least one embodiment, a geometry / fixed-function pipeline 1614 can be included in place of the geometry / fixed-function pipeline 1636 within the fixed-function block 1630 and can include similar logic units.

[0302] In at least one embodiment, the graphics core 1600C includes additional fixed-function logic 1616, which can include various fixed-function acceleration logics for use by the graphics core 1600C. In at least one embodiment, the additional fixed-function logic 1616 includes an additional geometry pipeline for use only in position shading. In position shading only, there are at least two geometry pipelines, and in the full geometry pipeline and culling pipeline within the geometry / fixed-function pipelines 1614, 1636, it is the additional geometry pipeline that can be included in the additional fixed-function logic 1616. In at least one embodiment, the culling pipeline is a trimmed version of the full geometry pipeline. In at least one embodiment, the full pipeline and the culling pipeline can execute different instances of an application, each instance having a separate environment. In at least one embodiment, position shading only can hide the long culling runs of discarded triangles, so that in some cases shading can be completed earlier. In at least one embodiment, the culling pipeline logic in the additional fixed-function logic 1616 can execute the position shader in parallel with the main application and generate critical results faster than the full pipeline because the culling pipeline fetches and masks the position attributes of vertices without performing rasterization and rendering pixels to the frame buffer. In at least one embodiment, the culling pipeline can use the generated critical results to calculate the visibility information for all triangles, regardless of whether these triangles are culled. In at least one embodiment, the full pipeline (which can be referred to as the replay pipeline in this case) can consume the visibility information to skip the culled triangles to only shade the visible triangles that are ultimately passed to the rasterization stage.

[0303] In at least one embodiment, the additional fixed function logic 1616 may further include machine learning acceleration logic, such as fixed function matrix multiplication logic, for implementing optimizations including those for machine learning training or inference.

[0304] In at least one embodiment, a set of execution resources is included within each graphics sub-core 1601A - 1601F, which can be used to perform graphics, media, and compute operations in response to requests from the graphics pipeline, media pipeline, or shader program. In at least one embodiment, the graphics sub-cores 1601A - 1601F include multiple EU arrays 1602A - 1602F, 1604A - 1604F, thread dispatch and inter-thread communication (TD / IC) logic 1603A - 1603F, 3D (e.g., texture) samplers 1605A - 1605F, media samplers 1606A - 1606F, shader processors 1607A - 1607F, and shared local memory (SLM) 1608A - 1608F. Each of the EU arrays 1602A - 1602F, 1604A - 1604F contains multiple execution units, which are general-purpose graphics processing units capable of servicing graphics, media, or compute operations, performing floating-point and integer / fixed-point logical operations, including graphics, media, or compute shader programs. In at least one embodiment, the TD / IC logic 1603A - 1603F performs local thread dispatch and thread control operations for the execution units within the sub-core and facilitates communication between the threads executing on the execution units of the sub-core. In at least one embodiment, the 3D samplers 1605A - 1605F can read data related to textures or other 3D graphics into memory. In at least one embodiment, the 3D sampler can read texture data differently based on the configured sampling state and texture format associated with a given texture. In at least one embodiment, the media samplers 1606A - 1606F can perform similar read operations based on the type and format associated with the media data. In at least one embodiment, each graphics sub-core 1601A - 1601F may alternatively include a unified 3D and media sampler. In at least one embodiment, the threads executing on the execution units within each sub-core 1601A - 1601F can utilize the shared local memory 1608A - 1608F within each sub-core to enable the threads executing within a thread group to use a common pool of on-chip memory for execution.

[0305] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. As described herein in connection with Figure 6B and / or Figure 6CProvide details regarding inference and / or training logic 615. In at least one embodiment, some or all of inference and / or training logic 615 may be incorporated into graphics processor 1610. In at least one embodiment, in at least one embodiment, the training and / or inference techniques described herein may be used in Figure 16B one or more ALUs embodied in graphics processor 1612, graphics microcontroller 1638, geometric and fixed function pipelines 1614 and 1636, or other logic in Figure 6B and / or Figure 6C shown. Additionally, in at least one embodiment, the inference and / or training operations described herein may be completed using logic other than the

[0306] Figures 16D - 16E logic shown. In at least one embodiment, weight parameters may be stored in on-chip or off-chip memory and / or registers (shown or not shown) that configure the ALUs of graphics processor 1600C to perform one or more of the machine learning algorithms, neural network architectures, use cases, or training techniques introduced herein. Figure 16D An embodiment is shown in which thread execution logic 1600D is used. Figure 16E Exemplary internal details of an execution unit are shown according to at least one embodiment.

[0307] As Figure 16DAs shown, in at least one embodiment, the thread execution logic 1600D includes a shader processor 1602, a thread dispatcher 1604, an instruction cache 1606, a scalable execution unit array including a plurality of execution units 1608A - 1608N, one or more samplers 1610, a data cache 1612, and a data port 1614. In at least one embodiment, the scalable execution unit array can be dynamically scaled, for example, based on the computational requirements of the workload, by enabling or disabling one or more execution units (e.g., any one of execution units 1608A, 1608B, 1608C, 1608D, through 1608N - 1 and 1608N). In at least one embodiment, the scalable execution units are interconnected by an interconnect structure linked to each execution unit. In at least one embodiment, the thread execution logic 1600D includes one or more connections to memory (such as system memory or cache memory) through the instruction cache 1606, the data port 1614, the sampler 1610, and one or more of the execution units 1608A - 1608N. In at least one embodiment, each execution unit (e.g., 1608A) is an independent programmable general - purpose computing unit capable of executing multiple simultaneous hardware threads while processing multiple data elements in parallel for each thread. In at least one embodiment, the array of execution units 1608A - 1608N can be scaled to include any number of individual execution units.

[0308] In at least one embodiment, the execution units 1608A - 1608N are mainly used to execute shader programs. In at least one embodiment, the shader processor 1602 can process various shader programs and dispatch execution threads associated with the shader programs via the thread dispatcher 1604. In at least one embodiment, the thread dispatcher 1604 includes logic for arbitrating thread initialization requests from the graphics and media pipelines and instantiating the requested threads on one or more of the execution units 1608A - 1608N. In at least one embodiment, in at least one embodiment, the geometry pipeline can dispatch vertex, tessellation, or geometry shaders to the thread execution logic for processing. In at least one embodiment, the thread dispatcher 1604 can also process runtime thread generation requests from executing shader programs.

[0309] In at least one embodiment, execution units 1608A - 1608N support an instruction set that includes native support for many standard 3D graphics shader instructions, enabling shader programs in graphics libraries (such as Direct3D and OpenGL) to execute with minimal translation. In at least one embodiment, the execution units support vertex and geometry processing (e.g., vertex programs, geometry programs, vertex shaders), pixel processing (e.g., pixel shaders, fragment shaders), and general - purpose processing (e.g., compute and media shaders). In at least one embodiment, each of the execution units 1608A - 1608N includes one or more arithmetic - logic units (ALUs) capable of performing multiple - issue single - instruction multiple - data (SIMD), and multithreaded operation enables an efficient execution environment despite higher - latency memory accesses. In at least one embodiment, each hardware thread within each execution unit has a dedicated high - bandwidth register file and associated independent thread state. In at least one embodiment, execution is multiple - issue per clock to the pipeline, and the pipeline is capable of integer, single - precision, and double - precision floating - point operations, SIMD branch functions, logical operations, transcendental operations, and other operations. In at least one embodiment, when waiting for data from one of memory or a shared function, dependency logic within the execution units 1608A - 1608N causes the waiting threads to sleep until the requested data is returned. In at least one embodiment, when the waiting threads are sleeping, hardware resources can be dedicated to processing other threads. In at least one embodiment, during the latency associated with vertex - shader operations, the execution units can perform operations on pixel shaders, fragment shaders, or another type of shader program (including a different vertex shader).

[0310] In at least one embodiment, each of the execution units 1608A - 1608N operates on an array of data elements. In at least one embodiment, the multiple data elements are the "execution size" or the number of lanes of an instruction. In at least one embodiment, an execution lane is a logical unit for the execution of data - element access, masking, and flow control within an instruction. In at least one embodiment, the multiple lanes can be independent of the multiple physical arithmetic - logic units (ALUs) or floating - point units (FPUs) for a particular graphics processor. In at least one embodiment, the execution units 1608A - 1608N support integer and floating - point data types.

[0311] In at least one embodiment, the execution unit instruction set includes SIMD instructions. In at least one embodiment, various data elements may be stored in registers as packed data types, and the execution unit will process the various elements based on the data size of those elements. In at least one embodiment, when operating on a 256-bit wide vector, the 256 bits of the vector are stored in a register, and the execution unit operates on the vector as four separate 64-bit packed data elements (quadword (QW) sized data elements), eight separate 32-bit packed data elements (doubleword (DW) sized data elements), sixteen separate 16-bit packed data elements (word (W) sized data elements), or thirty-two separate 8-bit data elements (byte (B) sized data elements). However, in at least one embodiment, different vector widths and register sizes are possible.

[0312] In at least one embodiment, one or more execution units may be combined into fused execution units 1609A - 1609N having thread control logic (1607A - 1607N) for executing threads for the fused EUs. In at least one embodiment, multiple EUs may be combined into an EU group. In at least one embodiment, each EU in the fused EU group may be configured to execute a separate SIMD hardware thread. The number of EUs in the fused EU group may vary according to the respective embodiments. In at least one embodiment, each EU may execute various SIMD widths, including but not limited to SIMD8, SIMD16, and SIMD32. In at least one embodiment, each fused graphics execution unit 1609A - 1609N includes at least two execution units. In at least one embodiment, in at least one embodiment, the fused execution unit 1609A includes a first EU 1608A, a second EU 1608B, and thread control logic 1607A common to the first EU 1608A and the second EU 1608B. In at least one embodiment, the thread control logic 1607A controls the threads executed on the fused graphics execution unit 1609A, thereby allowing each EU within the fused execution units 1609A - 1609N to execute using a common instruction pointer register.

[0313] In at least one embodiment, one or more internal instruction caches (e.g., instruction cache 1606) are included in the thread execution logic 1600D to cache thread instructions for the execution units. In at least one embodiment, one or more data caches (e.g., data cache 1612) are included to cache thread data during thread execution. In at least one embodiment, a sampler 1610 is included to provide texture sampling for 3D operations and media sampling for media operations. In at least one embodiment, the sampler 1610 includes specialized texture or media sampling functionality to process texture or media data during the sampling process before providing the sampled data to the execution units.

[0314] During execution, in at least one embodiment, the graphics and media pipeline sends thread initiation requests to the thread execution logic 1600D via the thread generation and dispatch logic. In at least one embodiment, once a set of geometric objects has been processed and rasterized into pixel data, pixel processor logic (e.g., pixel shader logic, fragment shader logic, etc.) within the shader processor 1602 is invoked to further compute output information and cause the results to be written to an output surface (e.g., color buffer, depth buffer, stencil buffer, etc.). In at least one embodiment, the pixel shader or fragment shader computes the values of various vertex attributes to be interpolated over the rasterized objects. In at least one embodiment, the pixel processor logic within the shader processor 1602 then executes the pixel or fragment shader program provided by the application programming interface (API). In at least one embodiment, to execute the shader program, the shader processor 1602 dispatches threads to the execution units (e.g., execution unit 1608A) via the thread dispatcher 1604. In at least one embodiment, the shader processor 1602 uses the texture sampling logic in the sampler 1610 to access texture data in a texture map stored in memory. In at least one embodiment, arithmetic operations on the texture data and the input geometric data compute pixel color data for each geometric fragment, or discard one or more pixels for further processing.

[0315] In at least one embodiment, the data port 1614 provides a memory access mechanism for the thread execution logic 1600D to output the processed data to memory for further processing on the graphics processor output pipeline. In at least one embodiment, the data port 1614 includes or is coupled to one or more cache memories (e.g., data cache 1612) to cache data for memory access via the data port.

[0316] As Figure 16EAs shown, in at least one embodiment, the graphics execution unit 1608 may include an instruction fetch unit 1637, a general register file array (GRF) 1624, an architecture register file array (ARF) 1626, a thread arbiter 1622, a dispatch unit 1630, a branch unit 1632, a set of SIMD floating-point units (FPU) 1634, and in at least one embodiment, a set of dedicated integer SIMD ALUs 1635. In at least one embodiment, the GRF 1624 and the ARF 1626 include a set of general register files and architecture register files associated with each simultaneously active hardware thread that may be in the graphics execution unit 1608. In at least one embodiment, each thread architecture state is maintained in the ARF 1626, while data used during thread execution is stored in the GRF 1624. In at least one embodiment, the execution state of each thread, including the instruction pointer of each thread, may be saved in a thread-specific register in the ARF 1626.

[0317] In at least one embodiment, the graphics execution unit 1608 has an architecture that is a combination of simultaneous multithreading (SMT) and fine-grained interleaved multithreading (IMT). In at least one embodiment, the architecture has a modular configuration that can be fine-tuned at design time based on the target number of simultaneous threads and the number of registers per execution unit, where the execution unit resources are logically allocated for executing multiple simultaneous threads.

[0318] In at least one embodiment, the graphics execution unit 1608 may co-issue multiple instructions, each of which may be a different instruction. In at least one embodiment, the thread arbiter 1622 of the graphics execution unit thread 1608 may dispatch instructions to one of the dispatch unit 1630, the branch unit 1632, or the SIMD FPU 1632 for execution. In at least one embodiment, each execution thread may access 128 general registers in the GRF 1624, where each register may store 32 bytes and may be accessed as a SIMD 8-element vector of 32-bit data elements. In at least one embodiment, each execution unit thread may access 4KB in the GRF 1624, although the embodiments are not limited thereto and more or fewer register resources may be provided in other embodiments. In at least one embodiment, although the number of threads per execution unit may also vary according to the embodiment, up to seven threads may be executed simultaneously. In at least one embodiment in which seven threads may access 4KB, the GRF 1624 may store a total of 28KB. In at least one embodiment, flexible addressing modes may allow registers to be addressed together to efficiently form wider registers or represent strided rectangular block data structures.

[0319] In at least one embodiment, memory operations, sampler operations, and other longer-latency system communications are scheduled via "send" instructions executed by the messaging sending unit 1630. In at least one embodiment, dispatching branch instructions to a dedicated branch unit 1632 facilitates SIMD divergence and eventual convergence.

[0320] In at least one embodiment, the graphics execution unit 1608 includes one or more SIMD floating-point units (FPUs) 1634 to perform floating-point operations. In at least one embodiment, one or more FPUs 1634 also support integer computations. In at least one embodiment, one or more FPUs 1634 can perform up to M 32-bit floating-point (or integer) operations SIMD, or up to 2M 16-bit integer or 16-bit floating-point operations SIMD. In at least one embodiment, at least one of the one or more FPUs provides extended mathematical capabilities to support high-throughput transcendental mathematical functions and double-precision 64-bit floating point. In at least one embodiment, there is also a set of 8-bit integer SIMD ALUs 1635, and can be specifically optimized to perform operations related to machine learning computations.

[0321] In at least one embodiment, an array of multiple instances of the graphics execution unit 1608 can be instantiated in a graphics sub-core grouping (e.g., sub-slice). In at least one embodiment, the execution unit 1608 can execute instructions across multiple execution channels. In at least one embodiment, each thread executed on the graphics execution unit 1608 executes on a different channel.

[0322] The inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. Details regarding the inference and / or training logic 615 are provided below in conjunction with Figure 6B and / or Figure 6C In at least one embodiment, part or all of the inference and / or training logic 615 can be incorporated into the execution logic 1600D. Additionally, in at least one embodiment, logic other than that shown in Figure 6B and / or Figure 6C can be used to complete the inference and / or training operations described herein. In at least one embodiment, the weight parameters can be stored in on-chip or off-chip memory and / or registers (shown or not shown), which configure the ALU of the execution logic 1600D to perform one or more machine learning algorithms, neural network architectures, use cases, or training techniques introduced herein.

[0323] Figure 17AShows a parallel processing unit (“PPU”) 1700A according to at least one embodiment. In at least one embodiment, the PPU 1700A is configured with machine-readable code that, if executed by the PPU 1700A, causes the PPU 1700A to perform some or all of the processes and techniques described throughout this disclosure. In at least one embodiment, the PPU 1700A is a multi-threaded processor implemented on one or more integrated circuit devices and utilizes multi-threading as a latency hiding technique designed to process computer-readable instructions (also referred to as machine-readable instructions or simply instructions) that are executed in parallel on multiple threads. In at least one embodiment, a thread refers to an execution thread and is an instance of a set of instructions configured to be executed by the PPU 1700A. In at least one embodiment, the PPU 1700A is a graphics processing unit (“GPU”), and the graphics processing unit is configured to implement a graphics rendering pipeline for processing three-dimensional (“3D”) graphics data in order to generate two-dimensional (“2D”) image data for display on a display device (such as a liquid crystal display (“LCD”) device). In at least one embodiment, the PPU 1700A is used to perform computations such as linear algebra operations and machine learning operations. Figure 17A An example parallel processor is shown for illustrative purposes only and should be construed as a non-limiting example of a processor architecture envisioned within the scope of this disclosure, and any suitable processor may be employed to supplement and / or replace it.

[0324] In at least one embodiment, one or more PPU 1700A are configured to accelerate high-performance computing (“HPC”), data center, and machine learning applications. In at least one embodiment, the PPU 1700A is configured to accelerate deep learning systems and applications, including the following non-limiting examples: autonomous vehicle platforms, deep learning, high-precision speech, image, text recognition systems, intelligent video analysis, molecular simulation, drug discovery, disease diagnosis, weather forecasting, big data analysis, astronomy, molecular dynamics simulation, financial modeling, robotics, factory automation, real-time language translation, online search optimization, and personalized user recommendations, etc.

[0325] In at least one embodiment, the PPU 1700A includes, but is not limited to, an input / output (“I / O”) unit 1706, a front-end unit 1710, a scheduler unit 1712, a work distribution unit 1714, a hub 1716, a crossbar (“Xbar”) 1720, one or more general processing clusters (“GPCs”) 1718, and one or more partition units (“memory partition units”) 1722. In at least one embodiment, the PPU 1700A is connected to a host processor or other PPU 1700A via one or more high-speed GPU interconnects (“GPU interconnects”) 1708. In at least one embodiment, the PPU 1700A is connected to a host processor or other peripheral devices via an interconnect 1702. In one embodiment, the PPU 1700A is connected to a local memory including one or more memory devices (“memory”) 1704. In at least one embodiment, the memory device 1704 includes, but is not limited to, one or more dynamic random access memory (“DRAM”) devices. In at least one embodiment, one or more DRAM devices are configured and / or configurable as a high-bandwidth memory (“HBM”) subsystem, and multiple DRAM dies are stacked within each device.

[0326] In at least one embodiment, the high-speed GPU interconnect 1708 may refer to a wire-based multi-channel communication link used by the system for scaling and includes one or more PPU 1700As in combination with one or more central processing units (“CPUs”), supporting cache coherence between the PPU 1700A and the CPU and CPU master control. In at least one embodiment, the high-speed GPU interconnect 1708 transfers data and / or commands to other units of the PPU 1700A, such as one or more copy engines, a video encoder, a video decoder, a power management unit, and / or other components that may not be explicitly shown in Figure 17A it.

[0327] In at least one embodiment, the I / O unit 1706 is configured to receive from the host processor via the interconnect 1702 ( Figure 17A(not shown in the figure) to send and receive communications (e.g., commands, data). In at least one embodiment, I / O unit 1706 communicates directly with the host processor via interconnect 1702 or through one or more intermediate devices (such as a memory bridge). In at least one embodiment, I / O unit 1706 may communicate with one or more other processors (such as one or more PPU 1700A) via interconnect 1702. In at least one embodiment, I / O unit 1706 implements a Peripheral Component Interconnect Express ("PCIe") interface for communication over a PCIe bus. In at least one embodiment, I / O unit 1706 implements an interface for communicating with external devices.

[0328] In at least one embodiment, I / O unit 1706 decodes packets received via interconnect 1702. In at least one embodiment, at least some of the packets represent commands configured to cause PPU 1700A to perform various operations. In at least one embodiment, I / O unit 1706 sends the decoded commands to various other units of PPU 1700A as specified by the commands. In at least one embodiment, the commands are sent to the front-end unit 1710 and / or sent to the hub 1716 or other units of PPU 1700A, such as one or more copy engines, video encoders, video decoders, power management units, etc. ( Figure 17A (not explicitly shown in the figure). In at least one embodiment, I / O unit 1706 is configured to route communications between various logical units of PPU 1700A.

[0329] In at least one embodiment, a program executed by the host processor encodes a command stream in a buffer that provides a workload to PPU 1700A for processing. In at least one embodiment, the workload includes instructions and data to be processed by those instructions. In at least one embodiment, the buffer is an area in a memory that can be accessed (e.g., read / written) by both the host processor and PPU 1700A - the host interface unit may be configured to access the buffer in the system memory connected to interconnect 1702 via a memory request transmitted through I / O unit 1706 via interconnect 1702. In at least one embodiment, the host processor writes the command stream to the buffer and then sends a pointer indicating the start of the command stream to PPU 1700A, such that the front-end unit 1710 receives one or more command stream pointers and manages one or more command streams, reads commands from the command streams and forwards the commands to various units of PPU 1700A.

[0330] In at least one embodiment, the front-end unit 1710 is coupled to a scheduler unit 1712 that configures various GPCs 1718 to process tasks defined by one or more command streams. In at least one embodiment, the scheduler unit 1712 is configured to track state information related to the various tasks managed by the scheduler unit 1712, where the state information may indicate which GPC 1718 a task is assigned to, whether the task is active or inactive, the priority associated with the task, and so on. In at least one embodiment, the scheduler unit 1712 manages multiple tasks executed on one or more GPCs 1718.

[0331] In at least one embodiment, the scheduler unit 1712 is coupled to a work distribution unit 1714 that is configured to dispatch tasks for execution on the GPCs 1718. In at least one embodiment, the work distribution unit 1714 tracks the multiple scheduled tasks received from the scheduler unit 1712 and the work distribution unit 1714 manages a pool of pending tasks and a pool of active tasks for each GPC 1718. In at least one embodiment, the pool of pending tasks includes multiple time slots (e.g., 32 time slots) that contain tasks assigned to be processed by a particular GPC 1718; the pool of active tasks may include multiple time slots (e.g., 4 time slots) for tasks to be actively processed by the GPC 1718 such that as one of the GPCs 1718 finishes executing a task, that task is evicted from the active task pool of the GPC 1718, and one of the other tasks from the pool of pending tasks is selected and scheduled for execution on the GPC 1718. In at least one embodiment, if an active task is idle on the GPC 1718, e.g., while waiting for a data dependency resolution, the active task is evicted from the GPC 1718 and returned to the pool of pending tasks while another task from the pool of pending tasks is selected and scheduled for execution on the GPC 1718.

[0332] In at least one embodiment, the work distribution unit 1714 communicates with one or more GPCs 1718 via an XBar 1720. In at least one embodiment, the XBar 1720 is an interconnect network that couples many of the units of the PPU 1700A to other units of the PPU 1700A and may be configured to couple the work distribution unit 1714 to a particular GPC 1718. In at least one embodiment, other units of one or more PPUs 1700A may also be connected to the XBar 1716 via a hub 1716.

[0333] In at least one embodiment, tasks are managed by a scheduler unit 1712 and assigned to one of the GPCs 1718 by a work distribution unit 1714. The GPC 1718 is configured to process the tasks and produce results. In at least one embodiment, the results can be consumed by other tasks in the GPC 1718, routed to a different GPC 1718 via the XBar 1716, or stored in the memory 1704. In at least one embodiment, the results can be written to the memory 1704 by a partitioning unit 1722, which implements a memory interface for writing data to or reading data from the memory 1704. In at least one embodiment, the results can be transmitted to another PPU 1704 or CPU via the high-speed GPU interconnect 1708. In at least one embodiment, the PPU 1700A includes, but is not limited to, U partitioning units 1722, where the number of partitioning units 1722 is equal to the number of separate and distinct memory devices 1704 coupled to the PPU 1700A. In at least one embodiment, the partitioning unit 1722 will be described in more detail below in conjunction with Figure 17C The partitioning unit 1722 will be described in more detail.

[0334] In at least one embodiment, a host processor executes a driver core that implements an application programming interface (API) that enables one or more applications executing on the host processor to schedule operations for execution on the PPU 1700A. In one embodiment, multiple compute applications are executed simultaneously by the PPU 1700A, and the PPU 1700A provides isolation, quality of service (“QoS”), and independent address spaces for the multiple compute applications. In at least one embodiment, an application generates instructions (e.g., in the form of API calls) that cause the driver core to generate one or more tasks for execution by the PPU 1700A, and the driver core outputs the tasks to one or more streams to be processed by the PPU 1700A. In at least one embodiment, each task includes one or more associated thread groups, which may be referred to as warps. In at least one embodiment, a warp includes multiple associated threads (e.g., 32 threads) that can be executed in parallel. In at least one embodiment, cooperating threads may refer to multiple threads that include instructions for performing a task and exchanging data via shared memory, in conjunction with Figure 17C Threads and cooperating threads are described in more detail according to at least one embodiment.

[0335] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. Below in conjunction with Figure 6B and / or Figure 6CProvide details regarding inference and / or training logic 615. In at least one embodiment, a deep learning application processor is used to train a machine learning model (such as a neural network) to predict or infer information provided to PPU 1700A. In at least one embodiment, PPU 1700A is used to infer or predict information based on a trained machine learning model (e.g., a neural network) that has been trained by another processor or system or PPU 1700A. In at least one embodiment, PPU 1700A can be used to execute one or more of the neural network use cases herein.

[0336] Figure 17B Illustrates a general processing cluster (“GPC”) 1700B according to at least one embodiment. In at least one embodiment, GPC 1700B is Figure 17A the GPC 1718. In at least one embodiment, each GPC 1700B includes, but is not limited to, multiple hardware units for processing tasks, and each GPC 1700B includes, but is not limited to, a pipeline manager 1702, a pre-raster operation unit (“PROP”) 1704, a raster engine 1708, a work distribution crossbar (“WDX”) 1716, a memory management unit (“MMU”) 1718, one or more data processing clusters (“DPC”) 1706, and any suitable combination of components.

[0337] In at least one embodiment, the operation of GPC 1700B is controlled by pipeline manager 1702. In at least one embodiment, pipeline manager 1702 manages the configuration of one or more DPCs 1706 to process tasks assigned to GPC 1700B. In at least one embodiment, pipeline manager 1702 configures at least one of the one or more DPCs 1706 to implement at least a portion of a graphics rendering pipeline. In at least one embodiment, DPC 1706 is configured to execute a vertex shader program on programmable stream multiprocessors (“SM”) 1714. In at least one embodiment, pipeline manager 1702 is configured to route data packets received from a work distribution unit to appropriate logic units within GPC 1700B, and in at least one embodiment, some data packets can be routed to fixed function hardware units in PROP 1704 and / or raster engine 1708, while other data packets can be routed to DPC 1706 for processing by primitive engine 1712 or SM 1714. In at least one embodiment, pipeline manager 1702 configures at least one of DPCs 1706 to implement a neural network model and / or a computational pipeline.

[0338] In at least one embodiment, the PROP unit 1704 is configured to route data generated by the raster engine 1708 and the DPC 1706 to the raster operation (“ROP”) units in the partition unit 1722, as described in more detail above in conjunction with Figure 17A Figure 17A . In at least one embodiment, the PROP unit 1704 is configured to perform optimizations for color blending, organize pixel data, perform address translation, and so on. In at least one embodiment, the raster engine 1708 includes, but is not limited to, a plurality of fixed-function hardware units configured to perform various raster operations, and in at least one embodiment, the raster engine 1708 includes, but is not limited to, a setup engine, a coarse raster engine, a culling engine, a clipping engine, a fine raster engine, a tile aggregation engine, and any suitable combination thereof. In at least one embodiment, the setup engine receives the transformed vertices and generates plane equations associated with the geometric primitives defined by the vertices; the plane equations are transmitted to the coarse raster engine to generate coverage information of the primitive (e.g., the x, y coverage range masks of the tile); the output of the coarse raster engine is transmitted to the culling engine, where the fragments associated with the primitives that fail the z-test are culled and transmitted to the clipping engine, where the fragments located outside the view frustum are clipped. In at least one embodiment, the clipped and culled fragments are passed to the fine raster engine to generate the attributes of the pixel fragments based on the plane equations generated by the setup engine. In at least one embodiment, the output of the raster engine 1708 includes the fragments that will be processed by any suitable entity (e.g., the fragment shader implemented within the DPC 1706).

[0339] In at least one embodiment, each DPC 1706 included in the GPC 1700B includes, but is not limited to, an M pipeline controller (“MPC”) 1710; a primitive engine 1712; one or more SMs 1714; and any suitable combination thereof. In at least one embodiment, the MPC 1710 controls the operation of the DPC 1706 and routes the packets received from the pipeline manager 1702 to the appropriate units in the DPC 1706. In at least one embodiment, the packets associated with the vertices are routed to the primitive engine 1712, and the primitive engine 1712 is configured to fetch the vertex attributes associated with the vertices from the memory; conversely, the data packets associated with the shader programs can be sent to the SM 1714.

[0340] In at least one embodiment, SM 1714 includes, but is not limited to, programmable streaming processors configured to process tasks represented by multiple threads. In at least one embodiment, SM 1714 is multi-threaded and configured to execute multiple threads (e.g., 32 threads) from a particular thread group simultaneously, and implements a single instruction, multiple data (“SIMD”) architecture, where each thread in a group of threads (e.g., a warp) is configured to process different data sets based on the same instruction set. In at least one embodiment, all threads in a thread group execute the same instruction. In at least one embodiment, SM 1714 implements a single instruction, multiple thread (“SIMT”) architecture, where each thread in a group of threads is configured to process different data sets based on the same instruction set, but where individual threads in a thread group are allowed to diverge during execution. In at least one embodiment, a program counter, call stack, and execution state are maintained for each warp, enabling concurrency between warps and serial execution within a warp when threads in the warp diverge. In another embodiment, a program counter, call stack, and execution state are maintained for each individual thread, enabling equal concurrency between all threads within and between warps. In at least one embodiment, an execution state is maintained for each individual thread, and threads that can converge and execute the same instruction in parallel can be executed for increased efficiency. At least one embodiment of SM 1714 is described in more detail below.

[0341] In at least one embodiment, MMU 1718 provides an interface between GPC 1700B and a memory partition unit (e.g., Figure 17A the partition unit 1722), and MMU 1718 provides virtual address to physical address translation, memory protection, and arbitration of memory requests. In at least one embodiment, MMU 1718 provides one or more translation lookaside buffers (“TLBs”) for performing virtual address to physical address translation in memory.

[0342] Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 615 are provided below in conjunction with Figure 6B and / or Figure 6C In at least one embodiment, a deep learning application processor is used to train a machine learning model (such as a neural network) to predict or infer information provided to GPC 1700B. In at least one embodiment, GPC 1700B is used to infer or predict information based on a machine learning model (e.g., a neural network) that has been trained by another processor or system or GPC 1700B. In at least one embodiment, GPC 1700B can be used to execute one or more of the neural network use cases described herein.

[0343] Figure 17C Shows the memory partition unit 1700C of a parallel processing unit (“PPU”) according to at least one embodiment. In at least one embodiment, the memory partition unit 1700C includes, but is not limited to, a raster operation (“ROP”) unit 1702; a secondary (“L2”) cache 1704; a memory interface 1706; and any suitable combination thereof. In at least one embodiment, the memory interface 1706 is coupled to a memory. In at least one embodiment, the memory interface 1706 can implement a 32, 64, 128, 1024-bit data bus, or a similar implementation for high-speed data transfer. In at least one embodiment, the PPU includes U memory interfaces 1706, one memory interface 1706 for each pair of partition units 1700C, where each pair of partition units 1700C is connected to a corresponding memory device. In at least one embodiment, in at least one embodiment, the PPU can be connected to up to Y memory devices, such as a high-bandwidth memory stack or a graphics double data rate version 5 synchronous dynamic random access memory (“GDDR5 SDRAM”).

[0344] In at least one embodiment, the memory interface 1706 implements a high-bandwidth memory second generation (“HBM2”) memory interface, and Y is equal to half of U. In at least one embodiment, the HBM2 memory stack is located on the same physical package as the PPU, providing a large amount of power and saving area compared to a GDDR5 SDRAM system. In at least one embodiment, each HBM2 stack includes, but is not limited to, four memory dies, and Y = 4. Each HBM2 stack includes two 128-bit channels per die for a total of 8 channels and a 1024-bit data bus width. In at least one embodiment, the memory supports single error correction double error detection (“SECDED”) error correction code (“ECC”) to protect data. In at least one embodiment, the ECC provides higher reliability for computational applications sensitive to data corruption.

[0345] In at least one embodiment, the PPU implements a multi-level memory hierarchy. In at least one embodiment, the memory partition unit 1700C supports unified memory to provide a single unified virtual address space for the central processing unit (“CPU”) and PPU memory, enabling data sharing between virtual memory systems. In at least one embodiment, the access frequency of the PPU to memory located on other processors is tracked to ensure that memory pages are moved to the physical memory of the PPU that accesses the pages more frequently. In at least one embodiment, the high-speed GPU interconnect 1708 supports address translation services, which allow the PPU to directly access the CPU's page table and provide full access to the CPU memory through the PPU.

[0346] In at least one embodiment, the copy engine transfers data between multiple PPUs or between a PPU and a CPU. In at least one embodiment, the copy engine can generate a page fault for an address not mapped to a page table, and then the memory partition unit 1700C services the page fault by mapping the address to the page table, after which the copy engine performs the transfer. In at least one embodiment, memory is fixed (in at least one embodiment, non-pageable) for multiple copy engine operations between multiple processors, thereby substantially reducing the available memory. In at least one embodiment, in the case of a hardware page fault, the address can be passed to the copy engine without regard to whether the memory page is resident, and the copy process is transparent.

[0347] According to at least one embodiment, data from Figure 17A memory 1704 or other system memory is fetched by the memory partition unit 1700C and stored in the L2 cache 1704, which is located on-chip and shared among various GPCs. In at least one embodiment, each memory partition unit 1700C includes, but is not limited to, at least a portion of the L2 cache associated with a corresponding memory device. In at least one embodiment, lower-level caches are implemented in various units within a GPC. In at least one embodiment, each SM 1714 can implement a level-one (“L1”) cache, where the L1 cache is private memory dedicated to a particular SM 1714 and fetches data from the L2 cache 1704 and stores it in each L1 cache for processing in the functional units of the SM 1714. In at least one embodiment, the L2 cache 1704 is coupled to the memory interface 1706 and the XBar 1720.

[0348] In at least one embodiment, the ROP unit 1702 performs graphics raster operations related to pixel colors, such as color compression, pixel blending, etc. In at least one embodiment, the ROP unit 1702 implements a depth test in conjunction with the raster engine 1708, receiving the depth of the sample position associated with the pixel fragment from the culling engine of the raster engine 1708. In at least one embodiment, for the corresponding depth test depth in the depth buffer at the sample position associated with the fragment. In at least one embodiment, if the fragment passes the depth test for the sample position, the ROP unit 1702 updates the depth buffer and sends the result of the depth test to the raster engine 1708. It will be appreciated that the number of partition units 1700C may be different from the number of GPCs, and thus, in at least one embodiment, each ROP unit 1702 may be coupled to each GPC. In at least one embodiment, the ROP unit 1702 tracks packets received from different GPCs and determines which result the result generated by the ROP unit 1702 is routed to via the XBar 1720.

[0349] Figure 17D Shows a streaming multiprocessor ("SM") 1700D according to at least one embodiment. In at least one embodiment, the SM 1700D is Figure 17BThe SM. In at least one embodiment, the SM 1700D includes, but is not limited to, an instruction cache 1702; one or more scheduler units 1704; a register file 1708; one or more processing cores (“cores”) 1710; one or more special function units (“SFUs”) 1712; one or more load / store units (“LSUs”) 1714; an interconnect network 1716; a shared memory / level 1 (“L1”) cache 1718; and any suitable combination thereof. In at least one embodiment, a work distribution unit schedules tasks to be executed on a general processing cluster (“GPC”) of a parallel processing unit (“PPU”), and each task is assigned to a specific data processing cluster (“DPC”) within the GPC, and if the task is associated with a shader program, the task is assigned to one of the SMs 1700D. In at least one embodiment, the scheduler unit 1704 receives tasks from the work distribution unit and manages the instruction scheduling for one or more thread blocks assigned to the SM 1700D. In at least one embodiment, the scheduler unit 1704 schedules thread blocks to execute as warps of parallel threads, where each thread block is assigned at least one warp. In at least one embodiment, each warp executes threads. In at least one embodiment, the scheduler unit 1704 manages multiple different thread blocks, assigns warps to different thread blocks, and then dispatches instructions from multiple different cooperative groups to various functional units (e.g., processing cores 1710, SFUs 1712, and LSUs 1714) within each clock cycle.

[0350] In at least one embodiment, a cooperative group may refer to a programming model for organizing groups of communication threads, which allows developers to express the granularity at which threads are communicating, enabling a richer and more effective parallel decomposition. In at least one embodiment, a cooperative launch API supports synchronization between thread blocks to execute parallel algorithms. In at least one embodiment, the programming model's applications provide a single, simple construct for synchronizing cooperative threads: a barrier for all threads across thread blocks (e.g., the syncthreads() function). However, in at least one embodiment, a programmer can define thread groups at a granularity smaller than that of a thread block and synchronize within the defined groups to achieve higher performance, design flexibility, and software reuse in the form of collective group-wide functional interfaces. In at least one embodiment, cooperative groups enable a programmer to explicitly define thread groups at sub-block (in at least one embodiment, as small as a single thread) and multi-block granularities and perform collective operations such as synchronizing the threads in a cooperative group. In at least one embodiment, the programming model supports clean composition across software boundaries, such that library and utility functions can synchronize safely in their native environments without having to make assumptions about convergence. In at least one embodiment, cooperative group primitives enable new patterns of cooperative parallelism, including but not limited to producer-consumer parallelism, opportunistic parallelism, and global synchronization across the entire thread block grid.

[0351] In at least one embodiment, a scheduling unit 1706 is configured to send instructions to one or more of the functional units, and the scheduler unit 1704 includes, but is not limited to, two scheduling units 1706 that enable two different instructions from the same warp to be scheduled in each clock cycle. In at least one embodiment, each scheduler unit 1704 includes a single scheduling unit 1706 or additional scheduling units 1706.

[0352] In at least one embodiment, each SM 1700D includes, but is not limited to, a register file 1708 in at least one embodiment, and this register file 1708 provides a set of registers for the functional units of the SM 1700D. In at least one embodiment, the register file 1708 is divided among each functional unit, so as to allocate a dedicated part of the register file 1708 to each functional unit. In at least one embodiment, the register file 1708 is divided among different warps executed by the SM 1700D, and the register file 1708 provides temporary storage for the operands of the data paths connected to the functional units. In at least one embodiment, each SM 1700D includes, but is not limited to, a plurality of L processing cores 1710 in at least one embodiment. In at least one embodiment, the SM 1700D includes, but is not limited to, a large number (e.g., 128 or more) of different processing cores 1710. In at least one embodiment, each processing core 1710 includes, but is not limited to, a fully pipelined, single-precision, double-precision, and / or mixed-precision processing unit, which includes, but is not limited to, a floating-point arithmetic logic unit and an integer arithmetic logic unit. In at least one embodiment, the floating-point arithmetic logic unit implements the IEEE 754-2008 standard for floating-point arithmetic. In at least one embodiment, the processing core 1710 includes, but is not limited to, 64 single-precision (32-bit) floating-point cores, 64 integer cores, 32 double-precision (64-bit) floating-point cores, and 8 tensor cores.

[0353] According to at least one embodiment, the tensor cores are configured to perform matrix operations. In at least one embodiment, one or more tensor cores are included in the processing core 1710. In at least one embodiment, the tensor cores are configured to perform deep learning matrix arithmetic, such as convolution operations for neural network training and inference. In at least one embodiment, each tensor core operates on a 4×4 matrix and performs matrix multiplication and accumulation operations D = A×B + C, where A, B, C, and D are 4×4 matrices.

[0354] In at least one embodiment, the matrix multiplication inputs A and B are 16-bit floating-point matrices, and the accumulation matrices C and D are 16-bit floating-point or 32-bit floating-point matrices. In at least one embodiment, the tensor core performs 32-bit floating-point accumulation operations on 16-bit floating-point input data. In at least one embodiment, the 16-bit floating-point multiplication uses 64 operations and obtains a full-precision product, which is then accumulated with other intermediate products using 32-bit floating-point addition for 4x4x4 matrix multiplication. In at least one embodiment, the tensor core is used to perform larger two-dimensional or higher-dimensional matrix operations composed of these smaller elements. In at least one embodiment, APIs (such as the CUDA 9 C++ API) expose specialized matrix load, matrix multiplication and accumulation, and matrix store operations to effectively use the tensor core from a CUDA-C++ program. In at least one embodiment, at the CUDA level, the warp-level interface assumes a 16×16 size matrix spanning all 32 warp threads.

[0355] In at least one embodiment, each SM 1700D includes, but is not limited to, M SFUs 1712 that perform special functions (e.g., attribute evaluation, reciprocal square root, etc.). In at least one embodiment, the SFU 1712 includes, but is not limited to, a tree traversal unit configured to traverse a hierarchical tree data structure. In at least one embodiment, the SFU 1712 includes, but is not limited to, a texture unit configured to perform texture mapping filtering operations. In at least one embodiment, the texture unit is configured to load a texture map (e.g., a 2D array of texture pixels) from memory and sample the texture map to produce a sampled texture value for use by a shader program executed by the SM 1700D. In at least one embodiment, the texture map is stored in the shared memory / L1 cache 1718. In at least one embodiment, according to at least one embodiment, the texture unit uses mip maps (e.g., texture maps with different levels of detail) to implement texture operations (such as filtering operations). In at least one embodiment, each SM 1700D includes, but is not limited to, two texture units.

[0356] In at least one embodiment, each SM 1700D...

Claims

1. A remedial system for threshold leakage in a data center liquid cooling system, comprising: A flow controller and a power controller within a power distribution unit (PDU), the flow controller and the power controller being adapted to receive inputs from a learning subsystem that executes a machine learning model, the learning subsystem being adapted to determine that at least one parameter associated with the data center liquid cooling system is outside a determined range, thereby indicating that a threshold leakage of coolant has occurred, the threshold leakage being associated with at least one computing component that operates within a normal temperature threshold and receives the coolant, the power controller being adapted to cause the at least one computing component to change its power state to reduce its dependence on the coolant, and the flow controller being adapted to cause a change in the flow rate of the coolant flowing to the at least one computing component. Wherein the learning subsystem is adapted to: Process parameter values associated with the at least one parameter using a plurality of neuron layers of the machine learning model; and After evaluating the parameter values, provide the input with the previously associated flow rate of the coolant and the previously associated power state of the at least one computing component.

2. The remedial system according to claim 1, wherein the threshold leakage indicates that while the at least one computing component operates within the normal temperature threshold and receives the coolant, a first quantity of coolant has incorrectly left the data center liquid cooling system, and wherein normal leakage is different from the threshold leakage in that it indicates at least a second quantity of coolant that exceeds the first quantity, the second quantity of coolant having incorrectly left the data center liquid cooling system such that the at least one computing component cannot operate within the normal temperature threshold and cannot receive the coolant to maintain the normal temperature threshold.

3. The remedial system according to claim 1, further comprising: At least one processor associated with the learning subsystem for controlling the flow controller and the power controller using the input, a first input in the input causing the shutdown of the at least one computing component, and a second input in the input causing the shutdown of the coolant.

4. The remedial system according to claim 3, further comprising: The at least one processor for enabling a load transfer subsystem to transfer the load associated with the at least one computing component before the at least one computing component is shut down, the load being transferred to at least one second computing component that receives the coolant or a second coolant that is not affected by the threshold leakage.

5. The remedial system according to claim 1, further comprising: At least one processor associated with the learning subsystem for determining that the threshold leakage of the coolant has occurred by determining that the pressure, flow rate, or temperature of the coolant entering and leaving the at least one computing component is outside the normal threshold of the determined range and within a warning range defined by at least an alarm threshold, at which alarm threshold the at least one computing component no longer operates within the normal temperature threshold and no longer receives the coolant to maintain the normal temperature threshold.

6. The remediation system according to claim 1, further comprising: A distributed control system, which includes one or more of the flow controller, the power controller, and at least a first part of the learning subsystem within the PDU, and includes one or more of an auxiliary flow controller, an auxiliary power controller, and a second part of the learning subsystem located in an auxiliary PDU; The distributed control system communicates the input from the learning subsystem with one or more of the auxiliary flow controller and the auxiliary power controller, and enables one or more of the auxiliary flow controller and the auxiliary power controller to cause at least one auxiliary computing component to change a second power state to reduce dependence on the coolant and cause a change in the flow rate of the coolant flowing to the at least one auxiliary computing component.

7. The remediation system according to claim 1, further comprising: The learning subsystem includes a processor associated with one or more of the flow controller of the PDU, the power controller of the PDU, the auxiliary flow controller of the auxiliary PDU, and the auxiliary power controller of the auxiliary PDU.

8. The remediation system according to claim 1, further comprising: The learning subsystem associated with at least one processor, for evaluating at least one parameter associated with the data center liquid cooling system using a moving range including the determined range, the moving range representing a parameter value of the at least one parameter, the parameter value being within a threshold of the at least one parameter when the at least one computing component operates within a normal temperature threshold and receives the coolant; And The learning subsystem, for providing a first input in the input to the power controller to cause the at least one computing component to change the power state to reduce dependence on the coolant, and providing a second input in the input to the flow controller to cause a change in the flow rate of the coolant flowing to the at least one computing component.

9. The remediation system according to claim 8, wherein the at least one parameter includes one or more of the following: The temperature of the coolant; The temperature of the at least one computing component or a first area including the at least one computing component; The temperature of the pipe carrying the coolant; The humidity or relative humidity of the first area or a second area including the flow controller; The flow rate of the coolant flowing to the first area; The flow rate of the coolant flowing from the first area; A proportional cooling response to the power drawn by the at least one computing component; and The fluid leakage rate of the coolant.

10. The remediation system according to claim 8, wherein the plurality of neuron levels have the parameter value and have a previously associated flow rate of the coolant and a previously associated power state of the at least one computing component.

11. A processor for remediating a threshold leak in a liquid cooling system, comprising: At least one logic unit for controlling a flow controller and a power controller within a power distribution unit (PDU), the flow controller and the power controller being for receiving inputs from a learning subsystem that executes a machine learning model, the learning subsystem being adapted to determine that at least one parameter associated with the liquid cooling system is outside a normal threshold of a determination range such that a threshold leak of coolant has occurred, the threshold leak being associated with at least one computing component that operates within a normal temperature threshold and receives the coolant, the power controller being for causing the at least one computing component to change its power state to reduce dependence on the coolant, and the flow controller being for causing a change in the flow rate of the coolant flowing to the at least one computing component, wherein the learning subsystem is for: processing parameter values associated with the at least one parameter using multiple neuron layers of the machine learning model; and after evaluating the parameter values, providing the input with a previously associated flow rate of the coolant and a previously associated power state of the at least one computing component.

12. The processor of claim 11, further comprising: the at least one logic unit associated with the learning subsystem for enabling a load transfer subsystem to transfer a load associated with the at least one computing component before the at least one computing component is shut down, the load being transferred to at least one second computing component that receives the coolant or a second coolant that is not affected by the threshold leak.

13. The processor of claim 11, further comprising: the at least one logic unit associated with the learning subsystem for determining that the threshold leak of the coolant has occurred by determining that the pressure, flow rate, or temperature of the coolant entering and leaving the at least one computing component is outside a normal threshold of the determination range and within a warning range defined by at least an alarm threshold, at which alarm threshold the at least one computing component no longer operates within the normal temperature threshold and no longer receives the coolant to maintain the normal temperature threshold.

14. The processor of claim 11, further comprising: an instruction output for transmitting the input from the learning subsystem to the flow controller and the power controller.

15. The processor of claim 11, further comprising: the at least one logic unit being adapted to receive parameter values from a parameter sensor associated with the liquid cooling system and functional information from a functional sensor of the at least one computing component, and being adapted to facilitate the change in the power state and the change in the flow rate of the coolant.

16. A processor for a liquid cooling system, comprising: At least one logic unit for training one or more neural networks having hidden layers of neurons to evaluate parameter values associated with at least one parameter of the liquid cooling system, performing the evaluation using the parameter values and using a previously associated flow rate of the coolant and a previously associated power state of at least one computing component, the evaluation for determining that a range of movement of the at least one parameter is outside a normal threshold and indicating that a threshold leak of the coolant has occurred, the threshold leak being associated with at least one computing component operating within a normal temperature threshold and receiving the coolant, and the one or more neural networks for providing an output associated with a first change in the power state of the at least one computing component and associated with a second change in the flow rate of the coolant flowing to the at least one computing component, the first change in the power state for reducing dependence on the coolant.

17. The processor of claim 16, further comprising: A learning subsystem of the at least one logic unit for executing a machine learning model to: Process parameter values associated with the at least one parameter using the previously associated flow rate of the coolant and the previously associated power state of the at least one computing component; And After evaluating the parameter values, provide the previously associated flow rate of the coolant and the previously associated power state to the output.

18. The processor of claim 16, further comprising: The at least one logic unit for outputting at least one instruction respectively associated with the first change in the power state and the second change in the flow rate of the coolant.

19. The processor of claim 16, further comprising: An instruction output for transmitting the output from the learning subsystem that executes the one or more neural networks to cause the first change in the power state of the at least one computing component and to cause the second change in the flow rate of the coolant to the at least one computing component.

20. The processor of claim 16, further comprising: The at least one logic unit adapted to receive parameter values from a parameter sensor associated with the liquid cooling system and functional information from a functional sensor of the at least one computing component, and adapted to facilitate respectively the first change in the power state and the second change in the flow rate of the coolant.

21. A remedial system for threshold leaks in a liquid cooling system, comprising: At least one processor for training one or more neural networks having a hidden layer of neurons to evaluate a parameter value associated with at least one parameter of the liquid cooling system, performing the evaluation using the parameter value and using a previously associated flow rate of the coolant and a previously associated power state of at least one computing component, the evaluation for determining that a movement range of the at least one parameter is outside a normal threshold and indicating that a threshold leak of the coolant has occurred, the threshold leak being associated with the at least one computing component operating within a normal temperature threshold and receiving the coolant, and the one or more neural networks providing an output associated with a first change in the power state of the at least one computing component and associated with a second change in the flow rate of the coolant flowing to the at least one computing component, the first change in the power state for reducing dependence on the coolant.

22. The remediation system of claim 21, further comprising: A learning subsystem of the at least one processor for executing a machine learning model to: Process a parameter value associated with the at least one parameter using the previously associated flow rate of the coolant and the previously associated power state of the at least one computing component; And After evaluating the parameter value, provide the previously associated flow rate of the coolant and the previously associated power state to the output.

23. The remediation system of claim 21, further comprising: At least one instruction for transmitting the output from the learning subsystem that executes the one or more neural networks to cause the first change in the power state of the at least one computing component and to cause the second change in the flow rate of the coolant to the at least one computing component.

24. The remediation system of claim 21, further comprising: The at least one processor unit for outputting at least one instruction respectively associated with the first change in the power state and the second change in the flow rate of the coolant.

25. The remediation system of claim 21, further comprising: The at least one processor adapted to receive a parameter value from a parameter sensor associated with the liquid cooling system and receive functional information from a functional sensor of the at least one computing component, and adapted to facilitate the first change in the power state and the second change in the flow rate of the coolant respectively.

26. A method for remediating a threshold leak in a liquid cooling system for a data center, comprising: Providing a flow controller and a power controller within a power distribution unit (PDU); Enabling a learning subsystem that executes a machine learning model to determine that at least one parameter associated with the data center liquid cooling system is outside a determined range and that a threshold leak of the coolant has occurred, the threshold leak being associated with at least one computing component operating within a normal temperature threshold and receiving the coolant; Enabling the flow controller and the power controller to receive an input from the learning subsystem; Causing the at least one computing component to change a power state by the power controller to reduce dependence on the coolant; And A change in the flow rate of the coolant flowing to the at least one computing component caused by the flow rate controller, wherein the learning subsystem processes parameter values associated with the at least one parameter using a plurality of neuron levels of the machine learning model, and after evaluating the parameter values, provides the previously associated flow rate of the coolant and the previously associated power state of the at least one computing component to the input.

27. The remedial method according to claim 26, wherein the threshold leakage indicates that while the at least one computing component is operating within the normal temperature threshold and receiving the coolant, a first quantity of the coolant is incorrectly leaving the data center liquid cooling system, and wherein the normal leakage is different from the threshold leakage in that at least a second quantity of the coolant exceeds the first quantity, the second quantity of the coolant incorrectly leaving the data center liquid cooling system such that the at least one computing component cannot operate within the normal temperature threshold and cannot receive the coolant to maintain the normal temperature threshold.

28. The remedial method according to claim 26, further comprising: using at least one processor associated with the learning subsystem to control the flow rate controller and the power controller using the input, wherein a first input in the input causes the shutdown of the at least one computing component, and a second input in the input causes the shutdown of the coolant.

29. The remedial method according to claim 26, further comprising: using at least one processor to enable a load transfer subsystem to transfer the load associated with the at least one computing component before the at least one computing component shuts down, the load being transferred to at least one second computing component that receives the coolant or a second coolant that is not affected by the threshold leakage.

30. The remedial method according to claim 26, further comprising: using at least one processor associated with the learning subsystem to determine that the threshold leakage of the coolant has occurred by determining that the pressure, flow rate, or temperature of the coolant flowing into and out of the at least one computing component is outside the normal threshold of the determined range and within a warning range defined by at least an alarm threshold, at which alarm threshold the at least one computing component no longer operates within the normal temperature threshold and no longer receives the coolant to maintain the normal temperature threshold.

Citation Information

Patent Citations

  • Leak detection and response system for liquid cooling of electronic racks of a data center

    CN110557925A

  • Leak detection with artificial intelligence

    WO2020033316A1