Smart reusable cooling system for mobile data centers
By adopting a combination of reusable refrigerant cooling subsystem and evaporative cooling subsystem in the data center cooling system, the problems of low cooling efficiency and high cost in the prior art are solved, and efficient and economical cooling effects are achieved to adapt to different cooling requirements.
Patent Information
- Application Number
- CN202180007634.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-07-08
- Filing Date
- 2021-06-30
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2041-06-30
AI Technical Summary
The existing data center cooling system cannot effectively meet the cooling needs of high thermal loads and high-density servers. The gas cooling system is inefficient, the liquid cooling system is costly and energy wasteful.
A multi-mode cooling system with a reusable refrigerant cooling subsystem is adopted, which includes an evaporative cooling subsystem and a reusable refrigerant cooling subsystem, which controls humidity through refrigerant circulation and provides independent cooling to meet different cooling requirements.
It achieves efficient and economical cooling effects, reduces operating costs, improves the cooling capacity and mobility of the data center, and adapts to different cooling requirements.
Smart Images

Figure CN114902821B_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This is a PCT application of U.S. Patent Application No. 16 / 923,971, filed on July 8, 2020. The disclosure of that application is incorporated herein by reference in its entirety for all purposes. Technical Field
[0003] At least one embodiment is directed to a cooling system for a data center. In at least one embodiment, an evaporative cooling subsystem provides blown air for cooling the data center, and a reusable refrigerant cooling subsystem controls humidity of the blown air in a first configuration and independently cools the data center in a second configuration. Background Art
[0004] Data center cooling systems typically use fans to circulate gas through server components. Some supercomputers or other high-capacity computers may use water or other cooling systems other than gas cooling systems to draw heat from server components or racks in the data center to an area outside the data center. The cooling system may include a chiller within the data center area. The area outside the data center may be a cooling tower or other external heat exchanger that receives heated coolant from the data center and disperses the heat to the environment (or external cooling medium) by forced gas or other means before the cooled coolant is recirculated back to the data center. In one example, the chiller and cooling tower together form a cooling facility that pulses in response to temperatures measured by external equipment applied to the data center. A gas cooling system by itself may not absorb enough heat to support effective or efficient cooling of a data center, and a liquid cooling system may not be able to economically meet the needs of a data center. BRIEF DESCRIPTION OF THE DRAWINGS
[0005] Various embodiments according to the present disclosure will be described with reference to the accompanying drawings, in which:
[0006] Figure 1 is a block diagram of an example data center having an improved cooling system as described in at least one embodiment;
[0007] Figure 2A is a block diagram illustrating mobile data center features of a cooling system incorporating a renewable refrigerant cooling subsystem according to at least one embodiment;
[0008] Figure 2B is a block diagram illustrating an evaporative cooling subsystem that benefits from a reusable refrigerant cooling subsystem according to at least one embodiment;
[0009] Figure 2Cis a block diagram illustrating data center level features of a cooling system incorporating an evaporative cooling subsystem and a reusable refrigerant cooling subsystem according to at least one embodiment;
[0010] Figure 3A is a block diagram illustrating rack-level and server-level features of a cooling system incorporating an evaporative cooling subsystem supported by a reusable refrigerant cooling subsystem, according to at least one embodiment;
[0011] Figure 3B is a block diagram illustrating rack-level and server-level features of a cooling system of a liquid-based cooling subsystem integrated with a multi-mode cooling subsystem according to at least one embodiment;
[0012] Figure 4A is a block diagram illustrating rack-level and server-level features of a cooling system integrating a reusable refrigerant cooling subsystem with a multi-mode cooling subsystem according to at least one embodiment;
[0013] Figure 4B is a block diagram illustrating rack-level and server-level features of a cooling system of a dielectric-based cooling subsystem integrated with a multi-mode cooling subsystem according to at least one embodiment;
[0014] Figure 5 According to at least one embodiment, a method for using or manufacturing Figure 2A-Figure 4B and Figures 6A-17D The process flow of the steps of the method of cooling the system;
[0015] Fig. 6A An example data center is shown where data from Figure 2A-Figure 5 at least one embodiment of;
[0016] Figure 6B , Figure 6C Inference and / or training logic for enabling and / or supporting a multi-mode cooling subsystem with a reusable refrigerant cooling subsystem according to various embodiments is shown, such as in Fig. 6A and reasoning and / or training logic used in at least one embodiment of the present disclosure;
[0017] Fig. 7A is a block diagram illustrating an exemplary computer system, which may be a system having interconnected devices and components, a system on a chip (SOC), or some combination thereof formed together with a processor, the processor may include an execution unit for executing instructions to support and / or implement a multi-mode cooling subsystem with a reusable refrigerant cooling subsystem as described herein, according to at least one embodiment;
[0018] Figure 7Bis a block diagram illustrating an electronic device for utilizing a processor to support and / or implement a multi-mode cooling subsystem having a reusable refrigerant cooling subsystem according to at least one embodiment;
[0019] Figure 7C A block diagram of an electronic device for utilizing a processor to support and / or implement a multi-mode cooling subsystem having a reusable refrigerant cooling subsystem according to at least one embodiment is shown;
[0020] Figure 8 Another exemplary computer system for implementing various processes and methods throughout a multi-mode cooling subsystem with a reusable refrigerant cooling subsystem for use with a reusable refrigerant cooling subsystem in accordance with at least one embodiment is shown;
[0021] Fig. 9A An exemplary architecture is shown in accordance with at least one embodiment of the present disclosure, wherein a GPU is communicatively coupled to a multi-core processor via a high-speed link to implement and / or support a multi-mode cooling subsystem having a reusable refrigerant cooling subsystem;
[0022] Fig. 9B shows additional details of the interconnection between a multi-core processor and a graphics acceleration module according to an exemplary embodiment;
[0023] Fig. 9C Another exemplary embodiment according to at least one embodiment of the present disclosure is shown, wherein an accelerator integrated circuit is integrated within a processor for implementing and / or supporting a multi-mode cooling subsystem with a reusable refrigerant cooling subsystem;
[0024] Fig.9D An exemplary accelerator integrated chip 990 for implementing and / or supporting a multi-mode cooling subsystem with a reusable refrigerant cooling subsystem according to at least one embodiment of the present disclosure is shown;
[0025] Fig.9E Additional details of an exemplary embodiment for implementing and / or supporting a shared model for a multi-mode cooling subsystem with a reusable refrigerant cooling subsystem according to at least one embodiment of the present disclosure are shown;
[0026] Fig.9F Additional details of an exemplary embodiment of a unified memory addressable by a common virtual memory address space for accessing physical processor memory and GPU memory to implement and / or support a multi-mode cooling subsystem with a reusable refrigerant cooling subsystem in accordance with at least one embodiment of the present disclosure are shown;
[0027] Fig. 10A An exemplary integrated circuit and associated graphics processor for a multi-mode cooling subsystem with a reusable refrigerant cooling subsystem according to embodiments described herein are shown;
[0028] Figures 10B-10C An exemplary integrated circuit and associated graphics processor for supporting and / or implementing a multi-mode cooling subsystem with a reusable refrigerant cooling subsystem according to at least one embodiment is shown;
[0029] Figures 10D-10E Additional exemplary graphics processor logic for supporting and / or implementing a multi-mode cooling subsystem with a reusable refrigerant cooling subsystem according to at least one embodiment is shown;
[0030] Fig.11A is a block diagram illustrating a computing system for supporting and / or implementing a multi-mode cooling subsystem with a reusable refrigerant cooling subsystem according to at least one embodiment;
[0031] Fig. 11B A parallel processor for supporting and / or implementing a multi-mode cooling subsystem with a reusable refrigerant cooling subsystem according to at least one embodiment is shown;
[0032] Fig. 11C is a block diagram of a partitioning unit according to at least one embodiment;
[0033] Fig.11D A graphics multiprocessor for a multi-mode cooling subsystem with a reusable refrigerant cooling subsystem is shown in accordance with at least one embodiment;
[0034] Fig.11E A graphics multiprocessor is shown in accordance with at least one embodiment;
[0035] Fig. 12A A multi-GPU computing system is shown in accordance with at least one embodiment;
[0036] Fig. 12B is a block diagram of a graphics processor according to at least one embodiment;
[0037] Fig.13 is a block diagram illustrating a microarchitecture for a processor, which may include logic circuitry for executing instructions, according to at least one embodiment;
[0038] Fig.14 A deep learning application processor according to at least one embodiment is shown;
[0039] Fig.15 shows a block diagram of a neuromorphic processor according to at least one embodiment;
[0040] Fig.16A is a block diagram of a processing system according to at least one embodiment;
[0041] Fig. 16B is a block diagram of a processor having one or more processor cores, an integrated memory controller, and an integrated graphics processor according to at least one embodiment;
[0042] Fig. 16C is a block diagram of the hardware logic of a graphics processor core according to at least one embodiment;
[0043] Figures 16D-16E Thread execution logic is shown that includes an array of processing elements of a graphics processor core in accordance with at least one embodiment.
[0044] Fig.17A illustrates a parallel processing unit according to at least one embodiment;
[0045] Fig. 17B illustrates a general processing cluster in accordance with at least one embodiment;
[0046] Fig. 17C A memory partitioning unit of a parallel processing unit according to at least one embodiment is shown; and
[0047] Fig.17D A streaming multiprocessor in accordance with at least one embodiment is shown. DETAILED DESCRIPTION
[0048] Considering the high heat load of today's computing components, gas cooling of high-density servers is inefficient and ineffective. However, these requirements may change or vary between the minimum and maximum values of different cooling requirements. Different cooling requirements also reflect the different heat dissipation functions of data centers. In at least one embodiment, the heat generated from components, servers and racks is cumulatively referred to as heat characteristics or cooling requirements, because the cooling requirements must fully meet the heat characteristics. In at least one embodiment, the heat characteristics or cooling requirements of the cooling system are the heat or cooling requirements generated by components, servers or racks associated with the cooling system and can be part of the components, servers and racks in the data center. For economic purposes, providing gas cooling or liquid cooling alone may not meet the cooling requirements of today's data centers. In addition, providing liquid cooling when the heat generated by the server reaches the minimum value of different cooling requirements will burden the components of the liquid cooling system and also waste energy requirements to enable the liquid cooling system. In addition, in a mobile data center, the ability to meet different cooling requirements is a challenge because there is a lack of space to support a cooling system that can provide a variety of different cooling media.
[0049] Therefore, the present disclosure is directed to the prospect of a cooling system having a reusable refrigerant cooling subsystem for one or more racks of a data center. One or more racks are mounted within a container or trailer, which also includes server racks of a mobile data center. The reusable refrigerant cooling subsystem has a reusable refrigerant cooling subsystem that controls humidity in an evaporative cooling subsystem in a first configuration and provides independent cooling for the data center in a second configuration. Therefore, the reusable refrigerant cooling subsystem has two or more different cooling systems, and the different cooling capacities provided by at least the evaporative cooling subsystem working with the reusable refrigerant cooling subsystem in the first configuration and the reusable refrigerant cooling subsystem configured to provide independent operation of at least two different cooling subsystems in the second configuration can be adjusted according to different cooling requirements of the data center.
[0050] In at least one embodiment, the additional cooling subsystem may provide at least two different cooling subsystems. In at least one embodiment, the cooling system of the present disclosure is a cooling system for cooling computing components such as a graphics processing unit (GPU), a central processing unit (CPU), or a switch component. These computing components are used in servers assembled in server trays on racks in a data center. As technological advances miniaturize computing components, server trays and racks accommodate more and more computing components, so more heat generated by each component needs to be dissipated compared to previous systems. However, there is also a situation where the calculation and participation of computing components that perform functions will cause more heat to be dissipated than when the computing components are idle. One problem solved by the present disclosure is that the basic cooling requirements of air-based cooling systems cannot be met due to the ineffectiveness of gas-based cooling systems and the clumsiness of evaporative cooling.
[0051] In at least one embodiment, a data center liquid cooling system is supported by a refrigeration device or system that may be expensive because it is designed with over-provisioning requirements. The over-provisioning requirements may be due to a lack of proper control to remove heat from a server (e.g., one or more components (e.g., GPUs, CPUs, etc.), from a collection of components in a server chassis, or from a collection of components in a rack. The efficient use of a high heat capacity liquid cooling system may be obscured by many intermediate features in the cooling system, including liquid coolers, pumps, coolant distribution units (CDUs), and heat exchangers. The reusable refrigerant cooling subsystem of the data center cooling system of the present disclosure is capable of providing cooling capabilities that perform at least dual functions, such as using high temperature facility water to control humidity and meet different cooling requirements. In at least one embodiment, the high temperature facility water can be used as a high heat component and system. The refrigerant cooling subsystem can be reused for controlling humidity of a gas-based subsystem that relies on evaporative cooling as well as for stand-alone cooling, while functioning as a reusable refrigerant cooling subsystem. The cooling system includes a dielectric-based cooling subsystem for immersion cooling; and an evaporative cooling subsystem, depending on the different cooling requirements. In addition to performing the dual purpose of controlling humidity of the evaporative cooling subsystem and providing stand-alone cooling, the reusable refrigerant cooling subsystem complements the evaporative cooling subsystem that relies on evaporative cooling, but also removes conventional to extreme heat through the additional use of a dielectric refrigerant that supports the dielectric-based cooling subsystem.
[0052] In at least one embodiment, the evaporative cooling subsystem provides blown air for cooling the data center. The blown air is provided according to a first thermal characteristic in the data center. In at least one embodiment, the first thermal characteristic is the minimum thermal characteristic satisfied by the evaporative cooling subsystem formed by the evaporative cooling subsystem. In addition to satisfying the minimum thermal characteristic of the data center, the evaporative cooling subsystem also reflects the most economical cooling subsystem in the data center cooling system. In at least one embodiment, the evaporative cooling subsystem operates or is activated simultaneously with the reusable refrigerant (or other reusable refrigerant) cooling subsystem to control the humidity of the blown air in a first configuration of the reusable refrigerant cooling subsystem. In at least one embodiment, the reusable refrigerant cooling subsystem can be independently activated to cool the data center in a second configuration and according to a second thermal characteristic of the data center. In at least one embodiment, in the second configuration, the evaporative cooling subsystem may not be in use.
[0053] In at least one embodiment, the reusable refrigerant cooling subsystem includes an evaporator coil associated with a condensing unit provided outside the data center. The phase of the refrigerant changes from liquid to vapor in the evaporator to reach a lower temperature in the circulation line (or after absorbing heat from the environment), thereby achieving reusable refrigerant cooling. Once the vapor is expanded while dissipating the lower temperature (or absorbing heat) and is consumed, it is compressed and condensed into a liquid phase in the condensing device and then sent back to the evaporator. In at least one embodiment, the condensing unit is as close to the secondary cooling loop as possible. In at least one embodiment, the condensing unit is also as close to the evaporative cooling system as possible to ensure timely and effective elimination of moisture in the evaporative cooling subsystem. In at least one embodiment, each secondary cooling loop based on liquid and reusable refrigerant cooling subsystem is adjacent in its respective pipelines and manifolds. In at least one embodiment, the condensing unit is located in the top area of the data center or on a container (of a mobile data center) and has direct access to the room or row manifold to provide refrigerant in the refrigerant cooling loop of the reusable refrigerant cooling subsystem, as well as for controlling moisture in the evaporative cooling subsystem. In at least one embodiment, the evaporator coil of the recyclable refrigerant cooling subsystem provides direct expansion refrigerant to the refrigeration cooling loop and coexists with a secondary cooling loop (also known as a hold temperature liquid cooling system) carrying coolant.
[0054] In at least one embodiment, the evaporative cooling subsystem includes its own condenser coil and its own evaporator coil to condition the blown air for the data center using refrigerant from the reusable refrigerant cooling subsystem. The refrigerant from the reusable refrigerant cooling subsystem enters the condenser coil of the evaporative cooling subsystem at high temperature and pressure and is cooled by the purge air flowing through the evaporative cooling subsystem. The high pressure refrigerant is now cold and in a liquid phase as it passes through the evaporator coil just before the blown air enters the data center. The moisture in the blown air condenses at the evaporator coil. The blown air is conditioned by removing moisture to at least a determined threshold and the blown air flows into the data center to meet a first thermal characteristic or cooling requirement of the data center. In at least one embodiment, the used refrigerant may be compressed by a compressor within the evaporative cooling subsystem before returning to the reusable refrigerant cooling subsystem. In at least one embodiment, the refrigerant of the evaporative subsystem circulates in the evaporative subsystem and is added from the reusable refrigerant cooling subsystem as needed.
[0055] A server box or server tray may include the above-described reference computing components, but multiple server boxes at the rack level or more racks benefit from the reusable refrigerant cooling subsystem of the present disclosure. The reusable refrigerant cooling subsystem, as well as the evaporative cooling subsystem, is able to meet the different cooling requirements of cooling subsystems installed in one rack or more racks as one form factor of the cooling system. This eliminates the additional cost of running a data center chiller, which is more efficient than the nominal ambient temperature or colder that the evaporative cooling system is sufficient for. In at least one embodiment, a micropump or other flow control device forms one or more flow controllers that are sensitive to the thermal requirements of at least one or more sensors in the server or rack for the reusable refrigerant cooling subsystem.
[0056] Figure 1 1 is a block diagram of an example data center 100 with an improved cooling system described in at least one embodiment. The data center 100 may be one or more spaces 102 with racks 110 and auxiliary equipment for accommodating one or more servers on one or more server trays. The data center 100 is supported by a cooling tower 104 located outside the data center 100. The cooling tower 104 dissipates heat within the data center 100 by acting on a primary cooling loop 106. In addition, a cooling distribution unit (CDU) 112 is used between the primary cooling loop 106 and the second or secondary cooling loop 108 so that heat can be extracted from the second or secondary cooling loop 108 to the primary cooling loop 106. In one aspect, the secondary cooling loop 108 is able to enter all the lines of the server tray as needed. The loops 106, 108 are illustrated as lines, but a person of ordinary skill will recognize that one or more pipeline features can be used. In one example, a flexible polyvinyl chloride (PVC) pipe can be used with associated pipelines to move fluid along each loop 106, 108. In at least one embodiment, one or more coolant pumps may be used to maintain a pressure differential within the loops 106 , 108 to move coolant based on temperature sensors at various locations, including within the room, within one or more racks 110 , and / or within server boxes or server trays within racks 110 .
[0057] In at least one embodiment, the coolant in the primary cooling loop 106 and the secondary cooling loop 108 can be at least water and an additive, such as ethylene glycol or propylene glycol. In operation, each primary and secondary cooling loop has its own coolant. In one aspect, the coolant in the secondary cooling loop can be dedicated to the components in the server tray or rack 110. The CDU 112 is capable of controlling the coolant in the loops 106, 108 independently or simultaneously. For example, the CDU can be adapted to control the flow rate so as to appropriately distribute one or more coolants to extract the heat generated in the rack 110. In addition, more manifolds 114 are provided from the secondary cooling loop 108 to enter each server tray and provide coolant to the electrical and / or computing components. In the present disclosure, electrical and / or computing components are interchangeably used to refer to heat generating components that benefit from the data center cooling system. The manifold 118 that constitutes part of the secondary cooling loop 108 can be referred to as a room manifold. In addition, the manifold 116 extending from the manifold 118 can also be part of the secondary cooling loop 108, but can be referred to as a row manifold. Manifold 114 enters the rack as part of the secondary cooling loop 108, but may be referred to as a rack cooling manifold. Additionally, row manifolds 116 extend along the rows in the data center 100 to all racks. The piping of the secondary cooling loop 108 (including manifolds 118, 116, and 114) may be improved by at least one embodiment of the present disclosure. An optional chiller 120 may be provided in the primary cooling loop within the data center 100 to support cooling prior to the cooling tower. Where additional loops are present in the primary control loop, a person of ordinary skill will recognize upon reading this disclosure that the additional loops provide cooling external to the rack and external to the secondary cooling loop; and may be used with the primary cooling loop of the present disclosure.
[0058] In at least one embodiment, in operation, heat generated within the server trays of the racks 110 can be transferred to the coolant leaving the racks 110 through the flexible tubes of the row manifolds 114 of the secondary cooling loop 108. Accordingly, the secondary coolant (in the secondary cooling loop 108) from the CDU 112 for cooling the racks 110 moves toward the racks 110. The secondary coolant from the CDU 112 flows from one side of the room manifold having the manifold 118 through the row manifold 116 to one side of the racks 110, and flows through one side of the server trays through the manifold 114. The used secondary coolant (or the exhausted secondary coolant carrying the heat of the computing components) flows out from the other side of the server trays (e.g., after passing through the server trays or components on the server trays, entering the left side of the rack, and then flowing out of the server trays from the right side of the rack). The spent second coolant from the server trays or racks 110 flows out of a different side (e.g., the outflow side) of the manifold 114 and moves to the parallel but also outflow side of the row manifold 116. From the row manifold 116, the second coolant moves in a parallel portion of the room manifold 118 toward the CDU 112 in an opposite direction to the incoming second coolant (which may also be refreshed second coolant).
[0059] In at least one embodiment, the spent second coolant exchanges heat with the primary coolant in the primary cooling loop 106 through the CDU 112. The spent second coolant is refreshed (e.g., relatively cooled compared to the temperature of the spent second coolant stage) and is ready to be circulated back to the computing components through the secondary cooling loop 108. Various flow and temperature control functions in the CDU 112 can control the heat exchanged from the spent second coolant or the flow of the second coolant into and out of the CDU 112. The CDU 112 can also control the flow of the primary coolant in the primary cooling loop 106. As a result, some components within the servers and racks do not receive the required coolant levels because the second or secondary loops typically provide coolant with their default temperature properties based in part on temperature sensors that can be within the servers and racks. In addition, due to space limitations in at least mobile data centers, the requirements herein meet multiple purposes of cooling and feature control through one or more cooling subsystems (e.g., through a reusable refrigerant cooling subsystem).
[0060] Figure 2A is a block diagram illustrating mobile data center features of a cooling system 200 including a reusable refrigerant cooling subsystem in a rack or multi-rack format, a multi-mode cooling subsystem 212A; 212B; 212C, according to at least one embodiment. Figure 2AA topology of containers or pods 204A-D of a mobile data center system or cooling system 200 may be provided in accordance with at least one embodiment, the containers or pods 204A-D being configured to implement an arrangement of circulation of a cooling medium associated with the mobile cooling system. In at least one embodiment, the mobile data center cooling system 200 is configured with a plurality of containers or pods 204A-D having racks 216 (referenced as a group of racks). In at least one embodiment, the racks 216 are cooled by one or more cooling mediums from a corresponding one of the container manifolds 206 (referenced as a container manifold having one or more manifolds for one or more different mediums). In at least one embodiment, the blown air may be delivered directly through ports of the containers 204A, B, C, or D or racks 216. In addition, the cooling medium flows from the container manifold 206 to a corresponding one of the row manifolds 208 (referenced as a group of row manifolds having inlet and outlet manifolds for one or more cooling mediums).
[0061] In at least one embodiment, each container manifold 216 extends across the perimeter of its respective container 204A-D. Thus, row manifolds 208 may not be required in at least one embodiment. In at least one embodiment, container manifolds 206 carry respective CDUs or couple to respective CDUs of multi-mode cooling subsystems 212A, B from respective one or more containers 204A-D. In at least one embodiment, the CDUs may reside in and be distributed within the cooling medium associated with one or more containers 204A-D, or each of the one or more containers 204A-D may have its own CDU (shown by reference numeral 212C, forming a multi-mode cooling subsystem with CDUs). In at least one embodiment, when the mobile data center features of the cooling system 200 include a multi-mode (or multiple modes) cooling subsystem using coolants, refrigerants, and blown air, the block with reference numerals 212A;B may be referred to as a multi-mode (or multiple modes) cooling subsystem and may include a CDU therein to handle features associated with a liquid-based cooling subsystem (e.g., a coolant-based cooling subsystem). Therefore, reference to a CDU within the multi-mode cooling subsystem 212A;B;C does not limit the functionality of the block shown in reference numerals 212A;B to CDU functionality, but refers to at least one function within the block with reference numerals 212A;B. In at least one embodiment, the block with reference numerals 212A;B;C may be referred to as a CDU, but may be a multi-mode cooling subsystem unless otherwise indicated.
[0062] In at least one embodiment, the cooling system 200 is located on one or more trailer beds. In at least one embodiment, the trailer bed or each container (e.g., container 204A, 204D) has space for infrastructure areas 202A, B. In at least one embodiment, the cooling system 200 is supported by a cooling tower 218 in the infrastructure area 202A on its own trailer bed, or shared between the infrastructure areas 202A, 202B of multiple containers or trailer beds. Therefore, in at least one embodiment, each container 204A-D and the infrastructure areas 202A, B are located on a separate trailer or a single trailer. In at least one embodiment, a single trailer bed can be located adjacent to each other at a deployment location. In at least one embodiment, the transportation of the cooling system 200 is sized according to a standard margin (e.g., a truck transportation margin of the Department of Transportation).
[0063] In at least one embodiment, each container 204A-D can be adapted to receive a primary cooling loop from a corresponding one or more CDUs of the multi-mode cooling subsystem 212A, B through respective external piping 214A; B. The corresponding one or more CDUs of the multi-mode cooling subsystem 212A, B are coupled to a secondary cooling loop to reach the rack 216. In at least one embodiment, the CDUs work together as part of a single multi-mode cooling subsystem 212A, B, where control of the different cooling media is provided by a distributed or centralized control system. The secondary cooling loop is associated with one or more external cooling towers 218B through the CDU, which is located between the primary and secondary cooling loops to allow a first cooling medium of the secondary cooling loop to heat exchange with a second cooling medium of the primary cooling loop. In at least one embodiment, the containers 204A-D can be adapted to receive a primary cooling loop indirectly via respective external piping 214A; B. The respective external piping 214A; B is part of the secondary cooling loop.
[0064] Although Figure 2AThe cooling tower 218B in the CDU is illustrated as a single tower, but each CDU may be associated with its own cooling tower. The primary cooling loop may have different configurations, for example, from the CDU to each group of racks in the container 204A, 204B; or from the CDU to the racks in the container 204A, where a separate CDU is coupled to the racks in the adjacent container 204B; or from the CDU to the racks in all containers 204A-D, with additional CDUs configured for redundancy in the event of a failure of the primary CDU. The primary cooling loop may include additional sub-loops, such as a loop with a row manifold 208, but the primary cooling loop may be fully extended to the row manifold by providing a single coolant through the container manifold 206 and the row manifold 208. In at least one embodiment, the CDU is positioned adjacent to and / or within one or more containers 204A-D (e.g., within an enclosed container area also referred to as the infrastructure area 202A; 202B). The cooling tower 218B may be located outside the container or enclosed area. In at least one embodiment, as also shown in the example of FIG. 2 , a cooling tower 218B can be located on top of one or more containers 204A-D. In at least one embodiment, one or more CDUs can be fully specified in a container to support racks of multiple containers. The CDU can be coupled to at least one cooling tower of a certain size located on top of the container. The combination of the cooling tower on top of the container and the CDU within the container forms a mobile cooling system that can be co-located with an adjacent trailer having only containers and racks.
[0065] In at least one embodiment, a condenser unit 218A having a compressor and a condenser is located outside of one or more containers 204A-D. In at least one embodiment, the condenser unit 218A can implement a reusable refrigerant cooling subsystem within a rack or multi-rack format multi-cooling subsystem of the present disclosure. The condenser unit 218A directly circulates the refrigerant required by the reusable refrigerant cooling subsystem within the rack or multi-rack format multi-cooling subsystem of the present disclosure through a line 210B that is independent of the coolant line 210A of the secondary cooling loop.
[0066] In at least one embodiment, the condenser unit 218A also circulates refrigerant directly to an evaporative cooling subsystem or a multi-rack format multi-cooling subsystem within a rack of the present disclosure. In at least one embodiment, each of the racks or multi-rack format multi-cooling subsystems is independent and meets the thermal characteristics or cooling requirements of the server racks associated with its container. In at least one embodiment, the condenser unit 218A is located at the top of the container. In at least one embodiment, the location of the condenser unit 218A is to ensure that it is as close as possible to a reusable refrigerant cooling subsystem that is in the form factor of one or more racks and is located within the container 204A; B; C or D. In at least one embodiment, the reusable refrigerant cooling subsystem is located in the infrastructure area 202A; 202B, adjacent to the CDU.
[0067] In at least one embodiment, the condenser unit 218A can be located at the top of the infrastructure area 202A and can circulate refrigerant directly to the evaporative cooling subsystem and directly to the reusable refrigerant cooling subsystem within the rack or the multi-rack format multi-cooling subsystem 212A. Therefore, in at least one embodiment, the block forming the reference numeral 212A is referred to as a multi-mode (or multiple mode) cooling subsystem. Figure 2C As shown, in at least one embodiment, the multi-mode (or multiple modes) cooling subsystem 212A is a rack or multi-rack format structure within a container or infrastructure area. In at least one embodiment, when the multi-mode (or multiple modes) cooling subsystem is used to cool a mobile or fixed data center, it can include a CDU therein to handle features associated with a liquid-based cooling subsystem (e.g., a coolant-based cooling subsystem).
[0068] In at least one embodiment, each of the rack or multi-rack format multi-cooling subsystems 212C is independent and meets the thermal characteristics or cooling requirements of the server rack associated with its container. In at least one embodiment, the condenser unit 218A is located at the top of the container. In at least one embodiment, the location of the condenser unit 218A is to ensure that it is as close as possible to the reusable refrigerant cooling subsystem, which can be the form factor of one or more racks and is located in the container 204A; B; C or D. In at least one embodiment, the reusable refrigerant cooling subsystem is located in the infrastructure area 202A; 202B, adjacent to the CDU in the multi-mode (or multiple modes) cooling subsystem 212A; 212B. In at least one embodiment, the refrigerant can reach the evaporative cooling subsystem and the reusable refrigerant cooling subsystem directly through the pipeline 210B, or can reach these subsystems through the pipeline in the container manifold 206.
[0069] In at least one embodiment, Figure 2AIt is also illustrated that multiple trailer beds (each carrying a container and infrastructure area or just a container) are adjacent to each other to enable coupling and sharing of cooling resources, such as the CDU, condenser unit 218A, and cooling tower 218B in cooling subsystems 212A; 212B. In at least one embodiment, the Figure 2A A single trailer bed with the ability to handle multiple containers or pods and cooling towers is shown in the configuration of FIG. 1 . In either case, the containers or pods 204A-D share cooling resources from at least two CDUs 212A;B. In at least one embodiment, each primary cooling loop from the cooling tower 218B terminates at a CDU within its own container 204A or other enclosure or within an adjacent container 204A;D. The secondary cooling loop extends through the container manifold 206 to provide coolant to the racks directly or indirectly through the row manifold 208. The coolant associated with the containers in the secondary cooling loop circulates in both CDUs of the respective container 204A;D, while the container 204B;C may have no CDU. Alternatively, in at least one embodiment, the secondary cooling loop may include additional CDUs in the container 204B;C, with the coolant associated with the CDU of the container 204A;D cooling the coolant associated with the CDU in the container 204B;C. The container manifolds 206 may be coupled to each other or to adjacent container manifolds via fluid couplings.
[0070] In at least one embodiment, the primary cooling loop terminates in a respective CDU of a respective container 204A-D. In at least one embodiment, the primary cooling loop extends from cooling tower 218B to a CDU in or near container 204A, wherein a split coupler allows coolant of the primary cooling loop to extend to the CDU in container 204B. In at least one embodiment, further piping of container manifold 206 enables the primary cooling loop of one cooling tower 218B to extend through the CDU of each container 204A-D, such as the CDU in cooling subsystem 212A. In this case, row manifold 208 forms a secondary cooling loop. Separately or concurrently, at least in this or one embodiment, the CDU in cooling subsystem 212B can be used to meet the cooling requirements of containers 204A-D through container manifold 206 forming a secondary cooling loop with row manifold 208, while cooling tower 218B is in the primary cooling loop with the CDU. This enables a redundancy plan in at least one embodiment with multiple cooling towers in the event that one of the cooling towers fails.
[0071] In at least one embodiment, Figure 2AThe configuration is implemented by at least one processor, which may include at least one logic unit having a trained neural network or other machine learning model that is trained to determine the cooling subsystem to meet the temperature requirements of a mobile data center deployed in response to customer needs, such as thermal characteristics or cooling requirements. The neural network or other machine learning model can be initially trained and can be further trained on an ongoing basis to improve its accuracy. At least one processor can be provided for training one or more neural networks having hidden layers of neurons for evaluating the temperature within a server or one or more racks in a data center using flow rates of different cooling media from at least an evaporative cooling subsystem and a reusable refrigerant cooling subsystem.
[0072] In at least one embodiment, the at least one processor can provide an output associated with at least two temperatures to facilitate movement of two cooling media (e.g., concurrently or independently, blown air and refrigerant) from the cooling system by controlling flow controllers associated with the two cooling media to meet a first thermal characteristic and a second thermal characteristic. Fig.14 and Fig.15 The at least one processor and the neural network training scheme in the embodiment of the present invention are further described. The training can also be used to evaluate the flow rate or flow rate of the two cooling media based in part on its correlation with the temperature requirement.
[0073] In at least one embodiment, the condenser unit 218A processes refrigerant returning from one or more of the refrigerant-based cooling subsystems in the multi-mode cooling subsystems 212A; B and C. In at least one embodiment, a separate condenser unit 218A may be provided to process refrigerant from more than one refrigerant-based cooling subsystem. In at least one embodiment, one or more refrigerant-based cooling subsystems are controlled by a centralized or distributed control system. In at least one embodiment, a condenser unit may be placed atop an infrastructure area to meet the refrigerant requirements of an evaporative cooling subsystem and one or more refrigerant-based cooling subsystems, all of which work together or independently through control of associated flow controllers of the evaporative cooling subsystem and one or more refrigerant-based cooling subsystems.
[0074] In at least one embodiment, at least one processor can be trained using previous or tested values of a thermal signature or cooling requirement, and a flow rate or flow rate of one or more cooling media. In at least one embodiment, at least one processor trained to associate a thermal signature or cooling requirement with a flow rate or flow rate of one or more cooling media can determine an output having instructions for one or more flow controllers. In at least one embodiment, one or more first flow controllers of the cooling system are instructed to maintain a first cooling medium associated with the evaporative cooling subsystem at a first flow rate and a second cooling medium associated with the reusable refrigerant cooling subsystem at a second flow rate. In at least one embodiment, the first cooling medium is flowing air and the second cooling medium is refrigerant. The first flow rate of the evaporative cooling subsystem is based in part on the first thermal signature. The second flow rate associated with the reusable refrigerant cooling subsystem is the flow rate required for the refrigerant to flow from the reusable refrigerant cooling subsystem to the components of the evaporative cooling subsystem to control the humidity of the blown air. In at least one embodiment, this setting of at least a second flow rate of the refrigerant to the reusable refrigerant cooling subsystem is associated with a first configuration of the reusable refrigerant cooling subsystem. Furthermore, in the first configuration, in order to control the humidity of the blown air, one or more first flow controllers are associated with at least a flow loop of the refrigerant to the evaporative cooling subsystem.
[0075] In at least one embodiment, one or more second flow controllers are used to maintain a second cooling medium associated with the reusable refrigerant cooling subsystem at a third flow rate. The third flow rate is based in part on the second thermal characteristic. In addition, one or more first flow controllers are deactivated based in part on the second thermal characteristic. In at least one embodiment, the third flow rate is at least part of a second configuration of the reusable refrigerant cooling subsystem. In the second configuration of the reusable refrigerant cooling subsystem, the refrigerant of the reusable refrigerant cooling subsystem flows to racks, servers, or data center areas that require cooling. Therefore, one or more second flow controllers may be different from one or more first flow controllers to circulate refrigerant from the reusable refrigerant cooling subsystem to the data center to cool components, servers, and racks in the data center. In at least one embodiment, a multi-directional flow controller may be used instead of the first and second flow controllers to adjust the refrigerant flow in the first configuration at a second flow rate for humidity control; and to adjust the refrigerant flow in the second configuration at a third flow rate for cooling in the data center.
[0076] When a thermal signature or cooling requirement (also referred to as a temperature requirement) associated with at least the racks 216 in the containers 204A-D is satisfied using the trained neural network, an output of the desired flow rate or flow rate of the refrigerant and other cooling media (such as blown air) required to satisfy the temperature requirement may be provided. In at least one embodiment, the output is an extrapolation of previous values or test values provided to control the flow rates of the various cooling media to satisfy the thermal signature or cooling requirement, and wherein the neural network determines that the error in achieving cooling is minimal in the configuration determined by the output. In at least one embodiment, the error is resolved by repeated back propagation, making the neural network more accurate in determining the response to the thermal signature or cooling requirement.
[0077] In at least one embodiment, the data center cooling system 200 includes at least one processor in its container 204A-D for continuously improving the neural network based in part on current information available from the deployed mobile data center cooling system 200. At least one processor executes the machine learning model to perform a function. One function is to process the temperature within a server or one or more racks 216 in the data center cooling system 200 using multiple neuron levels of the machine learning model. The multiple neuron levels of the machine learning model have temperatures and have previous associated flow rates of two cooling media for the temperatures. The temperatures and associated flow rates can come from previous actual cooling applications using the configuration of the cooling system 200, or from tests performed using the configuration of the cooling system 200. In at least one embodiment, another function of the machine learning model is to provide an output associated with the flow rates of the two cooling media to the flow controller. After evaluating the previous associated flow rates and previous associated temperatures of a single cooling medium of the two cooling media, the output is provided.
[0078] In at least one embodiment, the evaluation attempts to correlate (achieving zero error during training to reflect a fully classified machine learning model) one or more flow rates of a single cooling medium to achieve a temperature in the data center starting from an initial temperature that causes the single cooling medium to flow, to a second temperature that causes the flow to stagnate. In at least one embodiment, the cooling medium is blown air and a refrigerant, wherein the refrigerant is required to account for additional humidity that may be present in the blown air. The data center cooling system 200 can be adapted to operate at a maximum humidity of 60%, and the flow rate of the refrigerant is maintained at least along with the flow rate of the blown air to ensure that enough humidity is removed from the blown air to maintain the humidity in the data center at the adapted maximum limit.
[0079] In at least one embodiment, at least one processor of the cooling system 200 includes an instruction output for transmitting an output associated with a flow controller to achieve a first flow rate of a first separate cooling medium for satisfying a first thermal characteristic in the data center while maintaining a second flow rate of a second separate cooling medium to control the humidity of the first separate cooling medium. The instruction output also transmits an output with instructions to deactivate the first separate cooling medium and enable a third flow rate of the second separate cooling medium for satisfying a second thermal characteristic in the data center. In at least one embodiment, the instruction output of the processor is a pin of a connector bus or a ball of a ball grid array, which enables communication of the output in the form of a signal carrying the instruction. At least one processor is part of a distributed or centralized control system for controlling one or more flow controllers.
[0080] Figure 2B is a block diagram of an evaporative cooling subsystem 220 that benefits from a reusable refrigerant cooling subsystem according to at least one embodiment. In at least one embodiment, the evaporative cooling subsystem 220 draws supply air 222A from a data center (or other suitable source) using a supply air fan 224A for drawing the supply air. In at least one embodiment, the supply air then passes through a filter 226 to an indirect heat exchanger 232. In at least one embodiment, the indirect heat exchanger has a coil or perforated plate 232A. Purge air 240A is drawn through the indirect heat exchanger 232 by the action of a purge fan 224B, which blows the purge air out of the supply facility of the evaporative cooling subsystem 220 as exhaust air 240B.
[0081] In at least one embodiment, the evaporative cooling subsystem 220 includes a sump 234 with a suitable evaporative medium (e.g., facility water). A sump pump 236A is used to pump the evaporative medium through line 236B (or through an indirect heat exchanger) of a spray system 236C, which sprays the evaporative medium onto the coils or perforated plates 232A as the purge air 240A passes through and as the supply air 222A passes through. In at least one embodiment, the supply air 222A is mixed with the purge air 240A. In at least one embodiment, the supply air 222A is separated from the purge air 240A and flows through the coils or perforated plates 232A while absorbing humidity, while the purge air 240A alone acts to cool the coils or perforated plates 232A, thereby in turn cooling the supply air 222A as it passes through.
[0082] In at least one embodiment, the supply air 222A has some humidity due to the spraying of the evaporative medium 230, but is at least partially conditioned to cool the data center according to at least one thermal characteristic or cooling requirement of the data center. The air 238B flowing out of the indirect heat exchanger 232 can be provided to the data center as the blown air 222B. In at least one embodiment, the evaporative cooling subsystem 220 provides blown air to the data center according to a first thermal characteristic in the data center. In addition, the evaporative cooling subsystem 220 can include components (condenser 228, evaporator coil 242, compressor 246) that work with the reusable refrigerant cooling subsystem to control the humidity of the blown air in the first configuration of the reusable refrigerant cooling subsystem. In at least one embodiment, the reusable refrigerant cooling subsystem manages refrigerant through the condenser 228, which cools the refrigerant and transfers the refrigerant to the evaporator coil 242. In at least one embodiment, the refrigerant leaving the evaporator coil 242 is compressed by the direct cooling compressor 246 and returned to the reusable refrigerant cooling subsystem. In at least one embodiment, the compressed refrigerant is in a vapor state and fed back to the condenser 228 to exchange heat with the purge air 240A flowing through the indirect heat exchanger 232. The reusable refrigerant cooling subsystem can cool the data center using the refrigerant alone in a second configuration based on a second thermal characteristic in the data center.
[0083] In at least one embodiment, the refrigerant is at a first flow rate as it flows through the evaporator coil, which ensures that the evaporator coil 242 performs a dual function. In at least the first configuration, the first flow rate of the refrigerant through the evaporator coil 242 is to ensure that the coil is cool enough to allow moisture in the air 238B to condense on the evaporator coil 242. Thus, moisture is removed from the air 238B in the first configuration. In at least the second configuration, the second flow rate of the refrigerant enables it to meet the cooling requirements within the data center. In at least one embodiment, the air 238B is the blown air 222B returned to the data center for cooling the data center according to the first thermal characteristic or cooling requirement. This may occur when the air 238B has a determined moisture threshold within a range of limits suitable for a data center, such as 55-60%. In at least one embodiment, the refrigerant conditions the moisture content of the air 238B through the evaporator coil 242 so that the air 238B has a first moisture content, but the blown air 222B returned to the data center to cool the data center according to the first thermal characteristic or cooling requirement has a second moisture content lower than the first moisture content. This is to not affect the determined moisture threshold within the limit range applicable to the data center.
[0084] Figure 2C2 is a block diagram illustrating data center level features of a cooling system 250 including an evaporative cooling subsystem 278 and a reusable refrigerant cooling subsystem 282 within a multi-mode cooling subsystem 256 according to at least one embodiment. The cooling system 250 is for a data center 252. The reusable refrigerant cooling subsystem 282 includes support for humidity control of the evaporative cooling system 278. The multi-mode cooling subsystem 256 may also include a liquid-based cooling subsystem 280 and / or a dielectric cooling subsystem 284. The reusable refrigerant cooling subsystem 282 may have overlapping components with the evaporative cooling subsystem, and thus, as described with reference to FIG. Figure 2A , Figure 2B As described, it is suitable for two different configurations. In addition, the multi-mode cooling subsystem 256 can be rated for different cooling capacities because each of the subsystems 278-284 can have a different rated cooling capacity. In at least one embodiment, the evaporative cooling system 278 is rated for heating between 10 kilowatts (KW or kW) and 30KW, the liquid-based cooling subsystem 280 is rated for heating between 30KW and 60KW, the reusable refrigerant cooling subsystem 282 is rated for heating between 60KW and 100KW (in the cooling configuration, different from the humidity control configuration), and the dielectric cooling subsystem 282 is rated for heating between 100KW and 450KW.
[0085] In at least one embodiment, the heat is generated from one or more computing components in a server in a rack of a data center. The rating is a characteristic of the cooling subsystem that the cooling subsystem is able to handle (in at least one embodiment, by cooling or by dissipating) when it is within the range that the cooling subsystem can handle. Different cooling media can handle heat in different ways according to the ratings recorded for different cooling media. In at least one embodiment, for the reusable refrigerant cooling subsystem 282, the cooling medium is refrigerant; for the evaporative cooling subsystem 274, the cooling medium is blown air. The heat generated by one or more computing components, servers, or racks is different from the rating of the cooling subsystem's individual ability to handle the heat, but refers to the thermal characteristics or cooling requirements that the different cooling capacities of the cooling subsystem must meet.
[0086] In at least one embodiment, the rated cooling capacity of the multi-mode cooling subsystem 256 is different due to the different cooling media from different cooling subsystems 278-284, and the multi-mode cooling subsystem 256 can be adjusted within the minimum and maximum ranges of the different cooling capacities provided by the different cooling subsystems 278-284 to adapt to the different cooling requirements of the data center 252. In at least one embodiment, the different cooling media include air or gas (e.g., blown air), one or more different types of coolants (including water and glycol-based coolants), and one or more different types of refrigerants (including non-dielectric and dielectric refrigerants). In at least one embodiment, the multi-mode cooling subsystem 256 has a form factor of one or more racks 254 of the data center. In at least one embodiment, the multi-mode cooling subsystem 256 adopts the form factor of one or more racks and is installed near one or more server racks 254 to handle the different cooling requirements of the components and servers of the one or more server racks 254.
[0087] In at least one embodiment, one or more server racks 254 have different cooling requirements based in part on the different heat generated by the components of the server boxes or trays 270-276 within the one or more server racks 254. In at least one embodiment, multiple server boxes 270 have a nominal heat generated that corresponds to approximately 10KW to 30KW of heat generation, and the evaporative cooling subsystem 278 is activated to provide evaporative cooling, as shown by the directional arrow from the subsystem 278. The evaporative cooling is based on the blown air 278B provided from the evaporative cooling subsystem 278. The evaporative cooling subsystem 278 can share components with the reusable refrigerant cooling subsystem 282, so the blown air 278B provided to the server boxes 270 is cooled and humidity controlled air. The blown air 278B can be about Figure 2B The heat exchanger within the evaporative cooling subsystem 278 is part of an evaporative-based secondary cooling loop 292 that exchanges heat with a coolant-based primary cooling loop 288 associated with the CDU 286. In at least one embodiment, the coolant is facility water that is transferred to a sump, such as Figure 2B In at least one embodiment, a cooling facility consisting of at least one chiller unit 258, a cooling tower 262, and a pump 260 supports a coolant-based primary cooling loop 288.
[0088] In at least one embodiment, the cooling system of the present disclosure includes a refrigerant-based heat transfer subsystem to implement one or more of the reusable refrigerant cooling subsystem 282 and the medium cooling subsystem 284; and includes a liquid-based heat transfer subsystem to implement one or more of the evaporative cooling system 278 and the liquid-based cooling subsystem 280. In at least one embodiment, the refrigerant-based heat transfer subsystem includes at least a pipeline 278A that supports the flow of refrigerant to different cooling subsystems, and may include one or more evaporator coils, compressors, condensers, and expansion valves (some of which are in Figure 2B 294) to support the circulation of the refrigerant in the different cooling subsystems. In at least one embodiment, a dielectric refrigerant is used to implement the refrigerant-based cooling subsystem 282 and the medium cooling subsystem 284. This reduces the components required for the cooling system of the present disclosure, but supports a wide range of cooling capabilities to meet the different cooling requirements of the system. In addition, in at least one embodiment, a coolant-based secondary cooling loop 294 supports the liquid-based cooling subsystem 280.
[0089] In at least one embodiment, due to at least the above configuration, even if the server box 270 is air cooled by the evaporative cooling subsystem 278, the server tray 222 can be liquid cooled by the liquid-based cooling subsystem 280 (in at least one embodiment, using a coolant); and the server tray 279 can be immersion cooled via the medium-based cooling subsystem 284 (in at least one embodiment, using a medium refrigerant); and the server tray 276 can be refrigerant cooled (in at least one embodiment, using the same dielectric refrigerant or a separate refrigerant). In at least one embodiment, the refrigerant and medium cooling subsystems 282, 284 are directly coupled to at least one condensing unit 264 to meet the cooling requirements of the refrigerant through the refrigerant cooling loop 290.
[0090] In at least one embodiment, a coolant distributor (e.g., coolant distribution unit 286) provides coolant to both the evaporative cooling subsystem 278 and the liquid-based cooling subsystem 280. In at least one embodiment, for the liquid-based cooling subsystem 280, the primary coolant cools the secondary coolant that circulates through the coolant-based secondary cooling loop 294 to cool the components within the coolant-cooled server trays 272. Although the server trays 272 are referred to as coolant-cooled server trays, it should be understood that the coolant lines can be used for all server boxes 270-276 of the rack 254, as well as all racks of the data center 252, via at least the row manifold 266 and one or more rack manifolds 296A on the inlet side and one or more rack manifolds 296B on the outlet side.
[0091] In at least one embodiment, racks 254 (including one or more racks forming a cooling subsystem) are located on an elevated platform or floor 268, and row manifolds 266 may be a collection of manifolds therein to reach all server trays in available racks in a data center or select server trays in available racks in a data center. Similarly, in at least one embodiment, for refrigerant and immersion-based cooling subsystems 282, 284, one or more medium refrigerants or conventional refrigerants are circulated via a refrigerant-based secondary cooling loop 294 (which may extend from the primary cooling loop 240) to cool the components within the refrigerant and / or immersion-cooled server trays 279, 276. In at least one embodiment, although the server trays 279, 276 are referred to as refrigerant and / or immersion cooled server trays 279, 276, it should be understood that the refrigerant (or medium refrigerant) line can be used for all server boxes 270-276 of the rack 254 and all racks of the data center 252 through at least the row manifold 266, and one or more inlet side rack manifolds 296A and one or more outlet side rack manifolds 296B, respectively. The row manifold and the rack manifold are shown as a single line, but can be a set of manifolds for respectively conveying coolant, refrigerant, and medium refrigerant to reach all server trays in the available racks in the data center or select server trays in the available racks in the data center through one or more associated flow controllers. In this way, in at least one embodiment, all server trays can be able to individually request (or indicate for) the desired level of cooling from the multi-mode cooling subsystem 256.
[0092] In at least one embodiment, a first individual cooling subsystem of the multi-mode cooling subsystem 256, such as a coolant-based cooling subsystem 280, has associated first circuit components (e.g., coolant-based secondary cooling circuits in the row manifold 266 and the rack manifolds 296A, 296B) to direct a first cooling medium (e.g., coolant) to at least one first rack 254 of the data center 252 based on a first thermal characteristic or cooling requirement of the at least one first rack 254. In at least one embodiment, the multi-mode cooling subsystem 256 has associated inlet and outlet lines 298A to circulate one or more cooling media to and from appropriate manifolds of the row manifold 266. The first thermal characteristic or cooling requirement may be a temperature sensed by a sensor associated with the at least one first rack 254. In at least one embodiment, concurrently with the coolant-based cooling subsystem 280 cooling the at least one first rack 254 (or the servers therein), the refrigerant-based cooling subsystem 282 of the multi-mode cooling subsystem 256 has an associated second circuit component (e.g., a refrigerant-based second cooling circuit in the row manifold 266 and the rack manifolds 296A, 296B) to direct a second cooling medium (e.g., a refrigerant or a medium refrigerant) to at least one second rack of the data center based on a second thermal characteristic or a second cooling requirement of the at least one second rack. This type of different cooling mediums to meet different requirements of racks may also be used as described above. Figure 2C Rack 254 in the middle shows servers with different requirements in the same rack.
[0093] In at least one embodiment, the first thermal characteristic or first cooling requirement is within a first threshold of a minimum value of a different cooling capacity. In at least one embodiment, the minimum value of the different cooling capacity is provided by the evaporative cooling subsystem 278. In at least one embodiment, the second thermal characteristic or cooling requirement is within a second threshold of a maximum value of a different cooling capacity. In at least one embodiment, the maximum value of the different cooling capacity is provided by a medium-based cooling subsystem 284, which allows the computing component to be fully immersed in a dielectric to cool the computing component by direct convection cooling, as opposed to air, coolant, and refrigerant-based media that cool by radiation and / or conduction. In at least one embodiment, a cold plate associated with the computing component absorbs heat from the computing component and transfers the heat to a coolant or refrigerant when these types of cooling media are used. In at least one embodiment, when the cooling medium is air, the cold plate transfers the heat to air flowing through the cold plate (with associated heat sinks or directly through the cold plate).
[0094] In at least one embodiment, inlet and outlet lines 298B to rack manifolds 296A, 296B of one or more racks 254 enable individual server cassettes 270-276 associated with one or more rack manifolds 296A, 296B to receive different cooling media from respective cooling subsystems 280-284 of the multi-mode cooling subsystem 256. In at least one embodiment, one or more server cassettes 270-276 have or support two or more air distribution units, coolant distribution units, and refrigerant distribution units to enable different cooling media to address different cooling requirements of one or more servers. Figures 3A-4B Server-level features of such a server are described, but briefly, a coolant-based heat exchanger (at least in one embodiment, a coolant-based cold plate or coolant lines supported by a cold plate) can be used as a coolant distribution unit for the server tray 272; a refrigerant-based heat exchanger (at least in one embodiment, a refrigerant-based cold plate or a cold plate supported by refrigerant lines) can be used as a refrigerant distribution unit for the server, and one or more fans can be used as an air distribution unit for the server box 270.
[0095] In at least one embodiment, one or more sensors 278D are provided to monitor the air quality (or environment) within the data center 252, within the racks 254, or within the server boxes 270. In at least one embodiment, the one or more sensors 278D are disposed near a return air manifold or a supply air manifold associated with the evaporative cooling subsystem 278. In at least one embodiment, the one or more sensors 278D measure the temperature and humidity of the air within the data center 252, within the racks 254, or within the server boxes 270. In at least one embodiment, the one or more sensors provide input to a distributed or centralized control system, as described throughout this disclosure.
[0096] In at least one embodiment, the distributed or centralized control system includes at least one processor and a memory, the memory including instructions for causing the distributed or centralized control system to react to the input. In at least one embodiment, when the input indicates a first temperature reflecting a first thermal characteristic or a first cooling requirement, the distributed or centralized control system is able to determine that the blown air from the evaporative cooling subsystem 278 is the best cooling solution. Provide blown air 278B. The distributed or centralized control system still communicates with the sensor 278D to determine the moisture content in the blown air 278B. The sensor 278D indicates that there is moisture (possibly including a certain amount of moisture) in the return air from the evaporative cooling subsystem that forms the blown air 278B. The distributed or centralized control system is able to determine the effect of moisture on air quality based in part on the input (such as humidity) from the sensor 278D. In at least one embodiment, the input of moisture sensed at two different time intervals can be used as a basis for determining the moisture content in the return air. In at least one embodiment, the distributed or centralized control system is able to use information about the effect of moisture on air quality to determine which (if any) flow controller to open or close.
[0097] In at least one embodiment, the flow controller to be turned on or off is associated with both the evaporative cooling subsystem 278 and the reusable refrigerant cooling subsystem 282. In at least one embodiment, the first flow controller can be modified to control or maintain the flow rate of the first cooling medium, such as the blow-in air 278B from the evaporative cooling subsystem. In at least one embodiment, the second flow controller can be modified to control the flow rate of the second medium (such as the refrigerant on the pipeline 278A from the reusable refrigerant cooling subsystem 282), which will be used in the evaporative cooling subsystem 278 to control the moisture that may affect the air quality. In at least one embodiment, the components of the reusable refrigerant cooling subsystem 282 can be shared with the evaporative cooling subsystem 278 to provide refrigerant from the condenser to the inside of the evaporator coil. The moisture in the return air (if any) condenses on the evaporator coil and is removed before the return air enters the data center 252, the rack 254 or the server box 270 as the blow-in air. This reflects at least the first configuration of the reusable refrigerant cooling subsystem 282.
[0098] In at least one embodiment, when the input provided by the sensor 278D to the distributed or centralized control system indicates that the temperature within the data center 252, the rack 254, or the server box 270 is increasing, or is not responding sufficiently to the blown air provided by the evaporative cooling subsystem 278, the distributed or centralized control system can cause the first flow controller to capture activation and can cause the second flow controller or the third flow controller associated with the reusable refrigerant cooling subsystem 282 to provide refrigerant at a third flow rate that may be higher than the second flow rate. Since the refrigerant flows to the rack or server through the cooling line 298A, the row manifold 266, the rack line 298B, and the rack manifold 26A, 296B, the flow controller can be different from the flow controller used for the evaporative cooling subsystem line 278A. The refrigerant is caused to flow to the data center 252, the rack 254, or the server box 270 via the refrigerant line of the row manifold 298A to address the problem of temperature increase, which also reflects the second thermal characteristics or the second cooling requirements. In at least one embodiment, the first cooling medium can be maintained while the second cooling medium is introduced. The first cooling medium can be deactivated after a period of time while the second cooling medium continues to flow. In at least one embodiment, the second cooling medium can be caused to flow within the evaporator coil in response to the second thermal signature to perform the dual purpose of removing moisture from the data center 252, rack 254, or server box 270 and satisfying the second thermal signature by flowing through the necessary cooling components, such as cold plates associated with the components that require cooling. This reflects at least a second configuration of the reusable refrigerant cooling subsystem 282. In at least one embodiment, the flow rate of the refrigerant can be higher in the second configuration than in the first configuration because they meet different requirements of the reusable refrigerant cooling subsystem 282.
[0099] In at least one embodiment, due to the above features, when the evaporative cooling subsystem 278 is ineffective in reducing the heat generated in at least one server or rack of the data center, the reusable refrigerant cooling subsystem 282 is caused to operate in the second configuration. In addition, in at least one embodiment, the evaporative cooling subsystem 278 and the reusable refrigerant cooling subsystem 282 provide a determined range of cooling capabilities to support the mobility of the data center through at least different utilities provided by the reusable refrigerant cooling subsystem 282.
[0100] In at least one embodiment, the evaporative cooling subsystem 278 has associated first circuit components, such as Figure 2CThe supply air manifold 222D and the return air manifold 222C in the data center are associated with directing a first cooling medium to at least one first rack of the data center according to a first thermal characteristic or a first cooling requirement of the at least one first rack. In at least one embodiment, the reusable refrigerant cooling subsystem 282 has an associated second loop component, such as in the supply and return lines 298A, 298B and in the manifolds 266, 296A, 296B, directing a second cooling medium to at least one first rack of the data center according to a second thermal characteristic or a second cooling requirement of the at least one first rack. In at least one embodiment, the first thermal characteristic or the first cooling requirement is within a first threshold of different cooling capacities provided by the multi-mode cooling subsystem 256, which has at least evaporative and reusable refrigerant cooling subsystems 278, 282. In at least one embodiment, the second thermal characteristic or the second cooling requirement is within a second threshold of different cooling capacities. In at least one embodiment, the first threshold is a value between 8 and 9KW, because the evaporative cooling subsystem can handle heat generation between 10KW and 30KW. In at least one embodiment, the second threshold is between 31 and 59 KW because the reusable refrigerant cooling subsystem 282 is rated for heat generation between 60 KW and 100 KW. In at least one embodiment, the liquid-based cooling subsystem 280 can meet intermediate requirements between the evaporative and reusable refrigerant cooling subsystems. In at least one embodiment, the evaporative cooling subsystem remains activated along with the liquid-based cooling subsystem until the reusable refrigerant cooling subsystem is activated.
[0101] In at least one embodiment, at least one processor for a cooling system is disclosed, such as Fig.14 The processor 1400 in the embodiment of the present invention may use the neuron 1502 and its components implemented using circuits or logic, including one or more arithmetic logic units (ALUs), such as Fig.15 The at least one logic unit is adapted to control a flow controller associated with the reusable refrigerant cooling subsystem 282 and the evaporative cooling subsystem 278. The control is based in part on a first thermal characteristic in the data center and in part on a second thermal characteristic in the data center. The evaporative cooling subsystem 278 works with the reusable refrigerant cooling subsystem to control the moisture of the blown air 278B from the evaporative cooling subsystem in a first configuration of the reusable refrigerant cooling subsystem, as discussed throughout the present disclosure. The reusable refrigerant cooling subsystem 282 is adapted to cool the data center in a second configuration of the reusable refrigerant cooling subsystem. In at least one embodiment, at least one processor may provide an output with instructions to the flow controller to implement the first and second configurations. The cooling of the blown air 278B and the data center depends on the first thermal characteristic and the second thermal characteristic, respectively.
[0102] In at least one embodiment, a learning subsystem in the ALU is applied to evaluate the temperature within a server or one or more racks in a data center using flow rates of different cooling media. The learning subsystem provides outputs related to at least two temperatures, and facilitates the movement of two cooling media from a cooling system by controlling a flow controller using instructions. The flow controller is associated with the two cooling media to meet a first thermal characteristic and a second thermal characteristic.
[0103] In at least one embodiment, the learning subsystem executes a machine learning model. The machine learning model processes the temperature using multiple neuron levels of the machine learning model, and the machine learning model has the temperature and previously associated flow rates of the two cooling media. In at least one embodiment, the learning subsystem provides an output associated with the flow rates of the two cooling media to the flow controller. The output is provided after evaluating the previously associated flow rates and previously associated temperatures of a single cooling medium of the two cooling media.
[0104] In at least one embodiment, at least one processor includes an instruction output, such as a pin and ball, for transmitting an output associated with the flow controller. In at least one embodiment, the pump is adapted to receive the output from the at least one processor and is adapted to adjust an internal motor or valve to change the flow through the pump. The output enables a first flow rate of a first separate cooling medium for meeting a first thermal characteristic in the data center while maintaining a second flow rate of a second separate cooling medium to control the humidity of the first separate cooling medium. The output can simultaneously deactivate the first separate cooling medium and can enable a third flow rate of the second separate cooling medium for meeting the second thermal characteristic in the data center.
[0105] In at least one embodiment, at least one logic unit is adapted to receive a temperature value from a temperature sensor and a humidity value from a humidity sensor, both of which are implemented in sensor 278D in at least one embodiment. Thus, the temperature sensor and the humidity sensor are associated with a server or one or more racks and are adapted to concurrently facilitate movement of a first separate cooling medium and a second separate cooling medium to satisfy a first thermal characteristic of the data center and separately satisfy a second thermal characteristic of the data center. In at least one embodiment, sensor 278D is adapted to facilitate movement of the first and second separate cooling media through its communication with at least one processor within a centralized or distributed control system.
[0106] In at least one embodiment, at least one processor has at least one logic unit for training one or more neural networks having a hidden layer of neurons. The training is for one or more neural networks for evaluating the temperature, flow rate, and humidity associated with the reusable refrigerant cooling subsystem 282 and the evaporative cooling subsystem 278 based in part on a first thermal signature in the data center 252 and in part on a second thermal signature in the data center 252. The evaporative cooling subsystem 278 works in conjunction with the reusable refrigerant cooling subsystem 282 to control the humidity of the blown air 278B from the evaporative cooling subsystem 278 in the first configuration of the reusable refrigerant cooling subsystem. The reusable refrigerant cooling subsystem 282 is used to cool the data center 252 in the second configuration of the reusable refrigerant cooling subsystem 282. The cooling of the blown air 278B and the data center depends on the first thermal signature and the second thermal signature, respectively. The reusable refrigerant cooling subsystem 282 provides refrigerant directly to the cold plate of the rack, server, or component to provide cooling in the second configuration.
[0107] In at least one embodiment, at least one logic unit outputs at least one instruction associated with at least two temperatures to facilitate movement of at least two cooling media by controlling flow controllers associated with the at least two cooling media. The flow controllers facilitate concurrent flow of the two cooling media in a first configuration of the reusable refrigerant cooling subsystem and facilitate flow of a single one of the two cooling media in a second configuration of the reusable refrigerant cooling subsystem.
[0108] In at least one embodiment, at least one processor has at least one instruction output for transmitting an output associated with a flow controller. In at least one embodiment, the output transmitted to the flow controller implements a first flow rate of a first separate cooling medium for satisfying a first thermal characteristic in the data center while maintaining a second flow rate of a second separate cooling medium to control humidity of the first separate cooling medium. In at least one embodiment, the output transmitted to the flow controller deactivates the first separate cooling medium and enables a third flow rate of the second separate cooling medium for satisfying a second thermal characteristic in the data center.
[0109] Figure 3A300A and server level 300B features of a cooling system including an evaporative cooling subsystem in a multi-mode cooling subsystem are shown in accordance with at least one embodiment. In at least one embodiment, the evaporative cooling subsystem is capable of satisfying the entire rack 302 by blowing cold air across the entire rack 302 and circulating hot air 312 received from the opposite end of the rack 302. In at least one embodiment, one or more fans are provided in a duct or air manifold 318 to circulate cold air 314 (which may be cold air 310 at the inlet of the rack 302) from the evaporative cooling subsystem to a single server 304. Hot air 316 flows out of the other end of the server 304 and merges with the hot air 312 leaving the rack 302. Therefore, the server 304 is one or more servers 306 in the rack 302, has one or more components associated with the cold plate 308A-D, and requires a nominal cooling level. In order to better utilize the air cooling medium, the cold plate 308A-D may have associated heat sinks to achieve more heat dissipation.
[0110] In at least one embodiment, the components associated with the cold plate 308A-D are computing components that generate 50KW to 60KW of heat, and sensors associated with one or more servers or the entire rack can indicate the temperature (indicating a cooling requirement) to the multi-mode cooling subsystem, which can then determine a specific one of the cooling subsystems 278-284 to activate for one or more servers or the rack as a whole. In at least one embodiment, an indication of the temperature can be sent to the cooling subsystem 278-284, and the cooling subsystem is designated to handle a specific temperature that corresponds to the KW of heat generated (for that specific temperature), activating itself to handle the server or rack that needs cooling. This represents a distributed control system, where the control of activation or deactivation can be shared by processors of different subsystems 278-284. In at least one embodiment, a centralized control system can be implemented by one or more processors of the server 306 in the server rack 302. The one or more processors are adapted to execute the control requirements of the multi-mode cooling system, including by using artificial intelligence and / or machine learning to train the learning subsystem to recognize temperature indications and activate or deactivate the appropriate cooling subsystem 278-284. The distributed control system can also be adapted to use artificial intelligence and / or machine learning to train the learning subsystem, such as through a distributed neural network.
[0111] Figure 3B330A and server-level 330B features of a cooling system for a liquid-based cooling subsystem incorporating a multi-mode cooling subsystem according to at least one embodiment are shown. In at least one embodiment, the liquid-based cooling subsystem is a coolant-based cooling subsystem. The rack-level features 330A support or incorporate a liquid-based cooling subsystem, including a rack 332 having inlet and outlet coolant-based manifolds 358A, 358B, having inlet and outlet couplers or adapters 354, 356 associated therewith, and having separate flow controllers 362 associated therewith (two are referenced in the figure). The inlet and outlet couplers or adapters 354, 356 receive, direct, and send coolant 352 through one or more request (or instruction) servers 336; 304. In at least one embodiment, all servers 336 in a rack 332 can be addressed by their individual requests (or indications) for the same cooling requirement, or by a request (or indication) from the rack 332 on behalf of its servers 336, which can be addressed by a liquid-based cooling subsystem. The heated coolant 360 leaves the rack through a coupler or adapter 356. In at least one embodiment, each flow controller 362 can be a pump that is controlled (activated, deactivated, or maintained in a flow position) via input from a distributed or centralized control system. Further, in at least one embodiment, when a rack 332 makes a request for (or provides an indication of) a cooling requirement, all flow controllers 362 can be controlled simultaneously to address the cooling requirement, just as if the cooling requirement represented the cooling requirement of all servers 336 at the same time.
[0112] In at least one embodiment, the request or indication comes from a sensor associated with a rack or a server. There may be one or more sensors at different locations of a rack and different locations of a server. In at least one embodiment, sensors in a distributed configuration may send separate requests or indications of temperature. In at least one embodiment, the sensors operate in a centralized configuration and are adapted to send their requests or indications to one sensor or centralized control system that aggregates the requests or indications to determine a single action. In at least one embodiment, the single action may be to send one or more cooling media from a multi-mode cooling subsystem to address the request or indication associated with the single action. In at least one embodiment, this enables the current cooling system to send cooling media associated with a higher cooling capacity than the total cooling requirements of some servers in the rack, because the aggregated requests or indications may include at least some cooling requirements that are higher than the total cooling requirements.
[0113] In at least one embodiment, a pattern statistical metric is used as a set to indicate the number of servers with higher cooling requirements, which can be addressed instead of a pure average of the server cooling requirements. Therefore, it can be understood that when at least 50% of the servers in the rack have higher cooling requirements, a subsystem that provides higher cooling capacity (e.g., a refrigerant-based cooling subsystem) can be activated. In at least one embodiment, when at most 50% of the servers in the rack have higher cooling requirements, a subsystem that provides lower cooling capacity (e.g., a coolant-based cooling subsystem) can be activated. However, when multiple servers in the rack begin to request or indicate higher cooling requirements after the coolant-based cooling subsystem is activated for the rack, higher cooling capacity can be provided to address the shortfall. In at least one embodiment, the change in a particular cooling subsystem can be based in part on the amount of time that has passed since the first request or indication occurred and the cooling to higher or lower capacity occurs when the first request or indication persists or shows an associated temperature increase.
[0114] In at least one embodiment, the server 336 has a separate server-level manifold, such as the manifold 338 in the server 334 (also referred to as a server tray or server box). The manifold 338 of the server 334 includes one or more server-level flow controllers 340A, B for addressing the cooling requirements of individual one or more components. The server manifold 338 receives coolant from the rack manifold 358A (via an adapter or coupler 348) and is caused to distribute coolant (or associated server-level coolant in heat exchange with the rack-level coolant) to the components through the associated cold plates 342A-D based on the requirements of the computing components associated with the cold plates 342A-D. In at least one embodiment, a distributed or centralized control system associated with the multi-mode cooling subsystem is capable of controlling the rack-level flow controllers 362 and the server-level flow controllers 340A, B to facilitate the movement of coolant from the liquid-based cooling subsystem. Thus, in at least one embodiment, sensors on a first side of the server 334 having components associated with the first set of cold plates 342A, B may indicate or request a higher cooling requirement than sensors on a second side of the server 334 having components associated with the second set of cold plates 342C, D. Then, if coolant is sent as the cooling medium, the coolant can flow through the flow controller 340A and the inlet tube 344A, through the first set of cold plates 342A, B, out through the outlet tube 346A, and then back to the server manifold 338. The heated coolant can exit the server manifold through the adapter or coupler 350.
[0115] In at least one embodiment, the server manifold 338 and one or more row manifolds 358A, B include multiple sub-manifolds or lines and associated couplers to support other cooling media, including refrigerants (or medium refrigerants), but coolants (liquid-based media) are illustrated and discussed herein as part of this feature. In at least one embodiment, since the sensors on the second side of the server 334 having components associated with the second set of cold plates 342C, D do not indicate or request a higher cooling requirement than the sensors on the first side of the server 334 having components associated with the first set of cold plates 342A, B, the flow controller 340B remains closed or unchanged (if previously opened for a different cooling medium). There will be no change or no cooling medium in the inlet line 344B, the second set of cold plates 342C, D, and the outlet pipe 346B of the server manifold 338.
[0116] Figure 4A 4 is a block diagram illustrating rack-level 400A and server-level 400B features of a cooling system incorporating a reusable refrigerant cooling subsystem in a multi-mode cooling subsystem according to at least one embodiment. The rack-level features 400A support or incorporate a reusable refrigerant cooling subsystem, including a rack 402 having inlet and outlet refrigerant manifolds 428A, 428B, having associated inlet and outlet couplers or adapters 424, 426, and having associated individual flow controllers 432 (two are referenced in the figure). The inlet and outlet couplers or adapters 424, 426 receive, direct, and send refrigerant 422 through one or more requesting (or indicating) servers 406; 404. In at least one embodiment, all servers 406 in a rack 402 can be addressed by their individual requests (or indications) for the same cooling requirements, or by requests (or indications) from the rack 402 on behalf of its servers 406, which can be addressed by the reusable refrigerant cooling subsystem in the second configuration for the reusable refrigerant cooling subsystem. Liquid refrigerant may enter the rack manifold through coupler or adapter 424, and vapor refrigerant 430 may exit the rack through coupler or adapter 426. In at least one embodiment, each flow controller 432 may be a pump that is controlled (activated, deactivated, or maintained at a flow position) via input from a distributed or centralized control system, as described with reference to FIG. Figure 2A-Figure 3B Furthermore, in at least one embodiment, when a rack 402 issues a request for (or provides an indication of) a cooling requirement, all flow controllers 432 may be controlled simultaneously to address the cooling requirement, just as the cooling requirement simultaneously represents the cooling requirement of all servers 406. In at least one embodiment, the request or indication comes from a sensor associated with the rack that senses temperature and / or humidity within the rack.
[0117] In at least one embodiment, the request or indication comes from a sensor associated with a rack or server. Figure 3B In embodiments, there may be one or more sensors at different locations in the rack and at different locations of the servers. In at least one embodiment, the sensors in the distributed configuration may send individual requests or indications of temperature. In at least one embodiment, the sensors operate in a centralized configuration and are adapted to send their requests or indications to one sensor or to a centralized control system that aggregates the requests or indications to determine a single action. In at least one embodiment, the single action may be sending one or more cooling media from the reusable refrigerant cooling subsystem to address the request or indication associated with the single action. In at least one embodiment, this enables the current cooling system to send cooling media associated with a higher cooling capacity than the aggregated cooling requirements of some servers in the rack, because the aggregated requests or indications may include at least some cooling requirements that are higher than the aggregated cooling requirements.
[0118] In at least one embodiment, a mode statistical metric is used as a set to indicate the number of servers with higher cooling requirements, which can be processed instead of a pure average of the server cooling requirements. Therefore, it can be understood that when at least 50% of the servers in the rack have higher cooling requirements, a subsystem that provides higher cooling capacity or capability (e.g., based on an immersion cooling subsystem) can be activated. However, servers requesting higher cooling capacity must also support higher cooling capacity. In at least one embodiment, servers supporting immersion cooling must have a sealed space in which a dielectric cooling medium (e.g., a dielectric refrigerant) is allowed to directly contact the components. In at least one embodiment, when up to 50% of the servers in the rack have higher cooling requirements, a subsystem that provides lower cooling capacity (e.g., a refrigerant or coolant-based cooling subsystem) can be activated. However, when multiple servers in the rack begin to request or indicate higher cooling requirements after activating a refrigerant or coolant-based cooling subsystem for the rack, higher cooling capacity can be provided to address the shortfall. In at least one embodiment, changes to a particular cooling subsystem may be based in part on the amount of time that has passed since a first request or indication occurred and cooling to a higher or lower capacity may occur when the first request or indication persists or shows an increase in associated temperature.
[0119] In at least one embodiment, server 406 has a separate server-level manifold, such as manifold 408 in server 404. Manifold 408 of server 404 includes one or more server-level flow controllers 410A, 410B for addressing the cooling requirements of the individual one or more components. Server manifold 408 receives liquid refrigerant from rack manifold 428A (through adapter or coupler 418) and distributes the refrigerant to the components through associated cold plates 412A-D based on the requirements of the computing components associated with the cold plates 412A-D. In at least one embodiment, a distributed or centralized control system associated with the reusable refrigerant cooling subsystem is capable of controlling rack-level flow controllers 432 and server-level flow controllers 410A, B to facilitate the movement of refrigerant from the reusable refrigerant cooling subsystem. Thus, in at least one embodiment, sensors on a first side of the server 404 having components associated with the first set of cold plates 412A, B may indicate or request a higher cooling requirement than sensors on a second side of the server 404 having components associated with the second set of cold plates 412C, D. Then, if refrigerant is sent as the cooling medium, the refrigerant can flow through the first set of cold plates 412A, B via the flow controller 410A and the inlet tube 414A, out via the outlet tube 416A, and then back to the server manifold 408. The gas phase refrigerant can leave the server manifold via the adapter or coupler 420.
[0120] In at least one embodiment, the server manifold 408 and one or more row manifolds 428A, B include multiple sub-manifolds or lines and associated couplers to support other cooling media, including coolant (or medium refrigerant), but refrigerant (reusable refrigerant medium) is described and discussed herein as part of this feature. In at least one embodiment, since the sensors of the components associated with the second set of cold plates 412C, D on the second side of the server 404 do not indicate or request a higher cooling requirement than the sensors of the components associated with the first set of cold plates 412A, B on the first side of the server 404, the flow controller 410B remains closed or unchanged (if previously opened for a different cooling medium). There will be no change or no cooling medium in the inlet tube 414B, the second set of cold plates 412C, D, and the outlet tube 416B of the server manifold 408.
[0121] Figure 4B450A and server-level 450B features of a cooling system that incorporates a dielectric-based cooling subsystem and a reusable refrigerant cooling subsystem in a multi-mode cooling subsystem according to at least one embodiment. In at least one embodiment, the dielectric-based cooling subsystem is a dielectric refrigerant-based cooling subsystem and is supported by a dielectric refrigerant from a reusable refrigerant cooling subsystem, which reduces the footprint of the multi-mode cooling subsystem. The dielectric-based cooling subsystem is capable of immersion cooling of computing components in one or more servers 456 in a rack 452. The rack-level features 450A that support or incorporate a dielectric-based cooling subsystem include a rack 452 having dielectric-based inlet and outlet manifolds 478A, 478B, having inlet and outlet couplers or adapters 474, 476 associated therewith, and having separate flow controllers 482 (two are referenced in the figure) associated therewith. In at least one embodiment, when the dielectric cooling medium is a dielectric refrigerant, a cooling assembly associated with a refrigerant-based cooling subsystem can be used to support a medium-based cooling subsystem. Inlet and outlet couplers or adapters 474, 476 receive, direct, and send dielectric cooling medium 472 via one or more request (or instruction) servers 456; 454.
[0122] In at least one embodiment, all servers 456 in a rack 452 can be addressed by their individual requests (or indications) for the same cooling requirement, or by requests (or indications) from the rack 452 on behalf of its servers that can be addressed by the dielectric-based cooling subsystem. The dielectric cooling medium 480 exits the rack through a coupler or adapter 476. In at least one embodiment, individual flow controllers 483 can be pumps that are controlled (activated, deactivated, or maintained in flow position) by input from a distributed or centralized control system. In addition, in at least one embodiment, when a rack 452 issues a request for a cooling requirement (or provides an indication thereof), all flow controllers 482 can be controlled simultaneously to address the cooling requirement, as if the cooling requirement represents the cooling requirement of all servers 456 at the same time.
[0123] In at least one embodiment, the request or indication is issued by a sensor associated with a rack or server. There may be one or more sensors at different locations of a rack and different locations of a server. In at least one embodiment, in a distributed configuration, the sensors may send separate requests or indications of temperature. In at least one embodiment, the sensors operate in a centralized configuration and are adapted to send their requests or indications to one sensor or centralized control system that aggregates the requests or indications to determine a single action. In at least one embodiment, the single action may be sending one or more cooling media from a reusable refrigerant cooling subsystem to address the request or indication associated with the single action. In at least one embodiment, this enables the current cooling system to send cooling media associated with a cooling capacity that is higher than the aggregated cooling requirements of some servers in the system, or the indication may include at least some cooling requirements that are higher than the aggregated cooling requirements.
[0124] In at least one embodiment, the pattern statistical metric is used as an aggregate to indicate the number of servers with higher cooling requirements, which can be addressed instead of a pure average of the cooling requirements of the servers. Therefore, it can be understood that a subsystem that provides higher cooling capacity (e.g., an immersion or refrigerant-based cooling subsystem) can be activated when at least 50% of the servers in the rack have higher cooling requirements. In at least one embodiment, a subsystem that provides lower cooling capacity (e.g., a refrigerant or coolant-based cooling subsystem) can be activated when up to 50% of the rack servers have higher cooling requirements. However, when multiple servers in the rack begin to request or indicate higher cooling requirements after the refrigerant or coolant-based cooling subsystem is activated for the rack, if the servers (and computing components) to be addressed are not available, higher cooling capacity (e.g., from a dielectric-based cooling subsystem) can be provided to address this shortfall. In at least one embodiment, the change in a particular cooling subsystem may be based in part on the amount of time that has passed since the first request or indication occurred, and when the first request or indication persists or shows an associated temperature increase, it changes to a higher or lower cooling capacity.
[0125] In at least one embodiment, server 456 has a separate server-level manifold, such as manifold 458 in server 454 (also referred to as a server tray or server box). Because in at least one embodiment, the servers require specific adaptations to support immersion cooling, such as being sealed. Manifold 458 of server 454 is adapted to immerse the server box for immersion cooling. Server manifold 338 receives dielectric-based cooling medium from rack manifold 478A (through adapter or coupler 468) and flows the dielectric-based cooling medium into server 454 as required by computing components associated with cold plates 462A-D, bringing the components into direct contact with the dielectric-based cooling medium. Therefore, in at least one embodiment, cold plates 462A-D may not be required. In at least one embodiment, a distributed or centralized control system associated with the reusable refrigerant cooling subsystem is capable of controlling a rack-level flow controller 482 to facilitate the movement of dielectric-based cooling medium from the media-based cooling subsystem. Thus, in at least one embodiment, sensors associated with cold plates 462A-D may indicate or request a higher cooling requirement than currently supported by any coolant, refrigerant, or evaporative cooling medium. The dielectric-based cooling medium can then flow to the servers 454 via appropriate flow controllers 482 and inlet adapters or couplers 468 and out of outlet adapters or couplers 470. Each inlet and outlet adapters or couplers 468; 470 includes a one-way valve so that the dielectric-based cooling medium flows into the server box through one adapter or coupler 468 and out of the other adapter or coupler 470.
[0126] In at least one embodiment, the server manifold 458 and one or more row manifolds 478A, B include multiple sub-manifolds or lines and associated couplers to support other cooling media, including coolants and refrigerants (if the dielectric cooling medium is different from the refrigerant), but the dielectric-based cooling medium portion is illustrated and discussed herein as being relevant to this feature. In at least one embodiment, the dielectric-based cooling medium can enter the rack manifold via a coupler or adapter 474, and the heated or used dielectric-based cooling medium 480 can exit the rack via a coupler or adapter 476.
[0127] Figure 5 According to at least one embodiment, it is possible to use or manufacture Figure 2A-Figure 4B and Figures 6A-17DThe process flow of steps of method 500 of a cooling system for cooling a data center. In at least one embodiment, the method for cooling a data center includes step 502 for providing an evaporative cooling subsystem to provide blown air to the data center based on a first thermal signature in the data center. Step 504 of method 500 provides a reusable refrigerant cooling subsystem that controls moisture of the blown air in a first configuration of the reusable refrigerant cooling subsystem. Step 506 determines that a second thermal signature is indicated using one or more sensors deployed in different servers or racks. When it is determined in step 508 that the second thermal signature requires more cooling capacity than the first thermal signature, the reusable refrigerant cooling subsystem is activated in step 510 to cool the data center in a second configuration based on the second thermal signature in the data center. Step 506 can continue to monitor the thermal signature of the server or rack via one or more sensors.
[0128] In at least one embodiment, step 502 is further supported by the following steps: directing blown air from the evaporative cooling subsystem to at least one first rack of the data center using a first loop assembly based on a first thermal characteristic or a first cooling requirement of the at least one first rack. In at least one embodiment, step 510 is further supported by the following steps: directing refrigerant from the reusable refrigerant cooling subsystem to at least one first rack of the data center using a second loop assembly based on a second thermal characteristic or a second cooling requirement of the at least one first rack.
[0129] In at least one embodiment, step 510 includes the further step of enabling the reusable refrigerant cooling subsystem to operate in a second configuration when the evaporative cooling subsystem is ineffective in reducing heat generated in at least one server or rack of the data center. In at least one embodiment, step 510 implements the second configuration of the reusable refrigerant cooling subsystem using an output from a flow controller of a centralized or distributed control system having at least one processor to provide an output.
[0130] In at least one embodiment, step 510 includes the further step of implementing one or more of an evaporative cooling subsystem, a reusable refrigerant cooling subsystem, and a dielectric-based cooling subsystem using a refrigerant-based heat transfer subsystem, each subsystem having a different cooling capacity for the data center. Figure 2C As shown, various piping and features are provided to enable refrigerant from a reusable refrigerant cooling subsystem to reach an evaporative cooling subsystem to control moisture in the blown air; to enable refrigerant to reach racks, servers, and components through a refrigerant manifold; and to enable refrigerant to reach an immersion cooling subsystem to immerse supported server boxes.
[0131] In at least one embodiment, a flow controller can be provided on the outlet side of a pipeline in any data center feature. In addition, a flow controller can be provided on both the inlet and outlet sides of a pipeline in any data center feature. Thus, when a flow controller is located on the outlet side, it performs an intake action rather than an ejection action. When two flow controllers work together, there is both an intake and an ejection action. Connecting flow controllers in series can achieve higher flow rates. For example, the configuration or adaptation of the flow controller can be determined by the requirements of the component.
[0132] In at least one embodiment, in at least one embodiment, Figure 3A The coolant-based secondary cooling loop of the embodiment facilitates the default or standard movement of coolant in the component-level cooling loop (e.g., via tubes 344A, 346A in server 334), or if the secondary cooling loop is directly available to the component, the default or standard movement of the second coolant. In at least one embodiment, the learning subsystem can be implemented by a deep learning application processor (e.g., Fig.14 The processor 1400 in FIG. 1 may be implemented and may use the neuron 1502 and components implemented using circuits or logic, including Fig.15 One or more arithmetic logic units (ALUs) are shown. Thus, the learning subsystem includes at least one processor for evaluating temperatures within servers of one or more racks having different cooling medium flow rates. The learning subsystem also provides outputs associated with at least two temperatures to facilitate movement of two cooling media of different cooling media by controlling flow controllers associated with the two cooling media to meet different cooling requirements.
[0133] Once the training is complete, the learning subsystem will be able to provide one or more outputs to the flow controller with instructions related to the flow rates of different cooling media. Figure 2A-Figure 5 As described, an output is provided after evaluating the previous relevant flow rate and the previous relevant temperature of a single cooling medium of different cooling mediums. In at least one embodiment, the temperature used in the learning subsystem can be a value inferred from the voltage and / or current value output by each component (e.g., a sensor), but the voltage and / or current value itself can be used to train the learning subsystem, as well as the required flow rate of different cooling mediums to bring one or more computing components, servers, racks, or related sensors to the required temperature. Alternatively, the temperature reflects one or more temperatures at which the cooling medium requires a necessary flow rate of the cooling medium that exceeds the temperature that the cooling medium has already provided to cool one or more computing components, servers, racks, or related sensors, and subsequent temperatures can be used in the learning subsystem to reflect the temperature at which the required flow rate of the cooling medium is no longer required.
[0134] Alternatively, in at least one embodiment, the temperature differences and subsequent temperatures of different cooling media from different cooling subsystems, together with the desired flow rates, are used to train the learning system to recognize when to activate and deactivate the associated flow controllers of the different cooling media. Once trained, the learning subsystem will be able to provide one or more outputs to control the associated flow controllers of the different cooling subsystems in response to the temperatures received from the temperature sensors associated with the different components, different servers, and different racks, through the device controller described elsewhere in this disclosure (also referred to herein as a centralized control system or a distributed control system).
[0135] In addition, the learning subsystem executes a machine learning model that processes temperatures collected from a previous application (possibly in a test environment) to control temperatures within different servers and racks and associated with different components. The collected temperatures may include (a) temperatures obtained at specific flow rates of different cooling media; (b) temperature differences obtained at specific flow rates of different cooling media over a specific time period; and (c) initial temperatures and corresponding flow rates of different media for keeping different servers, racks, or components operating optimally based on different cooling requirements. In at least one embodiment, the collected information may include a reference to a coolant or refrigerant type, or a compatibility type reflecting a desired or available cooling type (e.g., a server or rack with immersion cooling compatibility).
[0136] In at least one embodiment, the processing aspects of the deep learning subsystem may use the Fig.14 , Fig.15 The collected information is processed by the discussed features. In one example, the temperature processing uses multiple neuron levels of a machine learning model that is loaded with one or more of the collected temperature features described above and the corresponding flow rates of different cooling media. The learning subsystem performs training, which can be represented as evaluating the temperature change associated with the previous flow rate (or flow change) of each cooling medium based on the adjustments made to one or more flow controllers associated with the different cooling media. The neuron level can store values related to the evaluation process and can represent the association or correlation between the temperature change and the flow rate of each different cooling medium.
[0137] In at least one embodiment, after the learning subsystem completes training, it is able to determine the flow rates of different media required to achieve cooling to temperatures (or changes, such as temperature reductions) associated with cooling requirements of, for example, different servers, different racks, or different components in an application. Since the cooling requirements must be within the cooling capacity of the corresponding cooling subsystems, the learning subsystem is able to select cooling subsystems to be activated to meet the temperatures sensed from different servers, different racks, or different components. The collected temperatures and previously associated flow rates of different cooling media used to obtain the collected temperatures (or differences, illustratively) can be used by the learning subsystem to provide one or more outputs associated with the required flow rates of different cooling media to meet the different cooling requirements reflected by the temperatures (indicating reductions) of different servers, different racks, or different components compared to the current temperatures.
[0138] In at least one embodiment, the result of the learning subsystem is one or more outputs to flow controllers associated with different cooling subsystems that modify the flow rates of corresponding different cooling mediums in response to sensed temperatures from different servers, different racks, or different components. The modification of the flow rate enables a determined flow rate of the corresponding different cooling medium to reach the different servers, different racks, or different components that need to be cooled. The modified flow rate can be maintained until the temperature in the desired area reaches the temperature associated with the cooling medium flow rate known to the learning subsystem. In at least one embodiment, the modified flow rate can be maintained until the temperature in the area changes to a determined value. In at least one embodiment, the modified flow rate can be maintained until the temperature in the area reaches the rated temperature of the different servers, different racks, different components, or different cooling mediums.
[0139] In at least one embodiment, the equipment controller (also referred to as a centralized or distributed control system) includes at least one processor having at least one logic unit for controlling flow controllers associated with different cooling subsystems. In at least one embodiment, the equipment controller can be at least one processor within a data center, such as Fig. 7A The flow controller facilitates movement of the respective cooling medium associated with the respective cooling subsystem and facilitates cooling the zones in the data center in response to the temperatures sensed in the zones. In at least one embodiment, at least one processor is a multi-core processor (e.g., Fig. 9A In at least one embodiment, at least one logic unit may be adapted to receive a temperature value from a temperature sensor associated with a server or one or more racks and to facilitate movement of a first separate cooling medium and a second separate cooling medium of at least two cooling media.
[0140] In at least one embodiment, a processor (such as Fig. 9AThe processor cores of the multi-core processors 905, 906 in the data center may include a learning subsystem for evaluating the temperature of sensors at different locations in the data center (e.g., different locations associated with servers, racks, or even components within a server), using flow rates associated with at least two cooling media of a reusable refrigerant cooling subsystem, the multi-mode cooling subsystem having two or more evaporative cooling subsystems, liquid-based cooling subsystems, dielectric-based cooling subsystems, or reusable refrigerant cooling subsystems. The learning subsystem provides outputs, such as instructions associated with at least two temperatures, by controlling flow controllers associated with the at least two cooling media to facilitate movement of the at least two cooling media to meet different cooling requirements.
[0141] In at least one embodiment, the learning subsystem executes the machine learning model to process the temperature using multiple neuron levels of the machine learning model with different temperatures and flow rates of the cooling medium previously associated. The machine learning model may use Fig.15 The neuronal structure and Fig.14 The deep learning processor described in the invention is implemented. The machine learning model provides an output related to flow rate to one or more flow controllers, the output being derived from an evaluation of previous related flows. In addition, a command output of the processor, such as a pin of a connector bus or a ball of a ball grid array, enables the output to communicate with the one or more flow controllers to modify a first flow rate of a first cooling medium of at least two different cooling media while maintaining a second flow rate of a second separate cooling medium.
[0142] In at least one embodiment, the present disclosure relates to at least one processor or a system having at least one processor for a cooling system. The at least one processor includes at least one logic unit for training a neural network having a hidden layer of neurons for evaluating temperatures and flow rates associated with at least two cooling media of a reusable refrigerant cooling subsystem, the reusable refrigerant cooling subsystem including two or more evaporative cooling subsystems, liquid-based cooling subsystems, dielectric-based cooling subsystems, or reusable refrigerant cooling subsystems. The reusable refrigerant cooling subsystems are rated for different cooling capacities and can be adjusted according to different cooling requirements of a data center within a range of minimum and maximum values of different cooling capacities. The cooling system has a form factor of one or more racks of a data center.
[0143] As other parts of this disclosure and reference Figure 2A-Figure 5 , Fig.14 and Fig.15As described, training can be performed by a neural layer that is provided with temperature inputs for one or more zones in a data center, and associated flow rates of different cooling media from previous applications, possibly in a test environment. The temperatures can include a starting temperature (and the associated flow rate of each cooling medium applied to the starting temperature to bring the temperature to the rated temperature of one or more zones), a final temperature (after each cooling medium is applied, over each cooling medium already present in the zone, for a period of time at a certain flow rate), and the temperature difference obtained from the flow rates and usage time of the different cooling media. One or more of these features can be used to train the neural network to determine when to apply different cooling medium flows and / or when to stop different cooling medium flows for cooling, such as when the temperature sensed for a zone reaches the starting temperature, when the temperature sensed for a zone reaches the final temperature, and / or when the temperature sensed for a zone reflects suitability for that zone (e.g., the temperature is reduced to the rated temperature).
[0144] In at least one embodiment, at least one processor includes at least one logic unit, is configured or adapted to train a neural network, and is a multi-core processor. In at least one embodiment, at least one logic unit may be located within a processor core, and the processor core is capable of evaluating temperatures and flow rates associated with at least two cooling media of a reusable refrigerant cooling subsystem having two or more of an evaporative cooling subsystem, a liquid-based cooling subsystem, a dielectric-based cooling subsystem, or a reusable refrigerant cooling subsystem. The reusable refrigerant cooling subsystem may be rated for different cooling capacities and may be adjusted for different cooling needs of a data center in different minimum and maximum values of cooling capacities. The cooling system has a form factor of one or more racks of a data center.
[0145] In at least one embodiment, the present disclosure implements an integrated air, liquid, refrigeration and immersion cooling system that is a form factor for a single server rack or more racks and provides the required cooling through an evaporative cooling subsystem with a heat exchanger associated with a coolant, a liquid-based cooling subsystem of warm or cold water for chip cooling, a reusable refrigerant cooling subsystem that cools to a heat exchanger, and immersion blade cooling that reflects a dielectric-based cooling subsystem. The reusable refrigerant cooling subsystem in the rack is part of a cooling system that is capable of providing cooling by passing air through a cooled heat exchanger, but is also capable of providing cooling water from a primary liquid supply, and is also capable of providing a vapor compression cooling system for refrigerated cooling through an associated evaporator heat exchanger for one of the servers or the chips. In addition, a coolant or refrigerant with dielectric characteristics can be used in both the liquid-based cooling subsystem and the reusable refrigerant (or phase)-based cooling subsystem to eliminate the heat load from the immersion blade cooling servers in the rack.
[0146] Thus, in at least one embodiment, the present disclosure is an all-in-one rack cooling system capable of providing cooling from relatively low density racks with approximately 10KW to higher density cooling of approximately 30KW using an air-based cooling subsystem; from approximately 30KW to 60KW using a liquid-based cooling subsystem; from approximately 60KW to 100kW using a reusable refrigerant cooling subsystem, and from approximately 100KW to 450KW using a dielectric-based cooling subsystem for immersion cooling. The all-in-one rack is an all-in-one structure and reflects a footprint that can be easily deployed in almost all data centers.
[0147] Data Center
[0148] Fig. 6A It shows that you can use Figure 2A-5 In at least one embodiment, the data center 600 includes a data center infrastructure layer 610, a framework layer 620, a software layer 630, and an application layer 640. In at least one embodiment, for example Figure 2A-Figure 5In at least one embodiment described in the foregoing, features in the components of the cooling subsystem of the reusable refrigerant cooling subsystem can be performed within or in cooperation with the example data center 600. In at least one embodiment, the infrastructure layer 610, the framework layer 620, the software layer 630, and the application layer 640 can be provided in part or in whole by computing components on server trays located in the racks 210 of the data center 200. This enables the cooling system of the present disclosure to directly cool certain systems of computing components in an efficient and effective manner. In addition, various aspects of the data center, including the data center architecture layer 610, the framework layer 620, the software layer 630, and the application layer 640, can be used to support at least the above reference Figure 2A-Figure 5 Therefore, the reference Figures 6A-17D The discussion can be understood as applicable to implementing or supporting e.g. Figure 2A-Figure 5 The hardware and software features required for a reusable refrigerant cooling subsystem for a data center.
[0149] In at least one embodiment, Fig. 6A As shown, the data center infrastructure layer 610 may include a resource coordinator 612, group computing resources 614, and node computing resources ("node CRs") 616 (1)-616 (N), where "N" represents any complete positive integer. In at least one embodiment, the node CRs 616 (1)-616 (N) may include, but are not limited to, any number of central processing units ("CPUs") or other processors (including accelerators, field programmable gate arrays (FPGAs), graphics processors, etc.), memory devices (e.g., dynamic read-only memories), storage devices (e.g., solid-state or disk drives), network input / output ("NW I / O") devices, network switches, virtual machines ("VMs"), power modules, and cooling modules, etc. In at least one embodiment, one or more of the node CRs 616 (1)-616 (N) may be a server having one or more of the above-mentioned computing resources.
[0150] In at least one embodiment, the grouped computing resources 614 may include a separate grouping (not shown) of node CRs housed in one or more racks, or many racks (also not shown) housed in data centers at various geographic locations. The separate grouping of node CRs within the grouped computing resources 614 may include computing, networks, memory, or storage resources that can be configured or allocated to support groupings of one or more workloads. In at least one embodiment, several node CRs including a CPU or processor may be grouped in one or more racks to provide computing resources to support one or more workloads. In at least one embodiment, one or more racks may also include any number of power modules, cooling modules, and network switches in any combination.
[0151] In at least one embodiment, resource coordinator 612 may configure or otherwise control one or more nodes CR 616(1)-616(N) and / or grouped computing resources 614. In at least one embodiment, resource coordinator 612 may include a software design infrastructure ("SDI") management entity for data center 600. In at least one embodiment, resource coordinator 612 may include hardware, software, or some combination thereof.
[0152] In at least one embodiment, Fig. 6AAs shown, the framework layer 620 includes a job scheduler 622, a configuration manager 624, a resource manager 626, and a distributed file system 628. In at least one embodiment, the framework layer 620 may include a framework that supports software 632 of the software layer 630 and / or one or more applications 642 of the application layer 640. In at least one embodiment, the software 632 or the application 642 may include web-based service software or applications, such as services or applications provided by Amazon Web Services, Google Cloud, and Microsoft Azure. In at least one embodiment, the framework layer 620 may be, but is not limited to, a free and open source software network application framework, such as Apache SparkTM (hereinafter referred to as "Spark") that can utilize the distributed file system 628 for large-scale data processing (e.g., "big data"). In at least one embodiment, the job scheduler 622 may include a Spark driver to facilitate scheduling of workloads supported by various layers of the data center 600. In at least one embodiment, the configuration manager 624 may be able to configure different layers, such as the software layer 630 and the framework layer 620 including Spark and a distributed file system 628 for supporting large-scale data processing. In at least one embodiment, the resource manager 626 can manage cluster or group computing resources mapped to or allocated to support the distributed file system 628 and the job scheduler 622. In at least one embodiment, the cluster or group computing resources can include group computing resources 614 on the data center infrastructure layer 610. In at least one embodiment, the resource manager 626 can coordinate with the resource coordinator 612 to manage these mapped or allocated computing resources.
[0153] In at least one embodiment, software 632 included in software layer 630 may include software used by at least a portion of node CRs 616(1)-616(N), grouped computing resources 614, and / or distributed file system 628 of framework layer 620. One or more types of software may include, but are not limited to, Internet web page search software, email virus scanning software, database software, and streaming video content software.
[0154] In at least one embodiment, one or more applications 642 included in the application layer 640 may include one or more types of applications used by at least a portion of the node CRs 616(1)-616(N), the grouped computing resources 614, and / or the distributed file system 628 of the framework layer 620. The one or more types of applications may include, but are not limited to, any number of genomics applications, cognitive computing, and machine learning applications, including training or inference software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), or other machine learning applications used in conjunction with one or more embodiments.
[0155] In at least one embodiment, any of configuration manager 624, resource manager 626, and resource coordinator 612 can implement any number and type of self-modification actions based on any number and type of data acquired in any technically feasible manner. In at least one embodiment, the self-modification actions can relieve a data center operator of data center 600 from making potentially bad configuration decisions and can avoid underutilized and / or poorly performing portions of the data center.
[0156] In at least one embodiment, the data center 600 may include tools, services, software, or other resources to train one or more machine learning models or use one or more machine learning models to predict or infer information according to one or more embodiments of the present invention. In at least one embodiment, the machine learning model can be trained by calculating weight parameters according to the neural network architecture using the software and computing resources described above for the data center 600. In at least one embodiment, by using the weight parameters calculated by one or more training techniques herein, the resources described above with respect to the data center 600 can be used to infer or predict information using the trained machine learning model corresponding to one or more neural networks. As previously discussed, deep learning techniques can be used to support intelligent control of flow controllers in refrigerant-assisted cooling by monitoring the regional temperature of the data center. Any appropriate learning network and computing function of the data center 600 can be used to advance deep learning. Therefore, in this way, hardware in the data center can be used to simultaneously or concurrently support deep neural networks (DNNs), recursive neural networks (RNNs), or convolutional neural networks (CNNs). For example, once the network is trained and successfully evaluated to identify data in a subset or slice, the trained network can provide similar representative data to be used with the collected data.
[0157] In at least one embodiment, the data center 600 can use a CPU, an application-specific integrated circuit (ASIC), a GPU, an FPGA, or other hardware to use the above resources to perform training and / or reasoning. In addition, one or more of the above software and / or hardware resources can be configured as a service to allow users to train or perform information reasoning, such as pressure, flow rate, temperature, and location information or other artificial intelligence services.
[0158] Reasoning and training logic
[0159] Reasoning and / or training logic 615 may be used to perform reasoning and / or training operations associated with one or more embodiments. In at least one embodiment, reasoning and / or training logic 615 may be used in the system Fig. 6A Inference and / or training logic 615 may be used to reason or predict operations based at least in part on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases herein. In at least one embodiment, reasoning and / or training logic 615 may include, but is not limited to, hardware logic in which computing resources are dedicated or otherwise uniquely used in conjunction with weight values or other information corresponding to one or more layers of neurons within a neural network. In at least one embodiment, reasoning and / or training logic 615 may be used in conjunction with an application specific integrated circuit (ASIC), such as the one from Google. Processing unit from Graphcore TM Inference Processing Unit (IPU) from Intel Corp (e.g. "LakeCrest") processor.
[0160] In at least one embodiment, the reasoning and / or training logic 615 can be used in conjunction with central processing unit (CPU) hardware, graphics processing unit (GPU) hardware, or other hardware (e.g., field programmable gate array (FPGA)). In at least one embodiment, the reasoning and / or training logic 615 includes, but is not limited to, a code and / or data storage model that can be used to store code (e.g., graphics code), weight values, and / or other information, including bias values, gradient information, momentum values, and / or other parameters or hyperparameter information. In at least one embodiment, each code and / or data storage module is associated with a dedicated computing resource. In at least one embodiment, the dedicated computing resource includes computing hardware that also includes one or more ALUs that perform mathematical functions (e.g., linear algebra functions) only on information stored in the code and / or data storage module, and stores the results stored therefrom in an activated storage module of the reasoning and / or training logic 615.
[0161] Figure 6B , Figure 6CInference and / or training logic according to at least one embodiment is shown, such as in Fig. 6A The reasoning and / or training logic 615 is used to perform reasoning and / or training operations associated with at least one embodiment of the present disclosure. Figure 6B and / or Figure 6C Provides details about the inference and / or training logic 615. Distinguished from the computational hardware 602, 606 by the use of an arithmetic logic unit (ALU) Figure 6B and Figure 6C In at least one embodiment, each of computing hardware 602 and computing hardware 606 includes one or more ALUs that perform mathematical functions (e.g., linear algebraic functions) on information stored in code and / or data storage 601 and information in code and / or data storage 605, respectively, with the results stored in activation storage 690. Thus, unless otherwise specified, Figure 6B and Figure 6C may be substituted and used interchangeably.
[0162] In at least one embodiment, the inference and / or training logic 615 may include, but is not limited to, code and / or data storage 601 to store forward and / or output weights and / or input / output data and / or other parameters of neurons or layers of a neural network trained and / or used for inference in at least one embodiment. In at least one embodiment, the training logic 615 may include or be coupled to code and / or data storage 601 for storing graph code or other software to control timing and / or sequence, wherein weight and / or other parameter information is loaded to configure logic, including integer and / or floating point units (collectively referred to as arithmetic logic units (ALUs)). In at least one embodiment, the code (such as graph code) loads weight or other parameter information into the processor ALU based on the architecture of the neural network to which the code corresponds. In at least one embodiment, the code and / or data storage 601 stores weight parameters and / or input / output data for each layer of a neural network trained or used in conjunction with at least one embodiment during forward propagation of input / output data and / or weight parameters during training and / or inference using aspects of at least one embodiment. In at least one embodiment, any portion of code and / or data storage 601 may be included in other on-chip or off-chip data storage, including a processor's L1, L2, or L3 cache or system memory.
[0163] In at least one embodiment, any portion of code and / or data storage 601 may be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, code and / or code and / or data storage 601 may be cache memory, dynamic random addressable memory ("DRAM"), static random addressable memory ("SRAM"), non-volatile memory (e.g., flash memory), or other storage. In at least one embodiment, the choice of whether code and / or data storage 601 is internal or external to a processor, for example, or includes DRAM, SRAM, flash memory, or some other storage type, may depend on the available storage space on-chip or off-chip, the latency requirements of the training and / or inference functions being performed, the data batch size used in inference and / or training of a neural network, or some combination of these factors.
[0164] In at least one embodiment, the inference and / or training logic 615 may include, but is not limited to, code and / or data storage 605 to store backpropagation and / or output weights and / or input / output data corresponding to neurons or layers of a neural network trained and / or used for inference in at least one embodiment aspect of the neural network. In at least one embodiment, the code and / or data storage 605 stores weight parameters and / or input / output data for each layer of a neural network trained or used in conjunction with at least one embodiment during backpropagation of input / output data and / or weight parameters during training and / or inference using at least one embodiment. In at least one embodiment, the training logic 615 may include or be coupled to code and / or data storage 605 for storing graph code or other software to control the timing and / or order in which weights and / or other parameter information is loaded to configure logic including integer and / or floating point units (collectively referred to as arithmetic logic units (ALUs)).
[0165] In at least one embodiment, code (such as graph code) loads weights or other parameter information into the processor ALU based on the architecture of the neural network to which the code corresponds. In at least one embodiment, any portion of the code and / or data storage 605 may be included with other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory. In at least one embodiment, any portion of the code and / or data storage 605 may be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, the code and / or data storage 605 may be cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other storage. In at least one embodiment, the choice of whether the code and / or data storage 605 is internal or external to the processor, for example, including DRAM, SRAM, flash memory, or some other storage type, depends on whether the available storage is on-chip or off-chip, the latency requirements of the training and / or inference functions being performed, the batch size of data used in the inference and / or training of the neural network, or some combination of these factors.
[0166] In at least one embodiment, code and / or data storage 601 and code and / or data storage 605 may be separate storage structures. In at least one embodiment, code and / or data storage 601 and code and / or data storage 605 may be the same storage structure. In at least one embodiment, code and / or data storage 601 and code and / or data storage 605 may be partially the same storage structure and partially separate storage structures. In at least one embodiment, any portion of code and / or data storage 601 and code and / or data storage 605 may be included with other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory.
[0167] In at least one embodiment, the inference and / or training logic 615 may include, but is not limited to, one or more arithmetic logic units (“ALUs”) including integer and / or floating point units for performing logic and / or mathematical operations based at least in part on or as directed by training and / or inference code (e.g., graph code), the results of which may produce activations (e.g., output values from a layer or neuron within a neural network) stored in activation storage 690, which is a function of input / output and / or weight parameter data stored in code and / or data storage 601 and / or code and / or data storage 605. In at least one embodiment, activations are generated in response to executing instructions or other code, linear algebra and / or matrix-based mathematics performed by the ALU to generate activations stored in activation storage 690, where weight values stored in code and / or data storage 605 and / or in code and / or data storage 601 are used as operands with other values, such as bias values, gradient information, momentum values, or other parameters or hyperparameters, any or all of which may be stored in code and / or data storage 605 and / or code and / or data storage 601 or other on-chip or off-chip storage.
[0168] In at least one embodiment, one or more ALUs are included in one or more processors or other hardware logic devices or circuits, while in another embodiment, one or more ALUs may be outside of a processor or other hardware logic device or circuit that uses them (e.g., a coprocessor). In at least one embodiment, one or more ALUs may be included within an execution unit of a processor, or otherwise included in a group of ALUs accessible by an execution unit of a processor, which may be within the same processor or distributed between different processors of different types (e.g., a central processing unit, a graphics processing unit, a fixed function unit, etc.). In at least one embodiment, code and / or data storage 601, code and / or data storage 605, and activation storage 690 may be on the same processor or other hardware logic device or circuit, while in another embodiment, they may be on different processors or other hardware logic devices or circuits or some combination of the same and different processors or other hardware logic devices or circuits. In at least one embodiment, any portion of activation storage 690 may be included with other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory. Additionally, inference and / or training code may be stored with other code accessible to a processor or other hardware logic or circuitry and may be retrieved and / or processed using the processor's fetch, decode, schedule, execute, exit, and / or other logic circuitry.
[0169] In at least one embodiment, activation storage 690 may be cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other storage. In at least one embodiment, activation storage 690 may be completely or partially internal or external to one or more processors or other logic circuits. In at least one embodiment, activation storage 690 may be internal or external to a processor, for example, or include DRAM, SRAM, flash memory, or other storage types, depending on the storage available on-chip or off-chip, the latency requirements for performing training and / or inference functions, the batch size of data used in inferencing and / or training neural networks, or some combination of these factors. In at least one embodiment, Figure 6B The inference and / or training logic 615 shown in FIG. 6 may be used in conjunction with an application specific integrated circuit (“ASIC”), such as the ASIC from Google. Processing unit from Graphcore TM Inference Processing Unit (IPU) from Intel Corp (e.g., "Lake Crest") processor. In at least one embodiment, Figure 6B The illustrated inference and / or training logic 615 may be used in conjunction with central processing unit (“CPU”) hardware, graphics processing unit (“GPU”) hardware, or other hardware such as a field programmable gate array (“FPGA”).
[0170] In at least one embodiment, Figure 6C Inference and / or training logic 615 is shown, which may include, but is not limited to, hardware logic, wherein computing resources are dedicated or otherwise uniquely used in conjunction with weight values or other information corresponding to one or more layers of neurons within a neural network, in accordance with at least one various embodiment. In at least one embodiment, Figure 6C The inference and / or training logic 615 shown in FIG. 6 can be used in conjunction with an application specific integrated circuit (ASIC), such as the ASIC from Google. Processing unit from Graphcore TM Inference Processing Unit (IPU) from Intel Corp (e.g., "Lake Crest") processor. In at least one embodiment, Figure 6CThe inference and / or training logic 615 shown in can be used in conjunction with central processing unit (CPU) hardware, graphics processing unit (GPU) hardware, or other hardware (e.g., field programmable gate array (FPGA)). In at least one embodiment, the inference and / or training logic 615 includes, but is not limited to, code and / or data storage 601 and code and / or data storage 605, which can be used to store code (e.g., chart code), weight values, and / or other information, including bias values, gradient information, momentum values, and / or other parameter or hyperparameter information. Figure 6C In at least one embodiment shown in , each of code and / or data store 601 and code and / or data store 605 are associated with dedicated computing resources (eg, computing hardware 602 and computing hardware 606 ), respectively.
[0171] In at least one embodiment, each of the code and / or data stores 601 and 605 and the corresponding computing hardware 602 and 606 corresponds to a different layer of the neural network, such that activations from one "storage / compute pair 601 / 602" of the code and / or data store 601 and computing hardware 602 are provided as inputs to the next "storage / compute pair 605 / 606" of the code and / or data store 605 and computing hardware 606, so as to reflect the conceptual organization of the neural network. In at least one embodiment, each storage / compute pair 601 / 602 and 605 / 606 can correspond to more than one neural network layer. In at least one embodiment, additional storage / compute pairs (not shown) can be included in the inference and / or training logic 615 after or in parallel with the storage / compute pairs 601 / 602 and 605 / 606.
[0172] Computer Systems
[0173] Fig. 7A A block diagram of an exemplary computer system 700A is shown in accordance with at least one embodiment, the exemplary computer system may be a system of interconnected devices and components, a system on a chip (SOC), or some combination formed with a processor, the processor may include an execution unit to execute instructions to support and / or enable intelligent control of a multi-mode cooling subsystem with a reusable refrigerant cooling subsystem as described herein. In at least one embodiment, the computer system 700A may include, but is not limited to, components (such as processor 702) to execute algorithms for process data using execution units including logic in accordance with the present disclosure (such as the embodiments herein). In at least one embodiment, the computer system 700A may include a processor such as an Intel® processor available from Intel Corporation of Santa Clara, California. Processor family, XeonTM, XScaleTM and / or StrongARMTM, CoreTM or Nervana TM microprocessor, but other systems (including PCs with other microprocessors, engineering workstations, set-top boxes, etc.) may also be used. In at least one embodiment, computer system 700B may execute a version of the WINDOWS operating system available from Microsoft Corporation of Redmond, Washington, although other operating systems (such as UNIX and Linux), embedded software, and / or graphical user interfaces may also be used.
[0174] In at least one embodiment, the exemplary computer system 700A may combine components 110-116 (from Figure 1 ) to support processing aspects for intelligent control of the reusable refrigerant cooling subsystem. For at least this reason, in one embodiment, Fig. 7A The system is shown as comprising interconnected hardware devices or "chips", while in other embodiments, Fig. 7A An exemplary system-on-chip SoC may be shown. In at least one embodiment, Fig. 7A The devices shown in the figure may be interconnected with a proprietary interconnect, a standardized interconnect (e.g., PCIe), or some combination thereof. In at least one embodiment, one or more components of computer system 700B are interconnected using a compute express link (CXL) interconnect. Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments, for example, as previously described with respect to Fig. 6A -C discussed below. Fig. 6A -C provides details about the reasoning and / or training logic 615. In at least one embodiment, the reasoning and / or training logic 615 may be implemented in the system Fig. 7A for inferring or predicting operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures, or neural network use cases herein.
[0175] Embodiments may be used in other devices, such as handheld devices and embedded applications. Some examples of handheld devices include cellular phones, Internet Protocol (Internet Protocol) devices, digital cameras, personal digital assistants ("PDAs"), and handheld PCs. In at least one embodiment, the embedded application may include a microcontroller, a digital signal processor ("DSP"), a system on a chip, a network computer ("NetPC"), a set-top box, a network hub, a wide area network ("WAN") switch, or any other system that can execute one or more instructions according to at least one embodiment.
[0176] In at least one embodiment, the computer system 700A may include, but is not limited to, a processor 702, which may include, but is not limited to, one or more execution units 708 to perform machine learning model training and / or reasoning according to the techniques described herein. In at least one embodiment, the computer system 700A is a single-processor desktop or server system, but in another embodiment, the computer system 700A may be a multi-processor system. In at least one embodiment, the processor 702 may include, but is not limited to, a complex instruction set computer ("CISC") microprocessor, a reduced instruction set computing ("RISC") microprocessor, a very long instruction word ("VLIW") microprocessor, a processor that implements an instruction set combination, or any other processor device, such as a digital signal processor. In at least one embodiment, the processor 702 may be coupled to a processor bus 710, which may transmit data signals between the processor 702 and other components in the computer system 700A.
[0177] In at least one embodiment, processor 702 may include, but is not limited to, level 1 ("L1") internal cache memory ("cache") 704. In at least one embodiment, processor 702 may have a single internal cache or multiple levels of internal cache. In at least one embodiment, cache memory may reside external to processor 702. Other embodiments may also include a combination of internal and external caches, depending on the particular implementation and needs. In at least one embodiment, register file 706 may store different types of data in various registers, including, but not limited to, integer registers, floating point registers, status registers, and instruction pointer registers.
[0178] In at least one embodiment, an execution unit 708, including but not limited to logic to perform integer and floating point operations, is also located in the processor 702. In at least one embodiment, the processor 702 may also include a microcode ("ucode") read-only memory ("ROM") for storing microcode for certain macroinstructions. In at least one embodiment, the execution unit 708 may include logic for processing a packed instruction set 709. In at least one embodiment, by including the packed instruction set 709 in the instruction set of a general-purpose processor, and the associated circuitry to execute the instructions, packed data in the general-purpose processor 702 may be used to perform operations used by many multimedia applications. In one or more embodiments, many multimedia applications may be accelerated and executed more efficiently by using the full width of the processor's data bus to perform operations on the packed data, which may not require the transfer of smaller units of data on the processor's data bus to perform one or more operations one data element at a time.
[0179] In at least one embodiment, execution unit 708 may also be used in microcontrollers, embedded processors, graphics devices, DSPs, and other types of logic circuits. In at least one embodiment, computer system 700A may include, but is not limited to, memory 720. In at least one embodiment, memory 720 may be implemented as a dynamic random access memory ("DRAM") device, a static random access memory ("SRAM") device, a flash memory device, or other storage device. In at least one embodiment, memory 720 may store instructions 719 and / or data 721 represented by data signals that may be executed by processor 702.
[0180] In at least one embodiment, the system logic chip may be coupled to the processor bus 710 and the memory 720. In at least one embodiment, the system logic chip may include, but is not limited to, a memory controller hub ("MCH") 716, and the processor 702 may communicate with the MCH 716 via the processor bus 710. In at least one embodiment, the MCH 716 may provide a high bandwidth memory path 718 to the memory 720 for instruction and data storage and for storage of graphics commands, data, and textures. In at least one embodiment, the MCH 716 may initiate data signals between the processor 702, the memory 720, and other components in the computer system 700A, and bridge data signals between the processor bus 710, the memory 720, and the system I / O 722. In at least one embodiment, the system logic chip may provide a graphics port for coupling to a graphics controller. In at least one embodiment, the MCH 716 may be coupled to the memory 720 via a high bandwidth memory path 718 , and the graphics / video card 712 may be coupled to the MCH 716 via an Accelerated Graphics Port (“AGP”) interconnect 714 .
[0181] In at least one embodiment, the computer system 700A can use the system I / O 722 as a proprietary hub interface bus to couple the MCH 716 to an I / O controller hub ("ICH") 730. In at least one embodiment, the ICH 730 can provide direct connections to certain I / O devices through a local I / O bus. In at least one embodiment, the local I / O bus can include, but is not limited to, a high-speed I / O bus used to connect peripheral devices to the memory 720, the chipset, and the processor 702. Examples can include, but are not limited to, an audio controller 729, a firmware hub ("Flash BIOS") 728, a wireless transceiver 726, a data store 724, a traditional I / O controller 723 including a user input and keyboard interface, a serial expansion port 727 (e.g., a universal serial bus (USB) port), and a network controller 734. The data store 724 can include a hard drive, a floppy drive, a CD-ROM device, a flash memory device, or other mass storage device.
[0182] Figure 7B 7 is a block diagram illustrating an electronic device 700B for utilizing a processor 710 to support and / or enable intelligent control of a multi-mode cooling subsystem with a reusable refrigerant cooling subsystem as described herein, according to at least one embodiment. In at least one embodiment, the electronic device 700B may be, for example but not limited to, a notebook computer, a tower server, a rack server, a blade server, a laptop computer, a desktop computer, a tablet computer, a mobile device, a phone, an embedded computer, or any other suitable electronic device. In at least one embodiment, the exemplary electronic device 700B may incorporate one or more of the components that support aspects of processing a reusable refrigerant cooling subsystem.
[0183] In at least one embodiment, system 700B may include, but is not limited to, a processor 710 communicatively coupled to any suitable number or kind of components, peripherals, modules, or devices. In at least one embodiment, processor 710 is coupled using a bus or interface, such as an I2C bus, a system management bus ("SMBus"), a low pin count (LPC) bus, a serial peripheral interface ("SPI"), a high-definition audio ("HDA") bus, a serial advanced technology attachment ("SATA") bus, a universal serial bus ("USB") (versions 1, 2, 3), or a universal asynchronous receiver / transmitter ("UART") bus.
[0184] In at least one embodiment, Figure 7B A system is shown that includes interconnected hardware devices or "chips", while in other embodiments, Figure 7B An exemplary system on a chip ("SoC") may be shown. In at least one embodiment, Figure 7BThe devices shown in can be interconnected with a proprietary interconnect, a standardized interconnect (e.g., PCIe), or some combination thereof. In at least one embodiment, Figure 7B One or more components of the system are interconnected using Compute Express Link (CXL) interconnect lines.
[0185] In at least one embodiment, Figure 7B The display 724, the touch screen 725, the touch pad 730, the near field communication unit ("NFC") 745, the sensor hub 740, the thermal sensor 746, the fast chipset ("EC") 735, the trusted platform module ("TPM") 738, the BIOS / firmware / flash memory ("BIOS, FW Flash") 722, the DSP 760, the drive 720 (such as a solid state disk ("SSD") or a hard disk drive ("HDD")), the wireless local area network unit ("WLAN") 750, the Bluetooth unit 752, the wireless wide area network unit ("WWAN") 756, the global positioning system (GPS) unit 755, the camera ("USB 3.0 camera") 754 (such as a USB 3.0 camera) and / or the low power double data rate ("LPDDR") memory unit ("LPDDR3") 715 implemented in, for example, the LPDDR3 standard. Each of these components can be implemented in any suitable manner.
[0186] In at least one embodiment, other components may be communicatively coupled to the processor 710 via the following components. In at least one embodiment, the accelerometer 741, ambient light sensor ("ALS") 742, compass 743, and gyroscope 744 may be communicatively coupled to the sensor hub 740. In at least one embodiment, the thermal sensor 739, fan 737, keyboard 736, and touchpad 730 may be communicatively coupled to the EC 735. In at least one embodiment, the speaker 763, earphone 764, and microphone ("mic") 765 may be communicatively coupled to the audio unit ("audio codec and class D amplifier") 762, which in turn may be communicatively coupled to the DSP 760. In at least one embodiment, the audio unit 762 may include, for example, but not limited to, an audio encoder / decoder ("codec") and a class D amplifier. In at least one embodiment, the SIM card ("SIM") 757 may be communicatively coupled to the WWAN unit 756. In at least one embodiment, components such as the WLAN unit 750 and the Bluetooth unit 752 and the WWAN unit 756 may be implemented as a next generation form factor (NGFF).
[0187] Reasoning and / or training logic 615 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 6B and / or Figure 6CProvides details about the reasoning and / or training logic 615. In at least one embodiment, the reasoning and / or training logic 615 may be implemented in the system Figure 7B for use in reasoning or predicting operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures, or neural network use cases herein.
[0188] Figure 7C A computer system 700C is shown according to at least one embodiment, which is used to support and / or implement the intelligent control of the multi-mode cooling subsystem with a reusable refrigerant cooling subsystem described herein. In at least one embodiment, the computer system 700C includes, but is not limited to, a computer 771 and a USB disk 770. In at least one embodiment, the computer 771 may include, but is not limited to, any number and type of processors (not shown) and memories (not shown). In at least one embodiment, the computer 771 includes, but is not limited to, a server, a cloud instance, a laptop computer, and a desktop computer.
[0189] In at least one embodiment, the USB disk 770 includes, but is not limited to, a processing unit 772, a USB interface 774, and a USB interface logic 773. In at least one embodiment, the processing unit 772 can be any instruction execution system, device, or device capable of executing instructions. In at least one embodiment, the processing unit 772 can include, but is not limited to, any number and type of processing cores (not shown). In at least one embodiment, the processing unit or core 772 includes an application specific integrated circuit ("ASIC") that is optimized to perform any number and type of operations associated with machine learning. For example, in at least one embodiment, the processing core 772 is a tensor processing unit ("TPC") that is optimized to perform machine learning reasoning operations. In at least one embodiment, the processing core 772 is a visual processing unit ("VPU") that is optimized to perform machine vision and machine learning reasoning operations.
[0190] In at least one embodiment, the USB interface 774 can be any type of USB connector or USB socket. For example, in at least one embodiment, the USB interface 774 is a USB 3.0 Type-C socket for data and power. In at least one embodiment, the USB interface 774 is a USB 3.0 Type-A connector. In at least one embodiment, the USB interface logic 773 can include any number and type of logic that enables the processing unit 772 to connect to a device (e.g., computer 771) via the USB connector 774.
[0191] Reasoning and / or training logic 615 (e.g., regarding Figure 6B and Figure 6CThe invention is described in detail and is used to perform reasoning and / or training operations related to one or more embodiments. Figure 6B and Figure 6C Provides details about the reasoning and / or training logic 615. In at least one embodiment, the reasoning and / or training logic 615 can be used to Figure 7C In a system of the present invention, operations can be inferred or predicted based at least in part on weight parameters calculated using the neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0192] Figure 8 A further exemplary computer system 800 for implementing the various processes and methods of the reusable refrigerant cooling subsystem described throughout the present disclosure is shown in accordance with at least one embodiment. In at least one embodiment, the computer system 800 includes, but is not limited to, at least one central processing unit ("CPU") 802 connected to a communication bus 810 implemented using any suitable protocol, such as PCI ("Peripheral Component Interconnect"), Peripheral Component Interconnect Express ("PCI-Express"), AGP ("Accelerated Graphics Port"), HyperTransport, or any other bus or point-to-point communication protocol. In at least one embodiment, the computer system 800 includes, but is not limited to, a main memory 804 and control logic (e.g., implemented as hardware, software, or a combination thereof), and data may be stored in the main memory 804 in the form of random access memory ("RAM"). In at least one embodiment, a network interface subsystem ("network interface") 822 provides an interface to other computing devices and networks for receiving data from the computer system 800 and transmitting data to other systems.
[0193] In at least one embodiment, computer system 800 includes, but is not limited to, input device 808, parallel processing system 812, and display device 806, which may be implemented using cathode ray tubes ("CRT"), liquid crystal displays ("LCD"), light emitting diodes ("LED"), plasma displays, or other suitable display technologies. In at least one embodiment, user input is received from input device 808 (such as a keyboard, mouse, touch pad, microphone, and more). In at least one embodiment, each of the above modules may be located on a single semiconductor platform to form a processing system.
[0194] The reasoning and / or training logic 615 is used to perform reasoning and / or training operations associated with one or more embodiments, such as those previously described with respect to Fig. 6A -C discussed. The following combination Fig. 6A -C provides details about the reasoning and / or training logic 615. In at least one embodiment, the reasoning and / or training logic 615 may be implemented in the system Figure 8 In at least one embodiment, the inference and / or training logic 615 may be used in the system to perform inference or prediction operations based at least in part on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases herein. Figure 8 to perform inference or prediction operations based at least in part on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases herein.
[0195] Fig. 9A An exemplary architecture is shown in which multiple GPUs 910-913 are communicatively coupled to multiple multi-core processors 940-943 via high-speed links 905-906 (e.g., bus / point-to-point interconnect, etc.). In one embodiment, the high-speed links 940-943 support 4GB / s, 30GB / s, 80GB / s, or higher communication throughput. Various interconnect protocols can be used, including but not limited to PCIe 4.0 or 5.0 and NVLink 2.0.
[0196] Furthermore, in one embodiment, two or more GPUs 910-913 are interconnected via high-speed links 929-930, which may be implemented using the same or different protocols / links as used for high-speed links 940-943. Similarly, two or more multi-core processors 905-906 may be connected via high-speed link 928, which may be a symmetric multiprocessor (SMP) bus running at 20 GB / s, 30 GB / s, 120 GB / s, or higher. Alternatively, the same protocol / links may be used (e.g., via a common interconnect fabric) to accomplish the above. Fig. 9A All communications between the various system components shown in .
[0197] In one embodiment, each multi-core processor 905-906 is communicatively coupled to processor memory 901-902 via memory interconnects 926-927, respectively, and each GPU 910-913 is communicatively coupled to GPU memory 920-923 via GPU memory interconnects 950-953, respectively. The memory interconnects 926-927 and 950-953 may utilize the same or different memory access technologies. By way of example and not limitation, the processor memory 901-902 and the GPU memory 920-923 may be volatile memory, such as dynamic random access memory (DRAM) (including stacked DRAM), graphics DDR SDRAM (GDDR) (e.g., GDDR5, GDDR6), or high bandwidth memory (HBM), and / or may be non-volatile memory, such as 3D XPoint or Nano-Ram. In one embodiment, some portion of the processor memory 901-902 may be volatile memory, while another portion may be non-volatile memory (e.g., using a two-level memory (2LM) hierarchy).
[0198] As follows, although the various processors 905-906 and GPUs 910-913 may be physically coupled to specific memories 901-902, 920-923, respectively, a unified memory architecture may be implemented in which a virtual system address space (also referred to as an "effective address" space) is distributed among the various physical memories. In at least one embodiment, the processor memories 901-902 may each include 64GB of system memory address space, and the GPU memories 920-923 may each include 32GB of system memory address space (resulting in a total addressable memory size of 256GB in this example).
[0199] As discussed elsewhere in this disclosure, at least flow rate and associated temperature can be established for an intelligent learning system, such as a neural network system. Since the first level represents previous data, it also represents a smaller subset of data that can be used to improve the system by training the system. Testing and training can be performed in parallel using multiple processing units, making the intelligent learning system robust. For example, Fig. 9A When the intelligent learning system achieves convergence, the data points and the number of data points used to cause convergence are recorded. The data and data points can be used to fully control the reference, for example Figure 2A-Figure 5 A reusable refrigerant cooling subsystem and a multi-mode cooling subsystem are discussed.
[0200] Fig. 9BAdditional details for interconnection between multi-core processor 907 and graphics acceleration module 946 are shown according to an exemplary embodiment. Graphics acceleration module 946 may include one or more GPU chips integrated on a line card that is coupled to processor 907 via high-speed link 940. Optionally, graphics acceleration module 946 may be integrated with processor 907 on the same package or chip.
[0201] In at least one embodiment, the processor 907 shown includes multiple cores 960A-960D, each core having a translation lookaside buffer 961A-961D and one or more caches 962A-962D. In at least one embodiment, the cores 960A-960D may include various other components not shown for executing instructions and processing data. The caches 962A-962D may include level 1 (L1) and level 2 (L2) caches. In addition, one or more shared caches 956 may be included in the caches 962A-962D and shared by each group of cores 960A-960D. In at least one embodiment, one embodiment of the processor 907 includes 24 cores, each core having its own L1 cache, twelve shared L2 caches, and twelve shared L3 caches. In this embodiment, two adjacent cores share one or more L2 and L3 caches. The processor 907 and the graphics acceleration module 946 are connected to the system memory 914, which may include Fig. 9A Processor memory 901-902 in.
[0202] Coherence is maintained for data and instructions stored in the various caches 962A-962D, 956 and system memory 914 via inter-core communications over the coherence bus 964. In at least one embodiment, each cache may have cache coherence logic / circuitry associated therewith to communicate over the coherence bus 964 in response to detecting a read or write to a particular cache line. In one implementation, a cache snooping protocol is implemented over the coherence bus 964 to snoop cache accesses.
[0203] In at least one embodiment, the proxy circuit 925 communicatively couples the graphics acceleration module 946 to the coherence bus 964, thereby allowing the graphics acceleration module 946 to participate in a cache coherence protocol as a peer of the cores 960A-960D. In particular, in at least one embodiment, the interface 935 provides a connection to the proxy circuit 925 via a high-speed link 940 (e.g., a PCIe bus, NVLink, etc.), and the interface 937 connects the graphics acceleration module 946 to the link 940.
[0204] In one implementation, the accelerator integrated circuit 936 provides cache management, memory access, context management, and interrupt management services on behalf of multiple graphics processing engines 931, 932, N of the graphics acceleration module. The graphics processing engines 931, 932, N may each include a separate graphics processing unit (GPU). In at least one embodiment, the graphics processing engines 931, 932, N may selectively include different types of graphics processing engines within a GPU, such as a graphics execution unit, a media processing engine (e.g., a video encoder / decoder), a sampler, and a blit engine. In at least one embodiment, the graphics acceleration module 946 may be a GPU having multiple graphics processing engines 931-932, N, or the graphics processing engines 931-932, N may be individual GPUs integrated on a common package, line card, or chip. As appropriate, the graphics processing engines 931-932, N may be integrated into a common package, line card, or chip. Fig. 9B The above determination of the reconstruction parameters and the reconstruction algorithm is performed in the GPU 931-N.
[0205] In one embodiment, the accelerator integrated circuit 936 includes a memory management unit (MMU) 939 for performing various memory management functions, such as virtual to physical memory translation (also known as effective to real memory translation), and also includes a memory access protocol for accessing the system memory 914. The MMU 939 may also include a translation lookaside buffer ("TLB") (not shown) for caching virtual / effective to physical / real address translations. In one implementation, the cache 938 may store commands and data for efficient access by the graphics processing engine 931-932, N. In at least one embodiment, the data stored in the cache 938 and the graphics memory 933-934, M may be kept consistent with the core caches 962A-962D, 956 and the system memory 914. As before, this task may be accomplished via proxy circuitry 925 acting on behalf of cache 938 and graphics memory 933-934, M (e.g., sending updates related to modifications / accesses of cache lines on processor caches 962A-962D, 956 to cache 938 and receiving updates from cache 938).
[0206] A set of registers 945 stores context data for threads executed by the graphics processing engines 931-932, N, and context management circuitry 948 manages thread contexts. In at least one embodiment, context management circuitry 948 may perform save and restore operations to save and restore contexts for various threads during context switches (e.g., where a first thread is saved and a second thread is stored so that the second thread can be executed by the graphics processing engine). In at least one embodiment, context management circuitry 948 may store current register values to a specified area in memory (e.g., identified by a context pointer) upon context switching. The register values may then be restored when returning to context. In one embodiment, interrupt management circuitry 947 receives and processes interrupts received from system devices.
[0207] In one implementation, the MMU 939 converts virtual / effective addresses from the graphics processing engine 931 into real / physical addresses in the system memory 914. One embodiment of the accelerator integrated circuit 936 supports multiple (e.g., 4, 8, 16) graphics accelerator modules 946 and / or other accelerator devices. The graphics accelerator module 946 can be dedicated to a single application executed on the processor 907, or can be shared between multiple applications. In one embodiment, a virtualized graphics execution environment is presented in which the resources of the graphics processing engines 931-932, N are shared with multiple applications or virtual machines (VMs). In at least one embodiment, resources can be subdivided into "slices" that are allocated to different VMs and / or applications based on processing requirements and priorities associated with VMs and / or applications.
[0208] In at least one embodiment, the accelerator integrated circuit 936 performs as a bridge to the system of graphics acceleration modules 946 and provides address translation and system memory cache services. In addition, the accelerator integrated circuit 936 can provide virtualization facilities for the host processor to manage virtualization, interrupts, and memory management of the graphics processing engines 931-932, N.
[0209] Since the hardware resources of the graphics processing engines 931-932, N are explicitly mapped to the real address space seen by the host processor 907, any host processor can directly address these resources using effective address values. In at least one embodiment, one function of the accelerator integrated circuit 936 is to physically separate the graphics processing engines 931-932, N so that they appear to the system as independent units.
[0210] In at least one embodiment, one or more graphics memories 933-934, M are respectively coupled to each graphics processing engine 931-932, N. The graphics memories 933-934, M store instructions and data, which are processed by each graphics processing engine 931-932, N. The graphics memories 933-934, M may be volatile memories, such as DRAM (including stacked DRAM), GDDR memories (e.g., GDDR5, GDDR6), or HBM, and / or may be non-volatile memories, such as 3DXPoint or Nano-Ram.
[0211] In one embodiment, to reduce data traffic on link 940, biasing techniques are used to ensure that data stored in graphics memory 933-934, M is data that is most frequently used by graphics processing engines 931-932, N, and that may not be used (at least not frequently) by cores 960A-960D. Similarly, the biasing mechanism attempts to keep data that is needed by a core (and may not be graphics processing engine 931-932, N) in cache 962A-962D, core 956, and system memory 914.
[0212] Fig. 9C Another exemplary embodiment is shown, in which an accelerator integrated circuit 936 is integrated into the processor 907 for enabling and / or supporting intelligent control of a multi-mode cooling subsystem with a reusable refrigerant cooling subsystem according to at least one embodiment disclosed herein. In at least this embodiment, the graphics processing engines 931-932, N communicate directly with the accelerator integrated circuit 936 via the interface 937 and the interface 935 (again, any form of bus or interface protocol can be used) through the high-speed link 940. The accelerator integrated circuit 936 can perform related Fig. 9B The operations described above are similar to the operations described above. However, due to its close proximity to the coherence bus 964 and caches 962A-962D, 956, it may have a higher throughput. At least one embodiment supports different programming models, including a dedicated process programming model (without graphics acceleration module virtualization) and a shared programming model (with virtualization), and the programming models may include a programming model controlled by the accelerator integrated circuit 936 and a programming model controlled by the graphics acceleration module 946.
[0213] In at least one embodiment, the graphics processing engines 931-932, N are dedicated to a single application or process under a single operating system. In at least one embodiment, a single application can funnel other application requests to the graphics processing engines 931-932, N, thereby providing virtualization within a VM / partition.
[0214] In at least one embodiment, the graphics processing engines 931-932, N can be shared by multiple VM / application partitions. In at least one embodiment, the sharing model can use a system hypervisor to virtualize the graphics processing engines 931-932, N to allow each operating system to access. For a single partition system without a hypervisor, the operating system owns the graphics processing engines 931-932, N. In at least one embodiment, the operating system can virtualize the graphics processing engines 931-932, N to provide access to each process or application.
[0215] In at least one embodiment, the graphics acceleration module 946 or individual graphics processing engines 931-932, N use a process handle to select a process element. In at least one embodiment, the process element is stored in the system memory 914 and can be addressed using the effective address to real address conversion technology of this article. In at least one embodiment, the process handle can be an implementation-specific value that is provided to the host process when registering its context with the graphics processing engine 931-932, N (i.e., calling system software to add the process element to the process element linked list). In at least one embodiment, the lower 16 bits of the process handle can be the offset of the process element in the process element linked list.
[0216] Fig.9D An exemplary accelerator integrated slice 990 for implementing and / or supporting intelligent control of a multi-mode cooling subsystem with a reusable refrigerant cooling subsystem according to at least one embodiment disclosed herein is shown. As used herein, a "slice" includes a specified portion of the processing resources of an accelerator integrated circuit 936. The application is an effective address space 982 in the system memory 914, which stores a process element 983. In at least one embodiment, the process element 983 is stored in response to a GPU call 981 from an application 980 executed on the processor 907. The process element 983 contains the process state of the corresponding application 980. The work descriptor (WD) 984 contained in the process element 983 can be a single job requested by the application, or can contain a pointer to a job queue. In at least one embodiment, the WD 984 is a pointer to a job request queue in the address space 982 of the application.
[0217] Graphics acceleration module 946 and / or individual graphics processing engines 931-932, N can be shared by all processes or a subset of processes in the system. In at least one embodiment, infrastructure for setting process state and sending WD 984 to graphics acceleration module 946 to start a job in a virtualized environment can be included.
[0218] In at least one embodiment, the dedicated process programming model is implementation specific. In this model, a single process owns a graphics acceleration module 946 or an individual graphics processing engine 931. When the graphics acceleration module 946 is owned by a single process, the hypervisor initializes the accelerator integrated circuit for the owned partition, and when the graphics acceleration module 946 is assigned, the operating system initializes the accelerator integrated circuit 936 for the owned process.
[0219] In operation, the WD acquisition unit 991 in the accelerator integrated slice 990 acquires the next WD 984, which includes an indication of the work to be completed by one or more graphics processing engines of the graphics acceleration module 946. Data from the WD 984 can be stored in registers 945 and used by the MMU 939, interrupt management circuits 947, and / or context management circuits 948, as shown. In at least one embodiment, an embodiment of the MMU 939 includes a segment / page roaming circuit for accessing a segment / page table 986 within the OS virtual address space 985. The interrupt management circuit 947 can handle interrupt events 992 received from the graphics acceleration module 946. In at least one embodiment, when performing graphics operations, the effective address 993 generated by the graphics processing engine 931-932, N is converted to a real address by the MMU 939.
[0220] In one embodiment, the same register set 945 is replicated for each graphics processing engine 931-932, N and / or graphics acceleration module 946, and the same register set 945 can be initialized by a hypervisor or operating system. Each of these replicated registers can be included in an accelerator integrated slice 990. Exemplary registers that can be initialized by a hypervisor are shown in Table 1.
[0221] Table 1 – Hypervisor Initialization Registers
[0222] 1 Slice Control Register 2 Real address (RA) dispatch process area pointer 3 Permission Mask Override Register 4 Interrupt vector table entry offset 5 Interrupt Vector Table Entry Limits 6 Status Register 7 Logical partition ID 8 Real Address (RA) Hypervisor Accelerator Utilizes Record Pointers 9 Storage Description Register
[0223] Example registers that may be initialized by the operating system are shown in Table 2.
[0224] Table 2 – Operating System Initialization Registers
[0225] 1 Process and thread identification 2 Effective Address (EA) context save / restore pointer 3 Virtual Address (VA) Accelerator Utilizes Record Pointers 4 Virtual Address (VA) Segment Table Pointer 5 Permission shielding 6 Job Descriptor
[0226] In at least one embodiment, each WD 984 is specific to a particular graphics acceleration module 946 and / or graphics processing engine 931-932, N. It contains all the information needed by the graphics processing engine 931-932, N to complete the work, or it can be a pointer to a memory location where the application has set up a command queue for the work to be done.
[0227] Fig.9E Additional details of an exemplary embodiment of a sharing model are shown. This embodiment includes a hypervisor real address space 998 in which a process element list 999 is stored. The hypervisor real address space 998 can be accessed via a hypervisor 996, which virtualizes a graphics acceleration module engine for an operating system 995.
[0228] In at least one embodiment, the shared programming model allows all processes or subsets of processes from all partitions or subsets of partitions in the system to use the graphics acceleration module 946. There are two programming models where the graphics acceleration module 946 is shared by multiple processes and partitions, namely, time-sliced sharing and graphics-directed sharing.
[0229] In this model, the hypervisor 996 owns the graphics acceleration module 946 and makes its functionality available to all operating systems 995. For the graphics acceleration module 946 to support virtualization through the hypervisor 996, the graphics acceleration module 946 may comply with the following requirements: (1) the application's job request must be autonomous (i.e., no state needs to be maintained between jobs), or the graphics acceleration module 946 must provide a context save and restore mechanism, (2) the graphics acceleration module 946 guarantees that the application's job request is completed within a specified amount of time, including any transition errors, or the graphics acceleration module 946 provides the ability to preempt job processing, and (3) fairness between graphics acceleration module 946 processes must be ensured when operating in a directed shared programming model.
[0230] In one embodiment, the application 980 is required to use the graphics acceleration module type, work descriptor (WD), permission mask register (AMR) value and context save / restore region pointer (CSRP) to make an operating system 995 system call. The graphics acceleration module type describes the target acceleration function for the system call. In at least one embodiment, the graphics acceleration module type can be a system-specific value. In at least one embodiment, the WD is formatted specifically for the graphics acceleration module 946 and can take the form of a graphics acceleration module 946 command, an effective address pointer to a user-defined structure, an effective address pointer to a command queue, or any other data structure describing the work to be completed by the graphics acceleration module 946. In at least one embodiment, the AMR value is the AMR state for the current process. In at least one embodiment, the value passed to the operating system is similar to the application that sets the AMR. If the implementation of the accelerator integrated circuit 936 and the graphics acceleration module 946 does not support the user permission mask override register (UAMOR), the operating system can apply the current UAMOR value to the AMR value before passing the AMR in the hypervisor call. In at least one embodiment, the hypervisor 996 may apply the current privilege mask overwrite register (AMOR) value before placing the AMR into the process element 983. In at least one embodiment, the CSRP is one of the registers 945 that contains the effective address of an area in the application's effective address space 982 for the graphics acceleration module 946 to save and restore context state. This pointer is used in at least one embodiment and is optional if state does not need to be saved between jobs or when a job is preempted. In at least one embodiment, the context save / restore area may be fixed system memory.
[0231] Upon receiving the system call, the operating system 995 may verify that the application 980 has been registered and granted permission to use the graphics acceleration module 946. The operating system 995 then calls the hypervisor 996 using the information shown in Table 3.
[0232] Table 3 – OS to hypervisor call parameters
[0233] 1 Work Descriptor (WD) 2 Access Mask Register (AMR) value (potentially masked) 3 Effective Address (EA) Context Save / Restore Region Pointer (CSRP) 4 Process ID (PID) and optional thread ID (TID) 5 Virtual Address (VA) Accelerator Usage Record Pointer (AURP) 6 Virtual address of the storage segment table pointer (SSTP) 7 Logical Interrupt Service Number (LISN)
[0234] Upon receiving the hypervisor call, the hypervisor 996 verifies that the operating system 995 has registered and is granted permission to use the graphics acceleration module 946. The hypervisor 996 then places the process element 983 into a linked list of process elements of the corresponding graphics acceleration module 946 type. The process element may include the information shown in Table 4.
[0235] Table 4 - Process element information
[0236] 1 Work Descriptor (WD) 2 The authority mask register (AMR) value (potentially masked). 3 Effective Address (EA) Context Save / Restore Region Pointer (CSRP) 4 Process ID (PID) and optional thread ID (TID) 5 Virtual Address (VA) Accelerator Usage Record Pointer (AURP) 6 Virtual address of the storage segment table pointer (SSTP) 7 Logical Interrupt Service Number (LISN) 8 Interrupt vector table, derived from the hypervisor call parameters 9 Status Register (SR) Value 10 Logical Partition ID (LPID) 11 Real Address (RA) Hypervisor Accelerator Using Record Pointers 12 Storage Descriptor Register (SDR)
[0237] In at least one embodiment, the hypervisor initializes the plurality of accelerator integrated slice 990 registers 945 .
[0238] like Fig.9F As shown, in at least one embodiment, a unified memory is used, and the unified memory can be addressed via a common virtual memory address space for accessing physical processor memories 901-902 and GPU memories 920-923. In this implementation, operations executed on GPUs 910-913 use the same virtual / effective memory address space to access processor memories 901-902, and vice versa, thereby simplifying programmability. In one embodiment, the first part of the virtual / effective address space is allocated to processor memory 901, the second part is allocated to the second processor memory 902, the third part is allocated to GPU memory 920, and so on. In at least one embodiment, the entire virtual / effective memory space (sometimes referred to as the effective address space) is thus distributed in each of processor memories 901-902 and GPU memories 920-923, thereby allowing any processor or GPU to access the memory using a virtual address mapped to any physical memory.
[0239] In one embodiment, bias / coherency management circuits 994A-994E within one or more MMUs 939A-939E ensure cache coherency between caches of one or more host processors (e.g., 905) and GPUs 910-913 and implement biasing techniques that indicate physical memory where certain types of data should be stored. Fig.9F Multiple instances of bias / consistency management circuits 994A- 994E are shown in , but bias / consistency circuits may be implemented within an MMU of one or more host processors 905 and / or within an accelerator integrated circuit 936 .
[0240] One embodiment allows the GPU additional memory 920-923 to be mapped as part of the system memory and accessed using shared virtual memory (SVM) technology, but without suffering from the performance defects associated with full system cache coherence. In at least one embodiment, the ability to access the GPU additional memory 920-923 as system memory without heavy cache coherence overhead provides a favorable operating environment for GPU offloading. This arrangement allows the host processor 905 software to set operands and access calculation results without the overhead of traditional I / O DMA data copying. Such traditional copies include driver calls, interrupts, and memory mapped I / O (MMIO) accesses, which are all less efficient than simple memory accesses. In at least one embodiment, the ability to access the GPU additional memory 920-923 without cache coherence overhead may be critical to the execution time of the offloaded calculation. For example, in the case of a large amount of streaming write memory traffic, the cache coherence overhead can significantly reduce the effective write bandwidth seen by the GPU 910-913. In at least one embodiment, the efficiency of operand setting, the efficiency of result access, and the efficiency of GPU calculation may play a role in determining the effectiveness of GPU offloading.
[0241] In at least one embodiment, the selection of GPU bias and host processor bias is driven by a bias tracker data structure. For example, a bias table can be used, which can be a page granular structure (e.g., controlled at the granularity of a memory page) that includes 1 or 2 bits per GPU attached memory page. In at least one embodiment, the bias table can be implemented in the stolen memory range of one or more GPU attached memories 920-923 with or without a bias cache in the GPU 910-913 (e.g., to cache frequently / recently used entries of the bias table). Alternatively, the entire bias table can be maintained within the GPU.
[0242] In at least one embodiment, before actually accessing the GPU memory, the bias table entry associated with each access to the GPU additional memory 920-923 is accessed, resulting in the following operations. Local requests from GPUs 910-913 that find their pages in the GPU bias are forwarded directly to the corresponding GPU memory 920-923. Local requests from GPUs that find their pages in the host bias are forwarded to processors 905 (e.g., through high-speed links as above). In one embodiment, the request from processor 905 to find the requested page in the host processor bias completes a request similar to a normal memory read. Alternatively, a request pointing to a GPU bias page can be forwarded to GPUs 910-913. In at least one embodiment, if the GPU is not currently using the page, the GPU can then migrate the page to the host processor bias. In at least one embodiment, the bias state of the page can be changed by a software-based mechanism, a hardware-assisted software-based mechanism, or in limited cases by a purely hardware-based mechanism.
[0243] One mechanism for changing the bias state employs an API call (e.g., OpenCL) that in turn calls the GPU's device driver, which in turn sends a message (or queues a command descriptor) to the GPU, directing the GPU to change the bias state and, in some migrations, performs a cache flush operation in the host. In at least one embodiment, the cache flush operation is used for migrations from host processor 905 bias to GPU bias, but not for the reverse migration.
[0244] In one embodiment, cache coherency is maintained by temporarily rendering GPU biased pages that cannot be cached by the host processor 905. To access these pages, the processor 905 may request access from the GPU 910, which may or may not immediately grant access. Therefore, in order to reduce communication between the processor 905 and the GPU 910, it is beneficial to ensure that the GPU biased pages are the pages required by the GPU and not the pages required by the host processor 905, and vice versa.
[0245] Reasoning and / or training logic 615 is used to perform one or more embodiments. Figure 6B and / or Figure 6C Details regarding the inference and / or training logic 615 are provided.
[0246] Fig. 10AAn exemplary integrated circuit and associated graphics processor according to various embodiments of the present invention are shown, which can be manufactured using one or more IP cores to support and / or implement a multi-mode cooling subsystem with a reusable refrigerant cooling subsystem described herein. In addition to the illustrations, other logic and circuits may be included in at least one embodiment, including additional graphics processors / cores, peripheral interface controllers, or general processor cores.
[0247] Fig. 10A 1 is a block diagram illustrating an exemplary system on a chip integrated circuit 1000A that may be manufactured using one or more IP cores according to at least one embodiment. In at least one embodiment, the integrated circuit 1000A includes one or more application processors 1005 (e.g., CPU), at least one graphics processor 1010, and may additionally include an image processor 1015 and / or a video processor 1020, any of which may be a modular IP core. In at least one embodiment, the integrated circuit 1000A includes peripheral or bus logic, which includes a USB controller 1025, a UART controller 1030, an SPI / SDIO controller 1035, and an I 2 S / I 2 C controller 1040. In at least one embodiment, the integrated circuit 1000A may include a display device 1045 coupled to one or more of a high-definition multimedia interface (HDMI) controller 1050 and a mobile industry processor interface (MIPI) display interface 1055. In at least one embodiment, storage may be provided by a flash memory subsystem 1060, including flash memory and a flash memory controller. In at least one embodiment, a memory interface may be provided via a memory controller 1065 for accessing SDRAM or SRAM memory devices. In at least one embodiment, some integrated circuits also include an embedded security engine 1070.
[0248] The reasoning and / or training logic 615 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 6B and / or Figure 6C Details are provided regarding the inference and / or training logic 615. In at least one embodiment, the inference and / or training logic 615 may be used in the integrated circuit 1000A to infer or predict operations based at least in part on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0249] Figures 10B-10CAn exemplary integrated circuit and associated graphics processor according to various embodiments of the present invention are shown, which can be manufactured using one or more IP cores to support and / or implement the multi-mode cooling subsystem with a reusable refrigerant cooling subsystem described herein. In addition to the illustrations, other logic and circuits may be included in at least one embodiment, including additional graphics processors / cores, peripheral interface controllers, or general processor cores.
[0250] Figures 10B-10C is a block diagram illustrating an exemplary graphics processor used within a SoC according to embodiments described herein to support and / or implement the multi-mode cooling subsystem with a reusable refrigerant cooling subsystem described herein. In one example, the graphics processor can be used in the intelligent control of the multi-mode cooling subsystem with a reusable refrigerant cooling subsystem because existing math engines can process multi-level neural networks faster. Fig. 10B An exemplary graphics processor 1010 of a system on a chip integrated circuit is shown, which may be manufactured using one or more IP cores, in accordance with at least one embodiment. Fig. 10C Another exemplary graphics processor 1040 of a system on a chip integrated circuit is shown, which can be manufactured using one or more IP cores according to at least one embodiment. In at least one embodiment, Fig. 10B The graphics processor 1010 is a low power graphics processor core. In at least one embodiment, Fig. 10C The graphics processor 1040 is a higher performance graphics processor core. In at least one embodiment, each graphics processor 1010, 1040 can be Fig. 10A A variant of graphics processor 1010.
[0251] In at least one embodiment, the graphics processor 1010 includes a vertex processor 1005 and one or more fragment processors 1015A-1015N (e.g., 1015A, 1015B, 1015C, 1015D to 1015N-1 and 1015N). In at least one embodiment, the graphics processor 1010 can execute different shader programs via separate logic, so that the vertex processor 1005 is optimized to perform operations for the vertex shader program, while one or more fragment processors 1015A-1015N perform fragment (e.g., pixel) shading operations for fragments or pixels or shader programs. In at least one embodiment, the vertex processor 1005 performs the vertex processing stage of the 3D graphics pipeline and generates primitives and vertex data. In at least one embodiment, the one or more fragment processors 1015A-1015N use the primitives and vertex data generated by the vertex processor 1005 to generate a frame buffer displayed on a display device. In at least one embodiment, one or more fragment processors 1015A-1015N are optimized to execute fragment shader programs as provided in the OpenGL API, which can be used to perform similar operations as pixel shader programs provided in the Direct 3D API.
[0252] In at least one embodiment, graphics processor 1010 additionally includes one or more memory management units (MMUs) 1020A-1020B, one or more caches 1025A-1025B, and one or more circuit interconnects 1030A-1030B. In at least one embodiment, one or more MMUs 1020A-1020B provide a mapping of virtual to physical addresses for graphics processor 1010, including for vertex processor 1005 and / or fragment processors 1015A-1015N, which may reference vertex or image / texture data stored in memory in addition to vertex or image / texture data stored in one or more caches 1025A-1025B. In at least one embodiment, one or more MMUs 1020A-1020B may synchronize with other MMUs within the system, including with Fig. 10A One or more MMUs associated with one or more application processors 1005, image processor 1015, and / or video processor 1020 enable each processor 1005-1020 to participate in a shared or unified virtual memory system. In at least one embodiment, one or more circuit interconnects 1030A-1030B enable graphics processor 1010 to connect to other IP cores within the SoC via an internal bus of the SoC or via a direct connection.
[0253] In at least one embodiment, graphics processor 1040 includes Fig. 10AOne or more MMUs 1020A-1020B, caches 1025A-1025B, and circuit interconnects 1030A-1030B of graphics processor 1010. In at least one embodiment, graphics processor 1040 includes one or more shader cores 1055A-1055N (e.g., 1055A, 1055B, 1055C, 1055D, 1055E, 1055F to 1055N-1 and 1055N), such as Fig. 10B As shown, it provides a unified shader core architecture in which a single core or type or core can execute all types of programmable shader code, including shader program code for implementing vertex shaders, fragment shaders, and / or compute shaders. In at least one embodiment, multiple shader cores can vary. In at least one embodiment, the graphics processor 1040 includes an inter-core task manager 1045 that acts as a thread dispatcher to dispatch execution threads to one or more shader cores 1055A-1055N and a blocking unit 1058 to accelerate tile-based rendering operations, in which rendering operations of a scene are subdivided in image space, for example, to exploit local spatial consistency within a scene or to optimize the use of internal caches.
[0254] Reasoning and / or training logic 615 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 6B and / or Figure 6C Provides details about the inference and / or training logic 615. In at least one embodiment, the inference and / or training logic 615 may be implemented in an integrated circuit. Fig. 10A and / or Fig. 10B for performing inference or prediction operations based at least in part on weight parameters computed using a neural network training operation, a neural network function or architecture, or a neural network use case herein.
[0255] Figures 10D-10E Additional exemplary graphics processor logic is shown according to embodiments described herein to support and / or implement the multi-mode cooling subsystem with reusable refrigerant cooling subsystem described herein. In at least one embodiment, Fig. 10D shows that it can be included in Fig. 10A The graphics core 1000D within the graphics processor 1010 of FIG. 1000B may be a graphics core 1000D of FIG. 1000D. Fig. 10C Unified shader cores 1055A-1055N are shown. Fig. 10B A highly parallel general purpose graphics processing unit ("GPGPU") 1030 suitable for deployment on a multi-chip module in at least one embodiment is shown.
[0256] In at least one embodiment, graphics core 1000D may include multiple slices 1001A-1001N or partitions of each core, and a graphics processor may include multiple instances of graphics core 1000D. In at least one embodiment, slices 1001A-1001N may include support logic including local instruction caches 1004A-1004N, thread schedulers 1006A-1006N, thread dispatchers 1008A-1008N, and a set of registers 1010A-1010N. In at least one embodiment, slices 1001A-1001N may include a set of additional function units (AFUs 1012A-1012N), floating point units (FPUs 1014A-1014N), integer arithmetic logic units (ALUs 109A-109N), address calculation units (ACUs 1013A-1013N), double precision floating point units (DPFPUs 1015A-1015N), and matrix processing units (MPUs 1017A-1017N).
[0257] In at least one embodiment, the FPU 1014A-1014N can perform single-precision (32-bit) and half-precision (16-bit) floating-point operations, while the DPFPU 1015A-1015N performs double-precision (64-bit) floating-point operations. In at least one embodiment, the ALU 1016A-1016N can perform variable-precision integer operations with 8-bit, 16-bit, and 32-bit precision, and can be configured for mixed-precision operations. In at least one embodiment, the MPU 1017A-1017N can also be configured for mixed-precision matrix operations, including half-precision floating-point operations and 8-bit integer operations. In at least one embodiment, the MPU 1017A-1010N can perform various matrix operations to accelerate machine learning application frameworks, including enabling general matrix-to-matrix multiplication (GEMM) to support acceleration. In at least one embodiment, the AFU 1012A-1012N can perform additional logical operations that are not supported by floating-point or integer units, including trigonometric operations (e.g., sine, cosine, etc.).
[0258] As discussed elsewhere in this disclosure, the reasoning and / or training logic 615 (at least in Figure 6B , Figure 6C Reference) is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 6B and / or Figure 6C Details are provided regarding inference and / or training logic 615. In at least one embodiment, inference and / or training logic 615 may be used in graphics core 1000D to infer or predict operations based at least in part on weight parameters computed using neural network training operations, neural network functions, and / or architectures or neural network use cases herein.
[0259] Fig.11A A block diagram of a computer system 1100A according to at least one embodiment is shown. In at least one embodiment, the computer system 1100A includes a processing subsystem 1101 having one or more processors 1102 and a system memory 1104, which communicates via an interconnect path that may include a memory hub 1105. In at least one embodiment, the memory hub 1105 may be a separate component within a chipset component, or may be integrated within one or more processors 1102. In at least one embodiment, the memory hub 1105 is coupled to an I / O subsystem 1111 via a communication link 1106. In one embodiment, the I / O subsystem 1111 includes an I / O hub 1107, which may enable the computer system 1100A to receive input from one or more input devices 1108. In at least one embodiment, the I / O hub 1107 may enable a display controller, which may be included in one or more processors 1102, to provide output to the one or more display devices 1110A. In at least one embodiment, the one or more display devices 1110A coupled to the I / O hub 1107 may include local, internal, or embedded display devices.
[0260] In at least one embodiment, the processing subsystem 1101 includes one or more parallel processors 1112 coupled to the memory hub 1105 via a bus or other communication link 1113. In at least one embodiment, the communication link 1113 can use any of a number of standard-based communication link technologies or protocols, such as but not limited to PCI Express, or can be a vendor-specific communication interface or communication structure. In at least one embodiment, the one or more parallel processors 1112 form a parallel or vector processing system in a computing concentration, and the system can include a large number of processing cores and / or processing clusters, such as a multi-integrated core (MIC) processor. In at least one embodiment, the one or more parallel processors 1112 form a graphics processing subsystem, which can output pixels to one of the one or more display devices 1110A coupled via the I / O hub 1107. In at least one embodiment, the one or more parallel processors 1112 can also include a display controller and a display interface (not shown) to enable direct connection to one or more display devices 1110B.
[0261] In at least one embodiment, a system storage unit 1114 can be connected to the I / O hub 1107 to provide a storage mechanism for the computer system 1100A. In at least one embodiment, an I / O switch 1116 can be used to provide an interface mechanism to enable connections between the I / O hub 1107 and other components, such as a network adapter 1118 and / or a wireless network adapter 1119 that can be integrated into one or more platforms, as well as various other devices that can be added through one or more additional devices 1120. In at least one embodiment, the network adapter 1118 can be an Ethernet adapter or another wired network adapter. In at least one embodiment, the wireless network adapter 1119 can include one or more of Wi-Fi, Bluetooth, near field communication (NFC), or other network devices including one or more radio devices.
[0262] In at least one embodiment, the computer system 1100A may include other components not explicitly shown, including USB or other port connections, optical storage drives, video capture devices, etc. Other components may also be connected to the I / O hub 1107. In at least one embodiment, the interconnection may be implemented using any suitable protocol, such as a PCI (Peripheral Component Interconnect) based protocol (e.g., PCI-Express) or other bus or point-to-point communication interface and / or protocol. Fig.11A The communication paths between the various components in the NV-Link high-speed interconnect or interconnect protocol.
[0263] In at least one embodiment, one or more parallel processors 1112 include circuits optimized for graphics and video processing, including, for example, video output circuits, and constitute a graphics processing unit (GPU). In at least one embodiment, one or more parallel processors 1112 include circuits optimized for general processing. In at least one embodiment, the components of the computer system 1100A can be integrated with one or more other system elements on a single integrated circuit. For example, in at least one embodiment, one or more parallel processors 1112, memory hub 1105, one or more processors 1102, and I / O hub 1107 can be integrated into a system-on-chip (SoC) integrated circuit. In at least one embodiment, the components of the computer system 1100A can be integrated into a single package to form a system-in-package (SIP) configuration. In at least one embodiment, at least a portion of the components of the computer system 1100A can be integrated into a multi-chip module (MCM), and the multi-chip module can be interconnected with other multi-chip modules into a modular computer system.
[0264] Reasoning and / or training logic 615 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 6B and / or Figure 6C Provide details about the reasoning and / or training logic 615. In at least one embodiment, the reasoning and / or training logic 615 may be Fig.11A for use in a system for performing inference or prediction operations based at least in part on weight parameters computed using a neural network training operation, a neural network function and / or architecture, or a neural network use case herein.
[0265] processor
[0266] Fig. 11B 100B according to at least one embodiment. In at least one embodiment, the various components of the parallel processor 1100B may be implemented using one or more integrated circuit devices, such as a programmable processor, an application specific integrated circuit (ASIC), or a field programmable gate array (FPGA). In at least one embodiment, the parallel processor 1100B shown is a processor according to an exemplary embodiment. Fig. 11B A variation of the one or more parallel processors 1112 is shown.
[0267] In at least one embodiment, parallel processor 1100B includes parallel processing unit 1102. In at least one embodiment, parallel processing unit 1102 includes I / O unit 1104, which enables communication with other devices, including other instances of parallel processing unit 1102. In at least one embodiment, I / O unit 1104 can be directly connected to other devices. In at least one embodiment, I / O unit 1104 is connected to other devices by using a hub or switch interface (e.g., memory hub 1105). In at least one embodiment, the connection between memory hub 1105 and I / O unit 1104 forms communication link 1113. In at least one embodiment, I / O unit 1104 is connected to host interface 1106 and memory crossbar switch 1116, wherein host interface 1106 receives commands for performing processing operations and memory crossbar switch 1116 receives commands for performing memory operations.
[0268] In at least one embodiment, when the host interface 1106 receives the command buffer via the I / O unit 1104, the host interface 1106 can direct work operations to execute those commands to the front end 1108. In at least one embodiment, the front end 1108 is coupled with a scheduler 1110, which is configured to distribute commands or other work items to the processing cluster array 1112. In at least one embodiment, the scheduler 1110 ensures that the processing cluster array 1112 is properly configured and in a valid state before assigning tasks to the processing cluster array 1112. In at least one embodiment, the scheduler 1110 is implemented by firmware logic executed on a microcontroller. In at least one embodiment, the microcontroller-implemented scheduler 1110 can be configured to perform complex scheduling and work distribution operations at coarse and fine granularity, thereby achieving fast preemption and context switching of threads executed on the processing array 1112. In at least one embodiment, the host software can prove the workload for scheduling on the processing array 1112 through one of the multiple graphics processing doorbells. In at least one embodiment, the workload may then be automatically distributed across the processing array 1112 by scheduler 1110 logic within a microcontroller that includes scheduler 1110 .
[0269] In at least one embodiment, the processing cluster array 1112 may include up to "N" processing clusters (e.g., cluster 1114A, cluster 1114B, through cluster 1114N). In at least one embodiment, each cluster 1114A-1114N of the processing cluster array 1112 may execute a large number of concurrent threads. In at least one embodiment, the scheduler 1110 may allocate work to the clusters 1114A-1114N of the processing cluster array 1112 using various scheduling and / or work allocation algorithms, which may vary depending on the workload generated by each program or type of calculation. In at least one embodiment, scheduling may be handled dynamically by the scheduler 1110, or may be partially assisted by compiler logic during the compilation of program logic configured to be executed by the processing cluster array 1112. In at least one embodiment, different clusters 1114A-1114N of the processing cluster array 1112 may be allocated to process different types of programs or to perform different types of calculations.
[0270] In at least one embodiment, processing cluster array 1112 can be configured to perform various types of parallel processing operations. In at least one embodiment, processing cluster array 1112 is configured to perform general-purpose parallel computing operations. In at least one embodiment, processing cluster array 1112 can include logic to perform processing tasks including filtering video and / or audio data, performing modeling operations including physics operations, and performing data transformations.
[0271] In at least one embodiment, processing cluster array 1112 is configured to perform parallel graphics processing operations. In at least one embodiment, processing cluster array 1112 may include additional logic to support the execution of such graphics processing operations, including but not limited to texture sampling logic to perform texture operations, as well as tessellation logic and other vertex processing logic. In at least one embodiment, processing cluster array 1112 may be configured to execute shader programs related to graphics processing, such as but not limited to vertex shaders, tessellation shaders, geometry shaders, and pixel shaders. In at least one embodiment, parallel processing units 1102 may transfer data from system memory via I / O units 1104 for processing. In at least one embodiment, during processing, the transferred data may be stored to on-chip memory (e.g., parallel processor memory 1122) during processing and then written back to system memory.
[0272] In at least one embodiment, when the parallel processing unit 1102 is used to perform graphics processing, the scheduler 1110 can be configured to divide the processing workload into tasks of approximately equal size to better distribute graphics processing operations to multiple clusters 1114A-1114N of the processing cluster array 1112. In at least one embodiment, portions of the processing cluster array 1112 can be configured to perform different types of processing. In at least one embodiment, a first portion can be configured to perform vertex shading and topology generation, a second portion can be configured to perform tessellation and geometry shading, and a third portion can be configured to perform pixel shading or other screen space operations to generate a rendered image for display, if necessary to simulate value control of a multi-mode cooling subsystem with a reusable refrigerant cooling subsystem. In at least one embodiment, intermediate data generated by one or more of the clusters 1114A-1114N can be stored in a buffer to allow the intermediate data to be transferred between the clusters 1114A-1114N for further processing.
[0273] In at least one embodiment, the processing cluster array 1112 can receive processing tasks to be performed via the scheduler 1110, which receives commands defining the processing tasks from the front end 1108. In at least one embodiment, the processing tasks can include an index of data to be processed, such as surface (patch) data, primitive data, vertex data, and / or pixel data, as well as state parameters and commands defining how to process the data (e.g., what program to execute). In at least one embodiment, the scheduler 1110 can be configured to obtain an index corresponding to a task, or can receive the index from the front end 1108. In at least one embodiment, the front end 1108 can be configured to ensure that the processing cluster array 1112 is configured to a valid state before starting a workload specified by an incoming command buffer (e.g., a batch-buffer, a push buffer, etc.).
[0274] In at least one embodiment, each of the one or more instances of parallel processing unit 1102 can be coupled to parallel processor memory 1122. In at least one embodiment, parallel processor memory 1122 can be accessed via memory crossbar switch 1116, which can receive memory requests from processing cluster array 1112 and I / O unit 1104. In at least one embodiment, memory crossbar switch 1116 can access parallel processor memory 1122 via memory interface 1118. In at least one embodiment, memory interface 1118 can include multiple partition units (e.g., partition unit 1120A, partition unit 1120B to partition unit 1120N), which can each be coupled to a portion of parallel processor memory 1122 (e.g., memory units). In at least one embodiment, the plurality of partition units 1120A-1120N are configured to be equal to the number of memory cells, such that the first partition unit 1120A has a corresponding first memory cell 1124A, the second partition unit 1120B has a corresponding memory cell 1124B, and the Nth partition unit 1120N has a corresponding Nth memory cell 1124N. In at least one embodiment, the number of partition units 1120A-1120N may not be equal to the number of memory devices.
[0275] In at least one embodiment, memory units 1124A-1124N may include various types of memory devices, including dynamic random access memory (DRAM) or graphics random access memory, such as synchronous graphics random access memory (SGRAM), including graphics double data rate (GDDR) memory. In at least one embodiment, memory units 1124A-1124N may also include 3D stacked memory, including but not limited to high bandwidth memory (HBM). In at least one embodiment, rendering targets such as frame buffers or texture maps may be stored across memory units 1124A-1124N, allowing partition units 1120A-1120N to write portions of each rendering target in parallel to efficiently use the available bandwidth of parallel processor memory 1122. In at least one embodiment, local instances of parallel processor memory 1122 may be excluded to facilitate a unified memory design that utilizes system memory in combination with local cache memory.
[0276] In at least one embodiment, any of the clusters 1114A-1114N of the processing cluster array 1112 can process data to be written to any memory unit 1124A-1124N within the parallel processor memory 1122. In at least one embodiment, the memory crossbar 1116 can be configured to transmit the output of each cluster 1114A-1114N to any partition unit 1120A-1120N or another cluster 1114A-1114N, and the cluster 1114A-1114N can perform other processing operations on the output. In at least one embodiment, each cluster 1114A-1114N can communicate with the memory interface 1118 through the memory crossbar 1116 to read from or write to various external storage devices. In at least one embodiment, memory crossbar switch 1116 has connections to memory interface 1118 to communicate with I / O unit 1104, and connections to local instances of parallel processor memory 1102, thereby enabling processing units within different processing clusters 1114A-1114N to communicate with system memory or other memory that is not local to parallel processing unit 1102. In at least one embodiment, memory crossbar switch 1116 may use virtual channels to separate traffic flows between clusters 1114A-1114N and partition units 1120A-1120N.
[0277] In at least one embodiment, multiple instances of parallel processing unit 1102 may be provided on a single plug-in card, or multiple plug-in cards may be interconnected. In at least one embodiment, different instances of parallel processing unit 1102 may be configured to interoperate, even if different instances have different numbers of processing cores, different numbers of local parallel processor memories, and / or other configuration differences. In at least one embodiment, some instances of parallel processing unit 1102 may include higher precision floating point units relative to other instances. In at least one embodiment, a system incorporating one or more instances of parallel processing unit 1102 or parallel processor 1100B may be implemented in a variety of configurations and form factors, including but not limited to desktop, laptop or handheld personal computers, servers, workstations, game consoles, and / or embedded systems.
[0278] Fig. 11C is a block diagram of a partition unit 1120 according to at least one embodiment. In at least one embodiment, the partition unit 1120 is Fig. 11B 1120N. In at least one embodiment, partition unit 1120 includes L2 cache 1121, frame buffer interface 1125, and ROP 1126 (raster operation unit). L2 cache 1121 is a read / write cache that is configured to perform load and store operations received from memory crossbar switch 1116 and ROP 1126. In at least one embodiment, L2 cache 1121 outputs read misses and urgent write-back requests to frame buffer interface 1125 for processing. In at least one embodiment, updates can also be sent to the frame buffer via frame buffer interface 1125 for processing. In at least one embodiment, frame buffer interface 1125 communicates with memory units (such as ROPs) in parallel processor memory. Fig. 11B interacts with one of the memory units 1124A-1124N (e.g., within parallel processor memory 1122).
[0279] In at least one embodiment, ROP 1126 is a processing unit that performs raster operations such as stenciling, z-testing, blending, and the like. In at least one embodiment, ROP 1126 then outputs processed graphics data that is stored in graphics memory. In at least one embodiment, ROP 1126 includes compression logic to compress depth or color data written to memory and decompress depth or color data read from memory. In at least one embodiment, the compression logic can be lossless compression logic that utilizes one or more of a variety of compression algorithms. The compression logic performed by ROP 1126 can vary based on the statistical characteristics of the data to be compressed. In at least one embodiment, incremental color compression is performed based on the depth and color data on a per-tile basis.
[0280] In at least one embodiment, ROP 1126 is included within each processing cluster (e.g., Fig. 11B In at least one embodiment, read and write requests for pixel data are transmitted through memory crossbar 1116 rather than pixel fragment data transmission. In at least one embodiment, the processed graphics data can be displayed on a display device (such as Fig.11A 110), routed by processor 1102 for further processing, or by Fig. 11B One of the processing entities within parallel processor 1100B is routed for further processing.
[0281] Fig.11D is a block diagram of a processing cluster 1114 within a parallel processing unit according to at least one embodiment. In at least one embodiment, a processing cluster is Fig. 11B In at least one embodiment, one or more processing clusters 1114 can be configured to execute many threads in parallel, where a "thread" refers to an instance of a specific program executed on a specific set of input data. In at least one embodiment, single instruction multiple data (SIMD) instruction issuance technology is used to support the parallel execution of a large number of threads without providing multiple independent instruction units. In at least one embodiment, single instruction multiple thread (SIMT) technology is used to support the parallel execution of a large number of synchronized threads, which uses a common instruction unit that is configured to issue instructions to a set of processing engines within each processing cluster.
[0282] In at least one embodiment, the operation of the processing cluster 1114 can be controlled by a pipeline manager 1132 that allocates processing tasks to SIMT parallel processors. In at least one embodiment, the pipeline manager 1132 Fig. 11BThe scheduler 1110 receives instructions and manages the execution of these instructions through the graphics multiprocessor 1134 and / or the texture unit 1136. In at least one embodiment, the graphics multiprocessor 1134 is an exemplary instance of a SIMT parallel processor. However, in at least one embodiment, various types of SIMT parallel processors of different architectures may be included in the processing cluster 1114. In at least one embodiment, one or more instances of the graphics multiprocessor 1134 may be included in the processing cluster 1114. In at least one embodiment, the graphics multiprocessor 1134 may process data, and the data crossbar 1140 may be used to distribute the processed data to one of a plurality of possible destinations (including other shader units). In at least one embodiment, the pipeline manager 1132 may facilitate the distribution of processed data by specifying the destination of the processed data to be distributed via the data crossbar 1140.
[0283] In at least one embodiment, each graphics multiprocessor 1134 within a processing cluster 1114 may include the same set of function execution logic (e.g., arithmetic logic units, load store units, etc.). In at least one embodiment, the function execution logic may be configured in a pipelined manner, where new instructions may be issued before previous instructions are completed. In at least one embodiment, the function execution logic supports a variety of operations, including integer and floating point arithmetic, comparison operations, Boolean operations, shifts, and calculations of various algebraic functions. In at least one embodiment, the same functional unit hardware may be utilized to perform different operations, and any combination of functional units may be present.
[0284] In at least one embodiment, the instructions transmitted to the processing cluster 1114 constitute threads. In at least one embodiment, a group of threads executed across a group of parallel processing engines is a thread group. In at least one embodiment, the thread group executes the program on different input data. In at least one embodiment, each thread in the thread group can be assigned to a different processing engine in the graphics multiprocessor 1134. In at least one embodiment, the thread group may include fewer threads than the number of processing engines in the graphics multiprocessor 1134. In at least one embodiment, when the number of threads included in the thread group is less than the number of processing engines, one or more processing engines may be idle during the cycle of the thread group being processed. In at least one embodiment, the thread group may also include more threads than the number of processing engines in the graphics multiprocessor 1134. In at least one embodiment, when the thread group includes more threads than the processing engines in the graphics multiprocessor 1134, processing can be performed in consecutive clock cycles. In at least one embodiment, multiple thread groups can be executed simultaneously on the graphics multiprocessor 1134.
[0285] In at least one embodiment, graphics multiprocessor 1134 includes internal cache memory to perform load and store operations. In at least one embodiment, graphics multiprocessor 1134 can abandon the internal cache and use cache memory (e.g., L1 cache 1148) within processing cluster 1114. In at least one embodiment, each graphics multiprocessor 1134 can also access partition units (e.g., Fig. 11B 1120N) that are shared between all processing clusters 1114 and can be used to transfer data between threads. In at least one embodiment, graphics multiprocessor 1134 can also access off-chip global memory, which can include one or more of local parallel processor memory and / or system memory. In at least one embodiment, any memory external to parallel processing unit 1102 can be used as global memory. In at least one embodiment, processing cluster 1114 includes multiple instances of graphics multiprocessor 1134, which can share common instructions and data that can be stored in L1 cache 1148.
[0286] In at least one embodiment, each processing cluster 1114 may include a memory management unit ("MMU") 1145 configured to map virtual addresses to physical addresses. In at least one embodiment, one or more instances of MMU 1145 may reside in Fig. 11B 1148. In at least one embodiment, the MMU 1145 includes a set of page table entries (PTEs) that are used to map virtual addresses to physical addresses of tiles and, in at least one embodiment, to cache line indices. In at least one embodiment, the MMU 1145 may include an address translation lookaside buffer (TLB) or a cache that may reside within the graphics multiprocessor 1134 or the L1 cache 1148 or the processing cluster 1114. In at least one embodiment, the physical addresses are processed to assign surface data access locality for efficient request interleaving between partition units. In at least one embodiment, the cache line index may be used to determine whether a request for a cache line is a hit or a miss.
[0287] In at least one embodiment, the processing clusters 1114 can be configured such that each graphics multiprocessor 1134 is coupled to a texture unit 1136 to perform texture mapping operations, operations to determine texture sample locations, read texture data, and filter texture data. In at least one embodiment, texture data is read from an internal texture L1 cache (not shown) or from an L1 cache within the graphics multiprocessor 1134, and texture data is retrieved from an L2 cache, local parallel processor memory, or system memory as needed. In at least one embodiment, each graphics multiprocessor 1134 outputs one or more processed tasks to a data crossbar 1140 to provide the processed tasks to another processing cluster 1114 for further processing or to store the processed one or more tasks in an L2 cache, local parallel processor memory, or system memory via a memory crossbar 1116. In at least one embodiment, a preROP 1142 (pre-raster operations unit) is configured to receive data from the graphics multiprocessor 1134, direct the data to a ROP unit, which can communicate with a partition unit (e.g., Fig. 11B In at least one embodiment, the PreROP 1142 unit may perform optimizations for color blending, organize pixel color data, and perform address translation.
[0288] The reasoning and / or training logic 615 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 6B and / or Figure 6C Details are provided regarding inference and / or training logic 615. In at least one embodiment, inference and / or training logic 615 may be used in graphics processing cluster 1114 to perform inference or prediction operations based at least in part on weight parameters computed using neural network training operations, neural network functions, and / or architectures or neural network use cases described herein.
[0289] Fig.11DA graphics multiprocessor 1134 is shown in accordance with at least one embodiment. In at least one embodiment, the graphics multiprocessor 1134 is coupled to a pipeline manager 1132 of a processing cluster 1114. In at least one embodiment, the graphics multiprocessor 1134 has an execution pipeline that includes, but is not limited to, an instruction cache 1152, an instruction unit 1154, an address mapping unit 1156, a register file 1158, one or more general purpose graphics processing unit (GPGPU) cores 1162, and one or more load / store units 1166. The one or more GPGPU cores 1162 and the one or more load / store units 1166 are coupled to a cache memory 1172 and a shared memory 1170 via a memory and cache interconnect 1168.
[0290] In at least one embodiment, the instruction cache 1152 receives a stream of instructions to be executed from the pipeline manager 1132. In at least one embodiment, the instructions are cached in the instruction cache 1152 and dispatched for execution by the instruction unit 1154. In one embodiment, the instruction unit 1154 can dispatch instructions as thread groups (e.g., warps), each thread group being assigned to a different execution unit within one or more GPGPU cores 1162. In at least one embodiment, the instructions can access any local, shared, or global address space by specifying an address within the unified address space. In at least one embodiment, the address mapping unit 1156 can be used to convert addresses in the unified address space into different memory addresses that can be accessed by one or more load / store units 1166.
[0291] In at least one embodiment, register file 1158 provides a set of registers for the functional units of graphics multiprocessor 1134. In at least one embodiment, register file 1158 provides temporary storage for operands for data paths connected to the functional units (e.g., GPGPU core 1162, load / store unit 1166) of graphics multiprocessor 1134. In at least one embodiment, register file 1158 is divided between each functional unit such that a dedicated portion of register file 1158 is allocated to each functional unit. In at least one embodiment, register file 1158 is divided between different warps being executed by graphics multiprocessor 1134.
[0292] In at least one embodiment, the GPGPU cores 1162 may each include a floating point unit (FPU) and / or an integer arithmetic logic unit (ALU) for executing instructions of the graphics multiprocessor 1134. The GPGPU cores 1162 may be similar in architecture or the architecture may be different. In at least one embodiment, the first portion of the GPGPU core 1162 includes a single-precision FPU and an integer ALU, while the second portion of the GPGPU core includes a double-precision FPU. In at least one embodiment, the FPU may implement the IEEE 754-2008 standard for floating-point arithmetic or enable variable-precision floating-point arithmetic. In at least one embodiment, the graphics multiprocessor 1134 may additionally include one or more fixed-function or special-function units to perform specific functions, such as copying rectangles or pixel blending operations. In at least one embodiment, one or more of the GPGPU cores may also include fixed or special-function logic.
[0293] In at least one embodiment, the GPGPU core 1162 includes SIMD logic capable of executing a single instruction to multiple sets of data. In one embodiment, the GPGPU core 1162 can physically execute SIMD4, SIMD8 and SIMD16 instructions, and logically execute SIMD1, SIMD2 and SIMD32 instructions. In at least one embodiment, the SIMD instructions for the GPGPU core can be generated by a shader compiler at compile time, or automatically generated when executing a program written and compiled for a single program multiple data (SPMD) or SIMT architecture. In at least one embodiment, multiple threads of a program configured for a SIMT execution model can be executed by a single SIMD instruction. For example, in at least one embodiment, eight SIMT threads that perform the same or similar operations can be executed in parallel by a single SIMD8 logic unit.
[0294] In at least one embodiment, the memory and cache interconnect 1168 is an interconnect network that connects each functional unit of the graphics multiprocessor 1134 to the register file 1158 and the shared memory 1170. In at least one embodiment, the memory and cache interconnect 1168 is a crossbar interconnect that allows the load / store unit 1166 to implement load and store operations between the shared memory 1170 and the register file 1158. In at least one embodiment, the register file 1158 can operate at the same frequency as the GPGPU core 1162, so that the latency of data transfer between the GPGPU core 1162 and the register file 1158 is very low. In at least one embodiment, the shared memory 1170 can be used to enable communication between threads executing on the functional units within the graphics multiprocessor 1134. In at least one embodiment, the cache memory 1172 can be used as, for example, a data cache to cache texture data communicated between the functional units and the texture unit 1136. In at least one embodiment, the shared memory 1170 can also be used as a program-managed cache. In at least one embodiment, in addition to automatically cached data stored in cache memory 1172, threads executing on GPGPU core 1162 may programmatically store data in shared memory.
[0295] In at least one embodiment, a parallel processor or GPGPU as described herein is communicatively coupled to a host / processor core to accelerate graphics operations, machine learning operations, pattern analysis operations, and various general purpose GPU (GPGPU) functions. In at least one embodiment, the GPU can be communicatively coupled to the host processor / core via a bus or other interconnect (e.g., a high-speed interconnect such as PCIe or NVLink). In at least one embodiment, the GPU can be integrated with the core on the same package or chip and communicatively coupled to the core via an internal processor bus / interconnect (inside the package or chip, in at least one embodiment). In at least one embodiment, regardless of the manner in which the GPU is connected, the processor core can assign work to the GPU in the form of a sequence of commands / instructions contained in a work descriptor. In at least one embodiment, the GPU then uses dedicated circuits / logic to efficiently process these commands / instructions.
[0296] The reasoning and / or training logic 615 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 6B and / or Figure 6CDetails are provided regarding inference and / or training logic 615. In at least one embodiment, inference and / or training logic 615 may be used in graphics multiprocessor 1134 to perform inference or prediction operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0297] 12A illustrates a multi-GPU computing system 1200A according to at least one embodiment. In at least one embodiment, the multi-GPU computing system 1200A may include a processor 1202 coupled to a plurality of general purpose graphics processing units (GPGPUs) 1206A-D via a host interface switch 1204. In at least one embodiment, the host interface switch 1204 is a PCI Express switch device that couples the processor 1202 to a PCI Express bus, and the processor 1202 may communicate with the GPGPUs 1206A-D via the PCI Express bus. The GPGPUs 1206A-D may be interconnected via a set of high-speed P2P GPU-to-GPU links 1216. In at least one embodiment, the GPU-to-GPU links 1216 are connected to each of the GPGPUs 1206A-D via a dedicated GPU link. In at least one embodiment, the P2P GPU links 1216 enable direct communication between each GPGPU 1206A-D without communicating through the host interface bus 1204 to which the processor 1202 is connected. In at least one embodiment, host interface bus 1204 remains available for system memory access or communication with other instances of multi-GPU computing system 1200A, e.g., via one or more network devices, with GPU-to-GPU traffic directed to P2P GPU link 1216. While in at least one embodiment, GPGPUs 1206A-D are connected to processor 1202 via host interface switch 1204, in at least one embodiment, processor 1202 includes direct support for P2P GPU link 1216 and can connect directly to GPGPUs 1206A-D.
[0298] Reasoning and / or training logic 615 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 6B and / or Figure 6C Details are provided regarding inference and / or training logic 615. In at least one embodiment, inference and / or training logic 615 may be used in multi-GPU computing system 1200A to perform inference or prediction operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures or neural network use cases described herein.
[0299] Fig. 12B is a block diagram of a graphics processor 1200B according to at least one embodiment. In at least one embodiment, graphics processor 1200B includes ring interconnect 1202, pipeline front end 1204, media engine 1237, and graphics cores 1280A-1280N. In at least one embodiment, ring interconnect 1202 couples graphics processor 1200B to other processing units, including other graphics processors or one or more general-purpose processor cores. In at least one embodiment, graphics processor 1200B is one of many processors integrated within a multi-core processing system.
[0300] In at least one embodiment, graphics processor 1200B receives batches of commands via ring interconnect 1202. In at least one embodiment, the incoming commands are interpreted by command streamer 1203 in pipeline front end 1204. In at least one embodiment, graphics processor 1200B includes scalable execution logic for performing 3D geometry processing and media processing via graphics cores 1280A-1280N. In at least one embodiment, for 3D geometry processing commands, command streamer 1203 provides commands to geometry pipeline 1236. In at least one embodiment, for at least some media processing commands, command streamer 1203 provides commands to video front end 1234, which is coupled to media engine 1237. In at least one embodiment, media engine 1237 includes video quality engine (VQE) 1230 for video and image post-processing, and multi-format encoding / decoding (MFX) 1233 engine for providing hardware accelerated media data encoding and decoding. In at least one embodiment, geometry pipeline 1236 and media engine 1237 each generate execution threads for thread execution resources provided by at least one graphics core 1280A.
[0301] In at least one embodiment, the graphics processor 1200B includes scalable thread execution resources featuring modular cores 1280A-1280N (sometimes referred to as core slices), each graphics core having multiple sub-cores 1250A-1250N, 1260A-1260N (sometimes referred to as core sub-slices). In at least one embodiment, the graphics processor 1200B can have any number of graphics cores 1280A. In at least one embodiment, the graphics processor 1200B includes a graphics core 1280A having at least a first sub-core 1250A and a second sub-core 1260A. In at least one embodiment, the graphics processor 1200B is a low-power processor having a single sub-core (e.g., 1250A). In at least one embodiment, the graphics processor 1200B includes multiple graphics cores 1280A-1280N, each graphics core including a group of first sub-cores 1250A-1250N and a group of second sub-cores 1260A-1260N. In at least one embodiment, each of the first sub-cores 1250A-1250N includes at least a first set of execution units 1252A-1252N and media / texture samplers 1254A-1254N. In at least one embodiment, each of the second sub-cores 1260A-1260N includes at least a second set of execution units 1262A-1262N and samplers 1264A-1264N. In at least one embodiment, each of the sub-cores 1250A-1250N, 1260A-1260N shares a set of shared resources 1270A-1270N. In at least one embodiment, the shared resources include shared cache memory and pixel operation logic.
[0302] Reasoning and / or training logic 615 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 6B and / or Figure 6C Details are provided regarding inference and / or training logic 615. In at least one embodiment, inference and / or training logic 615 may be used in graphics processor 1200B to perform inference or prediction operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0303] Fig.13is a block diagram of a microarchitecture for a processor 1300 that may include logic circuitry for executing instructions, according to an illustration of at least one embodiment. In at least one embodiment, the processor 1300 may execute instructions, including x86 instructions, ARM instructions, special instructions for an application-specific integrated circuit (ASIC), and the like. In at least one embodiment, the processor 1300 may include registers for storing packed data, such as 64-bit wide MMXTM registers in a microprocessor enabled with MMX technology by Intel Corporation of Santa Clara, California. In at least one embodiment, the MMX registers available in integer and floating point form may operate with packed data elements that accompany single instruction multiple data (“SIMD”) and streaming SIMD extension (“SSE”) instructions. In at least one embodiment, 128-bit wide XMM registers associated with SSE2, SSE3, SSE4, AVX, or higher (generally referred to as “SSEx”) technology may hold such packed data operands. In at least one embodiment, the processor 1300 may execute instructions to accelerate machine learning or deep learning algorithms, training, or reasoning.
[0304] In at least one embodiment, the processor 1300 includes an in-order front end ("front end") 1301 to fetch instructions to be executed and prepare instructions for later use in the processor pipeline. In at least one embodiment, the front end 1301 may include several units. In at least one embodiment, an instruction prefetcher 1326 fetches instructions from memory and provides the instructions to an instruction decoder 1328, which in turn decodes or interprets the instructions. For example, in at least one embodiment, the instruction decoder 1328 decodes the received instructions into one or more operations of so-called "microinstructions" or "microoperations" (also referred to as "micro-operations" or "microinstructions") that the machine can execute. In at least one embodiment, the instruction decoder 1328 parses the instructions into opcodes and corresponding data and control fields, which can be used by the microarchitecture to perform operations according to at least one embodiment. In at least one embodiment, the trace cache 1330 can assemble the decoded microinstructions into a program-ordered sequence or trace in the microinstruction queue 1334 for execution. In at least one embodiment, when trace cache 1330 encounters a complex instruction, microcode ROM 1332 provides the microinstructions necessary to complete the operation.
[0305] In at least one embodiment, some instructions may be converted into a single micro-operation, while other instructions may require several micro-operations to complete the entire operation. In at least one embodiment, if more than four micro-operations are required to complete an instruction, the instruction decoder 1328 may access the microcode ROM 1332 to execute the instruction. In at least one embodiment, the instruction may be decoded into a small number of micro-operations for processing at the instruction decoder 1328. In at least one embodiment, if multiple micro-operations are required to complete the operation, the instruction may be stored in the microcode ROM 1332. In at least one embodiment, the trace cache 1330 references the entry point programmable logic array ("PLA") to determine the correct micro-instruction pointer for reading the microcode sequence from the microcode ROM 1332 to complete one or more instructions according to at least one embodiment. In at least one embodiment, after the microcode ROM 1332 completes the micro-operation sequencing of the instruction, the front end 1301 of the machine may resume fetching micro-operations from the trace cache 1330.
[0306] In at least one embodiment, an out-of-order execution engine ("out-of-order engine") 1303 can prepare instructions for execution. In at least one embodiment, the out-of-order execution logic has multiple buffers to smooth and reorder the instruction flow to optimize performance as instructions go down the pipeline and are scheduled for execution. In at least one embodiment, the out-of-order execution engine 1303 includes, but is not limited to, an allocator / register renamer 1340, a memory microinstruction queue 1342, an integer / floating point microinstruction queue 1344, a memory scheduler 1346, a fast scheduler 1302, a slow / general purpose floating point scheduler ("slow / general purpose FP scheduler") 1304, and a simple floating point scheduler ("simple FP scheduler") 1306. In at least one embodiment, the fast scheduler 1302, the slow / general purpose floating point scheduler 1304, and the simple floating point scheduler 1306 are also collectively referred to as "microinstruction schedulers 1302, 1304, 1306". In at least one embodiment, the allocator / register renamer 1340 allocates machine buffers and resources required for each microinstruction to execute in sequence. In at least one embodiment, the allocator / register renamer 1340 renames logical registers into entries in the register file. In at least one embodiment, the allocator / register renamer 1340 also allocates entries for each microinstruction in one of the two microinstruction queues, the memory microinstruction queue 1342 for memory operations and the integer / floating point microinstruction queue 1344 for non-memory operations, in front of the memory scheduler 1346 and the microinstruction schedulers 1302, 1304, 1306. In at least one embodiment, the microinstruction schedulers 1302, 1304, 1306 determine when the microinstructions are ready to execute based on the readiness of their slave input register operand sources and the availability of the execution resource microinstructions that need to be completed. In at least one embodiment, the fast scheduler 1302 of at least one embodiment can be scheduled on each half of the main clock cycle, while the slow / general floating point scheduler 1304 and the simple floating point scheduler 1306 can be scheduled once per main processor clock cycle. In at least one embodiment, microinstruction schedulers 1302, 1304, 1306 arbitrate dispatch ports to schedule microinstructions for execution.
[0307] In at least one embodiment, execution block 1311 includes, but is not limited to, integer register file / bypass network 1308, floating point register file / bypass network ("FP register file / bypass network") 1310, address generation units ("AGUs") 1312 and 1314, fast arithmetic logic units ("fast ALUs") 1316 and 1318, slow arithmetic logic unit ("slow ALU") 1320, floating point ALU ("FP") 1322, and floating point move unit ("FP move") 1324. In at least one embodiment, integer register file / bypass network 1308 and floating point register file / bypass network 1310 are also referred to herein as "register files 1308, 1310". In at least one embodiment, AGUs 1312 and 1314, fast ALUs 1316 and 1318, slow ALU 1320, floating point ALU 1322, and floating point move unit 1324 are also referred to herein as "execution units 1312, 1314, 1316, 1318, 1320, 1322, and 1324." In at least one embodiment, execution block 1311 may include, but is not limited to, any number (including zero) and type of register files, bypass networks, address generation units, and execution units (in any combination).
[0308] In at least one embodiment, register networks 1308, 1310 may be arranged between microinstruction schedulers 1302, 1304, 1306 and execution units 1312, 1314, 1316, 1318, 1320, 1322, and 1324. In at least one embodiment, integer register file / bypass network 1308 performs integer operations. In at least one embodiment, floating point register file / bypass network 1310 performs floating point operations. In at least one embodiment, each of register networks 1308, 1310 may include, but is not limited to, a branch network that may bypass or forward a just completed result that has not yet been written to the register file to a new slave object. In at least one embodiment, register networks 1308, 1310 may communicate data with each other. In at least one embodiment, integer register file / bypass network 1308 may include, but is not limited to, two separate register files, one register file for low-order 32-bit data, and a second register file for high-order 32-bit data. In at least one embodiment, floating point register file / bypass network 1310 may include, but is not limited to, 128-bit wide entries, since floating point instructions typically have operands that are 64 to 128 bits wide.
[0309] In at least one embodiment, execution units 1312, 1314, 1316, 1318, 1320, 1322, 1324 can execute instructions. In at least one embodiment, register files 1308, 1310 store integer and floating point data operand values that microinstructions need to execute. In at least one embodiment, processor 1300 may include, but is not limited to, any number of execution units 1312, 1314, 1316, 1318, 1320, 1322, 1324 and combinations thereof. In at least one embodiment, floating point ALU 1322 and floating point move unit 1324 can perform floating point, MMX, SIMD, AVX and SSE or other operations, including specialized machine learning instructions. In at least one embodiment, floating point ALU 1322 may include, but is not limited to, a 64-bit by 64-bit floating point divider to perform division, square root and remainder micro-operations. In at least one embodiment, floating point hardware may be used to process instructions involving floating point values. In at least one embodiment, ALU operations can be passed to fast ALUs 1316, 1318. In at least one embodiment, fast ALUs 1316, 1318 can perform fast operations with an effective delay of half a clock cycle. In at least one embodiment, most complex integer operations go to slow ALU 1320 because slow ALU 1320 can include, but is not limited to, integer execution hardware for long-latency type operations, such as multipliers, shifts, flag logic, and branch processing. In at least one embodiment, memory load / store operations can be performed by AGUs 1312, 1314. In at least one embodiment, fast ALU 1316, fast ALU 1318, and slow ALU 1320 can perform integer operations on 64-bit data operands. In at least one embodiment, fast ALU 1316, fast ALU 1318, and slow ALU 1320 can be implemented to support various data bit sizes including sixteen, thirty-two, 128, 256, etc. In at least one embodiment, the floating point ALU 1322 and floating point move unit 1324 can be implemented to support a range of operands having bits of various widths. In at least one embodiment, the floating point ALU 1322 and floating point move unit 1324 can operate on 128-bit wide packed data operands in conjunction with SIMD and multimedia instructions.
[0310] In at least one embodiment, microinstruction schedulers 1302, 1304, 1306 schedule dependent operations before the parent load completes execution. In at least one embodiment, since microinstructions can be speculatively scheduled and executed in processor 1300, processor 1300 can also include logic for handling memory misses. In at least one embodiment, if the data load in the data cache misses, there may be dependent operations running in the pipeline, which temporarily prevents the scheduler from having the correct data. In at least one embodiment, a replay mechanism tracks and re-executes instructions that use incorrect data. In at least one embodiment, it may be necessary to replay dependent operations and independent operations may be allowed to complete. In at least one embodiment, the scheduler and replay mechanism of at least one embodiment of the processor may also be designed to capture instruction sequences for text string comparison operations.
[0311] In at least one embodiment, "register" may refer to an onboard processor storage location that can be used as part of an instruction to identify an operand. In at least one embodiment, registers may be those that can be used from outside the processor (from a programmer's perspective). In at least one embodiment, registers may not be limited to a particular type of circuit. On the contrary, in at least one embodiment, registers can store data, provide data, and perform the functions described herein. In at least one embodiment, the registers described herein can be implemented using a variety of different techniques by circuits within the processor, such as dedicated physical registers, physical registers dynamically allocated using register renaming, a combination of dedicated and dynamically allocated physical registers, etc. In at least one embodiment, integer registers store 32-bit integer data. The register file of at least one embodiment also includes eight multimedia SIMD registers for packaging data.
[0312] Reasoning and / or training logic 615 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 6B and / or Figure 6C Details are provided regarding the inference and / or training logic 615. In at least one embodiment, part or all of the inference and / or training logic 615 may be incorporated into the execution block 1311 and other memories or registers shown or not shown. For example, in at least one embodiment, the training and / or inference techniques described herein may use one or more ALUs shown in the execution block 1311. In addition, weight parameters may be stored in on-chip or off-chip memory and / or registers (shown or not shown) that configure the ALUs of the execution block 1311 to perform one or more of the machine learning algorithms, neural network architectures, use cases, or training techniques described herein.
[0313] Fig.14A deep learning application processor 1400 is shown in accordance with at least one embodiment. In at least one embodiment, the deep learning application processor 1400 uses instructions that, if executed by the deep learning application processor 1400, cause the deep learning application processor 1400 to perform some or all of the processes and techniques described throughout this disclosure. In at least one embodiment, the deep learning application processor 1400 is an application specific integrated circuit (ASIC). In at least one embodiment, the application processor 1400 performs matrix multiplication operations or is "hardwired" into hardware as a result of executing one or more instructions, or both. In at least one embodiment, the deep learning application processor 1400 includes, but is not limited to, processing clusters 1410(1)-1410(12), inter-chip links (“ICLs”) 1420(1)-1420(12), inter-chip controllers (“ICCs”) 1430(1)-1430(2), memory controllers (“Mem Ctrlr”) 1442(1)-1442(4), high bandwidth memory physical layer (“HBM PHY”) 1444(1)-1444(4), a management controller central processing unit (“management controller CPU”) 1450, serial peripheral interface, inter-integrated circuit, and general purpose input / output blocks (“SPI, I2C, GPIO”), peripheral component interconnect express controller and direct memory access block (“PCIe controller and DMA”) 1470, and sixteen-lane peripheral component interconnect express port (“PCI Express x 16”) 1480.
[0314] In at least one embodiment, the processing cluster 1410 may perform deep learning operations, including inference or prediction operations based on weight parameters calculated based on one or more training techniques, including those of this article. In at least one embodiment, each processing cluster 1410 may include, but is not limited to, any number and type of processors. In at least one embodiment, the deep learning application processor 1400 may include any number and type of processing clusters 1400. In at least one embodiment, the inter-chip link 1420 is bidirectional. In at least one embodiment, the inter-chip link 1420 and the inter-chip controller 1430 enable multiple deep learning application processors 1400 to exchange information, including activation information generated from executing one or more machine learning algorithms embodied in one or more neural networks. In at least one embodiment, the deep learning application processor 1400 may include any number (including zero) and type of ICL 1420 and ICC 1430.
[0315] In at least one embodiment, HBM2 1440 provides a total of 32GB of memory. HBM2 1440(i) is associated with both memory controller 1442(i) and HBM PHY 1444(i). In at least one embodiment, any number of HBM2 1440 can provide any type and total amount of high bandwidth memory and can be associated with any number (including zero) and type of memory controller 1442 and HBM PHY 1444. In at least one embodiment, SPI, I2C, GPIO 3360, PCIe controller 1460 and DMA 1470 and / or PCIe 1480 can be replaced with any number and type of blocks to implement any number and type of communication standards in any technically feasible manner.
[0316] Reasoning and / or training logic 615 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 6B and / or Figure 6C Details are provided regarding the inference and / or training logic 615. In at least one embodiment, the deep learning application processor is used to train a machine learning model (e.g., a neural network) to predict or infer information provided to the deep learning application processor 1400. In at least one embodiment, the deep learning application processor 1400 is used to infer or predict information based on a trained machine learning model (e.g., a neural network) that has been trained by another processor or system or by the deep learning application processor 1400. In at least one embodiment, the processor 1400 can be used to perform one or more of the neural network use cases described herein.
[0317] Fig.15 1 is a block diagram of a neuromorphic processor 1500 according to at least one embodiment. In at least one embodiment, the neuromorphic processor 1500 may receive one or more inputs from a source external to the neuromorphic processor 1500. In at least one embodiment, these inputs may be transmitted to one or more neurons 1502 within the neuromorphic processor 1500. In at least one embodiment, the neurons 1502 and their components may be implemented using circuits or logic including one or more arithmetic logic units (ALUs). In at least one embodiment, the neuromorphic processor 1500 may include, but is not limited to, thousands of instances of neurons 1502, although any suitable number of neurons 1502 may be used. In at least one embodiment, each instance of a neuron 1502 may include a neuron input 1504 and a neuron output 1506. In at least one embodiment, a neuron 1502 may generate an output that may be transmitted to the inputs of other instances of the neuron 1502. In at least one embodiment, the neuron input 1504 and the neuron output 1506 may be interconnected via a synapse 1508.
[0318] In at least one embodiment, the neurons 1502 and synapses 1508 may be interconnected such that the neuromorphic processor 1500 operates to process or analyze information received by the neuromorphic processor 1500. In at least one embodiment, the neuron 1502 may send an output pulse (or "trigger" or "spike") when the input received through the neuron input 1504 exceeds a threshold. In at least one embodiment, the neuron 1502 may sum or integrate the signal received at the neuron input 1504. For example, in at least one embodiment, the neuron 1502 may be implemented as a leaky integrate-trigger neuron, where if the sum (referred to as the "membrane potential") exceeds a threshold, the neuron 1502 may generate an output (or "trigger") using a transfer function such as a sigmoid or threshold function. In at least one embodiment, the leaky integrate-trigger neuron may sum the signal received at the neuron input 1504 into a membrane potential, and may apply an application attenuation factor (or leakage) to reduce the membrane potential. In at least one embodiment, a leaky integrate-and-trigger neuron may trigger if multiple input signals are received at the neuron input 1504 fast enough to exceed a threshold (in at least one embodiment, before the membrane potential decays too low to trigger). In at least one embodiment, the neuron 1502 may be implemented using circuitry or logic that receives inputs, integrates the inputs into a membrane potential, and decays the membrane potential. In at least one embodiment, the inputs may be averaged, or any other suitable transfer function may be used. In addition, in at least one embodiment, the neuron 1502 may include, but is not limited to, a comparator circuit or logic that generates an output spike at the neuron output 1506 when the result of applying the transfer function to the neuron input 1504 exceeds a threshold. In at least one embodiment, once the neuron 1502 triggers, it may ignore previously received input information by, for example, resetting the membrane potential to 0 or another suitable default value. In at least one embodiment, once the membrane potential is reset to 0, the neuron 1502 may resume normal operation after a suitable period of time (or recovery period).
[0319] In at least one embodiment, neurons 1502 may be interconnected via synapses 1508. In at least one embodiment, synapses 1508 may be operable to transmit a signal from an output of a first neuron 1502 to an input of a second neuron 1502. In at least one embodiment, a neuron 1502 may transmit information over more than one instance of synapse 1508. In at least one embodiment, one or more instances of a neuron output 1506 may be connected to an instance of a neuron input 1504 in the same neuron 1502 via an instance of synapse 1508. In at least one embodiment, an instance of a neuron 1502 that produces an output to be transmitted over an instance of synapse 1508 may be referred to as a "pre-synaptic neuron" relative to that instance of synapse 1508. In at least one embodiment, an instance of a neuron 1502 that receives an input transmitted through an instance of synapse 1508 may be referred to as a "post-synaptic neuron" relative to an instance of synapse 1508. In at least one embodiment, with respect to various instances of synapses 1508, because an instance of neuron 1502 can receive input from one or more instances of synapses 1508 and can also transmit output through one or more instances of synapses 1508, a single instance of neuron 1502 can be both a "pre-synaptic neuron" and a "post-synaptic neuron."
[0320] In at least one embodiment, neurons 1502 may be organized into one or more layers. Each instance of a neuron 1502 may have a neuron output 1506 that may fan out to one or more neuron inputs 1504 through one or more synapses 1508. In at least one embodiment, a neuron output 1506 of a neuron 1502 in a first layer 1510 may be connected to a neuron input 1504 of a neuron 1502 in a second layer 1512. In at least one embodiment, the layers 1510 may be referred to as "feed-forward layers". In at least one embodiment, each instance of a neuron 1502 in an instance of the first layer 1510 may fan out to each instance of a neuron 1502 in a second layer 1512. In at least one embodiment, the first layer 1510 may be referred to as a "fully connected feed-forward layer". In at least one embodiment, each instance of a neuron 1502 in each instance of the second layer 1512 may fan out to less than all instances of a neuron 1502 in a third layer 1514. In at least one embodiment, the second layer 1512 may be referred to as a "sparsely connected feed-forward layer". In at least one embodiment, neurons 1502 in the (same) second layer 1512 may fan out to neurons 1502 in multiple other layers, including also fanning out to neurons 1502 in the second layer 1512. In at least one embodiment, the second layer 1512 may be referred to as a "recurrent layer". In at least one embodiment, the neuromorphic processor 1500 may include, but is not limited to, any suitable combination of recurrent layers and feed-forward layers, including but not limited to sparsely connected feed-forward layers and fully connected feed-forward layers.
[0321] In at least one embodiment, the neuromorphic processor 1500 may include, but is not limited to, a reconfigurable interconnect architecture or a dedicated hardwired interconnect to connect the synapses 1508 to the neurons 1502. In at least one embodiment, the neuromorphic processor 1500 may include, but is not limited to, circuitry or logic that allows synapses to be assigned to different neurons 1502 as needed, depending on the neural network topology and neuron fan-in / fan-out. For example, in at least one embodiment, the synapses 1508 may be connected to the neurons 1502 using an interconnect architecture such as a network on a chip or through dedicated connections. In at least one embodiment, the synaptic interconnects and their components may be implemented using circuitry or logic.
[0322] Fig.16AA processing system according to at least one embodiment is shown. In at least one embodiment, system 1600A includes one or more processors 1602 and one or more graphics processors 1608, and can be a single processor desktop system, a multi-processor workstation system, or a server system with a large number of processors 1602 or processor cores 1607. In at least one embodiment, system 1600A is a processing platform incorporated within a system-on-chip (SoC) integrated circuit for use in a mobile, handheld, or embedded device.
[0323] In at least one embodiment, the system 1600A may include or be incorporated into a server-based gaming platform, including a gaming console, a mobile gaming console, a handheld gaming console, or an online gaming console for gaming and media consoles. In at least one embodiment, the system 1600A is a mobile phone, a smart phone, a tablet computing device, or a mobile Internet device. In at least one embodiment, the processing system 1600A may also include a wearable device coupled to or integrated in a wearable device, such as a smart watch wearable device, a smart glasses device, an augmented reality device, or a virtual reality device. In at least one embodiment, the processing system 1600A is a television or set-top box device having one or more processors 1602 and a graphical interface generated by one or more graphics processors 1608.
[0324] In at least one embodiment, one or more processors 1602 each include one or more processor cores 1607 to process instructions that, when executed, perform operations for system and user software. In at least one embodiment, each of the one or more processor cores 1607 is configured to process a specific instruction set 1609. In at least one embodiment, the instruction set 1609 can facilitate complex instruction set computing (CISC), reduced instruction set computing (RISC), or computing through very long instruction words (VLIW). In at least one embodiment, the processor cores 1607 can each process a different instruction set 1609, which can include instructions that help emulate other instruction sets. In at least one embodiment, the processor cores 1607 can also include other processing devices, such as digital signal processors (DSPs).
[0325] In at least one embodiment, the processor 1602 includes a cache memory 1604. In at least one embodiment, the processor 1602 may have a single internal cache or multiple levels of internal cache. In at least one embodiment, the cache memory is shared between various components of the processor 1602. In at least one embodiment, the processor 1602 also uses an external cache (e.g., a level 3 (L3) cache or a last level cache (LLC)) (not shown), which may be shared between the processor cores 1607 using known cache coherence techniques. In at least one embodiment, the processor 1602 additionally includes a register file 1606, which may include different types of registers (e.g., integer registers, floating point registers, status registers, and instruction pointer registers) for storing different types of data. In at least one embodiment, the register file 1606 may include general registers or other registers.
[0326] In at least one embodiment, one or more processors 1602 are coupled to one or more interface buses 1610 to transmit communication signals, such as address, data, or control signals, between the processor 1602 and other components in the system 1600A. In at least one embodiment, the interface bus 1610 can be a processor bus in one embodiment, such as a version of a direct media interface (DMI) bus. In at least one embodiment, the interface bus 1610 is not limited to a DMI bus, and can include one or more peripheral component interconnect buses (e.g., PCI, PCI Express), a memory bus, or other types of interface buses. In at least one embodiment, the processor 1602 includes an integrated memory controller 1616 and a platform controller hub 1630. In at least one embodiment, the memory controller 1616 facilitates communication between memory devices and other components of the processing system 1600A, while the platform controller hub (PCH) 1630 provides connections to input / output (I / O) devices through a local I / O bus.
[0327] In at least one embodiment, the memory device 1620 may be a dynamic random access memory (DRAM) device, a static random access memory (SRAM) device, a flash memory device, a phase change memory device, or have appropriate performance to be used as a processor memory. In at least one embodiment, the memory device 1620 may be used as a system memory of the processing system 1600A to store data 1622 and instructions 1621 for use when one or more processors 1602 execute applications or processes. In at least one embodiment, the memory controller 1616 is also coupled to an external graphics processor 1612 of at least one embodiment, which may communicate with one or more graphics processors 1608 in the processor 1602 to perform graphics and media operations. In at least one embodiment, the display device 1611 may be connected to the processor 1602. In at least one embodiment, the display device 1611 may include one or more of the internal display devices, such as in a mobile electronic device or laptop device or an external display device connected via a display interface (e.g., DisplayPort, etc.). In at least one embodiment, the display device 1611 may include a head mounted display (HMD), such as a stereoscopic display device used in virtual reality (VR) applications or augmented reality (AR) applications.
[0328] In at least one embodiment, the platform controller hub 1630 enables peripheral devices to be connected to the storage device 1620 and the processor 1602 via a high-speed I / O bus. In at least one embodiment, the I / O peripherals include, but are not limited to, an audio controller 1646, a network controller 1634, a firmware interface 1628, a wireless transceiver 1626, a touch sensor 1625, a data storage device 1624 (e.g., a hard drive, flash memory, etc.). In at least one embodiment, the data storage device 1624 can be connected via a storage interface (e.g., SATA) or via a peripheral bus, such as a peripheral component interconnect bus (e.g., PCI, PCIe). In at least one embodiment, the touch sensor 1625 can include a touch screen sensor, a pressure sensor, or a fingerprint sensor. In at least one embodiment, the wireless transceiver 1626 can be a Wi-Fi transceiver, a Bluetooth transceiver, or a mobile network transceiver, such as a 3G, 4G, or long-term evolution (LTE) transceiver. In at least one embodiment, the firmware interface 1628 enables communication with the system firmware and can be, for example, a unified extensible firmware interface (UEFI). In at least one embodiment, the network controller 1634 can enable a network connection to a wired network. In at least one embodiment, a high-performance network controller (not shown) is coupled to the interface bus 1610. In at least one embodiment, the audio controller 1646 is a multi-channel high-definition audio controller. In at least one embodiment, the processing system 1600A includes a legacy I / O controller 1640 for coupling legacy (e.g., Personal System 2 (PS / 2)) devices to the system 1600A. In at least one embodiment, the platform controller hub 1630 can also be connected to one or more universal serial bus (USB) controllers 1642, which connect input devices such as a keyboard and mouse 1643 combination, a camera 1644, or other USB input devices.
[0329] In at least one embodiment, instances of memory controller 1616 and platform controller hub 1630 may be integrated into a discrete external graphics processor, such as external graphics processor 1612. In at least one embodiment, platform controller hub 1630 and / or memory controller 1616 may be external to one or more processors 1602. In at least one embodiment, system 1600A may include external memory controller 1616 and platform controller hub 1630, which may be configured as a memory controller hub and a peripheral controller hub in a system chipset that communicates with processor 1602.
[0330] Reasoning and / or training logic 615 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 6B and / or Figure 6C1600A. In at least one embodiment, the training and / or inference techniques described herein may use one or more ALUs that are embodied in graphics processor 1612. In addition, in at least one embodiment, the inference and / or training operations described herein may use additional ALUs. Figure 6B and / or Figure 6C In at least one embodiment, the weight parameters may be stored in on-chip or off-chip memory and / or registers (shown or not shown) that configure the ALU of the graphics processor 1600A to execute one or more of the machine learning algorithms, neural network architectures, use cases, or training techniques described herein.
[0331] Fig. 16B is a block diagram of a processor 1600B having one or more processor cores 1602A-1602N, an integrated memory controller 1614, and an integrated graphics processor 1608, according to at least one embodiment. In at least one embodiment, the processor 1600B may include additional cores, up to and including the additional core 1602N represented by the dashed box. In at least one embodiment, each processor core 1602A-1602N includes one or more internal cache units 1604A-1604N. In at least one embodiment, each processor core may also have access to one or more shared cache units 1606.
[0332] In at least one embodiment, the internal cache units 1604A-1604N and the shared cache unit 1606 represent a cache memory hierarchy within the processor 1600B. In at least one embodiment, the cache memory units 1604A-1604N may include at least one level of instruction and data cache within each processor core and one or more levels of cache in a shared mid-level cache, such as level 2 (L2), level 3 (L3), level 4 (L4), or other levels of cache, where the highest level of cache before external memory is classified as LLC. In at least one embodiment, cache coherence logic maintains coherence between the various cache units 1606 and 1604A-1604N.
[0333] In at least one embodiment, the processor 1600B may also include a set of one or more bus controller units 1616 and a system agent core 1610. In at least one embodiment, the one or more bus controller units 1616 manage a set of peripheral buses, such as one or more PCI or PCIe buses. In at least one embodiment, the system agent core 1610 provides management functions for various processor components. In at least one embodiment, the system agent core 1610 includes one or more integrated memory controllers 1614 to manage access to various external memory devices (not shown).
[0334] In at least one embodiment, one or more processor cores 1602A-1602N include support for multiple threads simultaneously. In at least one embodiment, system agent core 1610 includes components for coordinating and operating cores 1602A-1602N during multithreaded processing. In at least one embodiment, system agent core 1610 may additionally include a power control unit (PCU) that includes logic and components for regulating one or more power states of processor cores 1602A-1602N and graphics processor 1608.
[0335] In at least one embodiment, the processor 1600B also includes a graphics processor 1608 for performing image processing operations. In at least one embodiment, the graphics processor 1608 is coupled to a shared cache unit 1606 and a system agent core 1610 including one or more integrated memory controllers 1614. In at least one embodiment, the system agent core 1610 also includes a display controller 1611 for driving the graphics processor output to one or more coupled displays. In at least one embodiment, the display controller 1611 may also be a separate module coupled to the graphics processor 1608 via at least one interconnect, or may be integrated within the graphics processor 1608.
[0336] In at least one embodiment, a ring-based interconnect unit 1612 is used to couple the internal components of the processor 1600B. In at least one embodiment, alternative interconnect units may be used, such as point-to-point interconnects, switched interconnects, or other technologies. In at least one embodiment, the graphics processor 1608 is coupled to the ring interconnect 1612 via an I / O link 1613.
[0337] In at least one embodiment, I / O link 1613 represents at least one of a variety of I / O interconnects, including packaged I / O interconnects that facilitate communication between various processor components and high-performance embedded memory modules 1618 (e.g., eDRAM modules). In at least one embodiment, each of processor cores 1602A-1602N and graphics processor 1608 uses embedded memory modules 1618 as a shared last level cache.
[0338] In at least one embodiment, the processor cores 1602A-1602N are homogeneous cores that execute a common instruction set architecture. In at least one embodiment, the processor cores 1602A-1602N are heterogeneous in terms of instruction set architecture (ISA), wherein one or more processor cores 1602A-1602N execute a common instruction set, while one or more other processor cores 1602A-1602N execute a subset or a different instruction set of the common instruction set. In at least one embodiment, the processor cores 1602A-1602N are heterogeneous in terms of microarchitecture, wherein one or more cores with relatively high power consumption are coupled with one or more power cores with lower power consumption. In at least one embodiment, the processor 1600B can be implemented on one or more chips or implemented as a SoC integrated circuit.
[0339] Reasoning and / or training logic 615 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 6B and / or Figure 6C 615. In at least one embodiment, some or all of the inference and / or training logic 615 may be incorporated into the processor 1600B. For example, in at least one embodiment, the training and / or inference techniques described herein may use one or more ALUs embodied in Fig.16A In addition, in at least one embodiment, the inference and / or training operations described herein may use the inference and / or training operations described herein. Figure 6B and / or Figure 6C In at least one embodiment, the weight parameters may be stored in on-chip or off-chip memory and / or registers (shown or not shown) that configure the ALU of the graphics processor 1600B to execute one or more machine learning algorithms, neural network architectures, use cases, or training techniques described herein.
[0340] Fig. 16Cis a block diagram of the hardware logic of a graphics processor core 1600C according to at least one embodiment herein. In at least one embodiment, graphics processor core 1600C is included in a graphics core array. In at least one embodiment, graphics processor core 1600C (sometimes referred to as a core slice) may be one or more graphics cores within a modular graphics processor. In at least one embodiment, graphics processor core 1600C is an example of a graphics core slice, and the graphics processor herein may include multiple graphics core slices based on a target power and performance envelope. In at least one embodiment, each graphics core 1600C may include a fixed function block 1630 coupled to multiple sub-cores 1601A-1601F, also referred to as a sub-slice, which includes modular blocks of general and fixed function logic.
[0341] In at least one embodiment, fixed function block 1630 includes a geometry fixed function pipeline 1636, which can be shared by all sub-cores in graphics processor 1600C, for example, in lower performance and / or lower power graphics processor implementations. In at least one embodiment, geometry and fixed function pipeline 1636 includes a 3D fixed function pipeline, a video front end unit, a thread generator and thread dispatcher, and a unified return buffer manager that manages a unified return buffer.
[0342] In at least one fixed embodiment, the fixed function block 1630 also includes a graphics SoC interface 1637, a graphics microcontroller 1638, and a media pipeline 1639. In at least one embodiment, the fixed graphics SoC interface 1637 provides an interface between the graphics core 1600C and other processor cores in the on-chip integrated circuit system. In at least one embodiment, the graphics microcontroller 1638 is a programmable subprocessor that can be configured to manage various functions of the graphics processor 1600C, including thread dispatching, scheduling, and preemption. In at least one embodiment, the media pipeline 1639 includes logic that helps decode, encode, pre-process, and / or post-process multimedia data including image and video data. In at least one embodiment, the media pipeline 1639 implements media operations via requests to calculation or sampling logic within the sub-cores 1601-1601F.
[0343] In at least one embodiment, the SoC interface 1637 enables the graphics core 1600C to communicate with a general application processor core (e.g., a CPU) and / or other components within the SoC, including memory hierarchy elements such as a shared last level cache, system RAM, and / or embedded on-chip or packaged DRAM. In at least one embodiment, the SoC interface 1637 may also enable communication with fixed-function devices within the SoC (e.g., a camera imaging pipeline), and enable the use and / or implementation of global memory atomics that can be shared between the graphics core 1600C and the CPU within the SoC. In at least one embodiment, the SoC interface 1637 may also implement power management controls for the graphics core 1600C and enable interfaces between the clock domain of the graphics core 1600C and other clock domains within the SoC. In at least one embodiment, the SoC interface 1637 enables command buffers to be received from a command stream converter and a global thread dispatcher, which is configured to provide commands and instructions to each of one or more graphics cores within the graphics processor. In at least one embodiment, commands and instructions may be dispatched to media pipeline 1639 when media operations are to be performed, or may be assigned to geometry and fixed function pipelines (e.g., geometry and fixed function pipeline 1636, geometry and fixed function pipeline 1614) when graphics processing operations are to be performed.
[0344] In at least one embodiment, the graphics microcontroller 1638 can be configured to perform various scheduling and management tasks for the graphics core 1600C. In at least one embodiment, the graphics microcontroller 1638 can perform graphics and / or computing workload scheduling on various graphics parallel engines within the execution unit (EU) arrays 1602A-1602F, 1604A-1604F in the sub-cores 1601A-1601F. In at least one embodiment, host software executing on the CPU core of the SoC including the graphics core 1600C can submit a workload to one of the multiple graphics processor doorbells, which invokes scheduling operations on the appropriate graphics engine. In at least one embodiment, the scheduling operations include determining which workload to run next, submitting the workload to the command stream converter, preempting existing workloads running on the engine, monitoring the progress of the workload, and notifying the host software when the workload is completed. In at least one embodiment, graphics microcontroller 1638 may also facilitate low power or idle states for graphics core 1600C, thereby providing graphics core 1600C with the ability to save and restore registers across low power state transitions within graphics core 1600C independent of the operating system and / or graphics driver software on the system.
[0345] In at least one embodiment, graphics core 1600C may have up to N modular sub-cores more or less than the sub-cores 1601A-1601F shown. For each set of N sub-cores, in at least one embodiment, graphics core 1600C may also include shared function logic 1610, shared and / or cache memory 1612, geometry / fixed function pipelines 1614, and additional fixed function logic 1616 to accelerate various graphics and compute processing operations. In at least one embodiment, shared function logic 1610 may include logic units (e.g., samplers, math and / or inter-thread communication logic) that may be shared by each of the N sub-cores within graphics core 1600C. In at least one embodiment, fixed, shared and / or cache memory 1612 may be the last level cache for the N sub-cores 1601A-1601F within graphics core 1600C, and may also be used as a shared memory accessible by multiple sub-cores. In at least one embodiment, geometry / fixed function pipeline 1614 may be included in place of geometry / fixed function pipeline 1636 within fixed function block 1630 and may include similar logic units.
[0346] In at least one embodiment, the graphics core 1600C includes additional fixed function logic 1616, which may include various fixed function acceleration logic for use by the graphics core 1600C. In at least one embodiment, the additional fixed function logic 1616 includes an additional geometry pipeline for use in position-only shading. In position-only shading, there are at least two geometry pipelines, and in the full geometry pipeline and culling pipeline within the geometry and fixed function pipelines 1614, 1636, it is an additional geometry pipeline that may be included in the additional fixed function logic 1616. In at least one embodimen...
Claims
1. A cooling system for a data center, include: an evaporative cooling subsystem for providing blown air to the data center based on a first thermal signature in the data center and cooperating with a reusable refrigerant cooling subsystem for controlling moisture of the blown air in a first configuration of the reusable refrigerant cooling subsystem, the reusable refrigerant cooling subsystem for cooling the data center in a second configuration and based on a second thermal signature in the data center.
2. The cooling system according to claim 1, further comprising: include: The reusable refrigerant cooling subsystem operates in the second configuration when the evaporative cooling subsystem is unable to effectively reduce heat generated in at least one server or rack of the data center.
3. The cooling system according to claim 1, further comprising: include: The evaporative cooling subsystem and the reusable refrigerant cooling subsystem located within the infrastructure area of the trailer are configured to provide a determined range of cooling capacities to support the mobility of the data center.
4. The cooling system according to claim 1, further comprising: include: A refrigerant-based heat transfer subsystem is provided to implement one or more of the evaporative cooling subsystem, the reusable refrigerant cooling subsystem, and a dielectric-based cooling subsystem, wherein each subsystem has a different cooling capacity for the data center.
5. The cooling system according to claim 4, further comprising: include: the evaporative cooling subsystem having an associated first loop component for directing a first cooling medium to at least one first rack of the data center based on a first thermal characteristic or a first cooling requirement of the at least one first rack; as well as The reusable refrigerant cooling subsystem has an associated second loop component for directing a second cooling medium to the at least one first rack of the data center based on a second thermal characteristic or a second cooling requirement of the at least one first rack.
6. The cooling system of claim 5, wherein the first thermal signature or first cooling requirement is within a first threshold of the different cooling capacities and the second thermal signature or second cooling requirement is within a second threshold of the different cooling capacities. 7 . The cooling system according to claim 5 , wherein the first cooling medium is the blown air, and the second cooling medium is a refrigerant.
8. The cooling system according to claim 1, further comprising: include: one or more humidity sensors associated with the data center for monitoring the humidity of the blown air; as well as One or more controllers having at least one processor and memory including instructions for execution on the at least one processor to cause the reusable refrigerant cooling subsystem to control the humidity of the blown air.
9. The cooling system according to claim 1, further comprising: include: One or more servers have an air distribution unit and a refrigerant distribution unit for enabling different cooling media to meet different cooling requirements of the one or more servers reflected by the first thermal signature and the second thermal signature.
10. The cooling system according to claim 1, further comprising: include: A learning subsystem comprising at least one processor for evaluating temperatures within a server or one or more racks in the data center using flow rates of different cooling media, and providing outputs associated with at least two temperatures to facilitate movement of two cooling media from the cooling system by controlling flow controllers associated with the two cooling media to meet the first thermal signature and the second thermal signature.
11. The cooling system according to claim 10, further comprising: include: a cold plate associated with a component in the server; the flow controller facilitating movement of the two cooling media through the cold plate or through the server; as well as The learning subsystem executes a machine learning model to: processing the temperature using a plurality of neuron levels of the machine learning model having the temperature and previously correlated flow rates of the two cooling media; and The output associated with the flow rates of the two cooling mediums is provided to the flow controller, the output being provided upon evaluation of the previous associated flow rates and previous associated temperatures of individual ones of the two cooling mediums.
12. The cooling system according to claim 1, further comprising: include: one or more first flow controllers to maintain a first cooling medium associated with the evaporative cooling subsystem at a first flow rate and a second cooling medium associated with the reusable refrigerant cooling subsystem at a second flow rate, the first flow rate being based in part on the first thermal characteristic; and One or more second flow controllers maintain a second cooling medium associated with the reusable refrigerant cooling subsystem at a third flow rate based in part on the second thermal characteristic, and the one or more first flow controllers are deactivated based in part on the second thermal characteristic.
13. At least one processor for cooling the system, include: At least one logic unit for controlling flow controllers associated with a reusable refrigerant cooling subsystem and an evaporative cooling subsystem based in part on a first thermal characteristic in a data center, and based in part on a second thermal characteristic in the data center, the evaporative cooling subsystem working in conjunction with the reusable refrigerant cooling subsystem to control moisture of blown air from the evaporative cooling subsystem in a first configuration of the reusable refrigerant cooling subsystem, and working in conjunction with the reusable refrigerant cooling subsystem to cool the data center in a second configuration of the reusable refrigerant cooling subsystem, the blown air and cooling of the data center being based on the first thermal characteristic and the second thermal characteristic, respectively.
14. The at least one processor of claim 13, further comprising: include: A learning subsystem for evaluating temperatures within a server or one or more racks in the data center using flow rates of different cooling media and providing outputs associated with at least two temperatures to facilitate movement of two cooling media from the cooling system by controlling flow controllers associated with the two cooling media to meet the first thermal signature and the second thermal signature.
15. The at least one processor of claim 14, further comprising: include: The learning subsystem executes a machine learning model to: processing the temperature using a plurality of neuron levels of the machine learning model having the temperature and previously correlated flow rates of the two cooling media; and The output associated with the flow rates of the two cooling mediums is provided to the flow controller, the output being provided upon evaluation of the previous associated flow rates and previous associated temperatures of individual ones of the two cooling mediums.
16. The at least one processor of claim 14, further comprising: include: an instruction output for transmitting the output associated with the flow controller to achieve a first flow rate for a first separate cooling medium for meeting a first thermal characteristic in the data center while maintaining a second flow rate for a second separate cooling medium to control the humidity of the first separate cooling medium; and to deactivate the first separate cooling medium and achieve a third flow rate for the second separate cooling medium for meeting the second thermal characteristic in the data center.
17. The at least one processor of claim 14, further comprising: include: The at least one logic unit is adapted to receive a temperature value from a temperature sensor and a humidity value from a humidity sensor, wherein the temperature sensor and the humidity sensor are associated with the server or the one or more racks, and is adapted to simultaneously facilitate movement of a first separate cooling medium and a second separate cooling medium to satisfy a first thermal characteristic of the data center and to satisfy the second thermal characteristic of the data center alone.
18. At least one processor for cooling the system, include: At least one logic unit for training one or more neural networks having hidden layers of neurons to evaluate temperature, flow rate, and humidity associated with a reusable refrigerant cooling subsystem and an evaporative cooling subsystem based in part on a first thermal signature in a data center and in part on a second thermal signature in the data center, the evaporative cooling subsystem working in conjunction with the reusable refrigerant cooling subsystem to control humidity of blown air from the evaporative cooling subsystem in a first configuration of the reusable refrigerant cooling subsystem, and working in conjunction with the reusable refrigerant cooling subsystem to cool the data center in a second configuration of the reusable refrigerant cooling subsystem, the blown air and cooling of the data center being based on the first thermal signature and the second thermal signature, respectively.
19. The at least one processor of claim 18, further comprising: include: The at least one logic unit is used to output at least one instruction associated with at least two temperatures to promote the movement of the at least two cooling media by controlling flow controllers associated with the at least two cooling media to promote simultaneous movement of the two cooling media in the first configuration of the reusable refrigerant cooling subsystem, and to promote the flow of one of the two cooling media alone in the second configuration of the reusable refrigerant cooling subsystem.
20. The at least one processor of claim 19, further comprising: include: At least one instruction output for transmitting an output associated with a flow controller to achieve a first flow rate of a first separate cooling medium for meeting the first thermal characteristic in the data center while maintaining a second flow rate of a second separate cooling medium to control the humidity of the first separate cooling medium; and deactivating the first separate cooling medium and enabling a third flow rate of the second separate cooling medium for meeting the second thermal characteristic in the data center.
21. The at least one processor of claim 18, further comprising: include: The at least one logic unit is adapted to receive a temperature value from a temperature sensor and a humidity value from a humidity sensor, wherein the temperature sensor and the humidity sensor are associated with a server or one or more racks in the data center, and is adapted to simultaneously facilitate movement of a first separate cooling medium and a second separate cooling medium to satisfy a first thermal characteristic of the data center and to satisfy a second thermal characteristic of the data center individually.
22. A data center cooling system, include: At least one processor for training one or more neural networks having hidden layers of neurons for evaluating temperatures and flow rates associated with at least two cooling media of a reusable refrigerant cooling subsystem, the reusable refrigerant cooling subsystem comprising at least an evaporative cooling subsystem and the reusable refrigerant cooling subsystem, the evaporative cooling subsystem and the reusable refrigerant cooling subsystem being rated for different cooling capacities and adjustable for different cooling requirements within minimum and maximum values for different cooling capacities of a data center, the cooling system also having a form factor of one or more racks of the data center.
23. The data center cooling system of claim 22, further comprising: include: The at least one processor is configured to output at least one instruction associated with at least two temperatures to facilitate movement of the at least two cooling media by controlling flow controllers associated with the at least two cooling media to facilitate simultaneous flow of the two cooling media in a first configuration of the reusable refrigerant cooling subsystem, and to facilitate flow of one of the two cooling media alone in a second configuration of the reusable refrigerant cooling subsystem.
24. The data center cooling system of claim 23, further comprising: include: At least one instruction output of the at least one processor is used to transmit an output associated with a flow controller to achieve a first flow rate of a first separate cooling medium for meeting a first thermal characteristic in the data center while maintaining a second flow rate of a second separate cooling medium to control the humidity of the first separate cooling medium; and deactivate the first separate cooling medium and enable a third flow rate of the second separate cooling medium for meeting a second thermal characteristic in the data center.
25. The data center cooling system of claim 22, further comprising: include: The at least one processor is adapted to receive a temperature value from a temperature sensor and a humidity value from a humidity sensor, the temperature sensor and the humidity sensor being associated with a server or one or more racks in the data center, and to simultaneously facilitate movement of a first separate cooling medium and a second separate cooling medium to satisfy a first thermal characteristic of the data center and to satisfy a second thermal characteristic of the data center individually.
26. A method for cooling a data center, include: An evaporative cooling subsystem is provided for providing blown air to the data center based on a first thermal characteristic in the data center and is coordinated with a reusable refrigerant cooling subsystem that controls moisture of the blown air in a first configuration of the reusable refrigerant cooling subsystem, the reusable refrigerant cooling subsystem being used to cool the data center in a second configuration based on a second thermal characteristic of the data center.
27. The method according to claim 26, further comprising: include: directing the blown air from the evaporative cooling subsystem to the at least one first rack of the data center using a first loop assembly and based on a first thermal characteristic or a first cooling requirement of the at least one first rack; as well as Refrigerant is directed from the reusable refrigerant cooling subsystem to the at least one first rack of the data center using a second loop assembly and based on a second thermal characteristic or a second cooling requirement of the at least one first rack.
28. The method according to claim 26, further comprising: include: The reusable refrigerant cooling subsystem is operated in the second configuration when the evaporative cooling subsystem is unable to effectively reduce heat generated in at least one server or rack of the data center.
29. The method according to claim 26, further comprising: include: One or more of the evaporative cooling subsystem, the reusable refrigerant cooling subsystem, and a dielectric-based cooling subsystem are implemented using a refrigerant-based heat transfer subsystem, wherein each subsystem has a different cooling capacity for the data center.
30. The method according to claim 26, further comprising: include: monitoring the humidity of the blown air using one or more humidity sensors associated with the data center; as well as The reusable refrigerant cooling subsystem is enabled to control the humidity of the blown air.
Citation Information
Patent Citations
Cooling system for electronic apparatus
CN102149267A
Evaporative cooling-mechanical refrigeration combined type energy-saving air-conditioning system for data room
CN105698314A
Trim cooling assembly for cooling electronic equipment
US10231358B1
Optimal controller for hybrid liquid-air cooling system of electronic racks of a data center
US10238011B1
Data center cooling with an air-side economizer and liquid-cooled electronics rack(s)
US20130021746A1