Smart integrated liquid cooling rack for data centers
By using controlled couplers and cooling manifold systems supported by intelligent interfaces in the data center, the problems of low cooling efficiency and reliability of high-density servers are solved, and efficient and reliable cooling effects are achieved, reducing the need for modifications to existing facilities.
Patent Information
- Application Number
- CN202180004788.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-02-21
- Filing Date
- 2021-02-19
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2041-02-19
AI Technical Summary
Existing data center cooling systems, especially high-density servers, are inefficient in air cooling methods, and liquid cooling systems may cause server components to fail, requiring large-scale modifications to the existing data center architecture.
The controlled coupler and cooling manifold system supported by intelligent interfaces are adopted to achieve accurate distribution and control of coolant, and the press-fit connection between the rigid or flexible cooling manifold in the rack is eliminated to eliminate overhead flexible pipes, reduce leakage risks, and optimize cooling effects with intelligent flow measurement/control system.
An efficient and reliable liquid cooling system is achieved, reducing modification needs for data center facility design, reducing downtime and cost, while improving cooling efficiency and reliability.
Smart Images

Figure CN114402707B_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This is a PCT application of U.S. Patent Application No. 16 / 798,216, filed February 21, 2020. The disclosure of this application is incorporated herein by reference in its entirety for all purposes. Technical Field
[0003] At least one embodiment relates to a cooling system for a data center. In at least one embodiment, a cooling manifold is equipped with a first controllable fluid coupler and one or more second controllable fluid couplers for receiving coolant and distributing the coolant to one or more server cooling manifolds within server trays of a rack in the data center, in accordance with various novel techniques described herein. Background Art
[0004] Data center cooling systems typically use fans to circulate air through the server components. Some supercomputers or other high-capacity computers may use water or other cooling systems rather than air cooling systems to draw heat from the server components or racks in the data center to an area outside the data center. The cooling system may include chillers within the data center area. The area outside the data center may be a cooling tower or other external heat exchanger that receives heated coolant from the data center and dissipates the heat to the environment (or external cooling medium) through forced air or other means before the cooled coolant is recirculated back into the data center. In one example, the chiller and cooling tower together form a cooling facility with a pump. Air cooling systems cannot absorb enough heat to support effective or efficient cooling of the data center, and liquid cooling systems can seriously damage server components or racks through electrical shorts, flooding or other problems. Summary of the Invention
[0005] In one aspect, a data center cooling system is described. The data center cooling system includes a cooling manifold located within a rack and comprising a first controllable fluid coupler for receiving coolant from a cooling circuit external to the rack and one or more second controllable fluid couplers for distributing the coolant to one or more server cooling manifolds within server trays of the rack.
[0006] In another aspect, a method of cooling a data center is described. The method includes providing a cooling manifold within a rack, coupling a first controllable fluid coupler of the cooling manifold to a cooling circuit external to the rack to receive a coolant, and coupling one or more second controllable fluid couplers of the cooling manifold to one or more server cooling manifolds within server trays of the rack to distribute the coolant within the rack. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] Various embodiments according to the present disclosure will be described with reference to the accompanying drawings, in which:
[0008] Figure 1 is a block diagram of an example data center having a cooling system subject to the improvements described in at least one embodiment;
[0009] Figure 2 is a block diagram of another example data center having an improved cooling system including a cooling manifold according to at least one embodiment;
[0010] Figure 3A 、 Figure 3B are different views of an example rack incorporating a cooling manifold according to at least one embodiment;
[0011] Figure 3C 、 Figure 3D is a detailed view of a manifold and coupling between a rack cooling manifold and a tray or plate that may be coupled to the rack cooling manifold in accordance with at least one embodiment;
[0012] Figure 4 is a feature diagram illustrating a portion of an internal rack assembly of an example server tray or an example cooling plate and example computing components for cooling using an example cooling system in accordance with at least one embodiment;
[0013] Figure 5 is usable for use or manufacture according to at least one embodiment Figure 2-17D A process flow of method steps for a cooling system;
[0014] Figure 6A An example data center is shown where data from Figure 2-5 at least one embodiment of;
[0015] Figure 6B 、 Figure 6C Inference and / or training logic for enabling and / or supporting intelligent cooling systems according to various embodiments is shown, such as in Figure 6A and reasoning and / or training logic used in at least one embodiment of the present disclosure;
[0016] Figure 7A is a block diagram illustrating an exemplary computer system, which may be a system having interconnected devices and components, a system on a chip (SOC), or some combination thereof, formed together with a processor that may include an execution unit for executing instructions to support and / or implement the intelligent cooling system described herein, in accordance with at least one embodiment;
[0017] Figure 7Bis a block diagram illustrating an electronic device for utilizing a processor to support and / or implement the smart cooling system described herein, according to at least one embodiment;
[0018] Figure 7C is a block diagram illustrating an electronic device for utilizing a processor to support and / or implement the smart cooling system described herein, according to at least one embodiment;
[0019] Figure 8 Another exemplary computer system for implementing various processes and methods for an intelligent cooling system described throughout this disclosure is shown in accordance with at least one embodiment;
[0020] Figure 9A An exemplary architecture is shown in accordance with at least one embodiment of the present disclosure, wherein a GPU is communicatively coupled to a multi-core processor via a high-speed link to implement and / or support an intelligent cooling system;
[0021] Figure 9B shows additional details of the interconnection between a multi-core processor and a graphics acceleration module according to an exemplary embodiment;
[0022] Figure 9C Another exemplary embodiment according to at least one embodiment of the present disclosure is shown, wherein an accelerator integrated circuit is integrated within a processor for implementing and / or supporting a smart cooling system;
[0023] Figure 9D An exemplary accelerator integrated chip 990 for implementing and / or supporting a smart cooling system according to at least one embodiment of the present disclosure is shown;
[0024] Figure 9E shows additional details of an exemplary embodiment for implementing and / or supporting a shared model for a smart cooling system in accordance with at least one embodiment of the present disclosure;
[0025] Figure 9F Additional details are shown for one exemplary embodiment of a unified memory addressable via a common virtual memory address space for accessing physical processor memory and GPU memory to implement and / or support an intelligent cooling system in accordance with at least one embodiment of the present disclosure;
[0026] Figure 10A An exemplary integrated circuit and associated graphics processor for a smart cooling system according to embodiments described herein are shown;
[0027] Figures 10B-10C An exemplary integrated circuit and associated graphics processor for supporting and / or implementing an intelligent cooling system in accordance with at least one embodiment is shown;
[0028] Figures 10D-10E Additional exemplary graphics processor logic for supporting and / or implementing a smart cooling system in accordance with at least one embodiment is shown;
[0029] Figure 11A is a block diagram illustrating a computing system for supporting and / or implementing an intelligent cooling system according to at least one embodiment;
[0030] Figure 11B A parallel processor for supporting and / or implementing an intelligent cooling system according to at least one embodiment is shown;
[0031] Figure 11C is a block diagram of a partitioning unit according to at least one embodiment;
[0032] Figure 11D A graphics multiprocessor for an intelligent cooling system according to at least one embodiment is shown;
[0033] Figure 11E A graphics multiprocessor is shown in accordance with at least one embodiment;
[0034] Figure 12A A multi-GPU computing system is shown in accordance with at least one embodiment;
[0035] Figure 12B is a block diagram of a graphics processor according to at least one embodiment;
[0036] Figure 13 is a block diagram illustrating a microarchitecture for a processor, which may include logic circuitry for executing instructions, according to at least one embodiment;
[0037] Figure 14 A deep learning application processor according to at least one embodiment is shown;
[0038] Figure 15 is a block diagram of a neuromorphic processor according to at least one embodiment;
[0039] Figure 16A is a block diagram of a processing system according to at least one embodiment;
[0040] Figure 16B is a block diagram of a processor having one or more processor cores, an integrated memory controller, and an integrated graphics processor according to at least one embodiment;
[0041] Figure 16C is a block diagram of the hardware logic of a graphics processor core according to at least one embodiment;
[0042] Figures 16D-16EThread execution logic including an array of processing elements of a graphics processor core is shown in accordance with at least one embodiment;
[0043] Figure 17A illustrates a parallel processing unit according to at least one embodiment;
[0044] Figure 17B illustrates a general processing cluster in accordance with at least one embodiment;
[0045] Figure 17C A memory partitioning unit of a parallel processing unit according to at least one embodiment is shown; and
[0046] Figure 17D A streaming multiprocessor in accordance with at least one embodiment is shown. DETAILED DESCRIPTION
[0047] Given the high heat demands caused by today's computing components, air cooling of high-density servers is inefficient and ineffective. Therefore, the present disclosure explores the prospects of using liquid coolants and related systems to cool computing components such as graphics processing units (GPUs), central processing units (CPUs), or switch components. These computing components are used to assemble servers in server trays on racks in data centers. As computing components are miniaturized due to technological advances, server trays and racks accommodate more and more computing components, and therefore each component needs to dissipate more of the generated heat compared to existing systems. One problem addressed in the present disclosure is the potential for data center system failures in the event of any liquid exploration of computing and electronic components. Failures can be costly and may require extensive changes to existing data center architectures. In the present disclosure, data center racks can be provided with controllable couplers supported by intelligent interfaces that facilitate intelligent cooling of data centers. Accordingly, intelligent distribution of coolant is achieved so that primary coolant or liquid or auxiliary coolant or liquid can be directly distributed to computing components in server trays and racks in a primary cooling circuit or a secondary cooling circuit.
[0048] In at least one embodiment, the present disclosure also relates to a rack having a rack cooling manifold associated therewith. The structure provides a built-in coolant distribution manifold, such as one or more rack cooling manifolds, which can be rigid or flexible and have inlet / outlet main couplers for receiving and returning coolant from a cooling circuit external to the rack. For example, the structure also includes a distribution coupler on the rack cooling manifold for distributing coolant from the rack cooling manifold to each server tray. The position and height of the main coupler and the distribution coupler allow coolant (whether from the primary cooling circuit or the secondary cooling circuit) to flow into the rack and the server tray. In addition, on the other hand, at least the distribution coupler of the rack cooling manifold mates with the mating coupler of the server tray to achieve a press-fit connection. The press-fit connection between the couplers provides a quick connection to deliver coolant from the cooling circuit external to the rack to the rack and the internal server trays. Further, the present disclosure uses an intelligent flow measurement / control system to provide a precise flow of coolant to each server tray within the rack.
[0049] In at least one embodiment, a rack cooling manifold for one rack may be provided with side couplers to allow the rack to be coupled to a second rack for cross-rack coolant delivery. In at least one embodiment, a number of row-level manifolds may be retained to provide a coolant channel from a room area to at least one rack, from which the side couplers may distribute the coolant to other racks. Thus, in another aspect, the row-level manifolds retained in the present disclosure may also be made into a built-in structure, lining the top of the rack and having side-to-side couplers (for side coupling along the rack aisle) and / or back-to-back couplers (for back coupling) and / or front-to-front couplers (for front coupling) so that the racks can be coupled to each other in rows and / or columns to pass coolant through the couplers. In addition, each rack may include at least one side-to-side (or side) coupler, back-to-back (or rear) coupler, or front-to-front (or front) coupler to at least allow or return coolant. For example, coolant may enter the current rack from an adjacent rack through at least one coupler, but may exit to the row manifold instead of being provided to another adjacent rack.
[0050] In at least one embodiment, the coupler can be a rigid structure with a press-fit feature for coupling and can be installed within and extend from the rack cooling manifold of at least one rack. In doing so, overhead flexible ducting can be eliminated, further reducing the potential for leaks. Furthermore, the side couplers enable several racks to be placed adjacent to each other and eliminate the need for additional manifolds to distribute coolant within the primary or auxiliary cooling circuits. Mechanical and / or electronic controls can be used in conjunction with an intelligent learning system to correlate temperature and coolant flow rate within the rack or within the primary and / or auxiliary cooling circuits. Furthermore, sensors can be used to measure the pressure and quality of the coolant within the cooling circuits. The previously mentioned coupler can include a pressure relief path to provide pressure relief when the server tray is removed from the coupler, for example, to increase pressure on the delivery side of the cooling circuit. The coupler can include a check valve to prevent leaks during coupling or decoupling of the server tray from the rack cooling manifold. These measurements can also be used to design racks with embedded control features to prevent leaks and control coolant distribution, thereby meeting any specifications for cooling certain computing components with minimal reliability risk or maintenance requirements.
[0051] Utilizing the above features, the present disclosure implements a cooling system for racks with server trays without extensive mechanical piping, distribution, metering, and control requirements. These requirements typically require extensive modifications to the design of data center facilities and often require downtime, impacting reliability and cost structures to achieve adequate cooling of high-density racks. Furthermore, because the couplers are controllable, as described in further examples below, the present disclosure enables remote / intelligent control of each coupler on the manifold and server tray to precisely control the temperature, pressure, flow rate, and chemistry of the coolant.
[0052] In at least one embodiment, the present disclosure addresses the above-mentioned and other problems in a fully internal liquid cooling system for electrical and computing components in a liquid-cooled data center. The present disclosure implements an intelligent, integrated liquid-cooled rack for a data center. The present disclosure provides a cooling manifold with associated press-fit tubing for modular or pluggable cooling of server trays within a rack. Furthermore, the coupler enables a pluggable cold plate header to, for example, transfer cooling liquid or coolant to a heat transfer element attached to the liquid-cooled electronic components. Thus, the present disclosure implements a cooling system in which there are no visible liquid cooling tubes. Instead, flexible tubing can be integrated into the pluggable liquid cooling unit via a rack cooling manifold having a coupler thereon. A press-fit connection is provided via the rack cooling manifold to couple to a row manifold or a room manifold. Furthermore, a similar press-fit connection can be provided on the server tray to couple to the rack cooling manifold and provide coolant to computing components (such as GPUs, CPUs, switches, or other heat sinks). The server tray can be formed as a mezzanine module to which various liquid cooling heat sink modules are attached to further dissipate heat from the computing and electronic components. In another embodiment, the row, rack, and server cooling manifolds must be rigid to better support the requirements of the press-fit connection, for example, so that the manifold does not yield when a server tray is pushed onto the manifold to complete a, for example, press-fit connection.
[0053] Figure 1 1 is a block diagram of an example data center 100 having a cooling system adapted to meet the requirements of at least one embodiment. The data center 100 may comprise one or more rooms 102 having racks 110 and auxiliary equipment to house one or more servers on one or more server trays. The data center 100 is supported by a cooling tower 104 located outside the data center 100. The cooling tower 104 removes heat from within the data center 100 by acting on a primary cooling loop 106. Furthermore, a cooling distribution unit (CDU) 112 is provided between the primary cooling loop 106 and a secondary cooling loop 108 to enable heat to be extracted from the secondary cooling loop 108 to the primary cooling loop 106. In one aspect, the secondary cooling loop 108 can be connected to the server trays via various piping as needed. The loops 106 and 108 are shown as line diagrams, but one of ordinary skill in the art will recognize that one or more piping features may be used. In one example, flexible polyvinyl chloride (PVC) tubing may be used with associated piping to move fluid along each of the loops 106 and 108. In at least one embodiment, one or more coolant pumps may be used to maintain a pressure differential within the circuits 106 , 108 to enable coolant movement.
[0054] In at least one embodiment, the coolant in the primary cooling loop 106 and the secondary cooling loop 108 can be at least water and an additive, such as ethylene glycol or propylene glycol. During operation, each primary cooling loop and the secondary cooling loop have their own coolant. In one aspect, the coolant in the secondary cooling loop can be specific to the components in the server tray or rack 110. Therefore, when a component is disconnected from the rack 110, the coolant in the secondary cooling loop 108 may need to be changed. The CDU 112 enables precise control of the coolant in loops 106 and 108, either independently or simultaneously. For example, the CDU can be adapted to control the flow rate so that the coolant is appropriately distributed to extract heat generated within the rack 110. Furthermore, more flexible tubing 114 is provided from the secondary cooling loop 108 to enter each server tray and provide coolant to the electrical and / or computing components. In this disclosure, the terms electrical and / or computing components are used interchangeably to refer to heat-generating components that benefit from the present data center cooling system. The tubing 118 forming part of the secondary cooling loop 108 may be referred to as a room manifold. Additionally, tubing 116 extending from tubing 118 may also be part of the second cooling circuit 108, but may be referred to as a row manifold. Tubing 114 enters the rack as part of the second cooling circuit 108, but may be referred to as a rack cooling manifold. Additionally, row manifold 116 extends to all racks along a row in the data center 100. The piping of the second cooling circuit 108, including manifolds 118, 116, and 114, may be improved by at least one embodiment of the present disclosure. An optional chiller 120 may be provided in the main cooling circuit within the data center 102 to support cooling prior to the cooling tower. To the extent that an additional circuit is present in the main control circuit, a person of ordinary skill reading this disclosure will recognize that the additional circuit provides cooling external to the racks and external to the second cooling circuit; and may be used in conjunction with the main cooling circuit in the present disclosure.
[0055] In at least one embodiment, during operation, heat generated within the server trays of rack 110 can be transferred to coolant that exits rack 110 via the flexible tubing of row manifold 114 of secondary cooling circuit 108. Specifically, secondary coolant from CDU 112 (in secondary cooling circuit 108) used to cool rack 110 moves toward rack 110. Second coolant from CDU 112 is transferred from one side of the room manifold, which has tubing 118, to one side of rack 110 via row manifold 116 and passes through one side of the server trays via tubing 114. Spent secondary coolant (or exiting secondary coolant, carrying heat from computing components) exits from the other side of the server trays (e.g., after circulating through the server trays or components on the server trays, entering the left side of the rack and exiting the server trays from the right side of the rack). Spent secondary coolant exiting the server trays or rack 110 exits from a different side (e.g., the exit side) of tubing 114 and moves to a parallel, but also exit side, of row manifold 116. From the row manifold 116, the spent second coolant moves in a parallel section of the room manifold 118 in a direction opposite to the incoming second coolant (which may also be refreshed second coolant) and toward the CDU 112. In the CDU 112, the spent second coolant exchanges its heat with the primary coolant in the primary cooling loop 106. The spent second coolant is refreshed (e.g., relatively cooled when compared to the temperature of the spent second coolant stage) and is ready to be circulated back to the computing components through the secondary cooling loop 108. Various flow and temperature control features in the CDU 112 enable control of the amount of heat exchanged from the spent second coolant or the flow of the second coolant into and out of the CDU 112. The CDU 112 is also capable of controlling the flow of the primary coolant in the primary cooling loop 106. Therefore, it is possible that the refreshed second coolant may not be fully cooled to its default temperature characteristics before being circulated to the racks 110.
[0056] Figure 2 is a block diagram of another example data center 200 having an improved cooling system according to at least one embodiment, the cooling system including a rigid cooling manifold (in Figure 3A -D). As shown, Figure 1In contrast, the primary cooling loop 206 exchanges heat with secondary cooling loops 208A, 208B via the CDU 212. The secondary cooling loops 208A, 208B include an inlet side 208A, which provides a path for fresh or cooled coolant from the CDU 212 into the racks 210, and an outlet side 208B, which provides a path for spent coolant (heated by the computing components) from the racks 210 to the CDU 212. Once heat is transferred from the spent coolant to the primary cooling loop via the CDU 212, the primary coolant of the primary cooling loop 206 exits the room 202 to exchange heat with the environment, for example, via the cooling tower 204 (or cooling facility). Alternatively, a greater degree of forced cooling than in the CDU can be applied to cool the primary coolant in the cooling tower 204. Furthermore, while the second cooling loops 208A, 208B are shown entering the racks 210, persons of ordinary skill in the art who read this disclosure will recognize that the racks internally provide a closed loop for the second coolant of the second cooling loops 208A, 208B. Furthermore, the second cooling loops 208A, 208B can utilize side couplers, front couplers, and rear couplers to transfer the second coolant between racks (across rows or columns) without the need for overhead flexible ducting (e.g., pipes or other components). Figure 1 These ducts may experience leaks into the racks. This is illustrated by duct lines 214 between racks 210 in data center 200. Duct lines 214 allow adjacent racks across a row or aisle to receive a second coolant.
[0057] In at least one embodiment, during operation, heat generated within the server trays of rack 210 can be dissipated via rigid or flexible rack cooling manifolds (in Figure 3A -D) is passed to a second coolant exiting the rack 210 to a room manifold that forms part of the second cooling loop 208A, 208B. Specifically, the second coolant from the CDU 212 for cooling the rack 210 moves toward the rack 210. When a rack cooling manifold is used, the manifold is secured within the rack to prevent movement and to support the press fit connection from moving away from, for example, the pressure of the connection. The couplers are rigid and can be used outside the rack so that adjacent racks can be coupled together. Coolant from the CDU 212 is passed from the inlet side 208A of the room manifold 208 to one or more racks 210. The room manifold 208 can also be a rigid manifold with its own distribution room coupler that locks into place with the rack coupler associated with the rack. Additionally, in operation, when a server tray requires cooling, the server tray is inserted into the rack such that its mating coupler locks into place with the distribution coupler of the rigid rack cooling manifold, as Figures 3A-3DAs shown. The locking or engagement of the couplers releases the check valves of the mating coupler and the distribution coupler so that the second coolant from the rigid rack cooling manifold inlet immediately flows through the inlet mating coupler, the server tray, the outlet mating coupler, and the rigid rack cooling manifold outlet. The second coolant that exits the server tray or at the outlet mating coupler may be referred to as spent second coolant. The spent second coolant is heated by the heat of the components within the server tray. The second coolant from the rigid rack cooling manifold outlet travels through the outlet side 208A of the chamber manifold 208 and flows to the CDU 212. The engagement or locking of the couplers may cause an audible and tactile response to notify the server tray that coolant is receiving coolant. When the other computing component is a fluid-cooled component, then the coolant to the server tray can remain within the mating coupler of the server tray until further coupling and passage are provided for the coolant to travel through the fluid-cooled components and return out of the server tray.
[0058] In at least one embodiment of operation, the coolant in the secondary cooling loop 208 undergoes heat exchange at the CDU 212 so that heat from the computing components of the server trays can be transferred to the primary coolant of the primary cooling loop 206. Thus, as the primary coolant of the primary cooling loop 206 moves to the cooling tower 204, the primary coolant is heated and is considered spent primary coolant. This spent primary coolant of the primary cooling loop 208 exchanges its heat with the environment through air cooling or further liquid cooling at a much higher temperature than that applied at the CDU 212. Thus, the spent primary coolant is refreshed (e.g., cooled relative to its temperature at the spent secondary coolant stage) and is ready to be circulated back to the CDU 212. Similarly, the spent secondary coolant of the secondary cooling loop 208 is refreshed after heat exchange with the primary cooling loop 206 and is recirculated back to the racks. Various flow and temperature control features within the CDU 212 enable control of the amount of heat exchanged from the spent secondary coolant, or the rate of heat exchange (e.g., by controlling the flow rate of the spent secondary coolant relative to the flow rate of the primary coolant). Furthermore, the CDU 212 can also control the movement of the refreshed secondary coolant in and out of the CDU 212 to achieve slower or faster heat removal from the secondary coolant. Alternatively, or concurrently, it is possible to control the movement and temperature of the secondary coolant by controlling the flow of the primary coolant in the primary cooling loop 206. In at least one embodiment, faster circulation of the primary coolant removes more heat from the spent secondary coolant in the secondary cooling loop 208. Consequently, the refreshed secondary coolant may or may not be fully cooled to its default temperature characteristics before being circulated to the racks 210.
[0059] Figure 3A 、 Figure 3Bare different views of example racks 300A, 300B incorporating rigid cooling manifolds within the racks according to at least one embodiment. In at least one embodiment, Figure 3A Rack 300A and Figure 3B The racks 300A and 300B may be front and rear views of the same rack, respectively. However, to illustrate that different types of racks are expected to work with the present disclosure, the racks 300A and 300B may not show all of the same features in the rear and front views. For example, the mezzanine module 302 is not shown in the rear view rack 300B. Figure 3A 300B, but the rack may not require a door. Alternatively, to maintain a constant or appropriate reference temperature, an aspect of the present disclosure may use doors on both the front and rear ends of the rack to maintain the cooling effect. For example, the racks 300A; 300B may be completely enclosed, or may be a frame structure that is stable enough to support various inserted trays or panels. In at least one embodiment, tracks or guides 316 are provided to receive the trays or panels. In at least one embodiment, a first main coupler 308 is provided on the first rack cooling manifold 304A at the rear of the rack 300A to receive a second coolant; and a second main coupler 310 is provided on the second rack cooling manifold 304B to return the second coolant after the second coolant has been used (e.g., heat absorbed from the computing components of the server trays in the racks 300A; 300B).
[0060] like Figure 3A 、 3B As shown, in at least one embodiment, the rack cooling manifolds 304A, 304B are rigid and located inside the rack. In addition, the rack cooling manifolds 304A, 304B are adapted to include a plurality of distribution couplers toward the front of the rack. Because the rack cooling manifolds 304A, 304B are rigid, when the server tray (or cooling plate or mezzanine module) is inserted into the track or rails 316, the server tray's mating coupler locks or mates with the appropriate distribution coupler - for example, one from the coolant inlet side (e.g., rack cooling manifold 304A) and the other from the coolant outlet side (e.g., rack cooling manifold 304B). The locking or mating provides audible or tactile feedback to notify the server tray that it is ready for use. The second coolant travels through the rack cooling manifolds 304A (rack side cooling manifold), 304D (under-rack cooling manifold), and 304B (rack side cooling manifold). The secondary coolant may exit the racks via the second primary coupler 310 at the rear of the racks 300A; 300B. Additionally, other paths may be provided, including an on-rack cooling manifold 304C.
[0061] In at least one embodiment, a mezzanine module 302 can be provided to slide within a track or rail 316 at the topmost or any other level of the rack 300A; 300B. The mezzanine module can include a cold plate or can include further couplers to distribute the first coolant within the rack 300A; 300B. In at least one embodiment, the mezzanine module is the only module in the rack that receives the second coolant. In such an embodiment, the mezzanine module and its cold plate can be adapted to extract heat from the entire rack and transfer the heat out of the rack via the second coolant. Additionally, the first coupler 312 shown along the side rack cooling manifold 304A serves as a coolant inlet valve for the server trays coupled thereto. Once the server trays are coupled to the inlet valve on one side and the corresponding outlet valve on the other side, the second coupler 314 shown along the side rack cooling manifold 304B serves as an outlet valve for the server trays. Thus, in operation, the secondary coolant enters the rack 300A; 300B via the first primary coupler 308 , passes through the connected server trays, the bottom rack cooling manifold 304D and / or the mezzanine module 302 , and exits via the second primary coupler 310 .
[0062] Figure 3C 、 Figure 3D is a detailed view of a rigid cooling manifold 300C and a coupling 300D between the rigid cooling manifold and a tray or plate that can be coupled to the rigid cooling manifold in accordance with at least one embodiment. The rack cooling manifold 300C can be Figure 3A 、 Figure 3B The side rack cooling manifold, lower rack cooling manifold, or upper rack cooling manifold shown in the example rack 300A; 300B. Each of the manifolds 304A; 304B shown may include a distribution (as an inlet or outlet) coupler 312, 314. Also refer to Figure 3A 、 Figure 3B The example racks 300A; 300B of FIG. 308 and 310 show the main couplers 308A, 310A. The main couplers 308A, 310A are shown as twist couplers, but may be similar to Figure 3D The couplers shown are press fit couplers. The distribution couplers 312, 314 can be press fit with the mating couplers of the server tray or plate. In at least one embodiment, the manifolds 304A, 304B are fixed to Figure 3A 、 Figure 3B The rear portion of the example bracket 300A; 300B.
[0063] In at least one embodiment, Figure 3A 、 Figure 3BAs shown in the example brackets 300A; 300B, the cooling manifold is also adapted to remain in place when the tray or plate is inserted along the track or rail. When the mating coupler of the tray or plate is pressed against the distribution couplers 312, 314, a leak-proof press fit is obtained. The completion of the second cooling circuit then enables the passage of the second coolant from the main coupler inlet 308A to the distribution coupler inlet 312, through the server or tray that is press-fitted on a corresponding one of the distribution coupler inlets 312, out of the distribution coupler outlet 314, and out of the main coupler outlet 310A. In addition, the main coupler outlet 310A is shown as having a side coupler that can be used to provide a connection to an adjacent rack or to provide a connection to the main coupler inlet 308A via an upper rack cooling manifold. In doing so, additional flexible tubing between adjacent racks and the room-level manifold is eliminated. For example, additional rack cooling manifolds (e.g., manifolds 304C, D) can be used to prevent pressure buildup within the second cooling circuit.
[0064] Figure 3D 300D shows a coupling between a rigid cooling manifold and a tray or plate within a rack, according to at least one embodiment, as shown in FIG. Figure 2 as discussed. Couplers 308A, 310A, 312, 314, 318 may include built-in check features. This enables one-way flow of fluid so that no leakage occurs when, for example, coupling 300D is removed. This also enables coupling 300D to enable the first or common coolant to flow immediately from the row manifold into the rack cooling manifold and further to the server trays, cooling plates, or mezzanine modules. Alternatively, the coupler includes a shutoff valve 326 that can be activated to close the flow of the second coolant from the rigid cooling manifold through the appropriate coupler before the tray or plate is removed from the rack. The shutoff valve 326 is shown on coupler 308B, but such a shutoff valve may be provided for all couplers, including the main coupler and the distribution coupler. In at least one embodiment, Figure 3D A manifold 304B, 304B is shown having a main inlet coupler 308B (with a shutoff valve 326) and a main outlet coupler 310B. Additionally, the main couplers 308B, 310B differ from the main couplers 308A, 308B in having slightly different requirements than a press fit. In at least one embodiment, the main couplers 308B, 310B may require a twist fit because these couplers may receive coolant from an existing piping system. However, the types of couplers are supported by different shutoff valves that can be used to control the flow of fluid from the rigid cooling manifold to the server trays—at the inlet, outlet, or both ends. Thus, the present disclosure is able to modify existing racks to function within existing data center architectures.
[0065] In at least one embodiment, Figure 3DAlso shown are the mating couplers 318 of the tray, cooling plate, or mezzanine module 310. The tray, cooling plate, or mezzanine module 310 may include a narrow, enclosed channel 316 for carrying fluid in a loop from the inlet manifold 304A, from the inlet mating coupler 318A to the outlet mating coupler 318B, and ultimately exiting via the outlet manifold 304B. The enclosed channel enables the drip-free and quick-connect functionality of the present disclosure through a press-fit connection. When the server tray 310 is ready for use with the rack of the present disclosure, the server tray 310 is aligned in the tracks or rails of the rack mentioned earlier and pushed into the rack. The mating couplers 318A, 318B are also aligned with the inlet and outlet distribution couplers 312, 314 of the rack cooling manifolds 304A, 304B. Internal rubber or silicone seals and ribs (shown as dashed lines within mating couplers 318A, 318B) can lock behind notches on the press-fit inlet and outlet distribution couplers 312, 314. Check valves associated with each distribution coupler and mating coupler open, allowing coolant to flow into the server tray 310. Alternatively, as previously described, shutoff valves are opened or closed to allow coolant to flow from the secondary cooling circuit of the rigid cooling manifold to the server trays or panels 310.
[0066] Figure 3D The intelligent aspects of the present disclosure are also shown. For example, electric, remote and intelligent valve control assemblies 320, 322, 328, 332 can be used with shutoff valve 326 via connector 324 to intelligently open or close valve 326. In at least one embodiment, as Figure 3D As shown, the valve may include an internal valve disc controlled by a handle. The handle may be rotated by the smart valve control assembly 320, 322, 328, 332 via activation of the motor 328, via instructions from the controller 332 using communication 330 to rotate the arm 322, thereby moving the handle from an open position to a closed position so that the valve disc can engage or disengage to restrict or allow coolant to flow through the valve disc. Alternatively, the control device may be internal to the valve 326 and may not be visible on the surface of the valve, but causes the valve disc to move in a manner similar to that described above. Further, although the valve 326 and the smart valve control assembly 320, 322, 328, 332 are shown as a coupler 308B, this combination is applicable to all couplers 308A, 310A, 312, 314, 318, without limitation.
[0067] In at least one embodiment, the intelligent control of the present disclosure uses intelligent valve control components 320, 322, 328, 332, which have or are coupled to a combination of sensors and further use at least some known data that references the flow rate, pressure, and temperature associated with the coolant in many different areas in which the coolant flows. For example, valve 326 may include a sensor for sensing the pressure or temperature of the coolant within the valve. The insertion of a mating coupler into a distribution coupler having valve 326 may cause the temperature of the coolant within the valve to change due to the hot server tray now associated with the mating coupler. The temperature change may trigger valve 326 to open and may cause coolant to flow into the server tray. Through this process, the flow of coolant can be automatic when the server tray is press-fit into the rack (e.g., the mating coupler is press-fit onto the distribution coupler).
[0068] In at least one embodiment, further intelligent control provided by smart valve control components 320, 322, 328, and 332 is used to establish reference information between the flow rate of coolant entering certain areas or zones within the rack and the temperature within those areas or zones. This reference information can also include pressure information sensed within the areas or zones. This reference information can be stored and used to train a neural network to identify the required opening in the valve disc to achieve a coolant flow rate within a specific area or zone, thereby further reducing the temperature within the specific area or zone. In at least one embodiment, neurons with inputs and outputs can be used in single-layer, multi-layer, full-loop, simple-loop, or competitive neural networks. Inputs to the first layer of neurons can be provided with adjusted pressure and flow rates, as well as different corresponding valve opening angle values (which can be translated into, for example, linear openings for disc valves). The purpose of the adjustments is to allow the various pressure and flow rate values to be represented within a fixed range. Error values can be provided to neurons in hidden or intermediate layers based in part on results from previous data, which can be temperature adjustments obtained for the inputs provided to the first layer of neurons. For example, a first layer of neurons can provide an output using an input value that is a combination of pressure and flow rate.When a sensor associated with a region or zone indicates a temperature change in that region or zone, the resulting temperature adjustment can be applied.
[0069] In at least one embodiment, the intelligence of the present system can then provide input flow rates and associated pressures to valves in a zone or area to restrict or allow coolant flow into the rigid cooling manifold or server tray associated with that zone or area. In at least one embodiment, reference to zones or areas is made to illustrate that the rigid cooling manifold (at multiple points within the rigid cooling manifold), the server trays (and each server tray or other areas within the server tray), and the couplers can all be maintained at different temperatures using the present system. In at least one embodiment, a portion of a server tray may be heated more than a different portion because that portion has a CPU or GPU. Sensors and the intelligent control of the present disclosure can ensure that coolant of the appropriate temperature flows to that portion of the server tray that is hotter at the appropriate flow rate, for example, so that the different portion is not cooled unnecessarily. In at least one embodiment, this can be accomplished by increasing the flow rate of the coolant so that it reaches the portion that needs to be cooled immediately. With reference to at least Figure 14 and Figure 15 Further details are provided for the intelligent control provided to the cooling system of the present disclosure.
[0070] Figure 4 FIG2 is a diagram illustrating portions of an internal rack assembly 400 including example server trays or example cooling plates 402 and 432 and example computing components 412 and 416 for cooling computing components 412 and 416 using an example cooling system, according to at least one embodiment. In at least one embodiment, the computing component can be a circuit board or component 412 or a single computing component 416, such as a GPU, having a connector 426. The circuit board 412 can include an area 428 for attaching computing components, such as one or more GPUs or CPUs 416. Furthermore, the example server tray or example cooling plate includes one or more of a bottom 402 and a top 432. When the circuit assembly is liquid-cooled, inlet and outlet couplers 422 and 424 can be present on the additional computing component 416. The inlet and outlet couplers 422 and 424 enable distribution of coolant from a secondary cooling circuit via the inlet and outlet mating couplers 408 and 406 of the server tray or cooling plate 402. The additional computing component can be inserted and secured to the circuit board or component 412 within the provided area 428. The circuit board or component 412 itself may include an inlet coupler 420 and an outlet coupler 418 for receiving and returning coolant from the second cooling loop. Thus, additional computing components 416 may be coupled to the mating couplers 408, 406, or to the inlet or outlet couplers 420, 418 of the circuit board or component 412.
[0071] In at least one embodiment, the server tray or cooling plate at the bottom 402 or top 432 may include enclosed channels 404; 436 (also shown as Figure 3D) which can circulate through the tray or plate. The enclosed channel 436 at the top 432 is coupled to the rack cooling manifold via server tray couplers (or mating couplers) 442, 434. Alternatively, the enclosed channel 436 can be a flexible duct within the server tray that is coupled to the rack cooling manifold via server couplers 438, 440 and to the compute component couplers 422, 424. The enclosed channel can be on the top 432 only and can be supported by fans 430 on the top rather than on the bottom 402, but can also be present on both parts. The enclosed channel on either part of the server tray or cooling plate can provide cooling via a coupler on the top that is directly coupled to a coupler on the compute component or additional compute component without further tubing connected to the compute component or additional compute component. However, the enclosed channel and further coupling to the compute component and additional compute component can exist simultaneously. When the tray 402 is a cooling plate, a fan 430 can be present. The fan 430 can provide additional forced air circulation for the coolant. The tray includes a compartment 410 for receiving a connector 426 .
[0072] Figure 5 is usable for use or manufacture according to at least one embodiment Figure 2-17D Process 500 steps flow through a method for providing a cooling system. In at least one embodiment of process 500, subprocess 502 provides a rigid cooling manifold within a rack. Subprocess 504 provides coupling of a first controllable fluid coupler for the rigid cooling manifold to a cooling circuit external to the rack for receiving coolant. It may be determined, via subprocess 506, that the rack of server trays requires cooling using the cooling system. In at least one embodiment, when it is determined that a computing component has changed and the new computing component requires coolant or cooling of a heated area, the process may alternatively provide such cooling, via subprocess 508. In subprocess 508, one or more second controllable fluid couplers for the rigid cooling manifold are coupled to one or more rigid server cooling manifolds within the rack's server trays. The coupling in subprocess 508 enables coolant distribution within the rack. Process 500 can be used with any rack or server tray with proprietary computing components, at least as determined in subprocess 506. Alternatively, if the components have not changed, the existing cooling system of subprocess 504 may be retained. Alternatively, process 500 may be applied to any data center to adapt an existing cooling system for smart cooling.
[0073] Data Center
[0074] Figure 6A An example data center 600 is shown where data from Figure 2-5In at least one embodiment, the data center 600 includes a data center infrastructure layer 610, a framework layer 620, a software layer 630, and an application layer 640. In at least one embodiment, such as Figure 2 As described, features in components 204-214 can be performed within or in conjunction with the example data center 600. In at least one embodiment, the infrastructure layer 610, framework layer 620, software layer 630, and application layer 640 can be provided in part or in whole by computing components located on server trays in racks 210 of the data center 200. This enables the cooling system of the present disclosure to cool computing components in an effective and efficient manner, reducing the chance of leaks, and allowing computing components to be replaced without downtime due to the use of smart cooling features in the cooling system. In addition, various aspects of the data center, including the data center infrastructure layer 610, framework layer 620, software layer 630, and application layer 640, can be used to support at least the above references. Figure 3D Therefore, the reference Figures 6A-17D The discussion can be understood as applicable to implementing or supporting e.g. Figure 2 The hardware and software features required for intelligent control of the cooling system of the data center 200.
[0075] In at least one embodiment, Figure 6A As shown, the data center infrastructure layer 610 may include a resource coordinator 612, group computing resources 614, and node computing resources ("node CRs") 616(1)-616(N), where "N" represents any complete positive integer. In at least one embodiment, the node CRs 616(1)-616(N) may include, but are not limited to, any number of central processing units ("CPUs") or other processors (including accelerators, field programmable gate arrays (FPGAs), graphics processors, etc.), memory devices (e.g., dynamic read-only memories), storage devices (e.g., solid-state or disk drives), network input / output ("NW I / O") devices, network switches, virtual machines ("VMs"), power modules, and cooling modules. In at least one embodiment, one or more of the node CRs 616(1)-616(N) may be servers having one or more of the above-mentioned computing resources.
[0076] In at least one embodiment, the grouped computing resources 614 may include separate groups of node CRs housed in one or more racks (not shown), or may include many racks (also not shown) housed in data centers at various geographic locations. The separate groups of node CRs within the grouped computing resources 614 may include computing, networking, memory, or storage resources that may be configured or allocated to support groupings of one or more workloads. In at least one embodiment, several node CRs comprising a CPU or processor may be grouped in one or more racks to provide computing resources to support one or more workloads. In at least one embodiment, one or more racks may also include any number of power modules, cooling modules, and network switches, in any combination.
[0077] In at least one embodiment, resource coordinator 612 may configure or otherwise control one or more nodes CR 616(1)-616(N) and / or grouped computing resources 614. In at least one embodiment, resource coordinator 612 may comprise a software design infrastructure ("SDI") management entity for data center 600. In at least one embodiment, resource coordinator 612 may comprise hardware, software, or some combination thereof.
[0078] In at least one embodiment, Figure 6AAs shown, the framework layer 620 includes a job scheduler 622, a configuration manager 624, a resource manager 626, and a distributed file system 628. In at least one embodiment, the framework layer 620 may include a framework that supports software 632 of the software layer 630 and / or one or more applications 642 of the application layer 640. In at least one embodiment, the software 632 or the application 642 may include web-based service software or applications, such as those provided by Amazon Web Services, Google Cloud, and Microsoft Azure. In at least one embodiment, the framework layer 620 may include, but is not limited to, a free and open source software web application framework, such as Apache Spark™ (hereinafter referred to as "Spark"), which can utilize the distributed file system 628 for large-scale data processing (e.g., "big data"). In at least one embodiment, the job scheduler 622 may include a Spark driver to facilitate scheduling workloads supported by various layers of the data center 600. In at least one embodiment, the configuration manager 624 may be capable of configuring different layers, such as the software layer 630 and the framework layer 620, which includes Spark and a distributed file system 628 for supporting large-scale data processing. In at least one embodiment, the resource manager 626 can manage clustered or grouped computing resources that are mapped to or allocated to support the distributed file system 628 and the job scheduler 622. In at least one embodiment, the clustered or grouped computing resources can include grouped computing resources 614 on the data center infrastructure layer 610. In at least one embodiment, the resource manager 626 can coordinate with the resource coordinator 612 to manage these mapped or allocated computing resources.
[0079] In at least one embodiment, the software 632 included in the software layer 630 may include software used by at least a portion of the node CRs 616(1)-616(N), the grouped computing resources 614, and / or the distributed file system 628 of the framework layer 620. The one or more types of software may include, but are not limited to, Internet web page search software, email virus scanning software, database software, and streaming video content software.
[0080] In at least one embodiment, the one or more applications 642 included in the application layer 640 may include one or more types of applications used by at least a portion of the node CRs 616(1)-616(N), the grouped computing resources 614, and / or the distributed file system 628 of the framework layer 620. The one or more types of applications may include, but are not limited to, any number of genomics applications, cognitive computing, and machine learning applications, including training or inference software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), or other machine learning applications used in conjunction with one or more embodiments.
[0081] In at least one embodiment, any of configuration manager 624, resource manager 626, and resource coordinator 612 can implement any number and type of self-modification actions based on any number and type of data obtained in any technically feasible manner. In at least one embodiment, the self-modification actions can relieve a data center operator of data center 600 from making potentially poor configuration decisions and can avoid underutilized and / or poorly performing portions of the data center.
[0082] In at least one embodiment, data center 600 may include tools, services, software, or other resources to train one or more machine learning models or use one or more machine learning models to predict or infer information according to one or more embodiments herein. In at least one embodiment, a machine learning model can be trained by calculating weight parameters according to a neural network architecture using the software and computing resources described above with respect to data center 600. In at least one embodiment, using the weight parameters calculated using one or more training techniques herein, the resources described above with respect to data center 600 can be used to infer or predict information using a trained machine learning model corresponding to one or more neural networks. As previously discussed, deep learning techniques can be used to support intelligent control of valves in intelligent cooling systems. Any suitable learning network and computing capabilities of data center 600 can be used to facilitate deep learning. Thus, hardware in the data center can be used to simultaneously or concurrently support deep neural networks (DNNs), recurrent neural networks (RNNs), or convolutional neural networks (CNNs). For example, once a network is trained and successfully evaluated to identify data in a subset or slice, the trained network can provide similar representative data for use with the collected data.
[0083] In at least one embodiment, the data center 600 can use a CPU, an application-specific integrated circuit (ASIC), a GPU, an FPGA, or other hardware to use the above resources to perform training and / or reasoning. In addition, one or more of the above software and / or hardware resources can be configured as a service to allow users to train or perform information reasoning, such as pressure, flow rate, temperature, location information, or other artificial intelligence services.
[0084] Reasoning and training logic
[0085] Reasoning and / or training logic 615 may be used to perform reasoning and / or training operations associated with one or more embodiments. In at least one embodiment, reasoning and / or training logic 615 may be used to perform reasoning and / or training operations associated with one or more embodiments. Figure 6A Inference and / or training logic 615 is used to reason or predict operations based at least in part on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases herein. In at least one embodiment, reasoning and / or training logic 615 may include, but is not limited to, hardware logic in which computing resources are dedicated or otherwise uniquely used in conjunction with weight values or other information corresponding to one or more layers of neurons within a neural network. In at least one embodiment, reasoning and / or training logic 615 may be used in conjunction with an application specific integrated circuit (ASIC), such as an ASIC from Google. Processing unit from Graphcore TM Inference Processing Unit (IPU) from Intel Corp. (e.g. "LakeCrest") processor.
[0086] In at least one embodiment, the reasoning and / or training logic 615 can be used in conjunction with central processing unit (CPU) hardware, graphics processing unit (GPU) hardware, or other hardware (e.g., field programmable gate array (FPGA)). In at least one embodiment, the reasoning and / or training logic 615 includes, but is not limited to, a code and / or data storage model that can be used to store code (e.g., graphics code), weight values, and / or other information, including bias values, gradient information, momentum values, and / or other parameter or hyperparameter information. In at least one embodiment, each code and / or data storage module is associated with a dedicated computing resource. In at least one embodiment, the dedicated computing resource includes computing hardware that also includes one or more ALUs that perform mathematical functions (e.g., linear algebra functions) only on the information stored in the code and / or data storage module, and stores the results stored therefrom in an activated storage module of the reasoning and / or training logic 615.
[0087] Figure 6B 、 Figure 6CInference and / or training logic according to at least one embodiment is shown, such as in Figure 6A The reasoning and / or training logic 615 is used to perform reasoning and / or training operations associated with at least one embodiment of the present disclosure. Figure 6B and / or Figure 6C Provides details about the inference and / or training logic 615. Distinguished from the computational hardware 602, 606 by the use of an arithmetic logic unit (ALU) 610 Figure 6B and Figure 6C Inference and / or training logic 615. In at least one embodiment, each of computational hardware 602 and computational hardware 606 includes one or more ALUs that perform mathematical functions (e.g., linear algebraic functions) on information stored in code and / or data memory 601 and information stored in code and / or data memory 605, respectively, with the results stored in activation memory 620. Thus, unless otherwise specified, Figure 6B and Figure 6C may be substituted and used interchangeably.
[0088] In at least one embodiment, the inference and / or training logic 615 may include, but is not limited to, code and / or data storage 601 to store forward and / or output weights and / or input / output data and / or other parameters for neurons or layers of a neural network trained and / or used for inference in at least one embodiment. As previously described elsewhere in this disclosure, these layers may be used with a certain degree of compression. In at least one embodiment, the training logic 615 may include or be coupled to code and / or data storage 601 for storing graph code or other software to control timing and / or sequence, wherein weight and / or other parameter information is loaded to configure logic, including integer and / or floating point units (collectively referred to as arithmetic logic units (ALUs)). In at least one embodiment, the code (such as graph code) loads weight or other parameter information into the processor ALU based on the architecture of the neural network to which the code corresponds. In at least one embodiment, the code and / or data storage 601 stores input / output data and / or weight parameters during training and / or inference using aspects of at least one embodiment during forward propagation of the weight parameters for each layer of the neural network trained or used in conjunction with at least one embodiment. In at least one embodiment, any portion of code and / or data storage 601 may be included within other on-chip or off-chip data storage, including a processor's L1, L2, or L3 cache or system memory.
[0089] In at least one embodiment, any portion of code and / or data storage 601 may be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, code and / or code and / or data storage 601 may be cache memory, dynamically randomly addressable memory ("DRAM"), static randomly addressable memory ("SRAM"), non-volatile memory (e.g., flash memory), or other storage. In at least one embodiment, the choice of whether code and / or data storage 601 is internal or external to a processor, for example, or includes DRAM, SRAM, flash memory, or some other storage type, may depend on the available storage space on-chip or off-chip, the latency requirements of the training and / or inference functions being performed, the data batch size used in inference and / or training of the neural network, or some combination of these factors.
[0090] In at least one embodiment, the inference and / or training logic 615 may include, but is not limited to, code and / or data storage 605 for storing backpropagation and / or output weights and / or input / output data corresponding to neurons or layers of a neural network trained and / or used for inference in at least one embodiment. In at least one embodiment, the code and / or data storage 605 stores weight parameters and / or input / output data for each layer of the neural network trained or used in conjunction with at least one embodiment during backpropagation of input / output data and / or weight parameters during training and / or inference using at least one embodiment. In at least one embodiment, the training logic 615 may include or be coupled to code and / or data storage 605 for storing graph code or other software to control the timing and / or sequence in which weight and / or other parameter information is loaded to configure logic, which includes integer and / or floating point units (collectively referred to as arithmetic logic units (ALUs)).
[0091] In at least one embodiment, code (such as graph code) loads weights or other parameter information into the processor ALU based on the architecture of the neural network to which the code corresponds. In at least one embodiment, any portion of code and / or data storage 605 can be included with other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory. In at least one embodiment, any portion of code and / or data storage 605 can be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, code and / or data storage 605 can be cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other storage. In at least one embodiment, the choice of whether code and / or data storage 605 is internal or external to the processor, for example, including DRAM, SRAM, flash memory, or some other storage type, depends on whether the available storage is on-chip or off-chip, the latency requirements of the training and / or inference functions being performed, the batch size of data used in inference and / or training of the neural network, or some combination of these factors.
[0092] In at least one embodiment, code and / or data store 601 and code and / or data store 605 may be separate storage structures. In at least one embodiment, code and / or data store 601 and code and / or data store 605 may be the same storage structure. In at least one embodiment, code and / or data store 601 and code and / or data store 605 may be partially the same storage structure and partially separate storage structures. In at least one embodiment, any portion of code and / or data store 601 and code and / or data store 605 may be included with other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory.
[0093] In at least one embodiment, inference and / or training logic 615 may include, but is not limited to, one or more arithmetic logic units (“ALUs”) 610 , including integer and / or floating point units, for performing logical and / or mathematical operations based at least in part on or directed by training and / or inference code (e.g., graph code), the results of which may produce activations (e.g., output values from a layer or neuron within a neural network) stored in activation storage 620 , which are functions of input / output and / or weight parameter data stored in code and / or data storage 601 and / or code and / or data storage 605 . In at least one embodiment, activations are performed in response to executing instructions or other code, linear algebra and / or matrix-based mathematics performed by ALU 610 to generate activations stored in activation storage 620, wherein weight values stored in code and / or data storage 605 and / or in code and / or data storage 601 are used as operands with other values, such as bias values, gradient information, momentum values, or other parameters or hyperparameters, any or all of which may be stored in code and / or data storage 605 and / or code and / or data storage 601 or other on-chip or off-chip storage.
[0094] In at least one embodiment, one or more ALUs 610 are included in one or more processors or other hardware logic devices or circuits, while in another embodiment, one or more ALUs 610 may be external to the processor or other hardware logic devices or circuits that use them (e.g., coprocessors). In at least one embodiment, one or more ALUs 610 may be included within an execution unit of a processor, or otherwise included in a group of ALUs accessible by an execution unit of a processor, which may be within the same processor or distributed across different processors of different types (e.g., central processing units, graphics processing units, fixed function units, etc.). In at least one embodiment, code and / or data storage 601, code and / or data storage 605, and activation storage 620 may be on the same processor or other hardware logic device or circuit, while in another embodiment, they may be on different processors or other hardware logic devices or circuits, or some combination of the same and different processors or other hardware logic devices or circuits. In at least one embodiment, any portion of activation storage 620 may be included with other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory. Additionally, inference and / or training code may be stored with other code accessible to a processor or other hardware logic or circuitry and may be retrieved and / or processed using the processor's fetch, decode, schedule, execute, exit, and / or other logic circuitry.
[0095] In at least one embodiment, activation storage 620 may be cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other storage. In at least one embodiment, activation storage 620 may be completely or partially internal or external to one or more processors or other logic circuits. In at least one embodiment, whether activation storage 620 is internal or external to the processor, for example, or comprises DRAM, SRAM, flash memory, or other storage types, may be selected based on the storage available on or off chip, the latency requirements for performing training and / or inference functions, the batch size of data used in inferring and / or training neural networks, or some combination of these factors. In at least one embodiment, Figure 6B The inference and / or training logic 615 shown in FIG can be used in conjunction with an application specific integrated circuit (“ASIC”), such as the one from Google. Processing unit from Graphcore TM Inference Processing Unit (IPU) from Intel Corp. (e.g., "Lake Crest") processor. In at least one embodiment, Figure 6B The illustrated inference and / or training logic 615 may be used in conjunction with central processing unit (“CPU”) hardware, graphics processing unit (“GPU”) hardware, or other hardware such as a field programmable gate array (“FPGA”).
[0096] In at least one embodiment, Figure 6C Inference and / or training logic 615 is shown, which may include, but is not limited to, hardware logic, wherein computing resources are dedicated or otherwise uniquely used in conjunction with weight values or other information corresponding to one or more layers of neurons within a neural network, in accordance with at least one various embodiment. In at least one embodiment, Figure 6C The inference and / or training logic 615 shown in FIG can be used in conjunction with an application specific integrated circuit (ASIC), such as the ASIC from Google. Processing unit from Graphcore TM Inference Processing Unit (IPU) or from Intel Corp (e.g., "Lake Crest") processor. In at least one embodiment, Figure 6CThe inference and / or training logic 615 shown in can be used in conjunction with central processing unit (CPU) hardware, graphics processing unit (GPU) hardware, or other hardware such as a field programmable gate array (FPGA). In at least one embodiment, the inference and / or training logic 615 includes, but is not limited to, code and / or data storage 601 and code and / or data storage 605, which can be used to store code (e.g., graph code), weight values, and / or other information, including bias values, gradient information, momentum values, and / or other parameter or hyperparameter information. Figure 6C In at least one embodiment shown in FIG, code and / or data store 601 and code and / or data store 605 are each associated with dedicated computing resources (eg, computing hardware 602 and computing hardware 606), respectively.
[0097] In at least one embodiment, each of the code and / or data stores 601 and 605 and the corresponding computing hardware 602 and 606 corresponds to a different layer of a neural network, such that activations from one "storage / compute pair 601 / 602" of the code and / or data store 601 and computing hardware 602 are provided as input to the next "storage / compute pair 605 / 606" of the code and / or data store 605 and computing hardware 606, reflecting the conceptual organization of the neural network. In at least one embodiment, each storage / compute pair 601 / 602 and 605 / 606 can correspond to more than one neural network layer. In at least one embodiment, additional storage / compute pairs (not shown) can be included in the inference and / or training logic 615 after or in parallel with the storage / compute pairs 601 / 602 and 605 / 606.
[0098] Computer system
[0099] Figure 7A A block diagram of an exemplary computer system 700A is shown, which may be a system of interconnected devices and components, a system on a chip (SOC), or some combination thereof with a processor that may include an execution unit to execute instructions to support and / or implement the intelligent cooling system described herein, in accordance with at least one embodiment. In at least one embodiment, the computer system 700A may include, but is not limited to, components (such as a processor 702) to execute algorithms for processing data using execution units including logic in accordance with the present disclosure (such as the embodiments herein). In at least one embodiment, the computer system 700A may include a processor such as an Intel® processor available from Intel Corporation of Santa Clara, California. Processor family, XeonTM, XScaleTM and / or StrongARMTM, Core TM or Nervana TM microprocessor, but other systems may be used (including PCs with other microprocessors, engineering workstations, set-top boxes, etc.). In at least one embodiment, computer system 700B may execute a version of the WINDOWS operating system available from Microsoft Corporation of Redmond, Washington, although other operating systems (e.g., UNIX and Linux), embedded software, and / or graphical user interfaces may also be used.
[0100] In at least one embodiment, exemplary computer system 700A may incorporate components 110-116 (from Figure 1 ) to support processing aspects for the intelligent cooling system. For at least this reason, in one embodiment, Figure 7A The system is shown as comprising interconnected hardware devices or "chips", while in other embodiments, Figure 7A An exemplary system-on-chip SoC may be shown. In at least one embodiment, Figure 7A The devices shown in FIG. 7 may be interconnected with a proprietary interconnect, a standardized interconnect (e.g., PCIe), or some combination thereof. In at least one embodiment, one or more components of computer system 700B are interconnected using a Compute Express Link (CXL) interconnect. Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments, for example, as previously described with respect to FIG. Figure 6A -C discussed. Figure 6A -C provides details about the reasoning and / or training logic 615. In at least one embodiment, the reasoning and / or training logic 615 may be implemented in the system Figure 7A for performing inference or prediction operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures, or neural network use cases herein.
[0101] Embodiments may be used in other devices, such as handheld devices and embedded applications. Some examples of handheld devices include cellular phones, Internet Protocol (IP) devices, digital cameras, personal digital assistants ("PDAs"), and handheld PCs. In at least one embodiment, the embedded application may include a microcontroller, a digital signal processor ("DSP"), a system on a chip, a network computer ("NetPC"), a set-top box, a network hub, a wide area network ("WAN") switch, or any other system that can execute one or more instructions according to at least one embodiment.
[0102] In at least one embodiment, the computer system 700A may include, but is not limited to, a processor 702, which may include, but is not limited to, one or more execution units 708 to perform machine learning model training and / or inference according to the techniques described herein. In at least one embodiment, the computer system 700A is a single-processor desktop or server system, but in another embodiment, the computer system 700A may be a multi-processor system. In at least one embodiment, the processor 702 may include, but is not limited to, a complex instruction set computer ("CISC") microprocessor, a reduced instruction set computing ("RISC") microprocessor, a very long instruction word ("VLIW") microprocessor, a processor that implements an instruction set combination, or any other processor device, such as a digital signal processor. In at least one embodiment, the processor 702 may be coupled to a processor bus 710, which may transmit data signals between the processor 702 and other components in the computer system 700A.
[0103] In at least one embodiment, processor 702 may include, but is not limited to, level 1 ("L1") internal cache memory ("cache") 704. In at least one embodiment, processor 702 may have a single internal cache or multiple levels of internal cache. In at least one embodiment, the cache memory may reside external to processor 702. Other embodiments may include a combination of internal and external caches, depending on the specific implementation and requirements. In at least one embodiment, register file 706 may store different types of data in various registers, including, but not limited to, integer registers, floating point registers, status registers, and an instruction pointer register.
[0104] In at least one embodiment, an execution unit 708, including but not limited to logic to perform integer and floating point operations, is also located in the processor 702. In at least one embodiment, the processor 702 may also include a microcode ("ucode") read-only memory ("ROM") for storing microcode for certain macroinstructions. In at least one embodiment, the execution unit 708 may include logic for processing a packed instruction set 709. In at least one embodiment, by including the packed instruction set 709 in the instruction set of the general-purpose processor, and the associated circuitry to execute the instructions, operations used by many multimedia applications may be performed using packed data in the general-purpose processor 702. In one or more embodiments, many multimedia applications may be accelerated and executed more efficiently by using the full width of the processor's data bus to perform operations on the packed data, which may not require transferring smaller units of data across the processor's data bus to perform one or more operations one data element at a time.
[0105] In at least one embodiment, execution unit 708 may also be used in a microcontroller, an embedded processor, a graphics device, a DSP, and other types of logic circuits. In at least one embodiment, computer system 700A may include, but is not limited to, memory 720. In at least one embodiment, memory 720 may be implemented as a dynamic random access memory ("DRAM") device, a static random access memory ("SRAM") device, a flash memory device, or other storage device. In at least one embodiment, memory 720 may store instructions 719 and / or data 721 represented by data signals that may be executed by processor 702.
[0106] In at least one embodiment, a system logic chip can be coupled to the processor bus 710 and the memory 720. In at least one embodiment, the system logic chip can include, but is not limited to, a memory controller hub ("MCH") 716, and the processor 702 can communicate with the MCH 716 via the processor bus 710. In at least one embodiment, the MCH 716 can provide a high-bandwidth memory path 718 to the memory 720 for instruction and data storage, as well as for storage of graphics commands, data, and textures. In at least one embodiment, the MCH 716 can initiate data signals between the processor 702, the memory 720, and other components in the computer system 700A, and bridge data signals between the processor bus 710, the memory 720, and the system I / O 722. In at least one embodiment, the system logic chip can provide a graphics port for coupling to a graphics controller. In at least one embodiment, the MCH 716 may be coupled to the memory 720 via a high-bandwidth memory path 718 , and the graphics / video card 712 may be coupled to the MCH 716 via an Accelerated Graphics Port (“AGP”) interconnect 714 .
[0107] In at least one embodiment, computer system 700A may use system I / O 722 as a proprietary hub interface bus to couple MCH 716 to I / O controller hub ("ICH") 730. In at least one embodiment, ICH 730 may provide direct connection to certain I / O devices via a local I / O bus. In at least one embodiment, the local I / O bus may include, but is not limited to, a high-speed I / O bus used to connect peripheral devices to memory 720, chipset, and processor 702. Examples may include, but are not limited to, an audio controller 729, a firmware hub ("Flash BIOS") 728, a wireless transceiver 726, a data store 724, a legacy I / O controller 723 including a user input and keyboard interface, a serial expansion port 727 (e.g., a Universal Serial Bus (USB) port), and a network controller 734. The data store 724 may include a hard drive, a floppy disk drive, a CD-ROM device, a flash memory device, or other mass storage device.
[0108] Figure 7B is a block diagram illustrating an electronic device 700B for utilizing a processor 710 to support and / or enable the intelligent cooling system described herein, according to at least one embodiment. In at least one embodiment, the electronic device 700B may be, for example, but not limited to, a notebook computer, a tower server, a rack server, a blade server, a laptop computer, a desktop computer, a tablet computer, a mobile device, a phone, an embedded computer, or any other suitable electronic device. In at least one embodiment, the exemplary electronic device 700B may be combined with components 328, 332 (from Figure 3D ) to support processing aspects for an intelligent cooling system.
[0109] In at least one embodiment, system 700B may include, but is not limited to, a processor 710 communicatively coupled to any suitable number or kind of components, peripherals, modules, or devices. In at least one embodiment, processor 710 is coupled using a bus or interface, such as an I2C bus, a system management bus ("SMBus"), a low pin count (LPC) bus, a serial peripheral interface ("SPI"), a high-definition audio ("HDA") bus, a serial advanced technology attachment ("SATA") bus, a universal serial bus ("USB") (versions 1, 2, 3), or a universal asynchronous receiver / transmitter ("UART") bus. In at least one embodiment, Figure 7B shows a system comprising interconnected hardware devices or "chips", while in other embodiments, Figure 7B An exemplary system on a chip ("SoC") may be shown. In at least one embodiment, Figure 7BThe devices shown in can be interconnected with proprietary interconnects, standardized interconnects (e.g., PCIe), or some combination thereof. In at least one embodiment, Figure 7B One or more components of the system are interconnected using Compute Express Link (CXL) interconnect lines.
[0110] In at least one embodiment, Figure 7B The system may include a display 724, a touch screen 725, a touchpad 730, a near field communication unit ("NFC") 745, a sensor hub 740, a thermal sensor 746, a fast chipset ("EC") 735, a trusted platform module ("TPM") 738, a BIOS / firmware / flash memory ("BIOS, FW Flash") 722, a DSP 760, a drive 720 (e.g., a solid-state disk ("SSD") or a hard disk drive ("HDD")), a wireless local area network unit ("WLAN") 750, a Bluetooth unit 752, a wireless wide area network unit ("WWAN") 756, a global positioning system (GPS) unit 755, a camera ("USB 3.0 camera") 754 (e.g., a USB 3.0 camera), and / or a low-power double data rate ("LPDDR") memory unit ("LPDDR3") 715 implemented using, for example, the LPDDR3 standard. Each of these components may be implemented in any suitable manner.
[0111] In at least one embodiment, other components may be communicatively coupled to the processor 710 via the following components. In at least one embodiment, an accelerometer 741, an ambient light sensor ("ALS") 742, a compass 743, and a gyroscope 744 may be communicatively coupled to the sensor hub 740. In at least one embodiment, a thermal sensor 739, a fan 737, a keyboard 746, and a touchpad 730 may be communicatively coupled to the EC 735. In at least one embodiment, a speaker 763, an earpiece 764, and a microphone ("mic") 765 may be communicatively coupled to an audio unit ("audio codec and class-D amplifier") 762, which in turn may be communicatively coupled to the DSP 760. In at least one embodiment, the audio unit 764 may include, for example, but not limited to, an audio codec / decoder ("codec") and a class-D amplifier. In at least one embodiment, a SIM card ("SIM") 757 may be communicatively coupled to the WWAN unit 756. In at least one embodiment, components such as the WLAN unit 750 and the Bluetooth unit 752 and the WWAN unit 756 may be implemented as a next generation form factor (NGFF).
[0112] Reasoning and / or training logic 615 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 6B and / or Figure 6CProvides details about the reasoning and / or training logic 615. In at least one embodiment, the reasoning and / or training logic 615 may be implemented in the system Figure 7B for use in performing inference or prediction operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures, or neural network use cases herein.
[0113] Figure 7C A computer system 700C is shown, according to at least one embodiment, for supporting and / or implementing the intelligent cooling system described herein. In at least one embodiment, computer system 700C includes, but is not limited to, a computer 771 and a USB drive 770. In at least one embodiment, computer 771 may include, but is not limited to, any number and type of processors (not shown) and memory (not shown). In at least one embodiment, computer 771 includes, but is not limited to, a server, a cloud instance, a laptop computer, and a desktop computer.
[0114] In at least one embodiment, the USB disk 770 includes, but is not limited to, a processing unit 772, a USB interface 774, and USB interface logic 773. In at least one embodiment, the processing unit 772 can be any instruction execution system, device, or device capable of executing instructions. In at least one embodiment, the processing unit 772 can include, but is not limited to, any number and type of processing cores (not shown). In at least one embodiment, the processing unit or core 772 includes an application specific integrated circuit ("ASIC") that is optimized to perform any number and type of operations associated with machine learning. For example, in at least one embodiment, the processing core 772 is a tensor processing unit ("TPC") that is optimized to perform machine learning reasoning operations. In at least one embodiment, the processing core 772 is a vision processing unit ("VPU") that is optimized to perform machine vision and machine learning reasoning operations.
[0115] In at least one embodiment, USB interface 774 can be any type of USB connector or USB receptacle. For example, in at least one embodiment, USB interface 774 is a USB 3.0 Type-C receptacle for data and power. In at least one embodiment, USB interface 774 is a USB 3.0 Type-A connector. In at least one embodiment, USB interface logic 773 can include any number and type of logic that enables processing unit 772 to connect to a device (e.g., computer 771) via USB connector 774.
[0116] Reasoning and / or training logic 615 (e.g., regarding Figure 6B and Figure 6C ) is used to perform reasoning and / or training operations related to one or more embodiments. Figure 6B and Figure 6C Provides details about the reasoning and / or training logic 615. In at least one embodiment, the reasoning and / or training logic 615 may be used to Figure 7C In a system, an inference or prediction operation is performed based at least in part on weight parameters calculated using the neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0117] Figure 8 A further exemplary computer system 800 is shown for implementing the various processes and methods of the intelligent cooling system described throughout this disclosure, in accordance with at least one embodiment. In at least one embodiment, the computer system 800 includes, but is not limited to, at least one central processing unit ("CPU") 802 connected to a communication bus 810 implemented using any suitable protocol, such as PCI ("Peripheral Component Interconnect"), Peripheral Component Interconnect Express ("PCI-Express"), AGP ("Accelerated Graphics Port"), HyperTransport, or any other bus or point-to-point communication protocol. In at least one embodiment, the computer system 800 includes, but is not limited to, a main memory 804 and control logic (e.g., implemented as hardware, software, or a combination thereof), and data may be stored in the main memory 804 in the form of random access memory ("RAM"). In at least one embodiment, a network interface subsystem ("network interface") 822 provides an interface to other computing devices and networks for receiving data from the computer system 800 and transmitting data to other systems.
[0118] In at least one embodiment, computer system 800 includes, but is not limited to, input device 808, parallel processing system 812, and display device 806, which can be implemented using cathode ray tubes ("CRTs"), liquid crystal displays ("LCDs"), light emitting diodes ("LEDs"), plasma displays, or other suitable display technologies. In at least one embodiment, user input is received from input device 808 (such as a keyboard, mouse, touchpad, microphone, and more). In at least one embodiment, each of the aforementioned modules can be located on a single semiconductor platform to form a processing system.
[0119] Reasoning and / or training logic 615 is used to perform reasoning and / or training operations associated with one or more embodiments, such as those previously described with respect to Figure 6A -C discussed. The following combination Figure 6A -C provides details about the reasoning and / or training logic 615. In at least one embodiment, the reasoning and / or training logic 615 may be implemented in the system Figure 8In at least one embodiment, the inference and / or training logic 615 may be used in the system to perform inference or prediction operations based at least in part on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases herein. Figure 8 to perform inference or prediction operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures, or neural network use cases herein.
[0120] Figure 9A An exemplary architecture is shown in which multiple GPUs 910-913 are connected via a high-speed link
[0121] 905-906 (e.g., bus / point-to-point interconnect, etc.) are communicatively coupled to multiple multi-core processors 940-943. In one embodiment, high-speed links 940-943 support 4GB / s, 30GB / s, 80GB / s, or higher communication throughput. Various interconnect protocols can be used, including but not limited to PCIe 4.0 or 5.0 and NVLink 2.0.
[0122] Furthermore, in one embodiment, two or more GPUs 910-913 are interconnected via high-speed links 929-930, which may be implemented using the same or different protocols / links as used for high-speed links 940-943. Similarly, two or more multi-core processors 905-906 may be connected via high-speed link 928, which may be a symmetric multiprocessor (SMP) bus running at 20 GB / s, 30 GB / s, 120 GB / s, or higher. Alternatively, the same protocol / links may be used (e.g., via a common interconnect fabric) to accomplish this. Figure 9A All communications between the various system components shown in .
[0123] In one embodiment, each multi-core processor 905-906 is communicatively coupled to processor memory 901-902 via memory interconnects 926-927, respectively, and each GPU 910-913 is communicatively coupled to GPU memory 920-923 via GPU memory interconnects 950-953, respectively. Memory interconnects 926-927 and 950-953 can utilize the same or different memory access technologies. By way of example and not limitation, processor memory 901-902 and GPU memory 920-923 can be volatile memory, such as dynamic random access memory (DRAM) (including stacked DRAM), graphics DDR SDRAM (GDDR) (e.g., GDDR5, GDDR6), or high bandwidth memory (HBM), and / or can be non-volatile memory, such as 3D XPoint or Nano-Ram. In one embodiment, some portion of processor memory 901-902 can be volatile memory, while another portion can be non-volatile memory (e.g., using a two-level memory (2LM) hierarchy).
[0124] As described below, although the various processors 905-906 and GPUs 910-913 may each be physically coupled to a specific memory 901-902, 920-923, a unified memory architecture may be implemented in which a virtual system address space (also referred to as an "effective address" space) is distributed among the various physical memories. In at least one embodiment, the processor memories 901-902 may each include 64GB of system memory address space, and the GPU memories 920-923 may each include 32GB of system memory address space (resulting in a total addressable memory size of 256GB in this example).
[0125] As discussed elsewhere in this disclosure, at least one flow rate and pressure value is established for a first level intelligent learning system (such as a neural network system). Since the first level represents previous associated data, it also represents a smaller subset of data that can be used to improve the system by retraining the system. Testing and training can be performed in parallel using multiple processor units, making the intelligent learning system robust. Figure 9A When the intelligent learning system achieves convergence, the number of data points and the data in the data points that led to convergence are recorded. The data and data points can be used to control the intelligent cooling system, as shown in FIG. Figure 3D discussed.
[0126] Figure 9B907 and a graphics acceleration module 946. The graphics acceleration module 946 may include one or more GPU chips integrated on a line card that is coupled to the processor 907 via a high-speed link 940. Alternatively, the graphics acceleration module 946 may be integrated with the processor 907 in the same package or chip.
[0127] In at least one embodiment, the illustrated processor 907 includes a plurality of cores 960A-960D, each core having a translation lookaside buffer 961A-961D and one or more caches 962A-962D. In at least one embodiment, the cores 960A-960D may include various other components not shown for executing instructions and processing data. The caches 962A-962D may include level 1 (L1) and level 2 (L2) caches. In addition, one or more shared caches 956 may be included in the caches 962A-962D and shared by each group of cores 960A-960D. In at least one embodiment, one embodiment of the processor 907 includes 24 cores, each core having its own L1 cache, twelve shared L2 caches, and twelve shared L3 caches. In this embodiment, two adjacent cores share one or more L2 and L3 caches. The processor 907 and the graphics acceleration module 946 are connected to a system memory 914, which may include Figure 9A Processor memory 901-902 in.
[0128] Coherence is maintained for data and instructions stored in the various caches 962A-962D, 956 and system memory 914 via inter-core communication over the coherence bus 964. In at least one embodiment, each cache may have cache coherence logic / circuitry associated therewith to communicate over the coherence bus 964 in response to detecting a read or write to a particular cache line. In one implementation, a cache snooping protocol is implemented over the coherence bus 964 to snoop cache accesses.
[0129] In at least one embodiment, the proxy circuitry 925 communicatively couples the graphics acceleration module 946 to the coherence bus 964, thereby allowing the graphics acceleration module 946 to participate in a cache coherence protocol as a peer of the cores 960A-960D. In particular, in at least one embodiment, the interface 935 provides a connection to the proxy circuitry 925 via a high-speed link 940 (e.g., a PCIe bus, NVLink, etc.), and the interface 937 connects the graphics acceleration module 946 to the link 940.
[0130] In one implementation, the accelerator integrated circuit 936 provides cache management, memory access, context management, and interrupt management services on behalf of the multiple graphics processing engines 931, 932, N of the graphics acceleration module. The graphics processing engines 931, 932, N may each include a separate graphics processing unit (GPU). In at least one embodiment, the graphics processing engines 931, 932, N may selectively include different types of graphics processing engines within the GPU, such as a graphics execution unit, a media processing engine (e.g., a video encoder / decoder), a sampler, and a blit engine. In at least one embodiment, the graphics acceleration module 946 may be a GPU having multiple graphics processing engines 931-932, N, or the graphics processing engines 931-932, N may be individual GPUs integrated on a common package, line card, or chip. As the case may be, Figure 9B The above determination of the reconstruction parameters and the reconstruction algorithm is performed in the GPU 931-N.
[0131] In one embodiment, the accelerator integrated circuit 936 includes a memory management unit (MMU) 939 for performing various memory management functions, such as virtual to physical memory translation (also known as effective to real memory translation), and a memory access protocol for accessing system memory 914. The MMU 939 may also include a translation lookaside buffer ("TLB") (not shown) for caching virtual / effective to physical / real address translations. In one implementation, a cache 938 may store commands and data for efficient access by the graphics processing engines 931-932, N. In at least one embodiment, data stored in the cache 938 and graphics memory 933-934, M may be kept consistent with the core caches 962A-962D, 956 and system memory 914. As before, this task may be accomplished via proxy circuitry 925 acting on behalf of cache 938 and graphics memory 933-934, M (e.g., sending updates related to modifications / accesses of cache lines on processor caches 962A-962D, 956 to cache 938 and receiving updates from cache 938).
[0132] A set of registers 945 stores context data for threads executed by graphics processing engines 931-932, N, and context management circuitry 948 manages thread contexts. In at least one embodiment, context management circuitry 948 can perform save and restore operations to save and restore the context of each thread during a context switch (e.g., where a first thread is saved and a second thread is stored so that the second thread can be executed by the graphics processing engine). In at least one embodiment, context management circuitry 948 can store current register values to a designated area in memory (e.g., identified by a context pointer) upon a context switch. The register values can then be restored upon returning to context. In one embodiment, interrupt management circuitry 947 receives and processes interrupts received from system devices.
[0133] In one implementation, the MMU 939 converts virtual / effective addresses from the graphics processing engine 931 into real / physical addresses in the system memory 914. One embodiment of the accelerator integrated circuit 936 supports multiple (e.g., 4, 8, 16) graphics accelerator modules 946 and / or other accelerator devices. The graphics accelerator module 946 can be dedicated to a single application executing on the processor 907, or can be shared among multiple applications. In one embodiment, a virtualized graphics execution environment is presented in which the resources of the graphics processing engines 931-932, N are shared with multiple applications or virtual machines (VMs). In at least one embodiment, the resources can be subdivided into "slices" that are allocated to different VMs and / or applications based on processing requirements and priorities associated with the VMs and / or applications.
[0134] In at least one embodiment, the accelerator integrated circuit 936 acts as a bridge to the system for the graphics acceleration module 946 and provides address translation and system memory cache services. In addition, the accelerator integrated circuit 936 can provide virtualization facilities for the host processor to manage virtualization, interrupts, and memory management of the graphics processing engines 931-932, N.
[0135] Because the hardware resources of graphics processing engines 931-932, N are explicitly mapped into the real address space seen by host processor 907, any host processor can directly address these resources using effective address values. In at least one embodiment, one function of accelerator integrated circuit 936 is to physically separate graphics processing engines 931-932, N so that they appear as independent units to the system.
[0136] In at least one embodiment, one or more graphics memories 933-934, M are respectively coupled to each graphics processing engine 931-932, N. The graphics memories 933-934, M store instructions and data, which are processed by each graphics processing engine 931-932, N. The graphics memories 933-934, M can be volatile memory, such as DRAM (including stacked DRAM), GDDR memory (e.g., GDDR5, GDDR6), or HBM, and / or can be non-volatile memory, such as 3D XPoint or Nano-Ram.
[0137] In one embodiment, to reduce data traffic on link 940, a biasing technique is used to ensure that the data stored in graphics memories 933-934, M is the data most frequently used by graphics processing engines 931-932, N and is data that cores 960A-960D may not use (at least not frequently). Similarly, the biasing mechanism attempts to keep data needed by a core (and possibly not graphics processing engines 931-932, N) in caches 962A-962D, core 956, and system memory 914.
[0138] Figure 9C Another exemplary embodiment is shown in which an accelerator integrated circuit 936 is integrated into the processor 907 for enabling and / or supporting the intelligent cooling system according to at least one embodiment disclosed herein. In at least this embodiment, the graphics processing engines 931-932, N communicate directly with the accelerator integrated circuit 936 via the interface 937 and the interface 935 (again, any form of bus or interface protocol can be used) through the high-speed link 940. The accelerator integrated circuit 936 can perform operations related to Figure 9B The operations described above are similar to those described above. However, due to its close proximity to the coherence bus 964 and caches 962A-962D, 956, higher throughput is possible. At least one embodiment supports different programming models, including a dedicated process programming model (without graphics acceleration module virtualization) and a shared programming model (with virtualization). The programming models can include a programming model controlled by the accelerator integrated circuit 936 and a programming model controlled by the graphics acceleration module 946.
[0139] In at least one embodiment, the graphics processing engines 931-932, N are dedicated to a single application or process under a single operating system. In at least one embodiment, a single application can funnel other application requests to the graphics processing engines 931-932, N, thereby providing virtualization within a VM / partition.
[0140] In at least one embodiment, graphics processing engines 931-932,N can be shared by multiple VM / application partitions. In at least one embodiment, the sharing model can use a hypervisor to virtualize graphics processing engines 931-932,N to allow each operating system to access them. For a single-partition system without a hypervisor, the operating system owns graphics processing engines 931-932,N. In at least one embodiment, the operating system can virtualize graphics processing engines 931-932,N to provide access to each process or application.
[0141] In at least one embodiment, the graphics acceleration module 946 or individual graphics processing engines 931-932, N use a process handle to select a process element. In at least one embodiment, the process element is stored in the system memory 914 and can be addressed using the effective address to real address translation techniques described herein. In at least one embodiment, the process handle can be an implementation-specific value that is provided to the host process when registering its context with the graphics processing engine 931-932, N (i.e., calling system software to add the process element to the process element linked list). In at least one embodiment, the lower 16 bits of the process handle can be the offset of the process element in the process element linked list.
[0142] Figure 9D An exemplary accelerator integrated slice 990 for implementing and / or supporting an intelligent cooling system according to at least one embodiment disclosed herein is shown. As used herein, a "slice" comprises a designated portion of the processing resources of an accelerator integrated circuit 936. An application is an effective address space 982 in system memory 914 that stores process elements 983. In at least one embodiment, process elements 983 are stored in response to a GPU call 981 from an application 980 executing on a processor 907. Process elements 983 contain the process state of the corresponding application 980. A work descriptor (WD) 984 contained in process element 983 may be a single job requested by an application, or may contain a pointer to a job queue. In at least one embodiment, WD 984 is a pointer to a job request queue in the address space 982 of the application.
[0143] The graphics acceleration module 946 and / or the individual graphics processing engines 931-932, N can be shared by all processes or a subset of processes in the system. In at least one embodiment, an infrastructure for setting process state and sending WD 984 to the graphics acceleration module 946 to start a job in a virtualized environment can be included.
[0144] In at least one embodiment, a dedicated process programming model is implementation-specific. In this model, a single process owns either a graphics acceleration module 946 or an individual graphics processing engine 931. When a graphics acceleration module 946 is owned by a single process, the hypervisor initializes the accelerator integrated circuit for the owned partition, and when a graphics acceleration module 946 is assigned, the operating system initializes the accelerator integrated circuit 936 for the owned process.
[0145] In operation, the WD fetch unit 991 in the accelerator integrated slice 990 fetches the next WD 984, which includes an indication of work to be completed by one or more graphics processing engines of the graphics acceleration module 946. Data from the WD 984 can be stored in registers 945 and used by the MMU 939, interrupt management circuitry 947, and / or context management circuitry 948, as shown. In at least one embodiment, one embodiment of the MMU 939 includes segment / page roaming circuitry for accessing segment / page tables 986 within the OS virtual address space 985. The interrupt management circuitry 947 can process interrupt events 992 received from the graphics acceleration module 946. In at least one embodiment, when performing graphics operations, effective addresses 993 generated by the graphics processing engines 931-932, N are converted into real addresses by the MMU 939.
[0146] In one embodiment, the same register set 945 is replicated for each graphics processing engine 931-932, N, and / or graphics acceleration module 946, and the same register set 945 can be initialized by the hypervisor or operating system. Each of these replicated registers can be included in the accelerator integration slice 990. Example registers that can be initialized by the hypervisor are shown in Table 1.
[0147] Table 1 – Hypervisor Initialization Registers
[0148]
[0149]
[0150] Example registers that may be initialized by the operating system are shown in Table 2.
[0151] Table 2 – Operating System Initialization Registers
[0152] 1 Process and thread identification 2 Effective Address (EA) context save / restore pointer 3 Virtual Address (VA) Accelerator Utilizes Record Pointers 4 Virtual Address (VA) Segment Table Pointer 5 Permission blocking 6 Job Descriptor
[0153] In at least one embodiment, each WD 984 is specific to a particular graphics acceleration module 946 and / or graphics processing engine 931-932, N. It contains all the information necessary for the graphics processing engine 931-932, N to complete the work, or it may be a pointer to a memory location where the application has set up a command queue for the work to be done.
[0154] Figure 9E 1 shows additional details of an exemplary embodiment of a sharing model. This embodiment includes a hypervisor real address space 998 in which a process element list 999 is stored. The hypervisor real address space 998 can be accessed via the hypervisor 996, which virtualizes the graphics acceleration module engine for the operating system 995.
[0155] In at least one embodiment, the shared programming model allows all processes or a subset of processes from all partitions or a subset of partitions in the system to use the graphics acceleration module 946. There are two programming models in which the graphics acceleration module 946 is shared by multiple processes and partitions, namely, time-sliced sharing and graphics-directed sharing.
[0156] In this model, the hypervisor 996 owns the graphics acceleration module 946 and makes its functionality available to all operating systems 995. For the graphics acceleration module 946 to support virtualization through the hypervisor 996, the graphics acceleration module 946 may adhere to the following requirements: (1) the application's job requests must be autonomous (i.e., no state needs to be maintained between jobs), or the graphics acceleration module 946 must provide a context save and restore mechanism, (2) the graphics acceleration module 946 guarantees that the application's job requests are completed within a specified amount of time, including any transition errors, or the graphics acceleration module 946 provides the ability to preempt job processing, and (3) fairness between graphics acceleration module 946 processes must be ensured when operating in a directed shared programming model.
[0157] In one embodiment, an application 980 is required to make an operating system 995 system call using a graphics acceleration module type, a work descriptor (WD), an authority mask register (AMR) value, and a context save / restore region pointer (CSRP). The graphics acceleration module type describes the target acceleration function for the system call. In at least one embodiment, the graphics acceleration module type can be a system-specific value. In at least one embodiment, the WD is formatted specifically for the graphics acceleration module 946 and can take the form of a graphics acceleration module 946 command, an effective address pointer to a user-defined structure, an effective address pointer to a command queue, or any other data structure describing work to be performed by the graphics acceleration module 946. In at least one embodiment, the AMR value is the AMR state for the current process. In at least one embodiment, the value passed to the operating system is similar to that of an application setting the AMR. If the implementation of the accelerator integrated circuit 936 and graphics acceleration module 946 does not support the User Authority Mask Override Register (UAMOR), the operating system can apply the current UAMOR value to the AMR value before passing the AMR in the hypervisor call. In at least one embodiment, the hypervisor 996 can apply the current privilege mask overwrite register (AMOR) value before placing the AMR into the process element 983. In at least one embodiment, the CSRP is one of the registers 945 that contains the effective address of an area in the application's effective address space 982 for the graphics acceleration module 946 to save and restore context state. This pointer is used in at least one embodiment but is optional if state does not need to be saved between jobs or when a job is preempted. In at least one embodiment, the context save / restore area can be fixed system memory.
[0158] Upon receiving the system call, the operating system 995 may verify that the application 980 has been registered and granted permission to use the graphics acceleration module 946. The operating system 995 then calls the hypervisor 996 using the information shown in Table 3.
[0159] Table 3 – OS to Hypervisor call parameters
[0160] 1 Work Descriptor (WD) 2 Access Mask Register (AMR) value (potentially masked) 3 Effective Address (EA) Context Save / Restore Region Pointer (CSRP) 4 Process ID (PID) and optional thread ID (TID) 5 Virtual Address (VA) Accelerator Usage Record Pointer (AURP) 6 Virtual address of the storage segment table pointer (SSTP) 7 Logical Interrupt Service Number (LISN)
[0161] Upon receiving the hypervisor call, the hypervisor 996 verifies that the operating system 995 has registered and been granted permission to use the graphics acceleration module 946. The hypervisor 996 then places the process element 983 into a linked list of process elements of the corresponding graphics acceleration module 946 type. The process element may include the information shown in Table 4.
[0162] Table 4 – Process element information
[0163]
[0164] In at least one embodiment, the hypervisor initializes the plurality of accelerator integrated slice 990 registers 945 .
[0165] like Figure 9F As shown, in at least one embodiment, a unified memory is used that can be addressed via a common virtual memory address space for accessing physical processor memories 901-902 and GPU memories 920-923. In this implementation, operations executed on GPUs 910-913 utilize the same virtual / effective memory address space to access processor memories 901-902, and vice versa, thereby simplifying programmability. In one embodiment, a first portion of the virtual / effective address space is allocated to processor memory 901, a second portion is allocated to second processor memory 902, a third portion is allocated to GPU memory 920, and so on. In at least one embodiment, the entire virtual / effective memory space (sometimes referred to as the effective address space) is thus distributed across each of processor memories 901-902 and GPU memories 920-923, allowing any processor or GPU to access the memory using a virtual address mapped to any physical memory.
[0166] In one embodiment, bias / coherency management circuitry 994A-994E within one or more MMUs 939A-939E ensures cache coherency between the caches of one or more host processors (e.g., 905) and GPUs 910-913 and implements biasing techniques that indicate the physical memory where certain types of data should be stored. Figure 9F Multiple instances of bias / coherence management circuits 994A- 994E are shown in , but bias / coherence circuits may be implemented within an MMU of one or more host processors 905 and / or within an accelerator integrated circuit 936 .
[0167] One embodiment allows GPU-attached memory 920-923 to be mapped as part of system memory and accessed using shared virtual memory (SVM) technology, without suffering the performance drawbacks associated with full system cache coherence. In at least one embodiment, the ability to access GPU-attached memory 920-923 as system memory without heavy cache coherence overhead provides a favorable operating environment for GPU offloading. This arrangement allows host processor 905 software to set operands and access computation results without the overhead of traditional I / O DMA data copies. Such traditional copies include driver calls, interrupts, and memory-mapped I / O (MMIO) accesses, all of which are less efficient than simple memory accesses. In at least one embodiment, the ability to access GPU-attached memory 920-923 without cache coherence overhead can be critical to the execution time of offloaded computations. For example, in situations with heavy streaming write-to-memory traffic, cache coherence overhead can significantly reduce the effective write bandwidth seen by GPUs 910-913. In at least one embodiment, the efficiency of operand setup, result access, and GPU computation can play a role in determining the effectiveness of GPU offloading.
[0168] In at least one embodiment, the selection of GPU bias and host processor bias is driven by a bias tracker data structure. For example, a bias table can be used, which can be a page-granular structure (e.g., controlled at the granularity of a memory page) that includes 1 or 2 bits per GPU-attached memory page. In at least one embodiment, the bias table can be implemented in the stolen memory range of one or more GPU-attached memories 920-923, with or without a bias cache in the GPUs 910-913 (e.g., to cache frequently / recently used entries in the bias table). Alternatively, the entire bias table can be maintained within the GPU.
[0169] In at least one embodiment, before actually accessing the GPU memory, the bias table entry associated with each access to the GPU attached memory 920-923 is accessed, resulting in the following operations. Local requests from GPUs 910-913 whose pages are found in the GPU bias are forwarded directly to the corresponding GPU memory 920-923. Local requests from the GPU whose pages are found in the host bias are forwarded to the processor 905 (e.g., via a high-speed link as above). In one embodiment, a request from processor 905 to find the requested page in the host processor bias completes a request similar to a normal memory read. Alternatively, a request to a GPU biased page can be forwarded to the GPUs 910-913. In at least one embodiment, if the GPU is not currently using the page, the GPU can subsequently migrate the page to the host processor bias. In at least one embodiment, the bias state of a page can be changed by a software-based mechanism, a hardware-assisted software-based mechanism, or in limited cases by a purely hardware-based mechanism.
[0170] One mechanism for changing the bias state employs an API call (e.g., OpenCL), which in turn calls the GPU's device driver, which in turn sends a message (or queues a command descriptor) to the GPU, directing the GPU to change the bias state and, in some migrations, performs a cache flush operation in the host. In at least one embodiment, the cache flush operation is used for migrations from host processor 905 bias to GPU bias, but not for the reverse migration.
[0171] In one embodiment, cache coherence is maintained by temporarily rendering GPU-biased pages that cannot be cached by the host processor 905. To access these pages, the processor 905 may request access from the GPU 910, which may or may not immediately grant access. Therefore, to reduce communication between the processor 905 and the GPU 910, it is beneficial to ensure that the GPU-biased pages are the pages required by the GPU and not the host processor 905, and vice versa.
[0172] Reasoning and / or training logic 615 is used to execute one or more embodiments. Figure 6B and / or Figure 6C Details regarding the inference and / or training logic 615 are provided.
[0173] Figure 10AAn exemplary integrated circuit and associated graphics processor according to various embodiments herein are shown, which can be manufactured using one or more IP cores to support and / or enable an intelligent cooling system. In addition to the illustrations, other logic and circuitry can be included in at least one embodiment, including additional graphics processors / cores, peripheral interface controllers, or general-purpose processor cores.
[0174] Figure 10A 1 is a block diagram illustrating an exemplary system on a chip integrated circuit 1000A that may be fabricated using one or more IP cores according to at least one embodiment. In at least one embodiment, the integrated circuit 1000A includes one or more application processors 1005 (e.g., CPUs), at least one graphics processor 1010, and may additionally include an image processor 1015 and / or a video processor 1020, any of which may be modular IP cores. In at least one embodiment, the integrated circuit 1000A includes peripheral or bus logic including a USB controller 1025, a UART controller 1030, an SPI / SDIO controller 1035, and an I / O controller. 2 S / I 2 C controller 1040. In at least one embodiment, integrated circuit 1000A may include a display device 1045 coupled to one or more of a High Definition Multimedia Interface (HDMI) controller 1050 and a Mobile Industry Processor Interface (MIPI) display interface 1055. In at least one embodiment, storage may be provided by a flash memory subsystem 1060, including flash memory and a flash memory controller. In at least one embodiment, a memory interface may be provided via a memory controller 1065 for accessing SDRAM or SRAM memory devices. In at least one embodiment, some integrated circuits also include an embedded security engine 1070.
[0175] The reasoning and / or training logic 615 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 6B and / or Figure 6C Details are provided regarding inference and / or training logic 615. In at least one embodiment, inference and / or training logic 615 may be used in integrated circuit 1000A to infer or predict operations based at least in part on weight parameters computed using neural network training operations, neural network functionality and / or architecture, or neural network use cases described herein.
[0176] Figures 10B-10CAn exemplary integrated circuit and associated graphics processor according to various embodiments herein are shown, which can be manufactured using one or more IP cores to support and / or implement an intelligent cooling system. In addition to the illustrations, other logic and circuitry can be included in at least one embodiment, including additional graphics processors / cores, peripheral interface controllers, or general-purpose processor cores.
[0177] Figures 10B-10C is a block diagram illustrating an exemplary graphics processor for use within a SoC to support and / or implement an intelligent cooling system according to embodiments described herein. In one example, a graphics processor can be used to perform simulation tests of an intelligent cooling system because existing math engines are able to process multi-level neural networks more quickly. Figure 10B An exemplary graphics processor 1010 of a system on a chip integrated circuit is shown, which may be manufactured using one or more IP cores, in accordance with at least one embodiment. Figure 10C Another exemplary graphics processor 1040 of a system on a chip integrated circuit is shown, which may be manufactured using one or more IP cores, in accordance with at least one embodiment. Figure 10B The graphics processor 1010 is a low power graphics processor core. In at least one embodiment, Figure 10C The graphics processor 1040 is a higher performance graphics processor core. In at least one embodiment, each graphics processor 1010, 1040 can be Figure 10A A variant of the graphics processor 1010.
[0178] In at least one embodiment, the graphics processor 1010 includes a vertex processor 1005 and one or more fragment processors 1015A-1015N (e.g., 1015A, 1015B, 1015C, 1015D through 1015N-1 and 1015N). In at least one embodiment, the graphics processor 1010 can execute different shader programs via separate logic, such that the vertex processor 1005 is optimized to perform operations for the vertex shader program, while the one or more fragment processors 1015A-1015N perform fragment (e.g., pixel) shading operations for the fragment or pixel or shader program. In at least one embodiment, the vertex processor 1005 performs the vertex processing stage of the 3D graphics pipeline and generates primitives and vertex data. In at least one embodiment, the one or more fragment processors 1015A-1015N use the primitives and vertex data generated by the vertex processor 1005 to generate a frame buffer for display on a display device. In at least one embodiment, one or more fragment processors 1015A-1015N are optimized to execute fragment shader programs as provided in the OpenGL API, which can be used to perform similar operations as pixel shader programs provided in the Direct 3D API.
[0179] In at least one embodiment, graphics processor 1010 additionally includes one or more memory management units (MMUs) 1020A-1020B, one or more caches 1025A-1025B, and one or more circuit interconnects 1030A-1030B. In at least one embodiment, one or more MMUs 1020A-1020B provide a mapping of virtual to physical addresses for graphics processor 1010, including for vertex processor 1005 and / or fragment processors 1015A-1015N, which may reference vertex or image / texture data stored in memory, in addition to vertex or image / texture data stored in one or more caches 1025A-1025B. In at least one embodiment, one or more MMUs 1020A-1020B may synchronize with other MMUs within the system, including with other MMUs. Figure 10A One or more MMUs associated with one or more application processors 1005, graphics processor 1015, and / or video processor 1020 enable each processor 1005-1020 to participate in a shared or unified virtual memory system. In at least one embodiment, one or more circuit interconnects 1030A-1030B enable graphics processor 1010 to connect to other IP cores within the SoC via an internal bus of the SoC or via a direct connection.
[0180] In at least one embodiment, graphics processor 1040 includes Figure 10AOne or more MMUs 1020A-1020B, caches 1025A-1025B, and circuit interconnects 1030A-1030B of graphics processor 1010. In at least one embodiment, graphics processor 1040 includes one or more shader cores 1055A-1055N (e.g., 1055A, 1055B, 1055C, 1055D, 1055E, 1055F through 1055N-1 and 1055N), such as Figure 10B As shown, it provides a unified shader core architecture in which a single core or type or core can execute all types of programmable shader code, including shader program code for implementing vertex shaders, fragment shaders, and / or compute shaders. In at least one embodiment, the number of shader cores can vary. In at least one embodiment, the graphics processor 1040 includes an inter-core task manager 1045 that acts as a thread dispatcher to dispatch execution threads to one or more shader cores 1055A-1055N and a tiling unit 1058 to accelerate tile-based rendering operations, in which rendering operations of a scene are subdivided in image space, for example, to exploit local spatial coherence within a scene or to optimize the use of internal caches.
[0181] Reasoning and / or training logic 615 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 6B and / or Figure 6C Provides details about the inference and / or training logic 615. In at least one embodiment, the inference and / or training logic 615 may be implemented in an integrated circuit. Figure 10A and / or Figure 10B for performing inference or prediction operations based at least in part on weight parameters computed using a neural network training operation, a neural network function or architecture, or a neural network use case herein.
[0182] Figures 10D-10E Additional exemplary graphics processor logic is shown to support and / or implement an intelligent cooling system according to embodiments described herein. In at least one embodiment, Figure 10D Shows that can be included in Figure 10A Graphics core 1000D within graphics processor 1010, and in at least one embodiment, may be such as Figure 10C Unified shader cores 1055A-1055N are shown. Figure 10B A highly parallel general purpose graphics processing unit ("GPGPU") 1030 suitable for deployment on a multi-chip module in at least one embodiment is shown.
[0183] In at least one embodiment, graphics core 1000D may include multiple slices 1001A-1001N, or partitions of each core, and a graphics processor may include multiple instances of graphics core 1000D. In at least one embodiment, slices 1001A-1001N may include support logic including local instruction caches 1004A-1004N, thread schedulers 1006A-1006N, thread dispatchers 1008A-1008N, and a set of registers 1010A-1010N. In at least one embodiment, slices 1001A-1001N may include a set of additional function units (AFUs 1012A-1012N), floating point units (FPUs 1014A-1014N), integer arithmetic logic units (ALUs 109A-109N), address calculation units (ACUs 1013A-1013N), double precision floating point units (DPFPUs 1015A-1015N), and matrix processing units (MPUs 1017A-1017N).
[0184] In at least one embodiment, the FPUs 1014A-1014N can perform single-precision (32-bit) and half-precision (16-bit) floating-point operations, while the DPFPUs 1015A-1015N perform double-precision (64-bit) floating-point operations. In at least one embodiment, the ALUs 1016A-1016N can perform variable-precision integer operations with 8-bit, 16-bit, and 32-bit precision, and can be configured for mixed-precision operations. In at least one embodiment, the MPUs 1017A-1017N can also be configured for mixed-precision matrix operations, including half-precision floating-point operations and 8-bit integer operations. In at least one embodiment, the MPUs 1017A-1010N can perform various matrix operations to accelerate machine learning application frameworks, including enabling support for accelerated general matrix-to-matrix multiplication (GEMM). In at least one embodiment, the AFUs 1012A-1012N can perform additional logical operations not supported by the floating-point or integer units, including trigonometric operations (e.g., sine, cosine, etc.).
[0185] As discussed elsewhere in this disclosure, the inference and / or training logic 615 (at least in Figure 6B 、 Figure 6C Reference) is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 6B and / or Figure 6C Details are provided regarding inference and / or training logic 615. In at least one embodiment, inference and / or training logic 615 may be used in graphics core 1000D to perform inference or prediction operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures, or neural network use cases herein.
[0186] Figure 11A A block diagram of a computer system 1100A according to at least one embodiment is shown. In at least one embodiment, computer system 1100A includes a processing subsystem 1101 having one or more processors 1102 and a system memory 1104, which communicates via an interconnect path that may include a memory hub 1105. In at least one embodiment, memory hub 1105 may be a separate component within a chipset assembly or may be integrated within one or more processors 1102. In at least one embodiment, memory hub 1105 is coupled to an I / O subsystem 1111 via a communication link 1106. In one embodiment, I / O subsystem 1111 includes an I / O hub 1107, which enables computer system 1100A to receive input from one or more input devices 1108. In at least one embodiment, I / O hub 1107 enables a display controller, which may be included in one or more processors 1102, to provide output to one or more display devices 1110A. In at least one embodiment, the one or more display devices 1110A coupled to the I / O hub 1107 may include local, internal, or embedded display devices.
[0187] In at least one embodiment, the processing subsystem 1101 includes one or more parallel processors 1112 coupled to the memory hub 1105 via a bus or other communication link 1113. In at least one embodiment, the communication link 1113 can use any of a number of standard-based communication link technologies or protocols, such as, but not limited to, PCI Express, or can be a vendor-specific communication interface or communication structure. In at least one embodiment, the one or more parallel processors 1112 form a parallel or vector processing system in a computational cluster, which can include a large number of processing cores and / or processing clusters, such as a multi-integrated core (MIC) processor. In at least one embodiment, the one or more parallel processors 1112 form a graphics processing subsystem that can output pixels to one of one or more display devices 1110A coupled via the I / O hub 1107. In at least one embodiment, the one or more parallel processors 1112 can also include a display controller and display interface (not shown) to enable direct connection to the one or more display devices 1110B.
[0188] In at least one embodiment, a system storage unit 1114 can be connected to the I / O hub 1107 to provide a storage mechanism for the computer system 1100A. In at least one embodiment, an I / O switch 1116 can be used to provide an interface mechanism to enable connections between the I / O hub 1107 and other components, such as a network adapter 1118 and / or a wireless network adapter 1119 that can be integrated into one or more platforms, as well as various other devices that can be added via one or more add-on devices 1120. In at least one embodiment, the network adapter 1118 can be an Ethernet adapter or another wired network adapter. In at least one embodiment, the wireless network adapter 1119 can include one or more of Wi-Fi, Bluetooth, near field communication (NFC), or other network devices including one or more radio devices.
[0189] In at least one embodiment, the computer system 1100A may include other components not explicitly shown, including USB or other port connections, optical storage drives, video capture devices, and the like, which may also be connected to the I / O hub 1107. In at least one embodiment, the interconnection may be implemented using any suitable protocol, such as a PCI (Peripheral Component Interconnect)-based protocol (e.g., PCI-Express) or other bus or point-to-point communication interface and / or protocol. Figure 11A Communication paths for various components in a chip, such as NV-Link high-speed interconnect or interconnect protocol.
[0190] In at least one embodiment, one or more parallel processors 1112 include circuits optimized for graphics and video processing, including, for example, video output circuitry, and constitute a graphics processing unit (GPU). In at least one embodiment, one or more parallel processors 1112 include circuits optimized for general-purpose processing. In at least one embodiment, the components of computer system 1100A can be integrated with one or more other system elements on a single integrated circuit. For example, in at least one embodiment, one or more parallel processors 1112, memory hub 1105, one or more processors 1102, and I / O hub 1107 can be integrated into a system-on-chip (SoC) integrated circuit. In at least one embodiment, the components of computer system 1100A can be integrated into a single package to form a system-in-package (SIP) configuration. In at least one embodiment, at least a portion of the components of computer system 1100A can be integrated into a multi-chip module (MCM), which can be interconnected with other multi-chip modules to form a modular computer system.
[0191] Reasoning and / or training logic 615 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 6B and / or Figure 6C Provides details about the reasoning and / or training logic 615. In at least one embodiment, the reasoning and / or training logic 615 may be Figure 11A for use in a system for performing inference or prediction operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures, or neural network use cases herein.
[0192] processor
[0193] Figure 11B FIG1 shows a parallel processor 1100B according to at least one embodiment. In at least one embodiment, the various components of the parallel processor 1100B may be implemented using one or more integrated circuit devices, such as a programmable processor, an application specific integrated circuit (ASIC), or a field programmable gate array (FPGA). In at least one embodiment, the parallel processor 1100B shown is a processor according to an exemplary embodiment. Figure 11B A variation of the one or more parallel processors 1112 is shown.
[0194] In at least one embodiment, parallel processor 1100B includes parallel processing unit 1102. In at least one embodiment, parallel processing unit 1102 includes an I / O unit 1104 that enables communication with other devices, including other instances of parallel processing unit 1102. In at least one embodiment, I / O unit 1104 can be directly connected to other devices. In at least one embodiment, I / O unit 1104 connects to other devices using a hub or switch interface (e.g., memory hub 1105). In at least one embodiment, the connection between memory hub 1105 and I / O unit 1104 forms a communication link 1113. In at least one embodiment, I / O unit 1104 is connected to a host interface 1106 and a memory crossbar switch 1116, where host interface 1106 receives commands for performing processing operations and memory crossbar switch 1116 receives commands for performing memory operations.
[0195] In at least one embodiment, when host interface 1106 receives command buffers via I / O unit 1104, host interface 1106 can direct work operations to execute those commands to front end 1108. In at least one embodiment, front end 1108 is coupled to scheduler 1110, which is configured to distribute commands or other work items to processing cluster array 1112. In at least one embodiment, scheduler 1110 ensures that processing cluster array 1112 is properly configured and in a valid state before distributing tasks to processing cluster array 1112. In at least one embodiment, scheduler 1110 is implemented by firmware logic executing on a microcontroller. In at least one embodiment, a microcontroller-implemented scheduler 1110 can be configured to perform complex scheduling and work distribution operations at both coarse and fine granularity, thereby enabling fast preemption and context switching of threads executing on processing array 1112. In at least one embodiment, host software can authenticate workloads for scheduling on processing array 1112 through one of multiple graphics processing doorbells. In at least one embodiment, the workload may then be automatically distributed across the processing array 1112 by scheduler 1110 logic within a microcontroller that includes scheduler 1110 .
[0196] In at least one embodiment, processing cluster array 1112 may include up to "N" processing clusters (e.g., cluster 1114A, cluster 1114B, through cluster 1114N). In at least one embodiment, each cluster 1114A-1114N of processing cluster array 1112 may execute a large number of concurrent threads. In at least one embodiment, scheduler 1110 may allocate work to clusters 1114A-1114N of processing cluster array 1112 using various scheduling and / or work distribution algorithms, which may vary depending on the workload generated by each program or computation type. In at least one embodiment, scheduling may be handled dynamically by scheduler 1110 or may be assisted in part by compiler logic during the compilation of program logic configured to be executed by processing cluster array 1112. In at least one embodiment, different clusters 1114A-1114N of processing cluster array 1112 may be assigned to process different types of programs or to perform different types of computations.
[0197] In at least one embodiment, processing cluster array 1112 can be configured to perform various types of parallel processing operations. In at least one embodiment, processing cluster array 1112 can be configured to perform general-purpose parallel computing operations. In at least one embodiment, processing cluster array 1112 can include logic to perform processing tasks including filtering video and / or audio data, performing modeling operations including physics operations, and performing data transformations.
[0198] In at least one embodiment, processing cluster array 1112 is configured to perform parallel graphics processing operations. In at least one embodiment, processing cluster array 1112 may include additional logic to support the execution of such graphics processing operations, including but not limited to texture sampling logic for performing texture operations, as well as tessellation logic and other vertex processing logic. In at least one embodiment, processing cluster array 1112 may be configured to execute shader programs related to graphics processing, such as, but not limited to, vertex shaders, tessellation shaders, geometry shaders, and pixel shaders. In at least one embodiment, parallel processing units 1102 may transfer data from system memory via I / O units 1104 for processing. In at least one embodiment, during processing, the transferred data may be stored in on-chip memory (e.g., parallel processor memory 1122) during processing and then written back to system memory.
[0199] In at least one embodiment, when parallel processing unit 1102 is used to perform graphics processing, scheduler 1110 can be configured to divide the processing workload into tasks of approximately equal size to better distribute graphics processing operations to multiple clusters 1114A-1114N of processing cluster array 1112. In at least one embodiment, portions of processing cluster array 1112 can be configured to perform different types of processing. In at least one embodiment, a first portion can be configured to perform vertex shading and topology generation, a second portion can be configured to perform tessellation and geometry shading, and a third portion can be configured to perform pixel shading or other screen-space operations to generate rendered images for display (if needed to simulate valve control of an intelligent cooling system). In at least one embodiment, intermediate data generated by one or more of clusters 1114A-1114N can be stored in a buffer to allow the intermediate data to be transferred between clusters 1114A-1114N for further processing.
[0200] In at least one embodiment, the processing cluster array 1112 can receive processing tasks to be executed via the scheduler 1110, which receives commands defining the processing tasks from the front end 1108. In at least one embodiment, the processing tasks can include an index of data to be processed, such as surface (patch) data, primitive data, vertex data, and / or pixel data, as well as state parameters and commands defining how to process the data (e.g., what program to execute). In at least one embodiment, the scheduler 1110 can be configured to obtain an index corresponding to a task, or can receive the index from the front end 1108. In at least one embodiment, the front end 1108 can be configured to ensure that the processing cluster array 1112 is configured in a valid state before starting a workload specified by an incoming command buffer (e.g., a batch buffer, a push buffer, etc.).
[0201] In at least one embodiment, each of one or more instances of parallel processing unit 1102 can be coupled to parallel processor memory 1122. In at least one embodiment, parallel processor memory 1122 can be accessed via memory crossbar 1116, which can receive memory requests from processing cluster array 1112 and I / O unit 1104. In at least one embodiment, memory crossbar 1116 can access parallel processor memory 1122 via memory interface 1118. In at least one embodiment, memory interface 1118 can include a plurality of partition units (e.g., partition unit 1120A, partition unit 1120B, through partition unit 1120N), which can each be coupled to a portion (e.g., a memory unit) of parallel processor memory 1122. In at least one embodiment, the plurality of partition units 1120A-1120N are configured to be equal to the number of memory cells, such that the first partition unit 1120A has a corresponding first memory cell 1124A, the second partition unit 1120B has a corresponding memory cell 1124B, and the Nth partition unit 1120N has a corresponding Nth memory cell 1124N. In at least one embodiment, the number of partition units 1120A-1120N may not be equal to the number of memory devices.
[0202] In at least one embodiment, memory units 1124A-1124N may include various types of memory devices, including dynamic random access memory (DRAM) or graphics random access memory, such as synchronous graphics random access memory (SGRAM), including graphics double data rate (GDDR) memory. In at least one embodiment, memory units 1124A-1124N may also include 3D stacked memory, including but not limited to high bandwidth memory (HBM). In at least one embodiment, render targets such as frame buffers or texture maps may be stored across memory units 1124A-1124N, allowing partition units 1120A-1120N to write portions of each render target in parallel to efficiently use the available bandwidth of parallel processor memory 1122. In at least one embodiment, local instances of parallel processor memory 1122 may be eliminated in favor of a unified memory design utilizing system memory in combination with local cache memory.
[0203] In at least one embodiment, any of the clusters 1114A-1114N in the processing cluster array 1112 can process data to be written to any memory unit 1124A-1124N within the parallel processor memory 1122. In at least one embodiment, the memory crossbar 1116 can be configured to transmit the output of each cluster 1114A-1114N to any partition unit 1120A-1120N or another cluster 1114A-1114N, which can perform other processing operations on the output. In at least one embodiment, each cluster 1114A-1114N can communicate with a memory interface 1118 via the memory crossbar 1116 to read from or write to various external storage devices. In at least one embodiment, memory crossbar 1116 has connections to memory interface 1118 for communicating with I / O unit 1104, as well as connections to local instances of parallel processor memory 1102, thereby enabling processing units within different processing clusters 1114A-1114N to communicate with system memory or other memory that is not local to parallel processing unit 1102. In at least one embodiment, memory crossbar 1116 may use virtual channels to separate traffic flows between clusters 1114A-1114N and partition units 1120A-1120N.
[0204] In at least one embodiment, multiple instances of parallel processing unit 1102 can be provided on a single plug-in card, or multiple plug-in cards can be interconnected. In at least one embodiment, different instances of parallel processing unit 1102 can be configured to interoperate with each other, even if the different instances have different numbers of processing cores, different amounts of local parallel processor memory, and / or other configuration differences. In at least one embodiment, some instances of parallel processing unit 1102 can include higher precision floating point units relative to other instances. In at least one embodiment, a system incorporating one or more instances of parallel processing unit 1102 or parallel processor 1100B can be implemented in a variety of configurations and form factors, including but not limited to desktop, laptop, or handheld personal computers, servers, workstations, game consoles, and / or embedded systems.
[0205] Figure 11C is a block diagram of a partition unit 1120 according to at least one embodiment. In at least one embodiment, the partition unit 1120 is Figure 11B 1120N。In at least one embodiment, the partition unit 1120 includes an L2 cache 1121, a frame buffer interface 1125 and an ROP 1126 (raster operation unit). The L2 cache 1121 is a read / write cache that is configured to perform load and store operations received from the memory crossbar switch 1116 and the ROP 1126. In at least one embodiment, the L2 cache 1121 outputs read misses and urgent write back requests to the frame buffer interface 1125 for processing. In at least one embodiment, updates can also be sent to the frame buffer via the frame buffer interface 1125 for processing. In at least one embodiment, the frame buffer interface 1125 communicates with memory units in the parallel processor memory (such as Figure 11B interacts with one of the memory units 1124A-1124N (e.g., within parallel processor memory 1122).
[0206] In at least one embodiment, ROP 1126 is a processing unit that performs raster operations such as stenciling, z-testing, blending, and the like. In at least one embodiment, ROP 1126 then outputs processed graphics data that is stored in graphics memory. In at least one embodiment, ROP 1126 includes compression logic to compress depth or color data written to memory and decompress depth or color data read from memory. In at least one embodiment, the compression logic can be lossless compression logic that utilizes one or more of a variety of compression algorithms. The compression logic implemented by ROP 1126 can vary based on the statistical characteristics of the data to be compressed. In at least one embodiment, incremental color compression is performed based on the depth and color data on a per-tile basis.
[0207] In at least one embodiment, ROP 1126 is included within each processing cluster (e.g., Figure 11B In at least one embodiment, read and write requests for pixel data are transmitted through the memory crossbar 1116 rather than through the pixel fragment data transfer. In at least one embodiment, the processed graphics data can be displayed on a display device such as a Figure 11A 110), is displayed on one of the one or more display devices 1110, is routed by the processor 1102 for further processing, or is displayed on one of the one or more display devices 1110 Figure 11B One of the processing entities within parallel processor 1100B is routed for further processing.
[0208] Figure 11D is a block diagram of a processing cluster 1114 within a parallel processing unit according to at least one embodiment. In at least one embodiment, a processing cluster is Figure 11B In at least one embodiment, one or more processing clusters 1114 can be configured to execute many threads in parallel, where a "thread" refers to an instance of a particular program executed on a particular set of input data. In at least one embodiment, single instruction multiple data (SIMD) instruction issuance technology is used to support the parallel execution of a large number of threads without providing multiple independent instruction units. In at least one embodiment, single instruction multiple thread (SIMT) technology is used to support the parallel execution of a large number of synchronized threads, which uses a common instruction unit that is configured to issue instructions to a group of processing engines within each processing cluster.
[0209] In at least one embodiment, the operation of the processing cluster 1114 can be controlled by a pipeline manager 1132 that assigns processing tasks to SIMT parallel processors. In at least one embodiment, the pipeline manager 1132 Figure 11BThe scheduler 1110 receives instructions and manages the execution of these instructions through the graphics multiprocessor 1134 and / or the texture unit 1136. In at least one embodiment, the graphics multiprocessor 1134 is an exemplary instance of a SIMT parallel processor. However, in at least one embodiment, various types of SIMT parallel processors with different architectures may be included within the processing cluster 1114. In at least one embodiment, one or more instances of the graphics multiprocessor 1134 may be included within the processing cluster 1114. In at least one embodiment, the graphics multiprocessor 1134 may process data, and the data crossbar 1140 may be used to distribute the processed data to one of multiple possible destinations (including other shader units). In at least one embodiment, the pipeline manager 1132 may facilitate the distribution of processed data by specifying the destination of the processed data to be distributed via the data crossbar 1140.
[0210] In at least one embodiment, each graphics multiprocessor 1134 within a processing cluster 1114 may include the same set of function execution logic (e.g., arithmetic logic unit, load-store unit, etc.). In at least one embodiment, the function execution logic may be configured in a pipelined manner, where new instructions may be issued before previous instructions have completed. In at least one embodiment, the function execution logic supports a variety of operations, including integer and floating-point arithmetic, comparison operations, Boolean operations, shifts, and calculations of various algebraic functions. In at least one embodiment, the same functional unit hardware may be utilized to perform different operations, and any combination of functional units may be present.
[0211] In at least one embodiment, instructions transmitted to processing cluster 1114 constitute threads. In at least one embodiment, a group of threads executed across a set of parallel processing engines is a thread group. In at least one embodiment, a thread group executes a program on different input data. In at least one embodiment, each thread within a thread group can be assigned to a different processing engine within graphics multiprocessor 1134. In at least one embodiment, a thread group can include fewer threads than the number of processing engines within graphics multiprocessor 1134. In at least one embodiment, when a thread group includes fewer threads than the number of processing engines, one or more processing engines may be idle during the processing of a loop by the thread group. In at least one embodiment, a thread group can also include more threads than the number of processing engines within graphics multiprocessor 1134. In at least one embodiment, when a thread group includes more threads than the number of processing engines within graphics multiprocessor 1134, processing can be performed within consecutive clock cycles. In at least one embodiment, multiple thread groups can be executed simultaneously on graphics multiprocessor 1134.
[0212] In at least one embodiment, the graphics multiprocessor 1134 includes internal cache memory to perform load and store operations. In at least one embodiment, the graphics multiprocessor 1134 can abandon the internal cache and use cache memory within the processing cluster 1114 (e.g., L1 cache 1148). In at least one embodiment, each graphics multiprocessor 1134 can also access a partition unit (e.g., Figure 11B 1120N) are shared across all processing clusters 1114 and can be used to transfer data between threads. In at least one embodiment, graphics multiprocessor 1134 can also access off-chip global memory, which can include one or more of local parallel processor memory and / or system memory. In at least one embodiment, any memory external to parallel processing unit 1102 can be used as global memory. In at least one embodiment, processing cluster 1114 includes multiple instances of graphics multiprocessor 1134, which can share common instructions and data, which can be stored in L1 cache 1148.
[0213] In at least one embodiment, each processing cluster 1114 may include a memory management unit ("MMU") 1145 configured to map virtual addresses to physical addresses. In at least one embodiment, one or more instances of MMU 1145 may reside in Figure 11B 1148 . In at least one embodiment, the MMU 1145 includes a set of page table entries (PTEs) that are used to map virtual addresses to physical addresses of tiles and, in at least one embodiment, to cache line indices. In at least one embodiment, the MMU 1145 may include an address translation lookaside buffer (TLB) or a cache that may reside within the graphics multiprocessor 1134 or L1 cache 1148 or processing cluster 1114. In at least one embodiment, the physical address is processed to assign surface data access locality for efficient request interleaving between partition units. In at least one embodiment, the cache line index may be used to determine whether a request for a cache line is a hit or a miss.
[0214] In at least one embodiment, the processing clusters 1114 can be configured such that each graphics multiprocessor 1134 is coupled to a texture unit 1136 to perform texture mapping operations, such as determining texture sample locations, reading texture data, and filtering texture data. In at least one embodiment, texture data is read from an internal texture L1 cache (not shown) or from an L1 cache within the graphics multiprocessor 1134, and texture data is retrieved from an L2 cache, local parallel processor memory, or system memory as needed. In at least one embodiment, each graphics multiprocessor 1134 outputs one or more processed tasks to a data crossbar 1140 to provide the processed tasks to another processing cluster 1114 for further processing or to store the processed tasks in an L2 cache, local parallel processor memory, or system memory via the memory crossbar 1116. In at least one embodiment, a preROP 1142 (pre-raster operations unit) is configured to receive data from the graphics multiprocessor 1134 and direct the data to a ROP unit, which can communicate with a partition unit (e.g., a partition unit) as described herein. Figure 11B In at least one embodiment, the PreROP 1142 unit may perform optimizations for color blending, organize pixel color data, and perform address translation.
[0215] Reasoning and / or training logic 615 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 6B and / or Figure 6C Details are provided regarding inference and / or training logic 615. In at least one embodiment, inference and / or training logic 615 can be used in graphics processing cluster 1114 to perform inference or prediction operations based at least in part on weight parameters computed using the neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0216] Figure 11EA graphics multiprocessor 1134 is shown in accordance with at least one embodiment. In at least one embodiment, the graphics multiprocessor 1134 is coupled to a pipeline manager 1132 of a processing cluster 1114. In at least one embodiment, the graphics multiprocessor 1134 has an execution pipeline that includes, but is not limited to, an instruction cache 1152, an instruction unit 1154, an address mapping unit 1156, a register file 1158, one or more general purpose graphics processing unit (GPGPU) cores 1162, and one or more load / store units 1166. The one or more GPGPU cores 1162 and the one or more load / store units 1166 are coupled to a cache memory 1172 and a shared memory 1170 via a memory and cache interconnect 1168.
[0217] In at least one embodiment, the instruction cache 1152 receives a stream of instructions to be executed from the pipeline manager 1132. In at least one embodiment, the instructions are cached in the instruction cache 1152 and dispatched for execution by the instruction unit 1154. In one embodiment, the instruction unit 1154 can dispatch instructions as thread groups (e.g., warps), assigning each thread group to a different execution unit within one or more GPGPU cores 1162. In at least one embodiment, instructions can access any local, shared, or global address space by specifying an address within the unified address space. In at least one embodiment, the address mapping unit 1156 can be used to convert addresses in the unified address space into different memory addresses that can be accessed by one or more load / store units 1166.
[0218] In at least one embodiment, register file 1158 provides a set of registers for the functional units of graphics multiprocessor 1134. In at least one embodiment, register file 1158 provides temporary storage for operands for the data paths of the functional units (e.g., GPGPU core 1162, load / store unit 1166) connected to graphics multiprocessor 1134. In at least one embodiment, register file 1158 is divided between each functional unit such that a dedicated portion of register file 1158 is allocated to each functional unit. In at least one embodiment, register file 1158 is divided between the different warps being executed by graphics multiprocessor 1134.
[0219] In at least one embodiment, the GPGPU cores 1162 may each include a floating-point unit (FPU) and / or an integer arithmetic logic unit (ALU) for executing instructions for the graphics multiprocessor 1134. The GPGPU cores 1162 may be architecturally similar or may differ in architecture. In at least one embodiment, a first portion of the GPGPU core 1162 includes a single-precision FPU and integer ALU, while a second portion of the GPGPU core includes a double-precision FPU. In at least one embodiment, the FPU may implement the IEEE 754-2008 standard for floating-point arithmetic or enable variable-precision floating-point arithmetic. In at least one embodiment, the graphics multiprocessor 1134 may additionally include one or more fixed-function or special-function units to perform specific functions, such as copying rectangles or pixel blending operations. In at least one embodiment, one or more of the GPGPU cores may also include fixed-function or special-function logic.
[0220] In at least one embodiment, the GPGPU core 1162 includes SIMD logic capable of executing a single instruction on multiple sets of data. In one embodiment, the GPGPU core 1162 can physically execute SIMD4, SIMD8, and SIMD16 instructions, and logically execute SIMD1, SIMD2, and SIMD32 instructions. In at least one embodiment, the SIMD instructions for the GPGPU core can be generated by a shader compiler at compile time, or automatically generated when executing a program written and compiled for a single program multiple data (SPMD) or SIMT architecture. In at least one embodiment, multiple threads of a program configured for a SIMT execution model can be executed by a single SIMD instruction. For example, in at least one embodiment, eight SIMT threads performing the same or similar operations can be executed in parallel by a single SIMD8 logic unit.
[0221] In at least one embodiment, the memory and cache interconnect 1168 is an interconnect network that connects each functional unit of the graphics multiprocessor 1134 to the register file 1158 and the shared memory 1170. In at least one embodiment, the memory and cache interconnect 1168 is a crossbar interconnect that allows the load / store unit 1166 to perform load and store operations between the shared memory 1170 and the register file 1158. In at least one embodiment, the register file 1158 can operate at the same frequency as the GPGPU core 1162, resulting in very low latency for data transfers between the GPGPU core 1162 and the register file 1158. In at least one embodiment, the shared memory 1170 can be used to enable communication between threads executing on the functional units within the graphics multiprocessor 1134. In at least one embodiment, the cache memory 1172 can be used, for example, as a data cache to cache texture data communicated between the functional units and the texture unit 1136. In at least one embodiment, the shared memory 1170 can also be used as a program-managed cache. In at least one embodiment, in addition to automatically cached data stored in cache memory 1172, threads executing on GPGPU core 1162 may programmatically store data in shared memory.
[0222] In at least one embodiment, a parallel processor or GPGPU as described herein is communicatively coupled to a host / processor core to accelerate graphics operations, machine learning operations, pattern analysis operations, and various general-purpose GPU (GPGPU) functions. In at least one embodiment, the GPU can be communicatively coupled to the host processor / core via a bus or other interconnect (e.g., a high-speed interconnect such as PCIe or NVLink). In at least one embodiment, the GPU can be integrated into the same package or chip as the core and communicatively coupled to the core via an internal processor bus / interconnect (inside the package or chip, in at least one embodiment). In at least one embodiment, regardless of the manner in which the GPU is connected, the processor core can assign work to the GPU in the form of a sequence of commands / instructions contained in a work descriptor. In at least one embodiment, the GPU then uses dedicated circuitry / logic to efficiently process these commands / instructions.
[0223] The reasoning and / or training logic 615 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 6B and / or Figure 6CDetails are provided regarding inference and / or training logic 615. In at least one embodiment, inference and / or training logic 615 may be used in graphics multiprocessor 1134 to perform inference or prediction operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0224] FIG12A illustrates a multi-GPU computing system 1200A according to at least one embodiment. In at least one embodiment, multi-GPU computing system 1200A may include a processor 1202 coupled to a plurality of general-purpose graphics processing units (GPGPUs) 1206A-D via a host interface switch 1204. In at least one embodiment, host interface switch 1204 is a PCI Express switch device that couples processor 1202 to a PCI Express bus, through which processor 1202 can communicate with GPGPUs 1206A-D. GPGPUs 1206A-D may be interconnected via a set of high-speed P2P GPU-to-GPU links 1216. In at least one embodiment, GPU-to-GPU links 1216 connect to each of GPGPUs 1206A-D via a dedicated GPU link. In at least one embodiment, P2P GPU links 1216 enable direct communication between each GPGPU 1206A-D without requiring communication through host interface bus 1204 to which processor 1202 is connected. In at least one embodiment, host interface bus 1204 remains available for system memory access or communication with other instances of multi-GPU computing system 1200A, e.g., via one or more network devices, with GPU-to-GPU traffic directed to P2P GPU link 1216. While in at least one embodiment, GPGPUs 1206A-D are connected to processor 1202 via host interface switch 1204, in at least one embodiment, processor 1202 includes direct support for P2P GPU link 1216 and can connect directly to GPGPUs 1206A-D.
[0225] Reasoning and / or training logic 615 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 6B and / or Figure 6C Details are provided regarding inference and / or training logic 615. In at least one embodiment, inference and / or training logic 615 can be used in multi-GPU computing system 1200A to perform inference or prediction operations based at least in part on weight parameters computed using the neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0226] Figure 12B FIG1 is a block diagram of a graphics processor 1200B according to at least one embodiment. In at least one embodiment, graphics processor 1200B includes a ring interconnect 1202, a pipeline front end 1204, a media engine 1237, and graphics cores 1280A-1280N. In at least one embodiment, ring interconnect 1202 couples graphics processor 1200B to other processing units, including other graphics processors or one or more general-purpose processor cores. In at least one embodiment, graphics processor 1200B is one of many processors integrated within a multi-core processing system.
[0227] In at least one embodiment, graphics processor 1200B receives batches of commands via ring interconnect 1202. In at least one embodiment, the incoming commands are interpreted by command streamer 1203 in pipeline front end 1204. In at least one embodiment, graphics processor 1200B includes scalable execution logic for performing 3D geometry processing and media processing via graphics cores 1280A-1280N. In at least one embodiment, for 3D geometry processing commands, command streamer 1203 provides commands to geometry pipeline 1236. In at least one embodiment, for at least some media processing commands, command streamer 1203 provides commands to video front end 1234, which is coupled to media engine 1237. In at least one embodiment, media engine 1237 includes a video quality engine (VQE) 1230 for video and image post-processing, and a multi-format encoding / decoding (MFX) 1233 engine for providing hardware-accelerated media data encoding and decoding. In at least one embodiment, the geometry pipeline 1236 and the media engine 1237 each generate execution threads for thread execution resources provided by at least one graphics core 1280A.
[0228] In at least one embodiment, graphics processor 1200B includes scalable thread execution resources featuring modular cores 1280A-1280N (sometimes referred to as core slices), each of which has multiple sub-cores 1250A-1250N, 1260A-1260N (sometimes referred to as core sub-slices). In at least one embodiment, graphics processor 1200B can have any number of graphics cores 1280A. In at least one embodiment, graphics processor 1200B includes graphics core 1280A having at least a first sub-core 1250A and a second sub-core 1260A. In at least one embodiment, graphics processor 1200B is a low-power processor having a single sub-core (e.g., 1250A). In at least one embodiment, graphics processor 1200B includes multiple graphics cores 1280A-1280N, each of which includes a set of first sub-cores 1250A-1250N and a set of second sub-cores 1260A-1260N. In at least one embodiment, each of the first sub-cores 1250A-1250N includes at least a first set of execution units 1252A-1252N and media / texture samplers 1254A-1254N. In at least one embodiment, each of the second sub-cores 1260A-1260N includes at least a second set of execution units 1262A-1262N and samplers 1264A-1264N. In at least one embodiment, each of the sub-cores 1250A-1250N, 1260A-1260N shares a set of shared resources 1270A-1270N. In at least one embodiment, the shared resources include shared cache memory and pixel operation logic.
[0229] Reasoning and / or training logic 615 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 6B and / or Figure 6C Details are provided regarding inference and / or training logic 615. In at least one embodiment, inference and / or training logic 615 may be used in graphics processor 1200B to perform inference or prediction operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0230] Figure 13is a block diagram illustrating a microarchitecture for a processor 1300, which may include logic circuitry for executing instructions, according to at least one embodiment. In at least one embodiment, processor 1300 may execute instructions including x86 instructions, ARM instructions, specialized instructions for application-specific integrated circuits (ASICs), and the like. In at least one embodiment, processor 1300 may include registers for storing packed data, such as the 64-bit-wide MMX™ registers in microprocessors enabled with MMX technology from Intel Corporation of Santa Clara, California. In at least one embodiment, MMX registers, available in integer and floating-point form, may operate with packed data elements with Single Instruction Multiple Data ("SIMD") and Streaming SIMD Extensions ("SSE") instructions. In at least one embodiment, 128-bit-wide XMM registers associated with SSE2, SSE3, SSE4, AVX, or later (generally referred to as "SSEx") technology may store such packed data operands. In at least one embodiment, processor 1300 may execute instructions to accelerate machine learning or deep learning algorithms, training, or inference.
[0231] In at least one embodiment, the processor 1300 includes an in-order front end ("front end") 1301 to fetch instructions to be executed and prepare the instructions for later use in the processor pipeline. In at least one embodiment, the front end 1301 may include several units. In at least one embodiment, an instruction prefetcher 1326 retrieves instructions from memory and provides the instructions to an instruction decoder 1328, which in turn decodes or interprets the instructions. For example, in at least one embodiment, the instruction decoder 1328 decodes the received instructions into one or more operations called "microinstructions" or "micro-operations" (also referred to as "micro-ops" or "micro-instructions") that the machine can execute. In at least one embodiment, the instruction decoder 1328 parses the instructions into an opcode and corresponding data and control fields, which can be used by the microarchitecture to perform the operations according to at least one embodiment. In at least one embodiment, the trace cache 1330 can assemble the decoded microinstructions into a program-ordered sequence or trace in the microinstruction queue 1334 for execution. In at least one embodiment, when trace cache 1330 encounters a complex instruction, microcode ROM 1332 provides the microinstructions necessary to complete the operation.
[0232] In at least one embodiment, some instructions may be converted into a single micro-op, while other instructions may require several micro-ops to complete the entire operation. In at least one embodiment, if more than four micro-ops are required to complete an instruction, the instruction decoder 1328 may access the microcode ROM 1332 to execute the instruction. In at least one embodiment, an instruction may be decoded into a smaller number of micro-ops for processing at the instruction decoder 1328. In at least one embodiment, if multiple micro-ops are required to complete the operation, the instruction may be stored in the microcode ROM 1332. In at least one embodiment, the trace cache 1330 references the entry point programmable logic array ("PLA") to determine the correct micro-op pointer for reading the microcode sequence from the microcode ROM 1332 to complete one or more instructions according to at least one embodiment. In at least one embodiment, after the microcode ROM 1332 completes the micro-op sequencing for the instruction, the front end 1301 of the machine may resume fetching micro-ops from the trace cache 1330.
[0233] In at least one embodiment, an out-of-order execution engine ("OOO engine") 1303 can prepare instructions for execution. In at least one embodiment, the OOO logic has multiple buffers to smooth and reorder the instruction flow to optimize performance as instructions flow down the pipeline and are scheduled for execution. In at least one embodiment, the OOO engine 1303 includes, but is not limited to, an allocator / register renamer 1340, a memory microinstruction queue 1342, an integer / floating-point microinstruction queue 1344, a memory scheduler 1346, a fast scheduler 1302, a slow / general purpose floating-point scheduler ("slow / general purpose FP scheduler") 1304, and a simple floating-point scheduler ("simple FP scheduler") 1306. In at least one embodiment, the fast scheduler 1302, the slow / general purpose floating-point scheduler 1304, and the simple floating-point scheduler 1306 are also collectively referred to as "microinstruction schedulers 1302, 1304, 1306." In at least one embodiment, the allocator / register renamer 1340 allocates the machine buffers and resources required for each microinstruction to execute in sequence. In at least one embodiment, the allocator / register renamer 1340 renames logical registers into entries in the register file. In at least one embodiment, the allocator / register renamer 1340 also allocates an entry for each microinstruction in one of two microinstruction queues: a memory microinstruction queue 1342 for memory operations and an integer / floating-point microinstruction queue 1344 for non-memory operations, preceding the memory scheduler 1346 and the microinstruction schedulers 1302, 1304, 1306. In at least one embodiment, the microinstruction schedulers 1302, 1304, 1306 determine when a microinstruction is ready to execute based on the readiness of its dependent input register operand sources and the availability of the execution resource microinstructions that need to be completed. In at least one embodiment, the fast scheduler 1302 of at least one embodiment can schedule every half of the main clock cycle, while the slow / general floating-point scheduler 1304 and the simple floating-point scheduler 1306 can schedule once per main processor clock cycle. In at least one embodiment, microinstruction schedulers 1302, 1304, 1306 arbitrate on dispatch ports to schedule microinstructions for execution.
[0234] In at least one embodiment, execution block 1311 includes, but is not limited to, integer register file / bypass network 1308, floating point register file / bypass network ("FP register file / bypass network") 1310, address generation units ("AGUs") 1312 and 1314, fast arithmetic logic units ("fast ALUs") 1316 and 1318, slow arithmetic logic unit ("slow ALU") 1320, floating point ALU ("FP") 1322, and floating point move unit ("FP move") 1324. In at least one embodiment, integer register file / bypass network 1308 and floating point register file / bypass network 1310 are also referred to herein as "register files 1308, 1310." In at least one embodiment, AGUs 1312 and 1314, fast ALUs 1316 and 1318, slow ALU 1320, floating-point ALU 1322, and floating-point move unit 1324 are also referred to herein as "execution units 1312, 1314, 1316, 1318, 1320, 1322, and 1324." In at least one embodiment, execution block 1311 may include, but is not limited to, any number (including zero) and type of register files, bypass networks, address generation units, and execution units (in any combination).
[0235] In at least one embodiment, register networks 1308 and 1310 may be arranged between microinstruction schedulers 1302, 1304, and 1306 and execution units 1312, 1314, 1316, 1318, 1320, 1322, and 1324. In at least one embodiment, integer register file / bypass network 1308 performs integer operations. In at least one embodiment, floating-point register file / bypass network 1310 performs floating-point operations. In at least one embodiment, each of register networks 1308 and 1310 may include, but is not limited to, a bypass network that can bypass or forward recently completed results that have not yet been written to the register file to new dependent objects. In at least one embodiment, register networks 1308 and 1310 may communicate data with each other. In at least one embodiment, integer register file / bypass network 1308 may include, but is not limited to, two separate register files: one register file for low-order 32-bit data and a second register file for high-order 32-bit data. In at least one embodiment, floating point register file / bypass network 1310 may include, but is not limited to, 128-bit wide entries, as floating point instructions typically have operands that are 64 to 128 bits wide.
[0236] In at least one embodiment, execution units 1312, 1314, 1316, 1318, 1320, 1322, and 1324 may execute instructions. In at least one embodiment, register files 1308 and 1310 store integer and floating-point data operand values required for microinstructions to execute. In at least one embodiment, processor 1300 may include, but is not limited to, any number of execution units 1312, 1314, 1316, 1318, 1320, 1322, and 1324, and any combination thereof. In at least one embodiment, floating-point ALU 1322 and floating-point move unit 1324 may perform floating-point, MMX, SIMD, AVX, SSE, or other operations, including specialized machine learning instructions. In at least one embodiment, floating-point ALU 1322 may include, but is not limited to, a 64-bit by 64-bit floating-point divider to perform division, square root, and remainder micro-operations. In at least one embodiment, floating-point hardware may be used to process instructions involving floating-point values. In at least one embodiment, ALU operations can be passed to fast ALUs 1316 and 1318. In at least one embodiment, fast ALUs 1316 and 1318 can perform fast operations with an effective latency of half a clock cycle. In at least one embodiment, most complex integer operations go to slow ALU 1320, as slow ALU 1320 may include, but is not limited to, integer execution hardware for long-latency operations, such as multipliers, shifts, flag logic, and branch processing. In at least one embodiment, memory load / store operations can be performed by AGUs 1312 and 1314. In at least one embodiment, fast ALU 1316, fast ALU 1318, and slow ALU 1320 can perform integer operations on 64-bit data operands. In at least one embodiment, fast ALU 1316, fast ALU 1318, and slow ALU 1320 can be implemented to support various data bit sizes, including 16, 32, 128, 256, and the like. In at least one embodiment, the floating point ALU 1322 and floating point shift unit 1324 can be implemented to support a range of operands having bits of various widths. In at least one embodiment, the floating point ALU 1322 and floating point shift unit 1324 can operate on 128-bit wide packed data operands in conjunction with SIMD and multimedia instructions.
[0237] In at least one embodiment, the microinstruction schedulers 1302, 1304, 1306 schedule dependent operations before the parent load completes execution. In at least one embodiment, because microinstructions can be speculatively scheduled and executed in processor 1300, processor 1300 can also include logic for handling memory misses. In at least one embodiment, if a data load misses in the data cache, there may be dependent operations running in the pipeline that temporarily prevent the scheduler from having the correct data. In at least one embodiment, a replay mechanism tracks and re-executes instructions that use incorrect data. In at least one embodiment, it may be necessary to replay dependent operations and allow independent operations to complete. In at least one embodiment, the scheduler and replay mechanism of at least one embodiment of the processor can also be designed to capture instruction sequences for text string comparison operations.
[0238] In at least one embodiment, "register" may refer to an on-board processor storage location that can be used as part of an instruction to identify an operand. In at least one embodiment, registers may be those that can be used from outside the processor (from a programmer's perspective). In at least one embodiment, registers may not be limited to a particular type of circuit. Instead, in at least one embodiment, registers can store data, provide data, and perform the functions described herein. In at least one embodiment, the registers described herein can be implemented by circuitry within the processor using a variety of different techniques, such as dedicated physical registers, physical registers dynamically allocated using register renaming, a combination of dedicated and dynamically allocated physical registers, and the like. In at least one embodiment, integer registers store 32-bit integer data. The register file of at least one embodiment also includes eight multimedia SIMD registers for packing data.
[0239] Reasoning and / or training logic 615 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 6B and / or Figure 6C Details are provided regarding the inference and / or training logic 615. In at least one embodiment, some or all of the inference and / or training logic 615 may be incorporated into the execution block 1311 and other memories or registers, shown or not shown. For example, in at least one embodiment, the training and / or inference techniques described herein may utilize one or more ALUs shown in the execution block 1311. Additionally, weight parameters may be stored in on-chip or off-chip memory and / or registers (shown or not shown) that configure the ALUs of the execution block 1311 to execute one or more of the machine learning algorithms, neural network architectures, use cases, or training techniques described herein.
[0240] Figure 14A deep learning application processor 1400 is shown in accordance with at least one embodiment. In at least one embodiment, the deep learning application processor 1400 uses instructions that, if executed by the deep learning application processor 1400, cause the deep learning application processor 1400 to perform some or all of the processes and techniques described throughout this disclosure. In at least one embodiment, the deep learning application processor 1400 is an application-specific integrated circuit (ASIC). In at least one embodiment, the application processor 1400 performs matrix multiplication operations or "hardwires" them into hardware as a result of executing one or more instructions, or both. In at least one embodiment, the deep learning application processor 1400 includes, but is not limited to, processing clusters 1410(1)-1410(12), inter-chip links (“ICLs”) 1420(1)-1420(12), inter-chip controllers (“ICCs”) 1430(1)-1430(2), memory controllers (“Mem Ctrlr”) 1442(1)-1442(4), high bandwidth memory physical layer (“HBM PHY”) 1444(1)-1444(4), a management controller central processing unit (“management controller CPU”) 1450, serial peripheral interface, inter-integrated circuit, and general purpose input / output blocks (“SPI, I2C, GPIO”), a peripheral component interconnect express controller and direct memory access block (“PCIe controller and DMA”) 1470, and a sixteen-lane peripheral component interconnect express port (“PCI Express x 16”) 1480.
[0241] In at least one embodiment, processing cluster 1410 can perform deep learning operations, including inference or prediction operations based on weight parameters calculated by one or more training techniques, including those described herein. In at least one embodiment, each processing cluster 1410 can include, but is not limited to, any number and type of processors. In at least one embodiment, deep learning application processor 1400 can include any number and type of processing clusters 1400. In at least one embodiment, inter-chip link 1420 is bidirectional. In at least one embodiment, inter-chip link 1420 and inter-chip controller 1430 enable multiple deep learning application processors 1400 to exchange information, including activation information generated from executing one or more machine learning algorithms embodied in one or more neural networks. In at least one embodiment, deep learning application processor 1400 can include any number (including zero) and type of ICL 1420 and ICC 1430.
[0242] In at least one embodiment, HBM2 1440 provides a total of 32GB of memory. HBM2 1440(i) is associated with both a memory controller 1442(i) and an HBM PHY 1444(i). In at least one embodiment, any number of HBM2 1440 can provide any type and total amount of high-bandwidth memory and can be associated with any number (including zero) and type of memory controllers 1442 and HBM PHYs 1444. In at least one embodiment, SPI, I2C, GPIO 3360, PCIe controller 1460, DMA 1470, and / or PCIe 1480 can be replaced with any number and type of blocks to implement any number and type of communication standards in any technically feasible manner.
[0243] Reasoning and / or training logic 615 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 6B and / or Figure 6C Details are provided regarding the inference and / or training logic 615. In at least one embodiment, the deep learning application processor is used to train a machine learning model (e.g., a neural network) to predict or infer information provided to the deep learning application processor 1400. In at least one embodiment, the deep learning application processor 1400 is used to infer or predict information based on a trained machine learning model (e.g., a neural network) that has been trained by another processor or system or by the deep learning application processor 1400. In at least one embodiment, the processor 1400 can be used to perform one or more of the neural network use cases described herein.
[0244] Figure 15 is a block diagram of a neuromorphic processor 1500 according to at least one embodiment. In at least one embodiment, the neuromorphic processor 1500 can receive one or more inputs from a source external to the neuromorphic processor 1500. In at least one embodiment, these inputs can be transmitted to one or more neurons 1502 within the neuromorphic processor 1500. In at least one embodiment, the neurons 1502 and their components can be implemented using circuitry or logic including one or more arithmetic logic units (ALUs). In at least one embodiment, the neuromorphic processor 1500 can include, but is not limited to, thousands of instances of neurons 1502, although any suitable number of neurons 1502 can be used. In at least one embodiment, each instance of a neuron 1502 can include a neuron input 1504 and a neuron output 1506. In at least one embodiment, a neuron 1502 can generate an output that can be transmitted to the inputs of other instances of the neuron 1502. In at least one embodiment, the neuron input 1504 and the neuron output 1506 can be interconnected via a synapse 1508.
[0245] In at least one embodiment, the neurons 1502 and synapses 1508 can be interconnected so that the neuromorphic processor 1500 operates to process or analyze information received by the neuromorphic processor 1500. In at least one embodiment, the neuron 1502 can send an output pulse (or "trigger" or "spike") when the input received through the neuron input 1504 exceeds a threshold. In at least one embodiment, the neuron 1502 can sum or integrate the signals received at the neuron input 1504. For example, in at least one embodiment, the neuron 1502 can be implemented as a leaky integrate-and-trigger neuron, where if the sum (referred to as the "membrane potential") exceeds a threshold, the neuron 1502 can generate an output (or "trigger") using a transfer function such as a sigmoid or threshold function. In at least one embodiment, the leaky integrate-and-trigger neuron can sum the signals received at the neuron input 1504 into a membrane potential and can apply an application attenuation factor (or leakage) to reduce the membrane potential. In at least one embodiment, a leaky integrate-and-trigger neuron may trigger if multiple input signals are received at neuron input 1504 quickly enough to exceed a threshold (in at least one embodiment, before the membrane potential decays too low to trigger). In at least one embodiment, neuron 1502 may be implemented using circuitry or logic that receives input, integrates the input into a membrane potential, and decays the membrane potential. In at least one embodiment, the inputs may be averaged, or any other suitable transfer function may be used. Furthermore, in at least one embodiment, neuron 1502 may include, but is not limited to, comparator circuitry or logic that generates an output spike at neuron output 1506 when the result of applying the transfer function to neuron input 1504 exceeds a threshold. In at least one embodiment, once neuron 1502 triggers, it may ignore previously received input information by, for example, resetting the membrane potential to 0 or another suitable default value. In at least one embodiment, once the membrane potential is reset to 0, neuron 1502 may resume normal operation after a suitable period of time (or recovery period).
[0246] In at least one embodiment, neurons 1502 can be interconnected via synapses 1508. In at least one embodiment, synapses 1508 can be operable to transmit a signal from an output of a first neuron 1502 to an input of a second neuron 1502. In at least one embodiment, a neuron 1502 can transmit information across more than one instance of synapse 1508. In at least one embodiment, one or more instances of a neuron output 1506 can be connected to an instance of a neuron input 1504 in the same neuron 1502 via an instance of synapse 1508. In at least one embodiment, an instance of a neuron 1502 that generates an output to be transmitted across an instance of synapse 1508 can be referred to as a "presynaptic neuron" relative to that instance of synapse 1508. In at least one embodiment, an instance of a neuron 1502 that receives an input transmitted across an instance of synapse 1508 can be referred to as a "postsynaptic neuron" relative to an instance of synapse 1508. In at least one embodiment, with respect to the various instances of synapses 1508, because an instance of neuron 1502 can receive input from one or more instances of synapses 1508 and can also transmit output through one or more instances of synapses 1508, a single instance of neuron 1502 can be both a "pre-synaptic neuron" and a "post-synaptic neuron."
[0247] In at least one embodiment, neurons 1502 can be organized into one or more layers. Each instance of a neuron 1502 can have a neuron output 1506 that can fan out to one or more neuron inputs 1504 via one or more synapses 1508.
[0248] In at least one embodiment, neuron outputs 1506 of neurons 1502 in a first layer 1510 may be connected to neuron inputs 1504 of neurons 1502 in a second layer 1512. In at least one embodiment, layer 1510 may be referred to as a "feed-forward layer." In at least one embodiment, each instance of neuron 1502 in an instance of first layer 1510 may fan out to each instance of neuron 1502 in second layer 1512. In at least one embodiment, first layer 1510 may be referred to as a "fully connected feed-forward layer." In at least one embodiment, each instance of neuron 1502 in each instance of second layer 1512 may fan out to fewer than all instances of neuron 1502 in third layer 1514. In at least one embodiment, second layer 1512 may be referred to as a "sparsely connected feed-forward layer." In at least one embodiment, neurons 1502 in the (same) second layer 1512 may fan out to neurons 1502 in multiple other layers, including neurons 1502 in second layer 1512. In at least one embodiment, second layer 1512 may be referred to as a "recurrent layer." In at least one embodiment, the neuromorphic processor 1500 may include, but is not limited to, any suitable combination of recurrent layers and feed-forward layers, including but not limited to sparsely connected feed-forward layers and fully connected feed-forward layers.
[0249] In at least one embodiment, the neuromorphic processor 1500 may include, but is not limited to, a reconfigurable interconnect architecture or a dedicated hardwired interconnect to connect synapses 1508 to neurons 1502. In at least one embodiment, the neuromorphic processor 1500 may include, but is not limited to, circuitry or logic that allows synapses to be assigned to different neurons 1502 as needed based on the neural network topology and neuron fan-in / fan-out. For example, in at least one embodiment, synapses 1508 may be connected to neurons 1502 using an interconnect architecture such as a network on a chip or through dedicated connections. In at least one embodiment, the synaptic interconnect and its components may be implemented using circuitry or logic.
[0250] Figure 16A 16 shows a processing system according to at least one embodiment. In at least one embodiment, system 1600A includes one or more processors 1602 and one or more graphics processors 1608, and can be a single-processor desktop system, a multi-processor workstation system, or a server system with a large number of processors 1602 or processor cores 1607. In at least one embodiment, system 1600A is a processing platform incorporated within a system-on-chip (SoC) integrated circuit for use in mobile, handheld, or embedded devices.
[0251] In at least one embodiment, the system 1600A may include or be incorporated into a server-based gaming platform, including a gaming console, a mobile gaming console, a handheld gaming console, or an online gaming console. In at least one embodiment, the system 1600A is a mobile phone, a smartphone, a tablet computing device, or a mobile internet device. In at least one embodiment, the processing system 1600A may also include a device coupled to or integrated into a wearable device, such as a smartwatch wearable device, a smart glasses device, an augmented reality device, or a virtual reality device. In at least one embodiment, the processing system 1600A is a television or set-top box device having one or more processors 1602 and a graphical interface generated by one or more graphics processors 1608.
[0252] In at least one embodiment, one or more processors 1602 each include one or more processor cores 1607 to process instructions that, when executed, perform operations for system and user software. In at least one embodiment, each of the one or more processor cores 1607 is configured to process a specific instruction set 1609. In at least one embodiment, the instruction set 1609 can facilitate complex instruction set computing (CISC), reduced instruction set computing (RISC), or computing via very long instruction words (VLIW). In at least one embodiment, the processor cores 1607 can each process a different instruction set 1609, which can include instructions that facilitate emulating other instruction sets. In at least one embodiment, the processor cores 1607 can also include other processing devices, such as a digital signal processor (DSP).
[0253] In at least one embodiment, processor 1602 includes cache memory 1604. In at least one embodiment, processor 1602 may have a single internal cache or multiple levels of internal cache. In at least one embodiment, cache memory is shared among various components of processor 1602. In at least one embodiment, processor 1602 also uses an external cache (e.g., a level 3 (L3) cache or a last level cache (LLC)) (not shown), which may be shared among processor cores 1607 using known cache coherence techniques. In at least one embodiment, processor 1602 further includes a register file 1606. The processor may include different types of registers (e.g., integer registers, floating point registers, status registers, and an instruction pointer register) for storing different types of data. In at least one embodiment, register file 1606 may include general purpose registers or other registers.
[0254] In at least one embodiment, one or more processors 1602 are coupled to one or more interface buses 1610 to transmit communication signals, such as address, data, or control signals, between the processors 1602 and other components in the system 1600A. In at least one embodiment, the interface bus 1610 can be a processor bus, such as a version of a Direct Media Interface (DMI) bus. In at least one embodiment, the interface bus 1610 is not limited to a DMI bus and can include one or more peripheral component interconnect buses (e.g., PCI, PCI Express), a memory bus, or other types of interface buses. In at least one embodiment, the processor 1602 includes an integrated memory controller 1616 and a platform controller hub 1630. In at least one embodiment, the memory controller 1616 facilitates communication between memory devices and other components of the processing system 1600A, while the platform controller hub (PCH) 1630 provides connections to input / output (I / O) devices via a local I / O bus.
[0255] In at least one embodiment, memory device 1620 may be a dynamic random access memory (DRAM) device, a static random access memory (SRAM) device, a flash memory device, a phase change memory device, or other suitable memory device for use as processor memory. In at least one embodiment, memory device 1620 may be used as system memory for processing system 1600A to store data 1622 and instructions 1621 for use when one or more processors 1602 execute applications or processes. In at least one embodiment, memory controller 1616 is also coupled to an external graphics processor 1612 of at least one embodiment, which may communicate with one or more graphics processors 1608 in processor 1602 to perform graphics and media operations. In at least one embodiment, display device 1611 may be connected to processor 1602. In at least one embodiment, display device 1611 may include one or more internal display devices, such as in a mobile electronic device or laptop, or an external display device connected via a display interface (e.g., DisplayPort). In at least one embodiment, the display device 1611 may include a head-mounted display (HMD), such as a stereoscopic display device used in virtual reality (VR) applications or augmented reality (AR) applications.
[0256] In at least one embodiment, the platform controller hub 1630 enables peripheral devices to connect to the storage device 1620 and the processor 1602 via a high-speed I / O bus. In at least one embodiment, the I / O peripherals include, but are not limited to, an audio controller 1646, a network controller 1634, a firmware interface 1628, a wireless transceiver 1626, a touch sensor 1625, and a data storage device 1624 (e.g., a hard drive, flash memory, etc.). In at least one embodiment, the data storage device 1624 can be connected via a storage interface (e.g., SATA) or via a peripheral bus, such as a peripheral component interconnect bus (e.g., PCI, PCIe). In at least one embodiment, the touch sensor 1625 can include a touch screen sensor, a pressure sensor, or a fingerprint sensor. In at least one embodiment, the wireless transceiver 1626 can be a Wi-Fi transceiver, a Bluetooth transceiver, or a mobile network transceiver, such as a 3G, 4G, or Long Term Evolution (LTE) transceiver. In at least one embodiment, the firmware interface 1628 enables communication with the system firmware and can be, for example, a Unified Extensible Firmware Interface (UEFI). In at least one embodiment, a network controller 1634 can enable network connectivity to a wired network. In at least one embodiment, a high-performance network controller (not shown) is coupled to the interface bus 1610. In at least one embodiment, the audio controller 1646 is a multi-channel high-definition audio controller. In at least one embodiment, the processing system 1600A includes a legacy I / O controller 1640 for coupling legacy (e.g., Personal System 2 (PS / 2)) devices to the system 1600A. In at least one embodiment, the platform controller hub 1630 can also be connected to one or more universal serial bus (USB) controllers 1642 that connect input devices such as a keyboard and mouse 1643 combination, a camera 1644, or other USB input devices.
[0257] In at least one embodiment, instances of memory controller 1616 and platform controller hub 1630 may be integrated into a discrete external graphics processor, such as external graphics processor 1612. In at least one embodiment, platform controller hub 1630 and / or memory controller 1616 may be external to one or more processors 1602. In at least one embodiment, system 1600A may include external memory controller 1616 and platform controller hub 1630, which may be configured as a memory controller hub and a peripheral controller hub in a system chipset that communicates with processor 1602.
[0258] Reasoning and / or training logic 615 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 6B and / or Figure 6C615. In at least one embodiment, some or all of the inference and / or training logic 615 may be incorporated into the graphics processor 1600A. In at least one embodiment, the training and / or inference techniques described herein may utilize one or more ALUs embodied in the graphics processor 1612. Additionally, in at least one embodiment, the inference and / or training operations described herein may utilize a number of ALUs other than the ALUs. Figure 6B and / or Figure 6C In at least one embodiment, the weight parameters may be stored in on-chip or off-chip memory and / or registers (shown or not shown), which configure the ALU of graphics processor 1600A to execute one or more of the machine learning algorithms, neural network architectures, use cases, or training techniques described herein.
[0259] Figure 16B is a block diagram of a processor 1600B having one or more processor cores 1602A-1602N, an integrated memory controller 1614, and an integrated graphics processor 1608, in accordance with at least one embodiment. In at least one embodiment, processor 1600B may include additional cores, up to and including additional core 1602N, which is represented by the dashed box. In at least one embodiment, each processor core 1602A-1602N includes one or more internal cache units 1604A-1604N. In at least one embodiment, each processor core may also have access to one or more shared cache units 1606.
[0260] In at least one embodiment, the internal cache units 1604A-1604N and the shared cache unit 1606 represent a cache memory hierarchy within the processor 1600B. In at least one embodiment, the cache memory units 1604A-1604N may include at least one level of instruction and data cache within each processor core and one or more levels of cache within a shared mid-level cache, such as level 2 (L2), level 3 (L3), level 4 (L4), or other levels of cache, with the highest level of cache before external memory being categorized as LLC. In at least one embodiment, cache coherence logic maintains coherence between the various cache units 1606 and 1604A-1604N.
[0261] In at least one embodiment, processor 1600B may also include a set of one or more bus controller units 1616 and a system agent core 1610. In at least one embodiment, the one or more bus controller units 1616 manage a set of peripheral buses, such as one or more PCI or PCIe buses. In at least one embodiment, the system agent core 1610 provides management functions for various processor components. In at least one embodiment, the system agent core 1610 includes one or more integrated memory controllers 1614 to manage access to various external memory devices (not shown).
[0262] In at least one embodiment, one or more processor cores 1602A-1602N include support for simultaneous multithreading. In at least one embodiment, system agent core 1610 includes components for coordinating and operating cores 1602A-1602N during multithreaded processing. In at least one embodiment, system agent core 1610 may additionally include a power control unit (PCU) that includes logic and components for regulating one or more power states of processor cores 1602A-1602N and graphics processor 1608.
[0263] In at least one embodiment, processor 1600B also includes a graphics processor 1608 for performing image processing operations. In at least one embodiment, graphics processor 1608 is coupled to a shared cache unit 1606 and a system agent core 1610 including one or more integrated memory controllers 1614. In at least one embodiment, system agent core 1610 also includes a display controller 1611 for driving the graphics processor output to one or more coupled displays. In at least one embodiment, display controller 1611 may also be a separate module coupled to graphics processor 1608 via at least one interconnect, or may be integrated within graphics processor 1608.
[0264] In at least one embodiment, a ring-based interconnect 1612 is used to couple the internal components of processor 1600B. In at least one embodiment, alternative interconnects may be used, such as point-to-point interconnects, switched interconnects, or other technologies. In at least one embodiment, graphics processor 1608 is coupled to ring interconnect 1612 via I / O link 1613.
[0265] In at least one embodiment, I / O link 1613 represents at least one of a variety of I / O interconnects, including an on-package I / O interconnect that facilitates communication between various processor components and a high-performance embedded memory module 1618 (e.g., an eDRAM module). In at least one embodiment, each of processor cores 1602A-1602N and graphics processor 1608 uses embedded memory module 1618 as a shared last-level cache.
[0266] In at least one embodiment, the processor cores 1602A-1602N are homogeneous cores that execute a common instruction set architecture. In at least one embodiment, the processor cores 1602A-1602N are heterogeneous in terms of instruction set architecture (ISA), wherein one or more processor cores 1602A-1602N execute a common instruction set, while one or more other processor cores 1602A-1602N execute a subset of the common instruction set or a different instruction set. In at least one embodiment, the processor cores 1602A-1602N are heterogeneous in terms of microarchitecture, wherein one or more cores with relatively high power consumption are coupled with one or more power cores with lower power consumption. In at least one embodiment, the processor 1600B can be implemented on one or more chips or as a SoC integrated circuit.
[0267] Reasoning and / or training logic 615 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 6B and / or Figure 6C 6. In at least one embodiment, some or all of the inference and / or training logic 615 may be incorporated into the processor 1600B. For example, in at least one embodiment, the training and / or inference techniques described herein may utilize one or more ALUs embodied in the processor 1600B. Figure 16A In addition, in at least one embodiment, the inference and / or training operations described herein may use a graphics core 1612, one or more processor cores 1602A-1602N, or other components. Figure 6B and / or Figure 6C In at least one embodiment, the weight parameters may be stored in on-chip or off-chip memory and / or registers (shown or not shown) that configure the ALU of the graphics processor 1600B to execute one or more machine learning algorithms, neural network architectures, use cases, or training techniques described herein.
[0268] Figure 16Cis a block diagram of the hardware logic of a graphics processor core 1600C according to at least one embodiment herein. In at least one embodiment, graphics processor core 1600C is included in a graphics core array. In at least one embodiment, graphics processor core 1600C (sometimes referred to as a core slice) can be one or more graphics cores within a modular graphics processor. In at least one embodiment, graphics processor core 1600C is an example of one graphics core slice, and the graphics processors herein can include multiple graphics core slices based on target power and performance envelopes. In at least one embodiment, each graphics core 1600C can include a fixed function block 1630 coupled to multiple sub-cores 1601A-1601F, also referred to as sub-slices, which include modular blocks of general-purpose and fixed-function logic.
[0269] In at least one embodiment, fixed function block 1630 includes a geometry fixed function pipeline 1636. For example, in lower performance and / or lower power graphics processor implementations, the geometry and fixed function pipeline 1636 can be shared by all sub-cores in graphics processor 1600C. In at least one embodiment, the geometry and fixed function pipeline 1636 includes a 3D fixed function pipeline, a video front end unit, a thread generator and thread dispatcher, and a unified return buffer manager that manages the unified return buffer.
[0270] In at least one fixed embodiment, fixed function block 1630 also includes a graphics SoC interface 1637, a graphics microcontroller 1638, and a media pipeline 1639. In at least one embodiment, fixed graphics SoC interface 1637 provides an interface between graphics core 1600C and other processor cores in the on-chip integrated circuit system. In at least one embodiment, graphics microcontroller 1638 is a programmable subprocessor that can be configured to manage various functions of graphics processor 1600C, including thread dispatching, scheduling, and preemption. In at least one embodiment, media pipeline 1639 includes logic that facilitates decoding, encoding, pre-processing, and / or post-processing of multimedia data, including image and video data. In at least one embodiment, media pipeline 1639 implements media operations via requests to computational or sampling logic within sub-cores 1601-1601F.
[0271] In at least one embodiment, SoC interface 1637 enables graphics core 1600C to communicate with a general-purpose application processor core (e.g., a CPU) and / or other components within the SoC, including memory hierarchy elements such as a shared last-level cache, system RAM, and / or embedded on-chip or packaged DRAM. In at least one embodiment, SoC interface 1637 may also enable communication with fixed-function devices within the SoC (e.g., a camera imaging pipeline) and enable the use and / or implementation of global memory atomics that can be shared between graphics core 1600C and the CPU within the SoC. In at least one embodiment, SoC interface 1637 may also implement power management controls for graphics core 1600C and enable interfaces between the clock domain of graphics core 1600C and other clock domains within the SoC. In at least one embodiment, SoC interface 1637 enables receiving command buffers from a command stream converter and a global thread dispatcher, which are configured to provide commands and instructions to each of one or more graphics cores within the graphics processor. In at least one embodiment, commands and instructions may be dispatched to the media pipeline 1639 when media operations are to be performed, or to the geometry and fixed function pipelines (e.g., geometry and fixed function pipeline 1636, geometry and fixed function pipeline 1614) when graphics processing operations are to be performed.
[0272] In at least one embodiment, graphics microcontroller 1638 can be configured to perform various scheduling and management tasks for graphics core 1600C. In at least one embodiment, graphics microcontroller 1638 can perform graphics and / or compute workload scheduling on various graphics parallel engines within execution unit (EU) arrays 1602A-1602F, 1604A-1604F in sub-cores 1601A-1601F. In at least one embodiment, host software executing on a CPU core of a SoC including graphics core 1600C can submit a workload to one of multiple graphics processor doorbells, which invokes scheduling operations on the appropriate graphics engine. In at least one embodiment, scheduling operations include determining which workload to run next, submitting the workload to the command stream converter, preempting existing workloads running on the engine, monitoring the progress of the workload, and notifying the host software when the workload is complete. In at least one embodiment, graphics microcontroller 1638 may also facilitate low power or idle states for graphics core 1600C, thereby providing graphics core 1600C with the ability to save and restore registers across low power state transitions within graphics core 1600C independent of the operating system and / or graphics driver software on the system.
[0273] In at least one embodiment, graphics core 1600C may have up to N modular sub-cores, more or less than the sub-cores 1601A-1601F shown. For each set of N sub-cores, in at least one embodiment, graphics core 1600C may also include shared function logic 1610, shared and / or cache memory 1612, geometry / fixed function pipelines 1614, and additional fixed function logic 1616 to accelerate various graphics and compute processing operations. In at least one embodiment, shared function logic 1610 may include logic units (e.g., samplers, math, and / or inter-thread communication logic) that may be shared by each of the N sub-cores within graphics core 1600C. In at least one embodiment, fixed, shared, and / or cache memory 1612 may serve as a last-level cache for the N sub-cores 1601A-1601F within graphics core 1600C and may also serve as shared memory accessible by multiple sub-cores. In at least one embodiment, geometry / fixed function pipeline 1614 may be included in place of geometry / fixed function pipeline 1636 within fixed function block 1630 and may include similar logic.
[0274] In at least one embodiment, graphics core 1600C includes additional fixed-function logic 1616, which may include various fixed-function acceleration logic for use with graphics core 1600C. In at least one embodiment, additional fixed-function logic 1616 includes an additional geometry pipeline for use in position-only shading. In position-only shading, there are at least two geometry pipelines, and within geometry and fixed-function pipelines 1614, 1636, there are at least two geometry pipelines, which are the full geometry pipeline and the culling pipeline, which may be included in additional fixed-function logic 1616. In at least one embodiment, the culling pipeline is a modified version of the full geometry pipeline. In at least one embodiment, the full pipeline and the culling pipeline can execute different instances of an application, each with a separate context. In at least one embodiment, position-only shading can hide long culling runs for discarded triangles, allowing shading to complete earlier in some cases. In at least one embodiment, the culling pipeline logic in the additional fixed function logic 1616 can execute position shaders in parallel with the main application and generate critical results faster than the full pipeline because the culling pipeline obtains and masks the position attributes of the vertices without having to perform rasterization and render the pixels to the frame buffer. In at least one embodiment, the culling pipeline can use the generated critical results to calculate visibility information for all triangles, regardless of whether they are culled. In at least one embodiment, the full pipeline (which in this case may be called a replay pipeline) can consume visibility information to skip culled triangles to mask only visible triangles that are ultimately passed to the rasterization stage.
[0275] In at least one embodiment, the additional fixed-function logic 1616 may also include machine learning acceleration logic, such as fixed-function matrix multiplication logic, to implement optimizations including for machine learning training or inference.
[0276] In at least one embodiment, a set of execution resources is included within each graphics sub-core 1601A-1601F that can be used to perform graphics, media, and compute operations in response to requests from the graphics pipeline, media pipeline, or shader programs. In at least one embodiment, the graphics sub-core 1601A-1601F includes multiple EU arrays 1602A-1602F, 1604A-1604F, thread dispatch and inter-thread communication (TD / IC) logic 1603A-1603F, 3D (e.g., texture) samplers 1605A-1605F, media samplers 1606A-1606F, shader processors 1607A-1607F, and shared local memory (SLM) 1608A-1608F. Each of EU arrays 1602A-1602F, 1604A-1604F includes multiple execution units, which are general-purpose graphics processing units capable of servicing graphics, media, or compute operations, executing floating-point and integer / fixed-point logic operations, including graphics, media, or compute shader programs. In at least one embodiment, TD / IC logic 1603A-1603F performs local thread dispatch and thread control operations for execution units within a sub-core and facilitates communication between threads executing on execution units within the sub-core. In at least one embodiment, 3D samplers 1605A-1605F can read texture or other 3D graphics-related data into memory. In at least one embodiment, 3D samplers can read texture data differently based on the configured sampling state and texture format associated with a given texture. In at least one embodiment, media samplers 1606A-1606F can perform similar read operations based on the type and format associated with the media data. In at least one embodiment, each graphics sub-core 1601A-1601F may alternatively include a unified 3D and media sampler. In at least one embodiment, threads executing on execution units within each sub-core 1601A-1601F may utilize shared local memory 1608A-1608F within each sub-core, enabling threads executing within a thread group to execute using a common pool of on-chip memory.
[0277] Reasoning and / or training logic 615 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 6B and / or Figure 6CDetails are provided regarding the inference and / or training logic 615. In at least one embodiment, some or all of the inference and / or training logic 615 may be incorporated into the graphics processor 1610. In at least one embodiment, the training and / or inference techniques described herein may be used in Figure 16B In addition, in at least one embodiment, the inference and / or training operations described herein may use the addition of Figure 6B and / or Figure 6C In at least one embodiment, the weight parameters may be stored in on-chip or off-chip memory and / or registers (shown or not shown), which configure the ALU of the graphics processor 1600C to execute one or more of the machine learning algorithms, neural network architectures, use cases, or training techniques described herein.
[0278] Figures 16D-16E Thread execution logic 1600D is shown for an array of processing elements comprising a graphics processor core, in accordance with at least one embodiment. Figure 16D At least one embodiment is shown in which thread execution logic 1600D is used. Figure 16E Illustrative internal details of an execution unit are shown in accordance with at least one embodiment.
[0279] like Figure 16DAs shown in FIG, in at least one embodiment, thread execution logic 1600D includes a shader processor 1602, a thread dispatcher 1604, an instruction cache 1606, a scalable execution unit array including a plurality of execution units 1608A-1608N, one or more samplers 1610, a data cache 1612, and a data port 1614. In at least one embodiment, the scalable execution unit array can be dynamically scaled by enabling or disabling one or more execution units (e.g., execution units 1608A, 1608B, 1608C, 1608D, 1608N-1, and 1608N), for example, based on the computational requirements of the workload. In at least one embodiment, the scalable execution units are interconnected via an interconnect structure that links to each execution unit. In at least one embodiment, thread execution logic 1600D includes one or more connections to memory (such as system memory or cache memory) through instruction cache 1606, data port 1614, sampler 1610, and one or more of execution units 1608A-1608N. In at least one embodiment, each execution unit (e.g., 1608A) is an independent programmable general-purpose computing unit capable of executing multiple simultaneous hardware threads while processing multiple data elements in parallel for each thread. In at least one embodiment, the array of execution units 1608A-1608N is scalable to include any number of individual execution units.
[0280] In at least one embodiment, execution units 1608A-1608N are primarily used to execute shader programs. In at least one embodiment, shader processor 1602 can process various shader programs and dispatch execution threads associated with the shader programs via thread dispatcher 1604. In at least one embodiment, thread dispatcher 1604 includes logic for arbitrating thread initialization requests from graphics and media pipelines and instantiating requested threads on one or more execution units 1608A-1608N. In at least one embodiment, in at least one embodiment, the geometry pipeline can dispatch vertex, tessellation, or geometry shaders to thread execution logic for processing. In at least one embodiment, thread dispatcher 1604 can also handle runtime thread generation requests from executing shader programs.
[0281] In at least one embodiment, execution units 1608A-1608N support an instruction set that includes native support for many standard 3D graphics shader instructions, enabling shader programs from graphics libraries such as Direct3D and OpenGL to execute with minimal translation. In at least one embodiment, the execution units support vertex and geometry processing (e.g., vertex programs, geometry programs, vertex shaders), pixel processing (e.g., pixel shaders, fragment shaders), and general-purpose processing (e.g., compute and media shaders). In at least one embodiment, each execution unit 1608A-1608N includes one or more arithmetic logic units (ALUs) capable of multi-issue single instruction, multiple data (SIMD) and multi-threaded operation, enabling an efficient execution environment despite higher latency memory accesses. In at least one embodiment, each hardware thread within each execution unit has a dedicated high-bandwidth register file and associated independent thread state. In at least one embodiment, execution is multiple issues per clock to the pipeline, which is capable of integer, single-precision, and double-precision floating-point operations, SIMD branching functions, logical operations, transcendental operations, and other operations. In at least one embodiment, while waiting for data from memory or one of the shared functions, dependency logic within execution units 1608A-1608N causes the waiting thread to sleep until the requested data is returned. In at least one embodiment, while the waiting thread is sleeping, hardware resources can be dedicated to processing other threads. In at least one embodiment, during the delay associated with vertex shader operations, the execution unit can perform operations on a pixel shader, a fragment shader, or another type of shader program (including a different vertex shader).
[0282] In at least one embodiment, each of execution units 1608A-1608N operates on an array of data elements. In at least one embodiment, the number of data elements is the "execution size" or number of lanes of an instruction. In at least one embodiment, an execution lane is the logic used to perform data element access, masking, and flow control within an instruction. In at least one embodiment, the number of lanes can be independent of the number of physical arithmetic logic units (ALUs) or floating point units (FPUs) used for a particular graphics processor. In at least one embodiment, execution units 1608A-1608N support integer and floating point data types.
[0283] In at least one embodiment, the execution unit instruction set includes SIMD instructions. In at least one embodiment, various data elements can be stored in registers as packed data types, and the execution unit will process various elements based on the data size of those elements. In at least one embodiment, in at least one embodiment, when operating on a 256-bit wide vector, the 256 bits of the vector are stored in registers, and the execution unit operates on the vector as four separate 64-bit packed data elements (quad word (QW) size data elements), eight separate 32-bit packed data elements (double word (DW) size data elements), sixteen separate 16-bit packed data elements (word (W) size data elements), or thirty-two separate 8-bit data elements (byte (B) size data elements). However, in at least one embodiment, different vector widths and register sizes are possible.
[0284] In at least one embodiment, one or more execution units can be combined into a fused execution unit 1609A-1609N having thread control logic (1607A-1607N) that executes for the fused EU. In at least one embodiment, multiple EUs can be merged into a EU group. In at least one embodiment, each EU in the fused EU group can be configured to execute a separate SIMD hardware thread. The number of EUs in the fused EU group can vary according to various embodiments. In at least one embodiment, each EU can execute various SIMD widths, including but not limited to SIMD8, SIMD16, and SIMD32. In at least one embodiment, each fused graphics execution unit 1609A-1609N includes at least two execution units. In at least one embodiment, in at least one embodiment, the fused execution unit 1609A includes a first EU 1608A, a second EU 1608B, and thread control logic 1607A shared by the first EU 1608A and the second EU 1608B. In at least one embodiment, thread control logic 1607A controls threads executing on fused graphics execution unit 1609A, allowing each EU within fused execution units 1609A-1609N to execute using a common instruction pointer register.
[0285] In at least one embodiment, one or more internal instruction caches (e.g., 1606) are included in thread execution logic 1600D to cache thread instructions for the execution units. In at least one embodiment, one or more data caches (e.g., 1612) are included to cache thread data during thread execution. In at least one embodiment, a sampler 1610 is included to provide texture sampling for 3D operations and media sampling for media operations. In at least one embodiment, sampler 1610 includes specialized texture or media sampling functionality to process texture or media data during the sampling process before providing the sampled data to the execution units.
[0286] During execution, in at least one embodiment, the graphics and media pipeline sends thread initiation requests to thread execution logic 1600D via thread generation and dispatch logic. In at least one embodiment, once a set of geometric objects has been processed and rasterized into pixel data, pixel processor logic (e.g., pixel shader logic, fragment shader logic, etc.) within shader processor 1602 is invoked to further calculate output information and cause the results to be written to output surfaces (e.g., color buffer, depth buffer, stencil buffer, etc.). In at least one embodiment, the pixel shader or fragment shader calculates the values of various vertex attributes to be interpolated across the rasterized objects. In at least one embodiment, the pixel processor logic within shader processor 1602 then executes a pixel or fragment shader program provided by an application programming interface (API). In at least one embodiment, to execute the shader program, shader processor 1602 dispatches threads to execution units (e.g., 1608A) via thread dispatcher 1604. In at least one embodiment, shader processor 1602 uses texture sampling logic in sampler 1610 to access texture data stored in texture maps in memory. In at least one embodiment, arithmetic operations on texture data and input geometry data calculate pixel color data for each geometry fragment, or discard one or more pixels for further processing.
[0287] In at least one embodiment, data port 1614 provides a memory access mechanism for thread execution logic 1600D to output processed data to memory for further processing on the graphics processor output pipeline. In at least one embodiment, data port 1614 includes or is coupled to one or more cache memories (e.g., data cache 1612) to cache data for memory access via the data port.
[0288] like Figure 16EAs shown, in at least one embodiment, graphics execution unit 1608 may include an instruction fetch unit 1637, a general register file array (GRF) 1624, an architectural register file array (ARF) 1626, a thread arbiter 1622, an issue unit 1630, a branch unit 1632, a set of SIMD floating point units (FPUs) 1634, and, in at least one embodiment, a set of dedicated integer SIMD ALUs 1635. In at least one embodiment, GRF 1624 and ARF 1626 include a set of general register files and architectural register files associated with each simultaneous hardware thread that can be active in graphics execution unit 1608. In at least one embodiment, per-thread architectural state is maintained in ARF 1626, while data used during thread execution is stored in GRF 1624. In at least one embodiment, the execution state of each thread, including the instruction pointer of each thread, may be maintained in thread-specific registers in ARF 1626.
[0289] In at least one embodiment, graphics execution unit 1608 has an architecture that is a combination of simultaneous multithreading (SMT) and fine-grained interleaved multithreading (IMT). In at least one embodiment, the architecture has a modular configuration that can be fine-tuned at design time based on the target number of simultaneous threads and the number of registers per execution unit, where execution unit resources are logically allocated for executing multiple simultaneous threads.
[0290] In at least one embodiment, graphics execution unit 1608 can collectively issue multiple instructions, each of which can be a different instruction. In at least one embodiment, the thread arbiter 1622 of a graphics execution unit thread 1608 can dispatch instructions to one of the issue unit 1630, branch unit 1632, or SIMD FPU 1632 for execution. In at least one embodiment, each execution thread can access 128 general purpose registers in GRF 1624, each of which can store 32 bytes and can be accessed as a SIMD 8-element vector of 32-bit data elements. In at least one embodiment, each execution unit thread can access 4KB of GRF 1624, although embodiments are not limited thereto and more or fewer register resources may be provided in other embodiments. In at least one embodiment, a maximum of seven threads can execute simultaneously, although the number of threads per execution unit may vary depending on the embodiment. In at least one embodiment where seven threads have access to 4KB, GRF 1624 can store a total of 28KB. In at least one embodiment, flexible addressing modes may allow registers to be addressed together to efficiently build wider registers or rectangular block data structures representing strides.
[0291] In at least one embodiment, memory operations, sampler operations, and other longer latency system communications are scheduled via "send" instructions executed by message passing send unit 1630. In at least one embodiment, dispatching branch instructions to a dedicated branch unit 1632 facilitates SIMD divergence and eventual convergence.
[0292] In at least one embodiment, the graphics execution unit 1608 includes one or more SIMD floating point units (FPUs) 1634 to perform floating point operations. In at least one embodiment, one or more FPUs 1634 also support integer computations. In at least one embodiment, one or more FPUs 1634 can SIMD perform up to M 32-bit floating point (or integer) operations, or SIMD perform up to 2M 16-bit integer or 16-bit floating point operations. In at least one embodiment, at least one of the one or more FPUs provides extended math capabilities to support high-throughput transcendental math functions and double-precision 64-bit floating point. In at least one embodiment, a set of 8-bit integer SIMD ALUs 1635 are also present and can be specifically optimized to perform operations related to machine learning computations.
[0293] In at least one embodiment, an array of multiple instances of graphics execution unit 1608 may be instantiated in graphics sub-core groupings (e.g., sub-slices). In at least one embodiment, execution unit 1608 may execute instructions across multiple execution lanes. In at least one embodiment, each thread executing on graphics execution unit 1608 executes on a different lane.
[0294] The reasoning and / or training logic 615 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 6B and / or Figure 6C Provides details about the reasoning and / or training logic 615. In at least one embodiment, some or all of the reasoning and / or training logic 615 may be incorporated into the execution logic 1600D. Additionally, in at least one embodiment, other than Figure 6B and / or Figure 6C In at least one embodiment, the weight parameters may be stored in on-chip or off-chip memory and / or registers (shown or not shown) that configure the ALUs of execution logic 1600D to execute one or more machine learning algorithms, neural network architectures, use cases, or training techniques described herein.
[0295] Figure 17AA parallel processing unit ("PPU") 1700A is shown in accordance with at least one embodiment. In at least one embodiment, PPU 1700A is configured with machine-readable code that, if executed by PPU 1700A, causes PPU 1700A to perform some or all of the processes and techniques described throughout this disclosure. In at least one embodiment, PPU 1700A is a multi-threaded processor implemented on one or more integrated circuit devices and utilizes multithreading as a latency hiding technique designed to process computer-readable instructions (also referred to as machine-readable instructions or simply instructions) executed in parallel on multiple threads. In at least one embodiment, a thread refers to an execution thread and is an instance of a group of instructions configured to be executed by PPU 1700A. In at least one embodiment, PPU 1700A is a graphics processing unit ("GPU") configured to implement a graphics rendering pipeline for processing three-dimensional ("3D") graphics data to generate two-dimensional ("2D") image data for display on a display device, such as a liquid crystal display ("LCD") device. In at least one embodiment, PPU 1700A is used to perform computations such as linear algebra operations and machine learning operations. Figure 17A The example parallel processor is shown for illustrative purposes only and should be construed as a non-limiting example of a processor architecture contemplated within the scope of the present disclosure, and any suitable processor may be employed in addition and / or in place thereof.
[0296] In at least one embodiment, one or more PPUs 1700A are configured to accelerate high-performance computing ("HPC"), data center, and machine learning applications. In at least one embodiment, the PPU 1700A is configured to accelerate deep learning systems and applications, including the following non-limiting examples: autonomous vehicle platforms, deep learning, high-precision speech, image, and text recognition systems, intelligent video analysis, molecular simulation, drug discovery, disease diagnosis, weather forecasting, big data analysis, astronomy, molecular dynamics simulation, financial modeling, robotics, factory automation, real-time language translation, online search optimization, and personalized user recommendations.
[0297] In at least one embodiment, PPU 1700A includes, but is not limited to, input / output ("I / O") units 1706, front-end units 1710, scheduler units 1712, work distribution units 1714, hubs 1716, crossbars ("Xbars") 1720, one or more general processing clusters ("GPCs") 1718, and one or more partitioning units ("memory partitioning units") 1722. In at least one embodiment, PPU 1700A is connected to a host processor or other PPUs 1700A via one or more high-speed GPU interconnects ("GPU interconnects") 1708. In at least one embodiment, PPU 1700A is connected to a host processor or other peripheral devices via interconnect 1702. In one embodiment, PPU 1700A is connected to local memory including one or more memory devices ("memory") 1704. In at least one embodiment, memory devices 1704 include, but are not limited to, one or more dynamic random access memory ("DRAM") devices. In at least one embodiment, one or more DRAM devices are configured and / or configurable as a high bandwidth memory ("HBM") subsystem with multiple DRAM dies stacked within each device.
[0298] In at least one embodiment, the high-speed GPU interconnect 1708 may refer to a wire-based, multi-lane communication link that a system uses to scale and includes one or more PPUs 1700A in conjunction with one or more central processing units ("CPUs"), supporting cache coherence between the PPUs 1700A and the CPUs and CPU mastering. In at least one embodiment, the high-speed GPU interconnect 1708 transmits data and / or commands to other units of the PPU 1700A, such as one or more copy engines, video encoders, video decoders, power management units, and / or other processors, via a hub 1716. Figure 17A Other components that may not be explicitly shown.
[0299] In at least one embodiment, I / O unit 1706 is configured to receive data from a host processor ( Figure 17A1700A). In at least one embodiment, I / O unit 1706 communicates with the host processor directly through interconnect 1702 or through one or more intermediate devices (e.g., a memory bridge). In at least one embodiment, I / O unit 1706 can communicate with one or more other processors (e.g., one or more PPUs 1700A) via interconnect 1702. In at least one embodiment, I / O unit 1706 implements a Peripheral Component Interconnect Express ("PCIe") interface for communicating over a PCIe bus. In at least one embodiment, I / O unit 1706 implements an interface for communicating with external devices.
[0300] In at least one embodiment, I / O unit 1706 decodes packets received via interconnect 1702. In at least one embodiment, at least some of the packets represent commands configured to cause PPU 1700A to perform various operations. In at least one embodiment, I / O unit 1706 sends the decoded commands to various other units of PPU 1700A as specified by the commands. In at least one embodiment, the commands are sent to front end unit 1710 and / or to hub 1716 or other units of PPU 1700A, such as one or more replication engines, video encoders, video decoders, power management units, etc. Figure 17A In at least one embodiment, I / O unit 1706 is configured to route communications between the various logical units of PPU 1700A.
[0301] In at least one embodiment, a program executed by a host processor encodes a command stream in a buffer that provides a workload to the PPU 1700A for processing. In at least one embodiment, the workload includes instructions and data to be processed by those instructions. In at least one embodiment, the buffer is an area in memory that is accessible (e.g., read / write) by both the host processor and the PPU 1700A—the host interface unit can be configured to access the buffer in system memory connected to the interconnect 1702 via memory requests transmitted via the I / O unit 1706 through the interconnect 1702. In at least one embodiment, the host processor writes the command stream into the buffer and then sends a pointer indicating the beginning of the command stream to the PPU 1700A, so that the front end unit 1710 receives pointers to one or more command streams and manages the one or more command streams, reading commands from the command streams and forwarding the commands to the various units of the PPU 1700A.
[0302] In at least one embodiment, the front end unit 1710 is coupled to a scheduler unit 1712 that configures the various GPCs 1718 to process tasks defined by one or more command streams. In at least one embodiment, the scheduler unit 1712 is configured to track state information related to the various tasks managed by the scheduler unit 1712, where the state information may indicate which GPC 1718 the task is assigned to, whether the task is active or inactive, a priority associated with the task, and the like. In at least one embodiment, the scheduler unit 1712 manages multiple tasks that execute on one or more GPCs 1718.
[0303] In at least one embodiment, scheduler unit 1712 is coupled to work distribution unit 1714, which is configured to dispatch tasks for execution on GPCs 1718. In at least one embodiment, work distribution unit 1714 tracks a plurality of scheduled tasks received from scheduler unit 1712 and manages a pending task pool and an active task pool for each GPC 1718. In at least one embodiment, the pending task pool includes a plurality of time slots (e.g., 16 time slots) containing tasks assigned to be processed by a particular GPC 1718; the active task pool may include a plurality of time slots (e.g., 4 time slots) for tasks actively being processed by GPC 1718, such that as a task in a GPC 1718 completes execution, the task is evicted from the active task pool of GPC 1718, and one of the other tasks is selected from the pending task pool and scheduled for execution on GPC 1718. In at least one embodiment, if an active task is idle on GPC 1718, such as while waiting for data dependencies to be resolved, the active task is evicted from GPC 1718 and returned to the pending task pool, while another task in the pending task pool is selected and scheduled for execution on GPC 1718.
[0304] In at least one embodiment, work distribution unit 1714 communicates with one or more GPCs 1718 via XBar 1720. In at least one embodiment, XBar 1720 is an interconnect network that couples many units of PPU 1700A to other units of PPU 1700A and can be configured to couple work distribution unit 1714 to a specific GPC 1718. In at least one embodiment, one or more other units of PPU 1700A can also be connected to XBar 1716 via hub 1716.
[0305] In at least one embodiment, tasks are managed by a scheduler unit 1712 and assigned to one of the GPCs 1718 by a work distribution unit 1714. The GPCs 1718 are configured to process tasks and produce results. In at least one embodiment, the results can be consumed by other tasks in the GPC 1718, routed to a different GPC 1718 via the XBar 1716, or stored in memory 1704. In at least one embodiment, the results can be written to the memory 1704 via a partition unit 1722, which implements a memory interface for writing data to or reading data from the memory 1704. In at least one embodiment, the results can be transferred to another PPU 1704 or a CPU via a high-speed GPU interconnect 1708. In at least one embodiment, the PPU 1700A includes, but is not limited to, U partition units 1722, where the partition units 1722 are equal to the number of separate and distinct memory devices 1704 coupled to the PPU 1700A. In at least one embodiment, the following will be combined with Figure 17C The partition unit 1722 is described in more detail.
[0306] In at least one embodiment, the host processor executes a driver core that implements an application programming interface (API) that enables one or more applications executing on the host processor to schedule operations for execution on the PPU 1700A. In one embodiment, multiple computing applications are executed simultaneously by the PPU 1700A, and the PPU 1700A provides isolation, quality of service ("QoS"), and independent address spaces for the multiple computing applications. In at least one embodiment, the application generates instructions (e.g., in the form of API calls) that cause the driver core to generate one or more tasks for execution by the PPU 1700A, and the driver core outputs the tasks to one or more streams processed by the PPU 1700A. In at least one embodiment, each task includes one or more related groups of threads, which may be referred to as warps. In at least one embodiment, a warp includes multiple related threads (e.g., 32 threads) that can be executed in parallel. In at least one embodiment, a cooperative thread may refer to multiple threads that include instructions for performing tasks and exchanging data through shared memory, in combination with Figure 17C Threads and cooperating threads are described in greater detail according to at least one embodiment.
[0307] The reasoning and / or training logic 615 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 6B and / or Figure 6CProvides details regarding inference and / or training logic 615. In at least one embodiment, the deep learning application processor is used to train a machine learning model (such as a neural network) to predict or infer information provided to the PPU 1700A. In at least one embodiment, the PPU 1700A is used to infer or predict information based on a trained machine learning model (e.g., a neural network) that has been trained by another processor or system or the PPU 1700A. In at least one embodiment, the PPU 1700A can be used to perform one or more of the neural network use cases described herein.
[0308] Figure 17B A general processing cluster ("GPC") 1700B is shown in accordance with at least one embodiment. In at least one embodiment, GPC 1700B is Figure 17A 1718. In at least one embodiment, each GPC 1700B includes, but is not limited to, multiple hardware units for processing tasks, and each GPC 1700B includes, but is not limited to, a pipeline manager 1702, a pre-raster operations unit ("PROP") 1704, a raster engine 1708, a work distribution crossbar ("WDX") 1716, a memory management unit ("MMU") 1718, one or more data processing clusters ("DPCs") 1706, and any suitable combination of components.
[0309] In at least one embodiment, the operation of GPC 1700B is controlled by pipeline manager 1702. In at least one embodiment, pipeline manager 1702 manages the configuration of one or more DPCs 1706 to process tasks assigned to GPC 1700B. In at least one embodiment, pipeline manager 1702 configures at least one of one or more DPCs 1706 to implement at least a portion of a graphics rendering pipeline. In at least one embodiment, DPC 1706 is configured to execute vertex shader programs on a programmable streaming multiprocessor ("SM") 1714. In at least one embodiment, pipeline manager 1702 is configured to route packets received from a work distribution unit to appropriate logic within GPC 1700B, and in at least one embodiment, some packets may be routed to fixed-function hardware units in PROP 1704 and / or raster engine 1708, while other packets may be routed to DPC 1706 for processing by primitive engine 1712 or SM 1714. In at least one embodiment, pipeline manager 1702 configures at least one of DPCs 1706 to implement a neural network model and / or a computational pipeline.
[0310] In at least one embodiment, PROP unit 1704 is configured to route data generated by raster engine 1708 and DPC 1706 to a raster operations ("ROP") unit in partition unit 1722 in at least one embodiment, in conjunction with Figure 17A Described in more detail. In at least one embodiment, PROP unit 1704 is configured to perform optimizations for color blending, organize pixel data, perform address translation, and the like. In at least one embodiment, raster engine 1708 includes, but is not limited to, a plurality of fixed-function hardware units configured to perform various raster operations, and in at least one embodiment, raster engine 1708 includes, but is not limited to, a setup engine, a coarse raster engine, a culling engine, a clipping engine, a fine raster engine, a tile aggregation engine, and any suitable combination thereof. In at least one embodiment, the setup engine receives transformed vertices and generates plane equations associated with the geometric primitives defined by the vertices; the plane equations are passed to the coarse raster engine to generate coverage information for the primitives (e.g., an x, y coverage mask for the tile); the output of the coarse raster engine is passed to the culling engine, where fragments associated with primitives that fail the z test are culled, and to the clipping engine, where fragments outside the viewing frustum are clipped. In at least one embodiment, the clipped and culled fragments are passed to the fine raster engine to generate properties for the pixel fragments based on the plane equations generated by the setup engine. In at least one embodiment, the output of raster engine 1708 includes fragments to be processed by any appropriate entity (e.g., by a fragment shader implemented within DPC 1706).
[0311] In at least one embodiment, each DPC 1706 included in a GPC 1700B includes, but is not limited to, an M-pipeline controller ("MPC") 1710; a primitive engine 1712; one or more SMs 1714; and any suitable combination thereof. In at least one embodiment, the MPC 1710 controls the operation of the DPC 1706, routing packets received from the pipeline manager 1702 to appropriate units within the DPC 1706. In at least one embodiment, packets associated with vertices are routed to the primitive engine 1712, which is configured to fetch vertex attributes associated with the vertices from memory; conversely, packets associated with shader programs may be sent to the SM 1714.
[0312] In at least one embodiment, SM 1714 includes, but is not limited to, a programmable streaming processor configured to process tasks represented by multiple threads. In at least one embodiment, SM 1714 is multithreaded and configured to simultaneously execute multiple threads (e.g., 32 threads) from a particular thread group and implements a single instruction, multiple data ("SIMD") architecture, in which each thread in a group of threads (e.g., a warp) is configured to process a different set of data based on the same instruction set. In at least one embodiment, all threads in a thread group execute the same instructions. In at least one embodiment, SM 1714 implements a single instruction, multiple thread ("SIMT") architecture, in which each thread in a group of threads is configured to process a different set of data based on the same instruction set, but in which individual threads in a thread group are allowed to diverge during execution. In at least one embodiment, a program counter, call stack, and execution state are maintained for each warp, thereby enabling concurrency between warps and serial execution within a warp when threads in the warp diverge. In another embodiment, a program counter, call stack, and execution state are maintained for each individual thread, thereby enabling equal concurrency between all threads within a warp and between warps. In at least one embodiment, execution state is maintained for each individual thread, and threads of the same instruction can be converged and executed in parallel to improve efficiency. At least one embodiment of SM 1714 is described in more detail below.
[0313] In at least one embodiment, the MMU 1718 is used between the GPC 1700B and the memory partition unit (e.g., Figure 17A The MMU 1718 provides an interface between the memory and the partition unit 1722, and provides virtual to physical address translation, memory protection, and arbitration of memory requests. In at least one embodiment, the MMU 1718 provides one or more translation lookaside buffers ("TLBs") for performing translation of virtual addresses to physical addresses in memory.
[0314] The reasoning and / or training logic 615 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 6B and / or Figure 6C Provides details regarding inference and / or training logic 615. In at least one embodiment, the deep learning application processor is used to train a machine learning model (such as a neural network) to predict or infer information provided to the GPC 1700B. In at least one embodiment, the GPC 1700B is used to infer or predict information based on a machine learning model (e.g., a neural network) that has been trained by another processor or system or the GPC 1700B. In at least one embodiment, the GPC 1700B can be used to perform one or more of the neural network use cases described herein.
[0315] Figure 17C A memory partition unit 1700C of a parallel processing unit ("PPU") is shown in accordance with at least one embodiment. In at least one embodiment, the memory partition unit 1700C includes, but is not limited to, a raster operations ("ROP") unit 1702; a level 2 ("L2") cache 1704; a memory interface 1706; and any suitable combination thereof. In at least one embodiment, the memory interface 1706 is coupled to a memory. In at least one embodiment, the memory interface 1706 can implement a 32-, 64-, 128-, 1024-bit data bus, or a similar implementation for high-speed data transfer. In at least one embodiment, the PPU includes U memory interfaces 1706, one memory interface 1706 for each pair of partition units 1700C, wherein each pair of partition units 1700C is connected to a corresponding memory device. In at least one embodiment, the PPU can be connected to up to Y memory devices, such as a high-bandwidth memory stack or graphics double data rate version 5 synchronous dynamic random access memory ("GDDR5 SDRAM").
[0316] In at least one embodiment, the memory interface 1706 implements a High Bandwidth Memory 2nd Generation ("HBM2") memory interface, and Y is equal to half of U. In at least one embodiment, the HBM2 memory stack is located on the same physical package as the PPU, providing significant power and area savings compared to GDDR5 SDRAM systems. In at least one embodiment, each HBM2 stack includes, but is not limited to, four memory dies, and Y=4, each HBM2 stack includes two 128-bit channels per die for a total of 8 channels and a data bus width of 1024 bits. In at least one embodiment, the memory supports single error correction double error detection ("SECDED") error correction code ("ECC") to protect data. In at least one embodiment, ECC provides higher reliability for computing applications that are sensitive to data corruption.
[0317] In at least one embodiment, the PPU implements a multi-level memory hierarchy. In at least one embodiment, the memory partition unit 1700C supports unified memory to provide a single unified virtual address space for the central processing unit ("CPU") and PPU memory, thereby enabling data sharing between virtual memory systems. In at least one embodiment, the frequency of PPU accesses to memory located on other processors is tracked to ensure that memory pages are moved to the physical memory of the PPU that accesses the pages more frequently. In at least one embodiment, the high-speed GPU interconnect 1708 supports address translation services that allow the PPU to directly access the CPU's page tables and provide full access to the CPU's memory through the PPU.
[0318] In at least one embodiment, the copy engine transfers data between multiple PPUs or between a PPU and a CPU. In at least one embodiment, the copy engine can generate a page fault for an address that is not mapped in a page table, and the memory partition unit 1700C then services the page fault, maps the address into a page table, and then the copy engine performs the transfer. In at least one embodiment, fixed (in at least one embodiment, non-pageable) memory is used for multiple copy engine operations between multiple processors, thereby substantially reducing the available memory. In at least one embodiment, in the event of a hardware page fault, the address can be passed to the copy engine without regard to whether it resides in the memory page, and the copy process is transparent.
[0319] According to at least one embodiment, Figure 17A Data from memory 1704 or other system memory is retrieved by memory partition unit 1700C and stored in L2 cache 1704, which is located on-chip and shared between the various GPCs. In at least one embodiment, each memory partition unit 1700C includes, but is not limited to, at least a portion of an L2 cache associated with a corresponding memory device. In at least one embodiment, lower-level caches are implemented in various units within a GPC. In at least one embodiment, each SM 1714 may implement a level 1 ("L1") cache, where the L1 cache is private memory dedicated to a particular SM 1714, and data is retrieved from L2 cache 1704 and stored in each L1 cache for processing within the functional units of SM 1714. In at least one embodiment, L2 cache 1704 is coupled to memory interface 1706 and XBar 1720.
[0320] In at least one embodiment, ROP unit 1702 performs graphics raster operations related to pixel color, such as color compression, pixel blending, and the like. In at least one embodiment, ROP unit 1702 performs depth testing in conjunction with raster engine 1708, receiving the depth of a sample location associated with a pixel fragment from the culling engine of raster engine 1708. In at least one embodiment, the depth is tested against the corresponding depth in the depth buffer associated with the sample location of the fragment. In at least one embodiment, if the fragment passes the depth test for the sample location, ROP unit 1702 updates the depth buffer and sends the result of the depth test to raster engine 1708. It will be appreciated that the number of partition units 1700C can differ from the number of GPCs, and therefore, in at least one embodiment, each ROP unit 1702 can be coupled to each GPC. In at least one embodiment, ROP unit 1702 tracks packets received from different GPCs and determines to which result generated by ROP unit 1702 via XBar 1720 to be routed.
[0321] Figure 17D Streaming multiprocessor ("SM") 1700D is shown in accordance with at least one embodiment. In at least one embodiment, SM 1700D is Figure 17BSM. In at least one embodiment, SM 1700D includes, but is not limited to, an instruction cache 1702; one or more scheduler units 1704; a register file 1708; one or more processing cores ("cores") 1710; one or more special function units ("SFUs") 1712; one or more load / store units ("LSUs") 1714; an interconnect network 1716; a shared memory / level 1 ("L1") cache 1718; and any suitable combination thereof. In at least one embodiment, a work distribution unit schedules tasks for execution on a general processing cluster ("GPC") of a parallel processing unit ("PPU"), with each task being assigned to a specific data processing cluster ("DPC") within the GPC, and if the task is associated with a shader program, the task is assigned to one of SMs 1700D. In at least one embodiment, scheduler unit 1704 receives tasks from the work distribution unit and manages instruction scheduling for one or more thread blocks assigned to SM 1700D. In at least one embodiment, the scheduler unit 1704 schedules thread blocks for execution as warps of parallel threads, where each thread block is assigned at least one warp. In at least one embodiment, each warp executes a thread. In at least one embodiment, the scheduler unit 1704 manages a plurality of different thread blocks, assigns warps to different thread blocks, and then dispatches instructions from a plurality of different cooperating groups to various functional units (e.g., processing cores 1710, SFUs 1712, and LSUs 1714) during each clock cycle.
[0322] In at least one embodiment, cooperative groups can refer to a programming model for organizing groups of communicating threads that allows developers to express the granularity at which threads are communicating, thereby enabling the expression of richer, more efficient decompositions of parallelism. In at least one embodiment, a cooperative launch API supports synchronization between thread blocks to execute parallel algorithms. In at least one embodiment, applications of the programming model provide a single, simple construct for synchronizing cooperating threads: a barrier across all threads of a thread block (e.g., the syncthreads() function). However, in at least one embodiment, programmers can define thread groups at a granularity smaller than a thread block and synchronize within the defined group to achieve higher performance, design flexibility, and software reuse in the form of a collective group-wide function interface. In at least one embodiment, cooperative groups enable programmers to explicitly define thread groups at sub-block (in at least one embodiment, down to a single thread) and multi-block granularity and perform collective operations, such as synchronizing threads within a cooperative group. In at least one embodiment, the programming model supports clean composition across software boundaries, allowing libraries and utility functions to safely synchronize in their local environment without making assumptions about convergence. In at least one embodiment, the cooperation group primitive enables new patterns of cooperative parallelism, including but not limited to producer-consumer parallelism, opportunistic parallelism, and global synchronization across an entire grid of thread blocks.
[0323] In at least one embodiment, the scheduling unit 1706 is configured to send instructions to one or more of the functional units, and the scheduler unit 1704 includes, but is not limited to, two scheduling units 1706 that enable two different instructions from the same warp to be scheduled per clock cycle. In at least one embodiment, each scheduler unit 1704 includes a single scheduling unit 1706 or additional scheduling units 1706.
[0324] In at least one embodiment, each SM 1700D includes, in at least one embodiment, but is not limited to, a register file 1708 that provides a set of registers for the functional units of the SM 1700D. In at least one embodiment, the register file 1708 is partitioned between each functional unit, such that each functional unit is allocated a dedicated portion of the register file 1708. In at least one embodiment, the register file 1708 is partitioned between the different warps executed by the SM 1700D, and the register file 1708 provides temporary storage for operands connected to the data paths of the functional units. In at least one embodiment, each SM 1700D includes, in at least one embodiment, but is not limited to, a plurality of L processing cores 1710. In at least one embodiment, the SM 1700D includes, but is not limited to, a large number (e.g., 128 or more) of different processing cores 1710. In at least one embodiment, each processing core 1710 includes, but is not limited to, a fully pipelined, single-precision, double-precision, and / or mixed-precision processing unit, including, but not limited to, a floating-point arithmetic logic unit and an integer arithmetic logic unit. In at least one embodiment, the floating-point arithmetic logic unit implements IEEE 802.11 for floating-point arithmetic.
[0325] In at least one embodiment, processing cores 1710 include, but are not limited to, 64 single-precision (32-bit) floating point cores, 64 integer cores, 32 double-precision (64-bit) floating point cores, and 8 tensor cores.
[0326] According to at least one embodiment, the tensor core is configured to perform matrix operations. In at least one embodiment, one or more tensor cores are included in processing core 1710. In at least one embodiment, the tensor core is configured to perform deep learning matrix arithmetic, such as convolution operations for neural network training and inference. In at least one embodiment, each tensor core operates on a 4×4 matrix and performs a matrix multiplication and accumulation operation D=A×B+C, where A, B, C, and D are 4×4 matrices.
[0327] In at least one embodiment, the matrix multiplication inputs A and B are 16-bit floating-point matrices, and the accumulation matrices C and D are 16-bit floating-point or 32-bit floating-point matrices. In at least one embodiment, the tensor core performs a 32-bit floating-point accumulation operation on the 16-bit floating-point input data. In at least one embodiment, the 16-bit floating-point multiplication uses 64 operations and results in a full-precision product, which is then accumulated with other intermediate products using 32-bit floating-point addition to perform a 4x4x4 matrix multiplication. In at least one embodiment, the tensor core is used to perform larger two-dimensional or higher-dimensional matrix operations composed of these smaller elements. In at least one embodiment, an API (such as the CUDA 9 C++ API) exposes specialized matrix load, matrix multiplication and accumulation, and matrix store operations to efficiently use the tensor cores from a CUDA-C++ program. In at least one embodiment, at the CUDA level, the warp-level interface assumes a 16×16 matrix size that spans all 32 warp threads.
[0328] In at least one embodiment, each SM 1700D includes, but is not limited to, M SFUs 1712 that perform specialized functions (e.g., attribute evaluation, reciprocal square root, etc.). In at least one embodiment, SFUs 1712 include, but are not limited to, tree traversal units configured to traverse a hierarchical tree data structure. In at least one embodiment, SFUs 1712 include, but are not limited to, texture units configured to perform texture map filtering operations. In at least one embodiment, the texture units are configured to load texture maps (e.g., 2D arrays of texels) from memory and sample the texture maps to generate sampled texture values for use by shader programs executed by the SM 1700D. In at least one embodiment, the texture maps are stored in shared memory / L1 cache 1718. In at least one embodiment, the texture units implement texture operations (such as filtering operations) using mip-maps (e.g., texture maps with different levels of detail). In at least one embodiment, each SM 1700D includes, but is not limited to, two texture units.
[0329] In at least one embodiment, each SM 1700D includes, but is not limited to, N LSUs 1714 that implement load and store operations between the shared memory / L1 cache 1718 and the register file 1708. In at least one embodiment, an interconnection network 1716 connects each functional unit to the register file 1708, and the LSUs 1714 connect to both the register file 1708 and the shared memory / L1 cache 1718. In at least one embodiment, the interconnection network 1716 is a crossbar switch that can be configured to connect any functional unit to any register in the register file 1708, and to connect the LSUs 1714 to memory locations in the register file 1708 and the shared memory / L1 cache 1718.
[0330] In at least one embodiment, shared memory / L1 cache 1718 is an array of on-chip memory that, in at least one embodiment, allows for data storage and communication between SM 1700D and primitive engines, as well as between threads within SM 1700D. In at least one embodiment, shared memory / L1 cache 1718 includes, but is not limited to, 128KB of storage capacity and is located in the path from SM 1700D to the partition unit. In at least one embodiment, shared memory / L1 cache 1718 is used to cache reads and writes in at least one embodiment. In at least one embodiment, one or more of shared memory / L1 cache 1718, L2 cache, and memory is a backing store.
[0331] In at least one embodiment, combining data cache and shared memory functionality into a single memory block provides improved performance for both types of memory accesses. In at least one embodiment, capacity is used by programs that do not use the shared memory or use it as a cache, for example, if the shared memory is configured to use half of its capacity, and texture and load / store operations can use the remaining capacity. According to at least one embodiment, integration within shared memory / L1 cache 1718 enables shared memory / L1 cache 1718 to function as a high-throughput pipeline for streaming data, while providing high-bandwidth and low-latency access to frequently reused data. In at least one embodiment, when configured for general-purpose parallel computing, a simpler configuration can be used compared to graphics processing. In at least one embodiment, the fixed-function graphics processing unit is bypassed, creating a simpler programming model. In at least one embodiment, in a general-purpose parallel computing configuration, the work distribution unit directly allocates and distributes blocks of threads to DPCs. In at least one embodiment, the threads in a block execute a common program, use unique thread IDs in computations to ensure that each thread generates unique results, use SM 1700D to execute the program and perform computations, use shared memory / L1 cache 1718 to communicate between threads, and use LSU 1714 to read and write global memory through shared memory / L1 cache 1718 and a memory partitioning unit. In at least one embodiment, when configured for general parallel computation, SM 1700D writes commands to scheduler unit 1704 that can be used to start new work on a DPC.
[0332] In at least one embodiment, the PPU is included in or coupled to a desktop computer, a laptop computer, a tablet computer, a server, a supercomputer, a smartphone (e.g., wireless, handheld device), a personal digital assistant ("PDA"), a digital camera, a vehicle, a head-mounted display, a handheld electronic device, etc. In at least one embodiment, the PPU is implemented on a single semiconductor substrate. In at least one embodiment, the PPU is included in a system-on-chip ("SoC") along with one or more other devices (e.g., an additional PPU, memory, a reduced instruction set computer ("RISC") CPU, one or more memory management units ("MMUs"), a digital-to-analog converter ("DAC"), etc.
[0333] In at least one embodiment, the PPU can be included on a graphics card that includes one or more storage devices. The graphics card can be configured to connect to a PCIe slot on a desktop computer motherboard. In at least one embodiment, the PPU can be an integrated graphics processing unit ("iGPU") included in a chipset on the motherboard.
[0334] Reasoning and / or training logic 615 is used to perform reasoning and / or training operations related to one or more embodiments. Figure 6B and / or Figure 6C Provides details regarding inference and / or training logic 615. In at least one embodiment, the deep learning application processor is used to train a machine learning model (such as a neural network) to predict or infer information provided to the SM 1700D. In at least one embodiment, the SM 1700D is used to infer or predict information based on a machine learning model (e.g., a neural network) that has been trained by another processor or system or by the SM 1700D. In at least one embodiment, the SM 1700D can be used to perform one or more of the neural network use cases described herein.
[0335] In at least one embodiment, a single semiconductor platform can refer to a single semiconductor-based integrated circuit or chip. In at least one embodiment, a multi-chip module with increased connectivity can be used, which emulates on-chip operations and provides substantial improvements over implementations utilizing a central processing unit ("CPU") and a bus. In at least one embodiment, the various modules can also be placed separately or in various combinations of semiconductor platforms, depending on user needs.
[0336] In at least one embodiment, a computer program in the form of machine-readable executable code or computer control logic algorithms is stored in main memory 4ee04 and / or auxiliary storage. According to at least one embodiment, if executed by one or more processors, the computer program enables system 4ee00 to perform various functions. In at least one embodiment, memory 4ee04, storage, and / or any other storage are possible examples of computer-readable media. In at least one embodiment, auxiliary storage can refer to any suitable storage device or system, such as a hard drive and / or removable storage drive, which represents a floppy disk drive, a tape drive, an optical drive, a digital versatile disk ("DVD") drive, a recording device, a universal serial bus ("USB") flash memory, etc. In at least one embodiment, the architecture and / or functionality of each of the previous figures is implemented in the context of a CPU 4ee02; a parallel processing system 4ee12; an integrated circuit capable of having at least some of the capabilities of two CPUs 4ee02; a parallel processing system 4ee12; a chipset (e.g., a group of integrated circuits designed to operate and sell as a unit to perform related functions, etc.); and any suitable combination of integrated circuits.
[0337] In at least one embodiment, the architecture and / or functionality of the various previous figures are implemented in the context of a general-purpose computer system, a circuit board system, a game console system dedicated for entertainment purposes, a dedicated system, etc. In at least one embodiment, the computer system 4ee00 can take the form of a desktop computer, a laptop computer, a tablet computer, a server, a supercomputer, a smartphone (e.g., wireless, handheld device), a personal digital assistant ("PDA"), a digital camera, a vehicle, a head-mounted display, a handheld electronic device, a mobile telephone device, a television, a workstation, a game console, an embedded system, and / or any other type of logic.
[0338] In at least one embodiment, the parallel processing system 4ee12 includes, but is not limited to, a plurality of parallel processing units ("PPUs") 4ee14 and associated memory 4ee16. In at least one embodiment, the PPUs 4ee14 are connected to a host processor or other peripheral device via an interconnect 4ee18 and a switch 4ee20 or multiplexer. In at least one embodiment, the parallel processing system 4ee12 distributes computational tasks across the parallelizable PPUs 4ee14, for example, as part of a distribution of computational tasks across multiple graphics processing unit ("GPU") thread blocks. In at least one embodiment, memory is shared and accessed (e.g., for read and / or write access) between some or all of the PPUs 4ee14, although such shared memory may incur a performance penalty relative to using local memory and registers resident on the PPUs 4ee ... using commands such as
[0339] __syncthreads()) to synchronize the operation of the PPU 4ee14, where all threads in a block (e.g., executing across multiple PPUs 4ee14) reach a certain code execution point before proceeding.
[0340] Other variations are within the spirit of the present disclosure. Thus, while the disclosed technology is susceptible to various modifications and alternative constructions, certain illustrated embodiments thereof are shown in the drawings and have been described above in detail. However, it should be understood that there is no intention to limit the disclosure to one or more specific forms disclosed, but on the contrary, the intention is to cover all modifications, alternative constructions, and equivalents falling within the spirit and scope of the present disclosure as defined by the appended claims.
[0341] Unless otherwise noted or clearly contradicted by the context, the use of the terms "a" and "an" and "the" and similar references in the context of describing the disclosed embodiments (particularly in the context of the appended claims) should be interpreted as covering the singular and plural, rather than as definitions of terms. Unless otherwise noted, the terms "include," "have," "include," and "contain" should be interpreted as open-ended terms (meaning "including but not limited to"). The term "connected" (when unmodified, refers to a physical connection) should be interpreted as partially or completely contained within, attached to, or connected together, even if there is some intervention. Unless otherwise noted herein, references to numerical ranges herein are intended only to be used as a shorthand method of referring to each individual value falling within the range, and each individual value is incorporated into the specification as if it were separately recited herein. Unless otherwise noted or contradicted by the context, the use of the term "set" (e.g., "item set") or "subset" should be interpreted as a non-empty set comprising one or more members. Furthermore, unless otherwise indicated or contradicted by context, a "subset" of a corresponding set does not necessarily mean a proper subset of the corresponding set, but rather a subset and a corresponding set may be equivalent.
[0342] Unless expressly indicated otherwise or clearly contradicted by context, conjunctions such as phrases of the form "at least one of A, B, and C" or "at least one of A, B, and C" are understood in context to generally refer to an item, clause, or the like, which may be A or B or C, or any non-empty subset of the set A, B, and C. For example, in the illustrative example of a set having three members, the conjunctions "at least one of A, B, and C" and "at least one of A, B, and C" refer to any of the following sets: {A}, {B}, {C}, {A, B}, {A, C}, {B, C}, {A, B, C}. Thus, such conjunctions may not be intended to imply that certain embodiments require the presence of at least one of A, at least one of B, and at least one of C. Additionally, unless expressly indicated otherwise or contradicted by context, "plurality" refers to plurality (e.g., "a plurality of items" refers to a plurality of items). The plural is at least two items, but may be more if expressly indicated or indicated by context. Further, unless stated otherwise or clear from context, "based on" means "based at least in part on" rather than "based solely on."
[0343] Unless otherwise indicated herein or clearly contradicted by the context, the operations of the processes described herein may be performed in any suitable order. In at least one embodiment, processes such as those herein (or variations thereof and / or combinations thereof) are performed under the control of one or more computer systems configured with executable instructions and are implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that are collectively executed on one or more processors by hardware or a combination thereof. In at least one embodiment, the code is stored on a computer-readable storage medium in the form of, for example, a computer program that includes a plurality of instructions that can be executed by one or more processors. In at least one embodiment, the computer-readable storage medium is a non-transitory computer-readable storage medium that excludes transient signals (e.g., propagated transient electrical or electromagnetic transmissions) but includes non-transitory data storage circuits (e.g., buffers, caches, and queues). In at least one embodiment, code (e.g., executable code or source code) is stored on a set of one or more non-transitory computer-readable storage media (or other memory for storing executable instructions) having executable instructions stored thereon, which, when executed by one or more processors of a computer system (in at least one embodiment, as a result of being executed), causes the computer system to perform the operations described herein. In at least one embodiment, the set of non-transitory computer-readable storage media includes a plurality of non-transitory computer-readable storage media, and one or more of the individual non-transitory storage media in the plurality of non-transitory computer-readable storage media lacks all of the code, but rather the plurality of non-transitory computer-readable storage media collectively store all of the code. In at least one embodiment, the executable instructions are executed so that different instructions are executed by different processors, for example, a non-transitory computer-readable storage medium stores instructions, and a main central processing unit ("CPU") executes some instructions, while a graphics processing unit ("GPU") executes other instructions. In at least one embodiment, different components of the computer system have separate processors, and different processors execute different subsets of instructions.
[0344] Thus, in at least one embodiment, a computer system is configured to implement one or more services that individually or collectively perform the operations of the processes herein, and such a computer system is configured with applicable hardware and / or software that enables the implementation of the operations. Furthermore, the computer system implementing at least one embodiment of the present disclosure is a single device, and in another embodiment is a distributed computer system comprising multiple devices operating in different ways such that the distributed computer system performs the operations herein and such that no single device performs all of the operations.
[0345] The use of any and all examples or exemplary language (e.g., "such as") provided herein is intended merely to better illuminate embodiments of the present disclosure and does not limit the scope of the disclosure unless otherwise required. No language in the specification should be construed as indicating any non-claimed element as essential to the practice of the disclosure.
[0346] All references, including publications, patent applications, and patents, cited herein are hereby incorporated by reference to the same extent as if each reference were individually and specifically indicated to be incorporated by reference and were set forth in its entirety herein.
[0347] In the description and claims, the terms "coupled" and "connected," along with their derivatives, may be used. It should be understood that these terms may not be intended as synonyms for each other. Rather, in specific examples, "connected" or "coupled" may be used to indicate that two or more elements are in direct or indirect physical or electrical contact with each other. "Coupled" may also mean that two or more elements are not in direct contact with each other, but still cooperate or interact with each other.
[0348] Unless expressly stated otherwise, it is understood that throughout this specification, references to “processing,” “computing,” “calculating,” “determining,” and the like refer to the actions and / or processes of a computer or computing system or similar electronic computing device that processes and / or converts data represented as physical quantities (e.g., electronic) in registers and / or memories of the computing system into other data similarly represented as physical quantities in the computing system's memories, registers, or other such information storage, transmission, or display devices.
[0349] In a similar manner, a "processor" may refer to any device or portion of a memory that processes electronic data from registers and / or memory and converts that electronic data into other electronic data that can be stored in registers and / or memory. As non-limiting examples, a "processor" may be a CPU or a GPU. A "computing platform" may include one or more processors. As used herein, a "software" process may include, for example, software and / or hardware entities that perform work over time, such as tasks, threads, and intelligent agents. Likewise, each process may refer to multiple processes to execute instructions continuously or intermittently, sequentially, or in parallel. The terms "system" and "method" may be used interchangeably herein, as long as a system may embody one or more methods, and a method may be considered a system.
[0350] In this document, reference may be made to obtaining, acquiring, receiving, or inputting analog or digital data into a subsystem, a computer system, or a computer-implemented machine. Obtaining, acquiring, receiving, or inputting analog and digital data can be accomplished in a variety of ways, such as by receiving data as parameters of a function call or a call to an application programming interface. In some implementations, the process of obtaining, acquiring, receiving, or inputting analog or digital data can be accomplished by transmitting data via a serial or parallel interface. In another implementation, the process of obtaining, acquiring, receiving, or inputting analog or digital data can be accomplished by transmitting data from a providing entity to an acquiring entity via a computer network. Reference may also be made to providing, outputting, transmitting, sending, or presenting analog or digital data. In various examples, the process of providing, outputting, transmitting, sending, or presenting analog or digital data can be accomplished by transmitting data as input or output parameters of a function call, an application programming interface, or an interprocess communication mechanism.
[0351] Although the above discussion sets forth example implementations of the described technology, other architectures may be used to implement the described functionality and are intended to fall within the scope of this disclosure. In addition, although specific responsibilities are defined above for discussion purposes, the various functions and responsibilities may be allocated and divided in different ways depending on the circumstances.
[0352] Furthermore, although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter claimed in the appended claims is not necessarily limited to the specific features or acts described. Rather, the specific features and acts are disclosed as example forms of implementing the claims.
Claims
1. A server rack, comprising: a first controllable fluid coupler for receiving coolant from a cooling circuit external to the server rack; a cooling manifold comprising one or more second controllable fluid couplers press-fitted to one or more couplers of one or more server cooling manifolds; and Temperature sensors within the server racks are configured to provide input to a learning subsystem of the data center cooling system, the learning subsystem being configured to execute a machine learning model to: processing the input using a plurality of neuron levels of the machine learning model having a priori temperature values and a priori associated flow values, and providing an output to the first controllable fluid coupler; The learning subsystem also includes at least one processor for transmitting a control signal to one or more of the first controllable fluid coupler and at least one of the one or more second controllable fluid couplers to adjust a flow rate of the coolant flowing in the server rack or one or more server trays of the server rack based on sensor data from one or more sensors. 2 . The server rack of claim 1 , wherein the first controllable fluid coupler modifies a restriction on the flow of the coolant from the cooling circuit or within the cooling manifold in response to the output.
3. The server rack according to claim 1, further comprising: temperature sensors within the racks for providing input to the learning subsystem of the data center cooling system; as well as The learning subsystem executes machine learning models to: processing the input using a plurality of neuron levels of the machine learning model having a priori temperature values and a priori associated flow values; and An output is provided to at least one of the one or more second controllable fluid couplers.
4. The server rack of claim 3, wherein at least one of the one or more second controllable fluid couplers modifies at least one restriction to the flow of the coolant in the one or more server cooling manifolds within the one or more server trays of the server rack.
5. The server rack according to claim 1, further comprising: one or more temperature sensors associated with the server rack or one or more server trays in the server rack for providing at least a first input to the learning subsystem of the data center cooling system; one or more flow sensors associated with the cooling circuit, within the cooling manifold, or within the one or more server trays, for providing at least a second input to the learning subsystem; and The learning subsystem executes machine learning models to: processing the first input and the second input using multiple levels of the machine learning model having prior temperature values and associated prior flow values; as well as An output is provided to one or more of the first controllable fluid coupler or at least one of the one or more second controllable fluid couplers. 6 . The server rack of claim 1 , wherein the one or more couplers of the one or more server cooling manifolds are to the one or more server trays in the server rack.
7. The server rack of claim 1, wherein the cooling manifold within the server rack has at least one side coupler for coupling adjacent racks in a forward, sideways, or rearward direction.
8. The server rack of claim 1 , wherein the one or more server cooling manifolds extend within the one or more server trays and include one or more mating couplers for coupling with one or more device couplers from one or more devices in the server trays.
9. A method of cooling a data center, comprising: providing a cooling manifold within the rack, and providing a first controllable fluid coupler to receive coolant from a cooling circuit external to the rack; press-fitting one or more second controllable fluid couplers of the cooling manifold with one or more couplers of one or more server cooling manifolds; providing input from temperature sensors within the racks to a learning subsystem of a data center cooling system; Use the learning subsystem to execute machine learning models to: processing the input using a plurality of neuron levels of the machine learning model having a prior temperature value and a prior associated flow value; as well as providing an output to the first controllable fluid coupler; and The learning subsystem, comprising at least a processor, transmits a control signal to one or more of at least one of the first controllable fluid coupler and the second controllable fluid coupler to adjust a flow rate of the coolant flowing in the rack or one or more server trays of the rack based on sensor data from one or more sensors.
10. The method according to claim 9, further comprising: Responsive to the output, a restriction on the flow of the coolant from the cooling circuit or within the cooling manifold is modified.
11. The method according to claim 9, further comprising: providing input to the learning subsystem of the data center cooling system from temperature sensors within the racks; as well as Use the learning subsystem to execute machine learning models to: processing the input using a plurality of neuron levels of the machine learning model having a priori temperature values and a priori associated flow values; and An output is provided to at least one of the one or more second controllable fluid couplers.
12. The method according to claim 11, further comprising: At least one restriction to the flow of the coolant in one or more server cooling manifolds within the one or more server trays of the rack is modified.
13. The method according to claim 9, further comprising: providing at least one first input to the learning subsystem of the data center cooling system from one or more temperature sensors associated with the rack or the one or more server trays of the rack; providing at least one second input to the learning subsystem from one or more flow sensors associated with the cooling circuit, within the cooling manifold, or within the one or more server trays; and Use the learning subsystem to execute machine learning models to: processing the first input and the second input using multiple levels of the machine learning model having prior temperature values and associated prior flow values; as well as An output is provided to one or more of the first controllable fluid coupler or at least one of the one or more second controllable fluid couplers.
14. The method of claim 9, wherein the one or more couplers of the one or more server cooling manifolds are of the one or more server trays in the rack.
15. The method according to claim 9, further comprising: At least one side coupler is provided for the cooling manifold within the rack for enabling the rack to be coupled to an adjacent rack in a forward, sideways or rearward direction.
16. The method of claim 9, wherein the one or more server cooling manifolds extend within the one or more server trays of the rack and include one or more mating couplers to couple with one or more device couplers from one or more devices in the one or more server trays.
Citation Information
Patent Citations
Manifolds having slidable dripless connectors
CN105793797A
Intelligent manifold assemblies for a light source, light sources including intelligent manifold assemblies, and methods of operating the same
CN107110484A
Data center without interline air conditioner and heat radiation system thereof
CN108012513A