Server unit with built-in stream distribution

By introducing server units with built-in flow distribution in the data center cooling system and utilizing integrated flow controllers and quick-connect connectors, the problems of multiple leakage points and difficult maintenance in the piping configuration are solved, achieving efficient management and simplified maintenance of the cooling system.

CN115297669BActive Publication Date: 2025-10-10NVIDIA CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210436766.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-05-04
Filing Date
2022-04-24
Publication Date
2025-10-10
Estimated Expiration
2042-04-24

AI Technical Summary

Technical Problem

The piping configuration of existing data center cooling systems has many leakage points, difficult and time-consuming maintenance, and is difficult to organize and maintain efficiently.

Method used

Adopt server units with built-in flow distribution, through integrated flow controllers and quick-connect connectors, to achieve efficient distribution and management of cooling fluid, reduce leak points and simplify maintenance processes.

Benefits of technology

It improves the reliability and maintenance efficiency of the cooling system, reduces the risk of leakage, and simplifies the installation and maintenance process of the cooling system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115297669B_ABST
    Figure CN115297669B_ABST
Patent Text Reader

Abstract

Server units with built-in flow distribution are disclosed, specifically configurations for cooling systems are disclosed. In at least one embodiment, one or more distribution manifolds are formed within at least one panel of a server unit housing for receiving liquid and directing the liquid to one or more outlets that are coupled to one or more cold plates.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] At least one embodiment relates to cooling systems. For example, at least one embodiment relates to systems and methods for operating a cooling system in a data center. BACKGROUND

[0002] Data center cooling systems can use water or other cooling fluids to remove heat from computing devices. Cooling systems can include multiple different manifold systems, such as row-level, rack-level, or server-level. Manifold systems include multiple piping runs to provide inlet and outlet flow between different flow loops and computing devices. Piping configurations can provide multiple different leak points, be difficult to organize and maintain, and can also take a significant amount of operational time to establish. BRIEF DESCRIPTION OF DRAWINGS

[0003] Figure 1 A data center cooling system is shown in accordance with at least one embodiment;

[0004] Figure 2 A rack assembly for a data center is shown in accordance with at least one embodiment;

[0005] Figure 3 A rack assembly for a data center is shown in accordance with at least one embodiment;

[0006] Figure 4A An integrated distribution manifold is shown in accordance with at least one embodiment;

[0007] Figure 4B An integrated distribution manifold is shown in accordance with at least one embodiment;

[0008] Figure 4C An integrated distribution manifold is shown in accordance with at least one embodiment;

[0009] Figure 5A And Figure 5B A mounting sequence for a panel for a server enclosure is shown in accordance with at least one embodiment;

[0010] Figure 5C And Figure 5D A mounting sequence for a panel for a server enclosure is shown in accordance with at least one embodiment;

[0011] Figure 6 A distributed system is shown in accordance with at least one embodiment;

[0012] Figure 7 An exemplary data center is shown in accordance with at least one embodiment;

[0013] Figure 8 A client-server network is shown in accordance with at least one embodiment;

[0014] Figure 9 illustrates a computer network according to at least one embodiment;

[0015] Figure 10A illustrates a networked computer system according to at least one embodiment;

[0016] Figure 10B illustrates a networked computer system according to at least one embodiment;

[0017] Figure 10C illustrates a networked computer system according to at least one embodiment;

[0018] Figure 11 illustrates one or more components of a system environment in which services may be provided as third-party network services according to at least one embodiment;

[0019] Figure 12 illustrates a cloud computing environment according to at least one embodiment;

[0020] Figure 13 illustrates a set of functional abstraction layers provided by a cloud computing environment in accordance with at least one embodiment;

[0021] Figure 14 shows a supercomputer at a chip level according to at least one embodiment;

[0022] Figure 15 illustrates a supercomputer at the rack module level according to at least one embodiment;

[0023] Figure 16 illustrates a supercomputer at the rack level according to at least one embodiment;

[0024] Figure 17 illustrates a supercomputer at an overall system level according to at least one embodiment;

[0025] Figure 18A Inference and / or training logic according to at least one embodiment is shown;

[0026] Figure 18B Inference and / or training logic according to at least one embodiment is shown;

[0027] Figure 19 illustrates the training and deployment of a neural network according to at least one embodiment;

[0028] Figure 20 shows the architecture of a network system according to at least one embodiment;

[0029] Figure 21shows the architecture of a network system according to at least one embodiment;

[0030] Figure 22 illustrates a control plane protocol stack according to at least one embodiment;

[0031] Figure 23 illustrates a user plane protocol stack according to at least one embodiment;

[0032] Figure 24 Components of a core network according to at least one embodiment are shown;

[0033] Figure 25 Components of a system supporting network functions virtualization (NFV) according to at least one embodiment are shown;

[0034] Figure 26 A processing system according to at least one embodiment is shown;

[0035] Figure 27 A computer system according to at least one embodiment is shown;

[0036] Figure 28 A system according to at least one embodiment is shown;

[0037] Figure 29 An exemplary integrated circuit according to at least one embodiment is shown;

[0038] Figure 30 A computing system according to at least one embodiment is shown;

[0039] Figure 31 An APU is shown according to at least one embodiment;

[0040] Figure 32 A CPU according to at least one embodiment is shown;

[0041] Figure 33 An exemplary accelerator integrated slice is shown in accordance with at least one embodiment;

[0042] Figures 34A-34B An exemplary graphics processor is shown in accordance with at least one embodiment;

[0043] Figure 35A A graphics core according to at least one embodiment is shown;

[0044] Figure 35B GPGPU according to at least one embodiment is shown;

[0045] Figure 36A A parallel processor according to at least one embodiment is shown;

[0046] Figure 36B illustrates a processing cluster according to at least one embodiment;

[0047] Figure 36C A graphics multiprocessor is shown in accordance with at least one embodiment;

[0048] Figure 37 A software stack for a programming platform according to at least one embodiment is shown;

[0049] Figure 38 According to at least one embodiment, Figure 37 CUDA implementation of the software stack;

[0050] Figure 39 According to at least one embodiment, Figure 37 ROCm implementation of the software stack;

[0051] Figure 40 According to at least one embodiment, Figure 37 OpenCL implementation of the software stack;

[0052] Figure 41 illustrates software supported by a programming platform according to at least one embodiment; and

[0053] Figure 42 According to at least one embodiment, a method for Figures 37-40 Compiled code that is executed on the programming platform. DETAILED DESCRIPTION

[0054] In at least one embodiment, the Figure 1A data center 100 with a cooling system 102 is shown. In at least one embodiment, the data center 100 can be one or more rooms 104 with racks 106 and ancillary equipment for housing one or more servers on one or more server trays. In at least one embodiment, the data center 100 is supported by a cooling tower 108 located outside of the data center 100. In at least one embodiment, the cooling tower 108 dissipates heat from within the data center 100 by acting on a primary cooling loop 110. In at least one embodiment, a cooling distribution unit (CDU) 112 is used between the primary cooling loop 110 and a secondary or auxiliary cooling loop 114 for enabling heat absorption from the secondary or auxiliary cooling loop 114 to the primary cooling loop 110. In at least one embodiment, in one aspect, the auxiliary cooling loop 114 can tap into individual lines that go to the server trays as needed. In at least one embodiment, the loops 110, 114 are shown as line graphs, but one of ordinary skill will recognize that one or more line features can be used. In at least one embodiment, flexible polyvinyl chloride (PVC) tubing can be used with the associated lines to move fluid along each provided loop 110, 114. In at least one embodiment, one or more coolant pumps can be used to maintain a pressure differential within the coolant loops 110, 114 to enable coolant to move according to temperature sensors in various locations, including in the room, in one or more racks 106, and / or in a server box or server tray within one or more racks 106.

[0055] In at least one embodiment, the coolant in the primary cooling loop 110 and the auxiliary cooling loop 114 can be at least water and an additive. In at least one embodiment, the additive can be ethylene glycol or propylene glycol. In at least one embodiment, each of the primary cooling loop and the auxiliary cooling loop can have its own coolant. In at least one embodiment, the coolant in the auxiliary cooling loop can be dedicated to the needs of the components in the server trays or associated racks 106. In at least one embodiment, the CDU 112 is capable of complex control of the coolant independently or simultaneously within the provided coolant loops 110, 114. In at least one embodiment, the CDU 112 can be adapted to control the flow rate of the coolant such that the coolant is properly distributed to absorb heat generated within the associated racks 106. In at least one embodiment, more flexible tubing 116 is provided from the auxiliary cooling loop 114 to enter each server tray to provide coolant to electrical components and / or computing components therein.

[0056] In at least one embodiment, the ducts 118 forming part of the auxiliary cooling loop 114 may be referred to as room manifolds. In at least one embodiment, additional ducts 120 may extend from the row manifold ducts 118 and may also be part of the auxiliary cooling loop 114, but may be referred to as row manifolds. In at least one embodiment, coolant ducts 122 enter the racks as part of the auxiliary cooling loop 114, but may be referred to as rack cooling manifolds within one or more racks. In at least one embodiment, the row manifolds 128 extend along the rows in the data center 100 to all racks. In at least one embodiment, chillers 124 may be provided in the primary cooling loop within the data center 100 to support cooling prior to the cooling tower. In at least one embodiment, for the present disclosure, additional cooling circuits that may be present in the primary control loop and provide cooling external to the racks and external to the auxiliary cooling circuit may be employed in conjunction with the primary cooling circuit and distinct from the auxiliary cooling circuit.

[0057] In at least one embodiment, during operation, heat generated within the server trays of a given rack 106 can be transferred via the flexible tubing of the row manifold 124 of the secondary cooling circuit 114 to coolant exiting one or more racks 106. In at least one embodiment, the secondary coolant (in the secondary cooling circuit 114) from the CDU 112 used to cool the given racks 106 moves toward the one or more racks 106 via the provided tubing. In at least one embodiment, the secondary coolant from the CDU 112 is transferred from one side of the room manifold having tubing 128 via the row manifold 120 to one side of the rack 106 and passes through one side of the server trays via different tubing 122. In at least one embodiment, the spent or returned secondary coolant (or heat-carrying secondary coolant exiting the computing components) exits the other side of the server trays (such as entering the left side of the rack and exiting the right side of the rack for the server trays after circulating through the server trays or components on the server trays). In at least one embodiment, the spent second coolant exiting the server trays or racks 106 comes out of a different side of the ducts 122, such as the exit side, and moves to the parallel, but also exit side, of the row manifold 120. In at least one embodiment, from the row manifold 120, the spent second coolant moves in a parallel portion of the room manifold 118 and travels toward the CDU 112 in an opposite direction from the incoming second coolant (which may also be refreshed second coolant).

[0058] In at least one embodiment, the cooling system 102 can be used with one or more additional cooling systems, such as an air-to-liquid heat exchanger, which uses an air flow (which can be a forced air flow) to remove heat from a liquid (such as a cooling fluid). In at least one embodiment, the heat exchanger can be positioned near the servers 106 to facilitate removing heat near the servers 106. In at least one embodiment, the air-to-liquid heat exchanger can provide supplemental cooling in addition to the cooling provided by the cooling system 102.

[0059] In at least one embodiment, server-level features 200 include a server tray or case 202. In at least one embodiment, server tray or case 202 includes a server manifold 204 intermediately coupled between provided cold plates 210A-D of the server tray or case 202 and a rack manifold, which can be considered a heat exchanger. In at least one embodiment, server tray or case 202 includes one or more cold plates 210A-D associated with one or more computing or data center components or devices 220A-D. In at least one embodiment, server manifold 204 is disposed within server tray or case 202. In at least one embodiment, server manifold 204 is external to server tray or case 202.

[0060] In at least one embodiment, one or more server-level cooling loops 214A, 214B can be provided between the server manifold 204 and one or more cold plates 210A-D. In at least one embodiment, each server-level cooling loop 214A, 214B includes an inlet line 210 and an outlet line 212. In at least one embodiment, when there are cold plates 210A, 210B configured in series, an intermediate line 216 can be provided. In at least one embodiment, one or more cold plates 210A-D can support different ports and channels for auxiliary coolant or different fluids for the auxiliary cooling loop. In at least one embodiment, a fluid for cooling, such as an auxiliary coolant, can be provided to the server manifold 204 via the provided inlet and outlet 206A, 206B. In at least one embodiment, a fluid for cooling can be provided to the server manifold 204 via the provided inlet and outlet 208A, 208B.

[0061] In at least one embodiment, the server tray 202 is an immersion cooled server tray that can be submerged in a fluid. In at least one embodiment, the fluid used for the immersion cooled server tray can be a dielectric engineered fluid that can be used in immersion cooled servers. In at least one embodiment, an auxiliary coolant or fluid or local coolant can be used to cool the engineered fluid. In at least one embodiment, the fluid or local coolant can be used to cool the engineered fluid when a primary cooling circuit associated with an auxiliary cooling circuit that circulates the auxiliary coolant has failed or is failing. In at least one embodiment, at least one cold plate thus has ports for the auxiliary cooling circuit and for the local cooling circuit and can support a local coolant that is activated in the event of a failure of the primary cooling circuit. In at least one embodiment, a top liquid to air heat exchanger with smart fan wall cooling can be used without the need for an auxiliary cooling circuit.

[0062] In at least one embodiment, at least one dual-cooling cold plate 210B, 250 can be configured to operate alongside conventional cold plates 210A, C, and D. In at least one embodiment, a three-dimensional (3D) exploded view (cold plate 250) provides internal details of at least some features that may be included in the dual-cooling cold plate 210B. In at least one embodiment, a tear-away view of a first section 250B of the cold plate 250 having microchannels 270 (also 270A) shows a different second section 250A having different microchannels 264. In at least one embodiment, a regular cold plate can have one set of microchannels 264, 270, rather than the two sets shown.

[0063] In at least one embodiment, the dual-cooled cold plate 250 has distinct paths 264, 270 (each also referred to as a microchannel) for auxiliary coolant in the auxiliary cooling circuit and fluid in the local cooling circuit. In at least one embodiment, the auxiliary coolant or fluid may not be dielectric in nature. In at least one embodiment, the auxiliary coolant or fluid may have the same or similar chemical properties and may be provided from the same source. In at least one embodiment, in the case of immersion-cooled servers, the fluid, which may be a dielectric engineered fluid, may be suitable for both cold plate applications and immersion-cooled server tray applications. In at least one embodiment, some of the microchannels 270 are paths provided by fins 270A or other such aspects that rise internally and perpendicularly from the base of the cold plate portion 250B and have gaps therebetween for coolant flow. In at least one embodiment, some of the microchannels 264 are fluid passages in different cold plate portions 250A of the cold plate 250.

[0064] In at least one embodiment, reference to cold plates and their dual cooling features can imply reference to cold plates that can support at least two types of cooling loops unless otherwise specified. In at least one embodiment, two types of cold plates receive a local coolant or fluid for cooling, but one type can support an auxiliary cooling loop and a local cooling loop. In at least one embodiment, a standard coolant such as facility water can be used in an auxiliary cooling loop.

[0065] In at least one embodiment, a fluid or local coolant can only support cold plate usage and can not be available for immersion cooling. In at least one embodiment, each type of cold plate receives a different fluid or local coolant and auxiliary coolant from a respective local cooling loop, or an auxiliary cooling loop or other cooling loop that interfaces with a main cooling loop. In at least one embodiment, in cases where different fluids such as coolants are used with different coolant distribution units (CDUs) of different auxiliary loops, then a local cooling loop can be adapted for dual cooling cold plates with a local cooling loop such that different passages can be used for each of the fluid or local coolant and for the different auxiliary coolant.

[0066] In at least one embodiment, dual cooling cold plate 250 is adapted to receive at least two types of fluids such as an auxiliary coolant and a fluid or local coolant and to keep the two types of fluids different from each other via their different ports 252, 272, 268, 262 and their different paths 264, 270 such as by different sections separated by a gasket and a plate such as in a gasketed cold plate. In at least one embodiment, each different path is a fluid path. In at least one embodiment, fluid such as a local coolant from a fluid or local coolant source and an auxiliary coolant can be provided simultaneously to address additional cooling needs.

[0067] In at least one embodiment, the dual-cooled cold plate 250 includes ports 252, 272 for receiving a fluid or local coolant into the cold plate 250 and for allowing the fluid or local coolant to flow out of the cold plate 250. In at least one embodiment, the dual-cooled cold plate 250 includes ports 268, 262 for receiving a supplemental coolant into the cold plate 250 and for allowing the supplemental coolant to flow out of the cold plate 250. In at least one embodiment, the ports 252, 272 may have valve covers 254, 260 that may be directional and pressure-controlled to enable the fluid or local coolant to flow through the cold plate 250. In at least one embodiment, a valve cover may be associated with all provided ports. In at least one embodiment, the provided valve covers 254, 260 are mechanical features of an associated flow controller that also have corresponding electronic features (such as at least one processor for executing instructions stored in an associated memory and controlling the mechanical features for the associated flow controller).

[0068] In at least one embodiment, each valve can be actuated by an electronic feature of an associated flow controller. In at least one embodiment, the electronic and mechanical features of the provided flow controller are integrated. In at least one embodiment, the electronic and mechanical features of the provided flow controller are physically distinct. In at least one embodiment, reference to the flow controller can be to one or more of the provided electronic and mechanical features, or a combination thereof, but at least to features that enable control of the flow of coolant or fluid through each cold plate or immersion-cooled server tray or box.

[0069] In at least one embodiment, the electronic feature of the provided flow controller receives the control signal and asserts control of the mechanical feature. In at least one embodiment, the electronic feature of the provided flow controller can be an actuator or other electronic component with similar electromechanical features. In at least one embodiment, a flow pump can be used as the flow controller. In at least one embodiment, an impeller, piston, or bellows can be the mechanical feature, and the motor and circuitry form the electronic feature of the provided flow controller.

[0070] In at least one embodiment, the circuitry of the provided flow controller can include a processor, memory, switches, sensors, and other components that collectively form the electronic features of the provided flow controller. In at least one embodiment, the provided ports 252, 262, 272, 268 of the provided flow controller are adapted to allow entry or exit of immersion fluid. In at least one embodiment, a flow controller 280 (capable of functioning as an expansion valve) can be associated with a fluid line 276 (also 256, 274) that allows fluid or local coolant to enter and exit the cold plate 210B. In at least one embodiment, other flow controllers can be similarly associated with the coolant lines 210, 216, 212 (also 266, 258) to allow auxiliary coolant to enter and exit the cold plate 210B.

[0071] In at least one embodiment, the fluid or local coolant enters the provided fluid lines 276 via dedicated fluid inlet lines 208A and outlet lines 208B. In at least one embodiment, the server manifold 204 is adapted to have channels therein (shown by dashed lines) to support different paths to the different fluid lines 276 (also 256, 274) and to any remaining loops 214A, B associated with the auxiliary coolant inlet and outlet lines 206A, B. In at least one embodiment, there can be multiple manifolds to support different fluid or local coolants and auxiliary coolants. In at least one embodiment, there can be multiple manifolds to support different inlets and outlets for each of the fluid or local coolant and auxiliary coolant. In at least one embodiment, if the fluid or local coolant is used alone without an auxiliary cooling loop, at least two different flows are enabled via the same fluid path (at least within the cold plate or server tray) to the fluid source and the coolant row manifold.

[0072] In at least one embodiment, a first flow can enable the auxiliary coolant to flow through one or more provided ports 252, 272 and associated pathways 270. In at least one embodiment, the dual cooling cold plate 250 can have isolated plate segments 250A, 250B that are flooded with fluid or local coolant and / or auxiliary coolant while being kept distinct from one another by gaskets or seals. In at least one embodiment, a second flow can enable the fluid or local coolant to flow through provided ports 268, 262 and associated pathways 264 through fins or microchannels 270A extending throughout the base of the cold plate segment 250B.

[0073] In at least one embodiment, flow controllers 278 may be associated with the fluid inlet 276 and outlet portions at the server manifold 204, rather than with the flow controllers 280 provided at the corresponding cold plates. In at least one embodiment, a first flow utilizes only fluid or local coolant and may be enabled when a fault is determined to have occurred in the auxiliary cooling loop or the primary cooling loop, such that the auxiliary coolant is unable to effectively absorb heat from at least one computing device. In at least one embodiment, the fault may be that the auxiliary coolant is not being adequately cooled by the CDU, and therefore, may not be able to absorb sufficient heat from the at least one computing device through its associated cold plate.

[0074] In at least one embodiment, Figure 1 The server-level feature 200 shown in FIG. 3 can be associated with a server tray or case 300 having built-in or integrated flow distribution. In at least one embodiment, the server-level feature 200 includes a server tray or case 302. In at least one embodiment, the server tray or case 302 includes an integrated or built-in server manifold 304, which is intermediately coupled between the provided cold plates 310A-D of the server tray or case 302 and the rack manifold, which can be considered a heat exchanger. In at least one embodiment, the server tray or case 302 includes one or more cold plates 310A-D associated with one or more computing or data center components or devices 320A-D. In at least one embodiment, the server manifold 304 is integrally formed into one or more panels forming the server tray or case 302. In at least one embodiment, the server manifold 304 includes one or more flow channels extending through one or more panels forming the server tray or case 302. In at least one embodiment, the server manifold 304 includes inlet and / or outlet flow channels for providing and receiving fluid from the cold plates 310A-D.

[0075] In at least one embodiment, one or more server-level cooling circuits 314A, 314B are integrated into one or more panels forming the server tray or box 302. In at least one embodiment, the server-level cooling circuits 314A, 314B. In at least one embodiment, the one or more server-level cooling circuits 314A, 314B may be provided between the server manifold 304 and one or more cold plates 310A-D. In at least one embodiment, each server-level cooling circuit 314A, 314B includes an inlet line 310 and an outlet line 312. In at least one embodiment, when cold plates 210A, 210B are configured in series, an intermediate line 316 may be provided. In at least one embodiment, one or more cold plates 310A-D may support different ports and channels for auxiliary coolant or different fluids for the auxiliary cooling circuits. In at least one embodiment, the fluid used for cooling, such as the auxiliary coolant, may be provided to the server manifold 304 via the provided inlet and outlet ports 306A, 306B. In at least one embodiment, fluid for cooling may be provided to the server manifold 304 via provided inlets and outlets 308A, 308B.

[0076] In at least one embodiment, the server-level cooling loops 314A, 314B are integrally formed within panels that form at least a portion of the server tray or box 302. In at least one embodiment, drop-down or tightly positioned connections are utilized to direct fluid into the cold plates 310A-310D. In at least one embodiment, the server manifold 304 includes internal channels that form at least a portion of the server-level cooling loops 314A, 314B. In at least one embodiment, the server manifold 304 includes outlets positioned adjacent to the associated inlets and outlets of the cold plates 310A-310D. In at least one embodiment, the server manifold 304 reduces piping or other items within the server tray or box 302 due to the close positioning of the outlets of the server-level cooling loops 314A, 314B with the associated inlets and outlets of the cold plates 310A-310D. In at least one embodiment, quick-connect couplings are utilized with the server manifold 304 at one or more locations, such as the inlets and outlets 306A, 306B, 308A, 308B. In at least one embodiment, outlets along the server-level cooling loops 314A, 314B also include quick-connect connectors. In at least one embodiment, the quick-connect, drip-free connectors are blind-mate connectors that couple together through a sliding or snapping action without external tools. In at least one embodiment, the quick-connect, drip-free connectors include self-aligning features to facilitate coupling. In at least one embodiment, the quick-connect, drip-free connectors include an indicator, such as an audible click or physical feedback (such as resistance or friction), to indicate that a connection is formed. In at least one embodiment, the quick-connect, drip-free connectors are arranged so that when the panel of the server tray or box 302 is moved to the closed position, the closing of the server tray or box 302 positions the mating connectors into alignment and forms a connection. In at least one embodiment, due to the alignment and / or feedback provided by the connection, the quick-connect, drip-free connectors are configured for installation without a direct line of sight to the connection. In at least one embodiment, the quick-connect, drip-free connectors can be tool-less connectors that require the removal of external tools (as described above) and also the removal of tools for internal connections within the server unit. In at least one embodiment, a quick-connect, no-drip connector blocks fluid flow until connected to a mating connector. In at least one embodiment, fluid is present in a line including a quick-connect, no-drip connector, but flow is blocked until coupled to a mating connector.

[0077] In at least one embodiment, the server tray 302 is an immersion cooled server tray that can be submerged in a fluid. In at least one embodiment, the fluid used for the immersion cooled server tray can be a dielectric engineered fluid that can be used in immersion cooled servers. In at least one embodiment, an auxiliary coolant or fluid or local coolant can be used to cool the engineered fluid. In at least one embodiment, the fluid or local coolant can be used to cool the engineered fluid in the event of a failure or malfunction of a primary cooling circuit associated with an auxiliary cooling circuit that circulates the auxiliary coolant. In at least one embodiment, at least one cold plate thus has ports for the auxiliary cooling circuit and for the local cooling circuit and is capable of supporting a local coolant that is activated in the event of a failure of the primary cooling circuit. In at least one embodiment, a top liquid to air heat exchanger with smart fan wall cooling can be used without the need for an auxiliary cooling circuit.

[0078] In at least one embodiment, the server-level cooling configuration 400 can be incorporated into one or more server units associated with a data center, such as Figure 4A As shown. In at least one embodiment, server unit 402 includes a body 404 formed by panels 406. In at least one embodiment, panels 406 include a top panel 408, a bottom panel 410, a first side panel 412, and a second side panel 414. In at least one embodiment, additional panels for the front and rear may be included. In at least one embodiment, cooling flow circuits 416A and 416B are integrated into panels 406 of body 404, such as top panel 408 as shown. In at least one embodiment, cooling flow circuits 416A and 416B may be circuitous or continue over at least a portion of the area of ​​top panel 408 and may include one or more bends. In at least one embodiment, cooling flow circuits 416A and 416B may include inlet and outlet circuits, with a single circuit operating as an inlet flow circuit and a single circuit operating as an outlet flow circuit. In at least one embodiment, both cooling flow circuits 416A and 416B may be inlet flow circuits. In at least one embodiment, both cooling flow circuits 416A and 416B may be outlet flow circuits. In at least one embodiment, cooling flow circuits 416A, 416B direct cooling fluid, which can be one or more types of liquids, among other options, toward or receive cooling fluid from a cold plate 418 associated with computing unit 420. In at least one embodiment, the cooling fluid can interact with cold plate 418 to receive heat generated by computing unit 420.

[0079] In at least one embodiment, the cooling flow loops 416A, 416B include a diameter 422 that can be equal to or greater than the diameter of an associated duct assembly that can be removed by utilizing the cooling system 400. In at least one embodiment, the cooling flow loops 416A, 416B can be designed to provide a minimum or threshold amount of fluid, such as a minimum volume of fluid. In at least one embodiment, the cooling flow loops 416A, 416B can replace plastic ducting using a distribution manifold within the server unit 402. In at least one embodiment, the cooling flow loops 416A, 416B are formed between a panel interior portion 424 and a panel exterior portion 426. In at least one embodiment, the panel interior portion 424 and the panel exterior portion 426 form a sealed environment, wherein a connecting portion couples the panel interior portion 424 to the panel exterior portion 426 to form a fluid-tight space between the portions 424, 426. In at least one embodiment, the cooling flow circuits 416A, 416B can be machined or otherwise positioned within a solid component, such as via an additive manufacturing process that creates space for channels associated with the cooling flow circuits 416A, 416B. In at least one embodiment, the cooling flow circuits 416A, 416B include structural supports 428 within the void 430 of the panel 406. In at least one embodiment, the cooling flow circuits 416A, 416B are formed from plastic tubing positioned within the void 430. In at least one embodiment, the cooling flow circuits 416A, 416B are formed from prefabricated components arranged within the void 430. In at least one embodiment, the void 430 is a sealed space such that seams between connected components prevent fluid from escaping the void 430. In at least one embodiment, the cooling flow circuits 416A, 416B are substantially symmetrical. In at least one embodiment, the cooling flow circuits 416A, 416B are different sizes, such as different diameters or different lengths. In at least one embodiment, the cooling flow circuits 416A, 416B have different routing patterns.

[0080] In at least one embodiment, the cooling flow circuits 416A, 416B provide cooling fluid to the cold plate 418 via corresponding quick-connect connectors. In at least one embodiment, the panel connectors 432A, 432B align with the plate connectors 434A, 434B to facilitate flow between the cooling flow circuits 416A, 416B and the cold plate 418. In at least one embodiment, the panel-to-plate connectors 432A, 432B, 434A, 434B are quick-connect type connectors, such as cone valves, which utilize one or more flow-restricting components (such as check valves, actuated valves, and other components) to block flow without a mating connector. In at least one embodiment, the panel-to-plate connectors 432A, 432B, 434A, 434B are drip-free connections, such that flow is blocked without a mating connection to reduce the possibility of leaks. In at least one embodiment, the panels and panel joints 432A, 432B, 434A, 434B include a coupling mechanism that may include one or more press fittings that provide feedback to the operator regarding installation, such as an audible sound indicating coupling, a friction response, or other types of feedback. In at least one embodiment, the corresponding panel joints and panel joints 432A, 432B, 434A, 434B are positioned to align when one or more panels 406 are in a closed position, such that closing the panels 406 forms a connection between the panels and panel joints 432A, 432B, 434A, 434B. In at least one embodiment, closing the panels 406 may include closing hinged panels, sliding the panels along tracks, or another type of connection. In at least one embodiment, additional components may be positioned between the panels and panel joints 432A, 432B, 434A, 434B, such as duct sections. In at least one embodiment, the flow splitters can be arranged so that a single panel header 432A, 432B can support multiple cold plates 418. In at least one embodiment, a direct connection is made between the panel and plate headers 432A, 432B, 434A, 434B.

[0081] In at least one embodiment, the cooling arrangement 450 can be built into a panel 406 that forms at least a portion of the server unit 402, such as a panel 406 that forms multiple portions of the body 404. In at least one embodiment, the cooling arrangement 450 can be considered to be a distribution manifold that directs and / or receives fluid from one or more locations. In at least one embodiment, the inlet 452, if built into the panel 406 (such as the top panel 408), can include a connector 432C. In at least one embodiment, the connector 432C is an inlet connector that can be a quick-connect type connector that acts as a one-way flow connector and can also be a drip-free connector. In at least one embodiment, the connector 432C can include one or more flow components, such as a check valve, an actuator valve, or other, for regulating the flow into the cooling flow circuit 416A.

[0082] In at least one embodiment, the cooling flow circuit 416A includes one or more flow controllers 454 along at least a portion of a flow path 456. In at least one embodiment, the flow path 456 includes various takeoffs, bypasses, and loops for directing flow according to one or more desired flow characteristics. In at least one embodiment, one or more sensors can be used to transmit signals to the flow controller 454 to allow, regulate, or block flow to various joints 432, which can be associated with mating joints 434 of the cold plate 418. In at least one embodiment, the one or more sensors can include leak sensors to block flow to leaking joints 432. In at least one embodiment, the one or more sensors can include flow sensors to determine whether a blockage or obstruction, such as fouling or material accumulation, has formed in the cooling flow circuit 416A. In at least one embodiment, the one or more sensors can include a pressure sensor to determine a pressure drop indicative of a blockage. In at least one embodiment, the one or more sensors can include a temperature sensor to indicate reduced cooling efficiency. In at least one embodiment, one or more controllers, which may include a memory and a processor, may receive signals from one or more sensors and send an alert or signal in response to the information from the one or more sensors.

[0083] In at least one embodiment, various flow configurations can be incorporated into flow path 456. In at least one embodiment, each joint 432 includes a bypass. In at least one joint, each joint 432 can be fluidically coupled to an adjacent joint 432 via one or more flow path segments. In at least one embodiment, a flow controller 454 is positioned at the junction where multiple flow path segments combine to regulate and direct fluid flow. In at least one embodiment, cooling flow circuit 416A can include a flow path extension 458, which can include one or more segments that loop through areas to increase the volume of cooling flow circuit 416A. In at least one embodiment, flow path extension 458 fills additional void space but may not be coupled to one or more joints 432. In at least one embodiment, flow path extension 458 increases the total volume of fluid received and / or distributed by cooling flow circuit 416A. In at least one embodiment, flow path extension 458 can be a circuitous configuration that increases volume, which can be related to length, while also reducing the number of bends or fittings to reduce line losses. In at least one embodiment, flow path extensions 458 can couple different groups of connectors 432 together. In at least one embodiment, adjacent connectors 432 may not have a flow path portion extending therebetween, but instead may be positioned on separate outlets. In at least one embodiment, various flow configurations can be combined to achieve different flow adjustments to different portions of panel 406.

[0084] In at least one embodiment, outlet 460 is coupled to cooling flow circuit 416B. In at least one embodiment, outlet 460 includes a quick-connect fitting 432D that may share one or more features with other fittings 432. In at least one embodiment, a continuous loop may extend between fittings 432. In at least one embodiment, adjacent fittings 432 may be coupled to each other via adjacent flow circuit segments. In at least one embodiment, a single outlet extends from fittings 432 to a common return line. In at least one embodiment, different configurations are incorporated into a single panel 406 to vary the flow characteristics for cooling.

[0085] In at least one embodiment, cooling fluid enters panel 406 and is distributed to cooling flow loop 416A, where it exits connector 432, for example, via a connection to mating connector 434, to provide cooling fluid to cold plate 418. In at least one embodiment, cooling fluid from cold plate 418 enters cooling flow loop 416B and is directed out of panel 406. In at least one embodiment, panel 406 acts as both a manifold for receiving fluid, distributing it to various locations, and returning it to a return manifold. In at least one embodiment, using panel 406 as a manifold can reduce the number of pipes or other connections that need to be formed and routed after installation. In at least one embodiment, using panel 406 as a manifold can reduce the number of leak points and, due to the use of drip-free quick-connect connectors, can reduce the number of leaks. In at least one embodiment, panel 406 can include multiple inlets 452 and / or multiple outlets 460 to provide a specific amount of cooling fluid. In at least one embodiment, panel 406 can include different numbers of inlets 452 and outlets 460, as well as cooling flow circuits 416A, 416B of different diameters, in order to adjust the amount of fluid, control the pressure, control the flow rate, or otherwise change the cooling characteristics. In at least one embodiment, one or more sensors can enable flow controller 454 to further adjust the flow characteristics, and therefore the cooling characteristics.

[0086] In at least one embodiment, one or more sensors 462 are positioned along the flow circuit 416 at the cold plate 418 within the panel 406 and at the computing unit 420. In at least one or more embodiments, the various sensor locations may include flow inlets, flow outlets, outlets, flow controllers, inlets, outlets, and other options. In at least one embodiment, the one or more sensors 462 may provide signals to the flow controller 454 to regulate the cooling flow throughout the flow circuit 416. In at least one embodiment, the one or more sensors 462 may include temperature sensors, flow sensors, pressure sensors, position sensors, leak sensors, or various other sensors. In at least one embodiment, flow sensor 462A may measure the flow rate through one or more sections of the flow circuit 416. In at least one embodiment, information from flow sensor 462A may be used to adjust the flow controller 454, such as by opening or closing a valve to increase or decrease the flow rate. In at least one embodiment, pressure sensor 462B may measure the pressure within one or more sections of the flow circuit 416, where a decrease in pressure may indicate a leak and an increase in pressure may indicate a blockage or flow restriction. In at least one embodiment, temperature sensor 462C can measure the temperature within one or more portions of flow circuit 416 to determine the cooling properties of the fluid within flow circuit 416. In at least one embodiment, temperature sensor 462C can be positioned at the inlet and outlet of cold plate 418 to determine the cooling efficiency or cooling load of cold plate 418. In at least one embodiment, position sensor 462D can determine the position of flow controller 454, for example, when flow controller 454 is a valve, to adjust the valve position. In at least one embodiment, leak sensor 462E can be positioned near sensitive components (such as computing unit 420) or within void 430 to detect leaks, which can be used to block or otherwise restrict flow to prevent damage to various components. In at least one embodiment, one or more alarms can be associated with one or more sensors 462 to provide information to an operator regarding flow conditions associated with unit 402.

[0087] In at least one embodiment, the server-level cooling configuration 400 can be incorporated into one or more server units associated with a data center, such as Figure 4Cshown. In at least one embodiment, server unit 402 includes a body 404 formed by panels 406, where each panel 406 includes at least one connection for facilitating the distribution of cooling fluid to one or more cold plates 418. In at least one embodiment, a cooling flow loop 416A extends to a panel junction 432A to distribute cooling fluid to a cold plate 318A via a connection at a plate junction 434A. In at least one embodiment, cooling flow loop 416A is positioned within a top panel 408 and can include one or more bends, flow controllers, or various leads to direct cooling fluid to cold plates 418. In at least one embodiment, heated fluid is transferred from plate junction 434B to plate junction 432B, which is coupled to a cooling flow loop 416B positioned within a bottom panel 410. In at least one embodiment, different portions of loops 416A, 416B can be arranged in different panels 406, such as a portion within top panel 408 and a portion within bottom panel 408.

[0088] In at least one embodiment, multiple junctions can be arranged to provide cooling fluid to cooling flow loop 416B. In at least one embodiment, cold plate 418B includes a plate junction 434C that provides an outlet flow to cooling flow loop 416B. In at least one embodiment, connections are formed when server unit 402 is assembled, such as when panels 406 are placed in position. In at least one embodiment, additional junctions 432 can also be incorporated into additional panels 406, such as first side panel 412 and second side panel 414. In at least one embodiment, additional cooling flow loops can be formed within other panels 406, providing various flow paths for cooling fluid to be distributed to cold plates 418. In at least one embodiment, junctions can be formed at predetermined locations based on an intended cold plate configuration, such that installation can reduce or remove the use of flexible tubing within server unit 402. In at least one embodiment, flexible tubing can be used to form one or more connections.

[0089] In at least one embodiment, cooling fluid connections are formed within server unit 402 at assembly, without the need for additional routing, tools, or connectors, such as Figure 5A and Figure 5BAs shown. In at least one embodiment, one or more panels 406 can be removed from the server unit 402. In at least one embodiment, removing one or more panels 406 can enable installation and configuration of internal components (such as the computing unit 420 and the cold plate 418). In at least one embodiment, assembly can include, for example, moving or sliding one or more panels 406 into place along a track. In at least one embodiment, one or more alignment features can be included so that the operator installing one or more panels 406 can do so without a direct line of sight to the connectors 432A, 432B. In at least one embodiment, installation can be performed in a rack that is higher than the operator, and therefore, the operator can utilize one or more features of the connectors 432A, 432B, 434A, 434B (such as the ability to act as a blind mate connector) for installation. In at least one embodiment, one or more panels 406 can be installed using a blind mate feature, wherein installation is performed with tactile or audible feedback so as to provide information to the operator installing one or more panels 406, for example, by receiving a feel of friction or hearing a click during installation. In at least one embodiment, the top panel 408 can slide or otherwise move axially over an interior portion of the server unit 402. In at least one embodiment, the panel connectors 432A, 432B extend axially below the top panel 408. In at least one embodiment, the panel connectors 432A, 432B can be flush with or recessed within the top panel 408. In at least one embodiment, the axial movement enables engagement of the panel connectors 432A, 432B and the panel connectors 434A, 434B. In at least one embodiment, one or more of the panels or panel connectors 432A, 432B, 434A, 434B include angled surfaces to facilitate engagement with mating connectors. In at least one embodiment, an additional force can be applied to engage between the connectors 432A, 432B, 434A, 434B. In at least one embodiment, the additional force can be substantially perpendicular to the axial direction of movement to position one or more panels 406. In at least one embodiment, positioning the top panel 408 on the cold plate 418 enables engagement between the panel and the plate joints 432A, 432B, 434A, 434B to facilitate cooling fluid flow to the cold plate 418 .

[0090] In at least one embodiment, cooling fluid connections are made within the server unit 402 during assembly without requiring additional routing, tools, or connectors, such as Figure 5C and Figure 5DAs shown. In at least one embodiment, one or more panels 406 can be removable or can be arranged to pivot or slide away from adjacent panels 406 forming the server unit 402. In at least one embodiment, removal or movement of one or more panels 406 can enable installation and configuration of internal components, such as the computing unit 420 and the cold plate 418. In at least one embodiment, assembly can include pivoting one or more panels 406 about hinges or pivot points. In at least one embodiment, the top panel 408 can be hinged relative to the first side panel 412 to expose the internal components of the server unit 402. In at least one embodiment, the panel joints 432A, 432B extend axially below the top panel 408. In at least one embodiment, the panel joints 432A, 432B can be flush with the top panel 408 or recessed in the top panel 408. In at least one embodiment, rotational movement about one or more pivot points enables engagement of the panel joints 432A, 432B and the plate joints 434A, 434B. In at least one embodiment, one or more of the panel or plate connectors 432A, 432B, 434A, 434B include angled surfaces to facilitate engagement with a mating connector. In at least one embodiment, positioning the top panel 408 on the cold plate 418 allows engagement between the panel and the plate connectors 432A, 432B, 434A, 434B to facilitate flow of cooling fluid to the cold plate 418.

[0091] Servers and Data Centers

[0092] The following figures illustrate, but are not limited to, exemplary network server and data center based systems that may be used to implement at least one embodiment.

[0093] Figure 6 A distributed system 600 is shown in accordance with at least one embodiment. In at least one embodiment, the distributed system 600 includes one or more client computing devices 602, 604, 606, and 608 configured to execute and operate client applications, such as network (web) browsers, proprietary clients, and / or variations thereof, over one or more networks 610. In at least one embodiment, a server 612 can be communicatively coupled to the remote client computing devices 602, 604, 606, and 608 via the network 610.

[0094] In at least one embodiment, the server 612 may be adapted to run one or more services or software applications, such as services and applications that can manage session activity for single sign-on (SSO) access across multiple data centers. In at least one embodiment, the server 612 may also provide other services or software applications that may include both non-virtualized and virtualized environments. In at least one embodiment, these services may be provided to users of the client computing devices 602, 604, 606, and / or 608 as web-based services or cloud services or under a software as a service (SaaS) model. In at least one embodiment, users operating the client computing devices 602, 604, 606, and / or 608 may, in turn, utilize one or more client applications to interact with the server 612 to utilize the services provided by these components.

[0095] In at least one embodiment, the software components 618, 620, and 622 of system 600 are implemented on server 612. In at least one embodiment, one or more components of system 600 and / or the services provided by these components can also be implemented by one or more of client computing devices 602, 604, 606, and / or 608. In at least one embodiment, a user operating a client computing device can then utilize one or more client applications to use the services provided by these components. In at least one embodiment, these components can be implemented in hardware, firmware, software, or a combination thereof. It should be understood that a variety of different system configurations are possible that can differ from the distributed system 600. Therefore, Figure 6 The illustrated embodiment is at least one embodiment of a distributed system for implementing an embodiment system and is not intended to be limiting.

[0096] In at least one embodiment, client computing devices 602, 604, 606, and / or 608 may include different types of computing systems. In at least one embodiment, client computing devices may include portable handheld devices (e.g., Cellular phones, computing tablets, personal digital assistants (PDAs), or wearable devices (e.g., Google head-mounted display), running software (such as Microsoft Windows ) and / or various mobile operating systems (such as iOS, Windows Phone, Android, BlackBerry 10, Palm OS and / or their variants). In at least one embodiment, the device can support different applications, such as different Internet-related applications, email, short message service (SMS) applications, and can use various other communication protocols. In at least one embodiment, the client computing device can also include a general-purpose personal computer, which in at least one embodiment includes a computer running various versions of Microsoft Apple and / or a personal computer and / or laptop computer running Linux operating system.

[0097] In at least one embodiment, the client computing device may be a computer running various commercially available or any workstation computer running any of the UNIX-like operating systems, including but not limited to various GNU / Linux operating systems, such as Google Chrome OS. In at least one embodiment, client computing devices may also include electronic devices capable of communicating over one or more networks 610, such as thin client computers, Internet-enabled gaming systems (e.g., with or without a PC), and gesture input devices for Microsoft Xbox gaming consoles), and / or personal messaging devices. Figure 6 The distributed system 600 in FIG. 6 is shown as having four client computing devices, but any number of client computing devices may be supported. Other devices (such as devices with sensors, etc.) may interact with the server 612 .

[0098] In at least one embodiment, the network 610 in the distributed system 600 can be any type of network capable of supporting data communications using any of a variety of available protocols, including but not limited to TCP / IP (Transmission Control Protocol / Internet Protocol), SNA (Systems Network Architecture), IPX (Internetwork Packet Exchange), AppleTalk, and / or variations thereof. In at least one embodiment, the network 610 can be a local area network (LAN), an Ethernet-based network, a token ring, a wide area network, the Internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a public switched telephone network (PSTN), an infrared network, a wireless network (e.g., a standard implemented in the Institute of Electrical and Electronics Engineers (IEEE) 802.11 protocol suite), a wireless network, a WLAN, a WLAN (e.g., a WLAN implemented in the Institute of Electrical and Electronics Engineers (IEEE) 802.11 protocol suite), a WLAN (e.g., a WLAN implemented in the Institute of Electrical and Electronics Engineers (IEEE) 802.11 protocol suite), a WLAN (e.g., a WLAN implemented in the Institute of Electrical and Electronics Engineers (IEEE) 802.11 protocol suite), a WLAN (e.g., a WLAN implemented in the Institute of Electrical and Electronics Engineers (IEEE) 802.11 protocol suite), a WLAN (e.g., a WLAN implemented in the Institute of Electrical and Electronics Engineers (IEE ... and / or any other wireless protocols), and / or any combination of these and / or other networks.

[0099] In at least one embodiment, the server 612 may be comprised of one or more general purpose computers, dedicated server computers (including, in at least one embodiment, PC (personal computer) servers, The server 612 may be composed of a plurality of servers (e.g., servers, mid-range servers, mainframe computers, rack servers, etc.), a server farm, a server cluster, or any other suitable arrangement and / or combination. In at least one embodiment, the server 612 may include one or more virtual machines running virtual operating systems or other computing architectures involving virtualization. In at least one embodiment, one or more flexible logical storage device pools may be virtualized to maintain virtual storage devices for the server. In at least one embodiment, the virtual network may be controlled by the server 612 using software-defined networking. In at least one embodiment, the server 612 may be adapted to run one or more services or software applications.

[0100] In at least one embodiment, the server 612 can run any operating system, and any commercially available server operating system. In at least one embodiment, the server 612 can also run any of a variety of additional server applications and / or mid-tier applications, including HTTP (Hypertext Transfer Protocol) servers, FTP (File Transfer Protocol) servers, CGI (Common Gateway Interface) servers, Server, database server and / or variants thereof.In at least one embodiment, exemplary database servers include, but are not limited to, those commercially available from Oracle, Microsoft, Sybase, IBM (International Business Machines), and / or variants thereof.

[0101] In at least one embodiment, server 612 may include one or more applications for analyzing and consolidating data feeds and / or event updates received from users of client computing devices 602, 604, 606, and 608. In at least one embodiment, data feeds and / or event updates may include, but are not limited to, data received from one or more third-party information sources and continuous data streams. feed, Updates or real-time updates, which may include real-time events related to sensor data applications, financial quoters, network performance measurement tools (e.g., network monitoring and business management applications), clickstream analysis tools, automobile traffic monitoring, and / or changes thereto. In at least one embodiment, server 612 may also include one or more applications for displaying data feeds and / or real-time events via one or more display devices of client computing devices 602, 604, 606, and 608.

[0102] In at least one embodiment, the distributed system 600 may further include one or more databases 614 and 616. In at least one embodiment, the database may provide a mechanism for storing information such as user interaction information, usage pattern information, adaptation rule information, and other information. In at least one embodiment, the databases 614 and 616 may reside in various locations. In at least one embodiment, one or more of the databases 614 and 616 may reside on a non-transitory storage medium local to the server 612 (and / or residing in the server 612). In at least one embodiment, the databases 614 and 616 may be remote from the server 612 and communicate with the server 612 via a network-based connection or a dedicated connection. In at least one embodiment, the databases 614 and 616 may reside in a storage area network (SAN). In at least one embodiment, any necessary files for performing the functions attributed to the server 612 may be appropriately stored locally on the server 612 and / or remotely. In at least one embodiment, the databases 614 and 616 may include relational databases, such as databases suitable for storing, updating, and retrieving data in response to SQL-formatted commands.

[0103] Figure 7 An exemplary data center 700 is shown in accordance with at least one embodiment. In at least one embodiment, data center 700 includes, but is not limited to, a data center infrastructure layer 710, a framework layer 720, a software layer 730, and an application layer 740.

[0104] In at least one embodiment, Figure 7 As shown, the data center infrastructure layer 710 may include a resource coordinator 712, grouped computing resources 714, and node computing resources ("node CRs") 716(1)-716(N), where "N" represents any complete positive integer. In at least one embodiment, the node CRs 716(1)-716(N) may include, but are not limited to, any number of central processing units ("CPUs") or other processors (including accelerators, field programmable gate arrays ("FPGAs"), graphics processors, etc.), memory devices (e.g., dynamic read-only memories), storage devices (e.g., solid-state drives or disk drives), network input / output ("NW I / O") devices, network switches, virtual machines ("VMs"), power modules, and cooling modules, etc. In at least one embodiment, one or more of the node CRs 716(1)-716(N) may be servers having one or more of the above-mentioned computing resources.

[0105] In at least one embodiment, the grouped computing resources 714 may include separate groups of node CRs housed in one or more racks (not shown), or may include many racks (also not shown) housed in data centers at various geographic locations. The separate groups of node CRs within the grouped computing resources 714 may include computing, networking, memory, or storage resources that may be configured or allocated to support groupings of one or more workloads. In at least one embodiment, several node CRs comprising a CPU or processor may be grouped in one or more racks to provide computing resources to support one or more workloads. In at least one embodiment, one or more racks may also include any number of power modules, cooling modules, and network switches in any combination.

[0106] In at least one embodiment, resource coordinator 712 may configure or otherwise control one or more nodes CR 716(1)-716(N) and / or grouped computing resources 714. In at least one embodiment, resource coordinator 712 may comprise a software design infrastructure ("SDI") management entity for data center 700. In at least one embodiment, resource coordinator 712 may comprise hardware, software, or some combination thereof.

[0107] In at least one embodiment, Figure 7As shown, framework layer 720 includes, but is not limited to, a job scheduler 732, a configuration manager 734, a resource manager 736, and a distributed file system 738. In at least one embodiment, framework layer 720 may include a framework that supports software 752 of software layer 730 and / or one or more applications 742 of application layer 740. In at least one embodiment, software 752 or applications 742 may include web-based service software or applications, such as those provided by Amazon Web Services, Google Cloud, and Microsoft Azure. In at least one embodiment, framework layer 720 may include, but is not limited to, a free and open source software web application framework, such as Apache Spark™ (hereinafter referred to as "Spark"), which can utilize distributed file system 738 for large-scale data processing (e.g., "big data"). In at least one embodiment, job scheduler 732 may include a Spark driver to facilitate scheduling workloads supported by various layers of data center 700. In at least one embodiment, configuration manager 734 may be capable of configuring different layers, such as software layer 730 and framework layer 720, which includes Spark and a distributed file system 738 for supporting large-scale data processing. In at least one embodiment, the resource manager 736 can manage clustered or grouped computing resources that are mapped to or allocated to support the distributed file system 738 and the job scheduler 732. In at least one embodiment, the clustered or grouped computing resources can include grouped computing resources 714 on the data center infrastructure layer 710. In at least one embodiment, the resource manager 736 can coordinate with the resource coordinator 712 to manage these mapped or allocated computing resources.

[0108] In at least one embodiment, the software 752 included in the software layer 730 may include software used by at least a portion of the node CRs 716(1)-716(N), the grouped computing resources 714, and / or the distributed file system 738 of the framework layer 720. The one or more types of software may include, but are not limited to, Internet web page search software, email virus scanning software, database software, and streaming video content software.

[0109] In at least one embodiment, the one or more applications 742 included in the application layer 740 may include one or more types of applications used by at least a portion of the node CRs 716(1)-716(N), the grouped computing resources 714, and / or the distributed file system 738 of the framework layer 720. The one or more types of applications may include, but are not limited to, CUDA applications, 5G network applications, artificial intelligence applications, data center applications, and / or variations thereof.

[0110] In at least one embodiment, any of configuration manager 734, resource manager 736, and resource coordinator 712 can implement any number and type of self-modification actions based on any number and type of data obtained in any technically feasible manner. In at least one embodiment, the self-modification actions can relieve a data center operator of data center 700 from making potentially poor configuration decisions and can avoid underutilized and / or poorly performing portions of the data center.

[0111] Figure 8 A client-server network 804 is shown formed by a plurality of interconnected network server computers 802, according to at least one embodiment. In at least one embodiment, each network server computer 802 stores data accessible to the other network server computers 802 and client computers 806 and networks 808 connected to the wide area network 804. In at least one embodiment, the configuration of the client-server network 804 can change over time as client computers 806 and one or more networks 808 are connected and disconnected from the network 804, and as one or more backbone server computers 802 are added to or removed from the network 804. In at least one embodiment, the client-server network includes client computers 806 and networks 808 when such client computers 806 and networks 808 are connected to the network server computers 802. In at least one embodiment, the term "computer" includes any device or machine capable of accepting data, applying a prescribed process to the data, and providing a result of the process.

[0112] In at least one embodiment, the client-server network 804 stores information accessible to the network server computer 802, the remote network 808, and the client computers 806. In at least one embodiment, the network server computer 802 is formed by a mainframe computer, a minicomputer, and / or a microcomputer, each having one or more processors. In at least one embodiment, the server computers 802 are linked together via wired and / or wireless transmission media (such as wires, fiber optic cables), and / or microwave transmission media, satellite transmission media, or other conductive, optical, or electromagnetic wave transmission media. In at least one embodiment, the client computers 806 access the network server computers 802 via similar wired or wireless transmission media. In at least one embodiment, the client computers 806 can be linked to the client-server network 804 using a modem and a standard telephone communication network. In at least one embodiment, alternative carrier systems (such as cable and satellite communication systems) can also be used to link to the client-server network 804. In at least one embodiment, other private or time-shared carrier systems can be used. In at least one embodiment, the network 804 is a global information network, such as the Internet. In at least one embodiment, the network is a private intranet that uses similar protocols to the Internet but with added security measures and restricted access controls.In at least one embodiment, the network 804 is a private or semi-private network that uses a proprietary communication protocol.

[0113] In at least one embodiment, client computer 806 is any end-user computer and may also be a mainframe computer, minicomputer, or microcomputer having one or more microprocessors. In at least one embodiment, server computer 802 may sometimes be used as a client computer to access another server computer 802. In at least one embodiment, remote network 808 may be a local area network, a network added to a wide area network via an independent service provider (ISP) for the Internet, or another group of computers interconnected via a wired or wireless transmission medium with a fixed or time-varying configuration. In at least one embodiment, client computer 806 may be linked to network 804 independently or via remote network 808 and access network 804.

[0114] Figure 9A computer network 908 connecting one or more computer machines is shown in accordance with at least one embodiment. In at least one embodiment, network 908 can be any type of electrically connected computer group, including, for example, the following networks: the Internet, an intranet, a local area network (LAN), a wide area network (WAN), or an interconnected combination of these network types. In at least one embodiment, connections within network 908 can be a remote modem, Ethernet (IEEE 802.3), Token Ring (IEEE 802.5), Fiber Distributed Datalink Interface (FDDI), Asynchronous Transfer Mode (ATM), or any other communications protocol. In at least one embodiment, computing devices linked to network can be desktops, servers, portable, handheld, set-top, personal digital assistants (PDAs), terminals, or any other desired type or configuration. In at least one embodiment, network-connected devices can vary widely in processing power, internal memory, and other performance depending on their functionality.

[0115] In at least one embodiment, communications within network and to or from computing devices connected to network can be wired or wireless. In at least one embodiment, network 908 can include, at least in part, the world-wide public Internet, which typically connects multiple users according to a client-server model according to Transmission Control Protocol / Internet Protocol (TCP / IP) specifications. In at least one embodiment, a client-server network is a dominant model for communication between two computers. In at least one embodiment, a client computer (“client”) issues one or more commands to a server computer (“server”). In at least one embodiment, a server fulfills client commands by accessing available network resources and returning information to a client according to client commands. In at least one embodiment, client computer systems and network resources residing on network servers are assigned network addresses for identification during communications between elements of a network. In at least one embodiment, communications from other network-connected systems to a server will include a network address of a relevant server / network resource as part of a communication, so that an appropriate destination for data / requests is identified as a recipient. In at least one embodiment, when network 908 includes the global Internet, network addresses are IP addresses in TCP / IP format, which can route data, at least in part, to an email account, website, or other Internet tool residing on a server. In at least one embodiment, information and services residing on network servers can be available to web browsers of client computers through a domain name (e.g., www.site.com), which maps to an IP address of a network server.

[0116] In at least one embodiment, a plurality of clients 902, 904, and 906 are connected to a network 908 via respective communication links. In at least one embodiment, each of these clients can access the network 908 via any desired form of communication, such as via a dial-up modem connection, a cable link, a digital subscriber line (DSL), a wireless or satellite link, or any other form of communication. In at least one embodiment, each client can communicate using any machine (e.g., a personal computer (PC), a workstation, a dedicated terminal, a personal data assistant (PDA), or other similar device) that is compatible with the network 908. In at least one embodiment, the clients 902, 904, and 906 may or may not be located in the same geographic area.

[0117] In at least one embodiment, multiple servers 910, 912, and 914 are connected to a network 918 to serve clients communicating with the network 918. In at least one embodiment, each server is typically a powerful computer or device that manages network resources and responds to client commands. In at least one embodiment, the servers include computer-readable data storage media, such as hard drives and RAM memory, that store program instructions and data. In at least one embodiment, servers 910, 912, and 914 run applications that respond to client commands. In at least one embodiment, server 910 may run a web server application that responds to client requests for HTML pages and may also run a mail server application that receives and routes emails. In at least one embodiment, other applications may also run on server 910, such as an FTP server or media server for streaming audio / video data to clients. In at least one embodiment, different servers may be dedicated to performing different tasks. In at least one embodiment, server 910 may be a dedicated web server that manages website-related resources for different users, while server 912 may be dedicated to providing email management. In at least one embodiment, other servers may be dedicated to media (audio, video, etc.), file transfer protocol (FTP), or a combination of any two or more services typically available or provided over a network. In at least one embodiment, each server may be in the same or different location as the other servers. In at least one embodiment, multiple servers may be present to perform mirroring tasks for users, thereby alleviating congestion or minimizing traffic directed to and from a single server. In at least one embodiment, servers 910, 912, 914 are under the control of a web hosting provider in the business of maintaining and delivering third-party content over network 918.

[0118] In at least one embodiment, a web hosting provider delivers services to two different types of clients. In at least one embodiment, one type, which may be referred to as a browser, requests content from servers 910, 912, 914, such as web pages, email messages, video clips, etc. In at least one embodiment, a second type, which may be referred to as a user, hires the web hosting provider to maintain network resources (such as a website) and make them available to the browser. In at least one embodiment, the user contracts with the web hosting provider to make available the memory space, processor capacity, and communication bandwidth required for the network resources they desire, depending on the amount of server resources they desire to utilize.

[0119] In at least one embodiment, in order for the web hosting provider to serve both clients, an application that manages the network resources hosted by the server must be appropriately configured. In at least one embodiment, the program configuration process involves defining a set of parameters that at least partially control the application's responses to browser requests and also at least partially define the server resources available to a particular user.

[0120] In one embodiment, intranet server 916 communicates with network 908 via a communication link. In at least one embodiment, intranet server 916 communicates with server manager 918. In at least one embodiment, server manager 918 includes a database of application configuration parameters used in servers 910, 912, 914. In at least one embodiment, a user modifies database 920 via intranet 916, and server manager 918 interacts with servers 910, 912, 914 to modify application parameters so that they match the contents of the database. In at least one embodiment, a user logs into intranet 916 by connecting to intranet 916 via computer 902 and entering authentication information such as a username and password.

[0121] In at least one embodiment, when a user wishes to log in to a new service or modify an existing service, the intranet server 916 authenticates the user and provides the user with an interactive screen display / control panel that allows the user to access configuration parameters for a particular application. In at least one embodiment, the user is presented with a plurality of modifiable text boxes that describe aspects of the configuration of the user's website or other network resources. In at least one embodiment, if the user desires to increase the memory space reserved for their website on the server, the user is provided with a field in which the user specifies the desired memory space. In at least one embodiment, in response to receiving this information, the intranet server 916 updates the database 920. In at least one embodiment, the server manager 918 forwards this information to the appropriate server and uses the new parameters during application operation. In at least one embodiment, the intranet server 916 is configured to provide the user with access to configuration parameters for hosted network resources (e.g., web pages, email, FTP sites, media sites, etc.) that the user has contracted with a web hosting service provider.

[0122] Figure 10A A networked computer system 1000A according to at least one embodiment is shown. In at least one embodiment, the networked computer system 1000A includes a plurality of nodes or personal computers ("PCs") 1002, 1018, 1020. In at least one embodiment, the personal computer or node 1002 includes a processor 1014, a memory 1016, a camera 1004, a microphone 1006, a mouse 1008, a speaker 1010, and a monitor 1012. In at least one embodiment, the PCs 1002, 1018, 1020 can each run one or more desktop servers, such as an internal network within a given company, or can be servers of a general-purpose network that is not limited to a particular environment. In at least one embodiment, each PC node of the network has one server, such that each PC node of the network represents a specific network server with a specific network URL address. In at least one embodiment, each server defaults to a default web page for users of that server, which itself can contain embedded URLs pointing to further subpages for that user on that server, or to other servers on the network or to pages on other servers.

[0123] In at least one embodiment, nodes 1002, 1018, 1020 and other nodes of the network are interconnected via a medium 1022. In at least one embodiment, the medium 1022 can be a communication channel such as an Integrated Services Digital Network ("ISDN"). In at least one embodiment, the various nodes of the networked computer system can be connected via various communication media, including a local area network ("LAN"), a plain old telephone line ("POTS") (sometimes referred to as a public switched telephone network ("PSTN")), and / or variations thereof. In at least one embodiment, the various nodes of the network can also constitute computer system users interconnected via a network such as the Internet. In at least one embodiment, each server on the network (operated from a particular node of the network at a given instance) has a unique address or identification within the network, which can be specified according to a URL.

[0124] In at least one embodiment, a plurality of multipoint conferencing units ("MCUs") can thus be used to transmit data to and from various nodes or "endpoints" of a conferencing system. In at least one embodiment, the nodes and / or MCUs can be interconnected via ISDN links or through a local area network ("LAN"), in addition to various other communication media (such as, nodes connected via the Internet). In at least one embodiment, the nodes of a conferencing system can generally be connected directly to a communication medium (such as a LAN) or through an MCU, and the conferencing system can include other nodes or elements, such as routers, servers, and / or variations thereof.

[0125] In at least one embodiment, processor 1014 is a general-purpose programmable processor. In at least one embodiment, the processor of a node of networked computer system 1000A may also be a dedicated video processor. In at least one embodiment, the various peripherals and components of a node (such as those of node 1002) may be different from those of other nodes. In at least one embodiment, node 1018 and node 1020 may be configured to be the same as or different from node 1002. In at least one embodiment, the node may be implemented on any suitable computer system other than a PC system.

[0126] Figure 10B10. A networked computer system 1000B is shown according to at least one embodiment. In at least one embodiment, system 1000B shows a network (such as LAN 1024) that can be used to interconnect various nodes that can communicate with each other. In at least one embodiment, attached to LAN 1024 are multiple nodes, such as PC nodes 1026, 1028, 1030. In at least one embodiment, the nodes can also be connected to the LAN via a network server or other device. In at least one embodiment, system 1000B includes other types of nodes or elements, including routers, servers, and nodes for at least one embodiment.

[0127] Figure 10C A networked computer system 1000C is shown in accordance with at least one embodiment. In at least one embodiment, system 1000C shows a WWW system with communications across a backbone communications network, such as the Internet 1032, which may be used to interconnect various nodes of the network. In at least one embodiment, the WWW is a set of protocols that operate on top of the Internet and allow graphical interface systems to operate on top of it to access information through the Internet. In at least one embodiment, attached to the Internet 1032 in the WWW are multiple nodes, such as PCs 1040, 1042, 1044. In at least one embodiment, the nodes interface with other nodes of the WWW through WWW HTTP servers, such as servers 1034, 1036. In at least one embodiment, PC 1044 may be a PC that forms a node of network 1032, and PC 1044 itself runs its server 1036, although for illustrative purposes only. Figure 10C PC 1044 and server 1036 are shown separately in FIG.

[0128] In at least one embodiment, the WWW is a distributed type of application characterized by WWW HTTP, the protocol of the WWW, which runs on top of the Internet's Transmission Control Protocol / Internet Protocol ("TCP / IP"). In at least one embodiment, the WWW can therefore be characterized by a set of protocols (i.e., HTTP) running on the Internet as its "backbone."

[0129] In at least one embodiment, a web browser is an application running on a node of the network in a WWW-compatible network system that allows users of a particular server or node to view such information and, therefore, to search for graphics and text-based files linked together using hypertext links embedded in documents or files available from servers on a network that understands HTTP. In at least one embodiment, when a user retrieves a given web page from a first server associated with a first node using another server on a network such as the Internet, the retrieved document may have different hypertext links embedded therein, and a local copy of the page is created locally on the retrieving user's machine. In at least one embodiment, when a user clicks on a hypertext link, the locally stored information associated with the selected hypertext link is typically sufficient to allow the user's machine to open a connection over the Internet to the server indicated by the hypertext link.

[0130] In at least one embodiment, more than one user can be coupled to each HTTP server via a LAN (such as LAN 1038, as shown with respect to WWW HTTP server 1034). In at least one embodiment, system 1000C may also include other types of nodes or elements. In at least one embodiment, the WWW HTTP server is an application running on a machine such as a PC. In at least one embodiment, each user can be considered to have a unique "server," as shown with respect to PC 1044. In at least one embodiment, a server can be considered to be a server, such as WWW HTTP server 1034, that provides access to the network for a LAN, or multiple nodes, or multiple LANs. In at least one embodiment, there are multiple users, each with a desktop PC or node on the network, each desktop PC potentially establishing a server for its user. In at least one embodiment, each server is associated with a specific network address or URL that, when accessed, provides a default web page for that user. In at least one embodiment, the web page may contain further links (embedded URLs) pointing to further subpages for that user on that server, or to other servers on the network or to pages on other servers on the network.

[0131] Cloud computing and services

[0132] The following figures illustrate, but are not limited to, exemplary cloud-based systems that may be used to implement at least one embodiment.

[0133] In at least one embodiment, cloud computing is a style of computing in which dynamically scalable and often virtualized resources are provided as services over the Internet. In at least one embodiment, users do not need knowledge, expertise, or control over the technical infrastructure that supports them; the technical infrastructure can be referred to as being "in the cloud." In at least one embodiment, cloud computing combines infrastructure as a service, platform as a service, software as a service, and other variations with a common theme of relying on the Internet to meet users' computing needs. In at least one embodiment, a typical cloud deployment (such as in a private cloud (e.g., an enterprise network)) or a data center (DC) in a public cloud (e.g., the Internet) can consist of thousands of servers (or alternatively, VMs), hundreds of Ethernet, Fibre Channel, or Fibre Channel over Ethernet (FCoE) ports, switching and storage infrastructure, etc. In at least one embodiment, the cloud can also consist of network service infrastructure, such as IPsec VPN concentrators, firewalls, load balancers, wide area network (WAN) optimizers, etc. In at least one embodiment, remote subscribers can securely access cloud applications and services by connecting via a VPN tunnel (e.g., an IPsec VPN tunnel).

[0134] In at least one embodiment, cloud computing is a model for enabling convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, servers, storage, applications, and services) that can be quickly provisioned and released with minimal management effort or service provider interaction.

[0135] In at least one embodiment, cloud computing is characterized by on-demand self-service, where consumers can automatically and unilaterally provision computing capabilities, such as server time and network storage, as needed, without requiring human interaction with each service provider. In at least one embodiment, cloud computing is characterized by broad network access, where capabilities are available over the network and accessed through standard mechanisms that facilitate use by heterogeneous thin or thick client platforms (e.g., mobile phones, laptops, and PDAs). In at least one embodiment, cloud computing is characterized by resource pooling, where a provider's computing resources are pooled to serve multiple consumers using a multi-tenant model, where different physical and virtual resources are dynamically signed up and reallocated based on consumer demand. In at least one embodiment, there is a sense of location independence, as consumers generally have no control or knowledge of the exact location of the provisioned resources, but may be able to specify the location at a higher level of abstraction (e.g., country, state, or data center).

[0136] In at least one embodiment, resources include storage, processing, memory, network bandwidth, and virtual machines. In at least one embodiment, cloud computing is characterized by rapid elasticity, where capacity can be quickly and elastically provisioned (in some cases automatically) to quickly scale down and quickly released to quickly scale up. In at least one embodiment, the capacity available for provisioning generally appears unlimited to the consumer and can be purchased in any quantity at any time. In at least one embodiment, cloud computing is characterized by metered services, where the cloud system automatically controls and optimizes resource usage by utilizing metering capabilities at some level of abstraction appropriate to the type of service (e.g., storage, processing, bandwidth, and active user accounts). In at least one embodiment, resource usage can be monitored, controlled, and reported, thereby providing transparency to both the provider and the consumer of the utilized service.

[0137] In at least one embodiment, cloud computing can be associated with a variety of services. In at least one embodiment, cloud software as a service (SaaS) can refer to a service that provides consumers with the ability to use a provider's applications running on a cloud infrastructure. In at least one embodiment, the applications can be accessed from various client devices through a thin client interface such as a web browser (e.g., web-based email). In at least one embodiment, the consumer does not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, storage, or even individual application capabilities, with the possible exception of limited user-specific application configuration settings.

[0138] In at least one embodiment, cloud platform as a service (PaaS) may refer to a service in which the capability provided to the consumer is to deploy consumer-created or acquired applications onto a cloud infrastructure, where these applications are created using programming languages ​​and tools supported by the provider. In at least one embodiment, the consumer does not manage or control the underlying cloud infrastructure, including networks, servers, operating systems, or storage, but does have control over the deployed applications and possibly the configuration of the application hosting environment.

[0139] In at least one embodiment, cloud infrastructure as a service (IaaS) can refer to a service in which the capabilities provided to the consumer are processing, storage, networking, and other basic computing resources upon which the consumer can deploy and run arbitrary software, which may include operating systems and applications. In at least one embodiment, the consumer does not manage or control the underlying cloud infrastructure, but rather has control over the operating system, storage, deployed applications, and possibly limited control over selected networking components (e.g., host firewalls).

[0140] In at least one embodiment, cloud computing can be deployed in different ways. In at least one embodiment, a private cloud may refer to cloud infrastructure that operates only for an organization. In at least one embodiment, a private cloud may be managed by an organization or a third party and may exist on-premises or off-premises. In at least one embodiment, a community cloud may refer to cloud infrastructure shared by several organizations and supporting a specific community with shared concerns (e.g., mission, security requirements, policies, and compliance considerations). In at least one embodiment, a community cloud may be managed by an organization or a third party and may exist on-premises or off-premises. In at least one embodiment, a public cloud may refer to cloud infrastructure that is available to the general public or a large industry group and owned by an organization providing cloud services. In at least one embodiment, a hybrid cloud may refer to a cloud infrastructure that is a composite of two or more clouds (private, community, or public) that remain unique entities but are bound together by standardized or proprietary technologies that enable data and application portability (e.g., cloud bursting for load balancing between clouds). In at least one embodiment, a cloud computing environment is service-oriented, focusing on statelessness, low coupling, modularity, and semantic interoperability.

[0141] Figure 11 One or more components of a system environment 1100 are shown, according to at least one embodiment, in which services may be provided as third-party network services. In at least one embodiment, the third-party network may be referred to as a cloud, a cloud network, a cloud computing network, and / or variations thereof. In at least one embodiment, the system environment 1100 includes one or more client computing devices 1104, 1106, and 1108, which can be used by users to interact with a third-party network infrastructure system 1102 that provides the third-party network services (which may be referred to as cloud computing services). In at least one embodiment, the third-party network infrastructure system 1102 may include one or more computers and / or servers.

[0142] It should be understood that Figure 11 The third-party network infrastructure system 1102 depicted in FIG may have other components in addition to those depicted. Further, Figure 11 In at least one embodiment, the third party network infrastructure system 1102 may have Figure 11 More or fewer components may be depicted, two or more components may be combined, or there may be a different configuration or arrangement of components.

[0143] In at least one embodiment, the client computing devices 1104, 1106, and 1108 can be configured to operate a client application, such as a web browser, a proprietary client application, or some other application that can be used by users of the client computing devices to interact with the third-party network infrastructure system 1102 to utilize services provided by the third-party network infrastructure system 1102. Although the exemplary system environment 1100 is shown with three client computing devices, any number of client computing devices can be supported. In at least one embodiment, other devices, such as devices with sensors, can interact with the third-party network infrastructure system 1102. In at least one embodiment, one or more networks 1110 can facilitate communication and data exchange between the client computing devices 1104, 1106, and 1108 and the third-party network infrastructure system 1102.

[0144] In at least one embodiment, the services provided by the third-party network infrastructure system 1102 may include a host of services available on-demand to users of the third-party network infrastructure system. In at least one embodiment, a variety of services may also be provided, including but not limited to online data storage and backup solutions, web-based email services, hosted office suites and document collaboration services, database management and processing, managed technical support services, and / or variations thereof. In at least one embodiment, the services provided by the third-party network infrastructure system may be dynamically scalable to meet the needs of its users.

[0145] In at least one embodiment, a specific instantiation of a service provided by the third-party network infrastructure system 1102 may be referred to as a "service instance." In at least one embodiment, generally, any service available to a user from a third-party network service provider system via a communication network (such as the Internet) is referred to as a "third-party network service." In at least one embodiment, in a public third-party network environment, the servers and systems that comprise the third-party network service provider system are different from the customer's own on-premises servers and systems. In at least one embodiment, the third-party network service provider system can host applications, and users can subscribe to and use the applications on demand via a communication network (such as the Internet).

[0146] In at least one embodiment, services within a computer network third-party network infrastructure may include protected computer network access to storage, hosted databases, hosted web servers, software applications, or other services provided to users by a third-party network provider. In at least one embodiment, services may include password-protected access to remote storage devices on the third-party network via the Internet. In at least one embodiment, services may include a hosted relational database and scripting language middleware engine based on a web service for private use by networked developers. In at least one embodiment, services may include access to an email software application hosted on a website of a third-party network provider.

[0147] In at least one embodiment, the third-party network infrastructure system 1102 may include a suite of application, middleware, and database service offerings delivered to customers in a self-service, subscription-based, elastically scalable, reliable, highly available, and secure manner. In at least one embodiment, the third-party network infrastructure system 1102 may also provide computing and analytical services related to "big data." In at least one embodiment, the term "big data" is generally used to refer to extremely large data sets that can be stored and manipulated by analysts and researchers to visualize, detect trends, and / or otherwise interact with the data. In at least one embodiment, big data and related applications can be hosted and / or manipulated by the infrastructure system at many levels and at varying scales. In at least one embodiment, dozens, hundreds, or thousands of processors linked in parallel may act on such data to render it or simulate external forces acting on the data or its representation. In at least one embodiment, these data sets may involve structured data (such as structured data organized in a database or otherwise according to a structured model) and / or unstructured data (e.g., emails, images, data blobs (binary large objects), web pages, complex event processing). In at least one embodiment, by leveraging the ability of embodiments to focus more (or fewer) computing resources on a target relatively quickly, third-party network infrastructure systems may be better available to perform tasks on large data sets based on demand from businesses, government agencies, research organizations, private individuals, groups of like-minded individuals or organizations, or other entities.

[0148] In at least one embodiment, the third-party network infrastructure system 1102 can be adapted to automatically provision, manage, and track customer subscriptions to services provided by the third-party network infrastructure system 1102. In at least one embodiment, the third-party network infrastructure system 1102 can provide third-party network services via different deployment models. In at least one embodiment, services can be provided under a public third-party network model, in which the third-party network infrastructure system 1102 is owned by the organization selling the third-party network services and makes the services available to the general public or businesses across various industries. In at least one embodiment, services can be provided under a private third-party network model, in which the third-party network infrastructure system 1102 operates solely for a single organization and can provide services to one or more entities within the organization. In at least one embodiment, third-party network services can also be provided under a community third-party network model, in which the third-party network infrastructure system 1102 and the services provided by the third-party network infrastructure system 1102 are shared by several organizations within a related community. In at least one embodiment, third-party network services can also be provided under a hybrid third-party network model, which is a combination of two or more different models.

[0149] In at least one embodiment, the services provided by the third-party network infrastructure system 1102 may include one or more services provided under the Software as a Service (SaaS) category, the Platform as a Service (PaaS) category, the Infrastructure as a Service (IaaS) category, or other service categories including hybrid services. In at least one embodiment, a customer may subscribe to one or more services provided by the third-party network infrastructure system 1102 via a subscription order. In at least one embodiment, the third-party network infrastructure system 1102 then performs processing to provide the services in the customer's subscription order.

[0150] In at least one embodiment, services provided by third party network infrastructure systems 1102 can include, without limitation, application services, platform services, and infrastructure services. In at least one embodiment, application services can be provided by third party network infrastructure systems via a SaaS platform. In at least one embodiment, a SaaS platform can be configured to provide third party network services that fall into the SaaS category. In at least one embodiment, a SaaS platform can provide the capability for customers to use applications, running on third party network infrastructure systems, that are built using an integrated development and deployment platform. In at least one embodiment, a SaaS platform can manage and control underlying software and infrastructure for providing the SaaS services. In at least one embodiment, by utilizing the services provided by a SaaS platform, customers can no longer have to worry about acquiring and managing the underlying hardware and software. In at least one embodiment, customers can obtain an application service without the need for customers to purchase, install, and manage software or hardware. In at least one embodiment, various different SaaS services can be provided. In at least one embodiment, this can include, without limitation, services for sales performance management, enterprise integration, and business flexibility that provide solutions for managing sales, aligning sales with customers, and improving business responsiveness, respectively.

[0151] In at least one embodiment, platform services can be provided by third party network infrastructure systems 1102 via a PaaS platform. In at least one embodiment, a PaaS platform can be configured to provide third party network services that fall into the PaaS category. In at least one embodiment, platform services can include, without limitation, services enabling organizations to combine existing applications with new applications built using the shared services provided by the platform, as well as the ability to establish new applications that leverage the shared services provided by the platform. In at least one embodiment, a PaaS platform can manage and control the underlying software and infrastructure for providing the PaaS services. In at least one embodiment, customers can obtain PaaS services provided by third party network infrastructure systems 1102 without the need for customers to purchase, install, and manage the underlying hardware and software.

[0152] In at least one embodiment, by utilizing the services provided by a PaaS platform, customers can use programming languages and tools supported by the third party network infrastructure system and also control deployed services. In at least one embodiment, platform services provided by a third party network infrastructure system can include database third party network services, middleware third party network services, and third party network services. In at least one embodiment, database third party network services can support a shared services deployment model that enables organizations to pool database resources and offer customers database as a service in the form of a database third party network. In at least one embodiment, middleware third party network services can provide customers with a platform for developing and deploying various business applications, and third party network services can provide customers with a platform to deploy applications in a third party network infrastructure system.

[0153] In at least one embodiment, various different infrastructure services can be provided by an IaaS platform in third party network infrastructure system. In at least one embodiment, infrastructure services facilitate the management and control of underlying computing resources, such as storage, networks, and other fundamental computing resources for customers utilizing services provided by SaaS and PaaS platforms.

[0154] In at least one embodiment, third party network infrastructure system 1102 can also include infrastructure resources 1130 for providing resources used to provide various services to customers of third party network infrastructure system. In at least one embodiment, infrastructure resources 1130 can include pre-integrated and optimized combinations of hardware, such as, for example, servers, storage, and networking resources, for

[0155] In at least one embodiment, resources in third party network infrastructure system 1102 can be shared by multiple users and dynamically re-allocated per demand. In at least one embodiment, resources can be allocated to users in different time zones. In at least one embodiment, third party network infrastructure system 1102 can enable a first set of users in a first time zone to utilize resources of third party network infrastructure system for a specified number of hours and subsequently enable reallocation of same resources to another set of users located in a different time zone, thereby maximizing resource utilization.

[0156] In at least one embodiment, a number of internal shared services 1132 can be provided that are shared by different components or modules of third party network infrastructure system 1102 for enabling services provided by third party network infrastructure system 1102. In at least one embodiment, these internal shared services can include, but are not limited to, security and identity services, integration services, enterprise repository services, enterprise manager services, virus scanning and white list services, high availability, backup and recovery services, services for enabling third party network support, email services, notification services, file transfer services, and / or variations thereof.

[0157] In at least one embodiment, third party network infrastructure system 1102 can provide comprehensive management of third party network services (e.g., SaaS, PaaS, and IaaS services) in third party network infrastructure system. In at least one embodiment, third party network management functionality can include the ability to provision, manage and track subscriptions of customers received by third party network infrastructure system 1102, and / or variations thereof.

[0158] In at least one embodiment, as Figure 11As shown, third-party network management functionality may be provided by one or more modules, such as an order management module 1120, an order coordination module 1122, an order provisioning module 1124, an order management and monitoring module 1126, and an identity management module 1128. In at least one embodiment, these modules may include or be provided using one or more computers and / or servers, which may be general-purpose computers, dedicated server computers, server farms, server clusters, or any other suitable arrangement and / or combination.

[0159] In at least one embodiment, at step 1134, a customer using a client device (such as client computing device 1104, 1106, or 1108) can interact with the third-party network infrastructure system 1102 by requesting one or more services provided by the third-party network infrastructure system 1102 and placing an order for a subscription to the one or more services provided by the third-party network infrastructure system 1102. In at least one embodiment, the customer can access a third-party network user interface (UI), such as third-party network UI 1112, third-party network UI 1114, and / or third-party network UI 1116, and place an order via these UIs. In at least one embodiment, the order information received by the third-party network infrastructure system 1102 in response to the customer placing the order can include information identifying the customer and one or more services provided by the third-party network infrastructure system 1102 to which the customer wishes to subscribe.

[0160] In at least one embodiment, at step 1136, the order information received from the customer can be stored in order database 1118. In at least one embodiment, if this is a new order, a new record can be created for the order. In at least one embodiment, order database 1118 can be one of several databases operated by third-party network infrastructure system 1118 and in conjunction with other system components.

[0161] In at least one embodiment, at step 1138, the order information may be forwarded to order management module 1120, which may be configured to perform billing and accounting functions related to the order, such as verifying the order and, upon verification, booking an order.

[0162] In at least one embodiment, at step 1140, information about the order can be transmitted to an order coordination module 1122 that is configured to coordinate the provisioning of services and resources for orders placed by customers. In at least one embodiment, order coordination module 1122 can use the services of an order provisioning module 1124 for provisioning. In at least one embodiment, order coordination module 1122 enables management of business processes associated with each order and applies business logic to determine whether an order should continue to be provisioned.

[0163] In at least one embodiment, at step 1142, upon receiving a new subscribed order, order coordination module 1122 sends a request to order provisioning module 1124 to allocate resources and configure resources needed to fulfill the subscribed order. In at least one embodiment, order provisioning module 1124 implements resource allocation for services ordered by customers. In at least one embodiment, order provisioning module 1124 provides a level of abstraction between third-party network infrastructure systems 1100 provided third-party network services and physical implementation layers used to provision resources for providing the requested services. In at least one embodiment, this enables order coordination module 1122 to be isolated from implementation details, such as whether services and resources are provisioned in real-time or pre-provisioned and only allocated / assigned upon request.

[0164] In at least one embodiment, at step 1144, once services and resources are provisioned, a notification can be sent to the subscribing customer indicating that the requested services are now ready for use. In at least one embodiment, information (e.g., a link) can be sent to the customer that enables the customer to begin using the requested services.

[0165] In at least one embodiment, at step 1146, orders for customers that are subscribed to can be managed and tracked by an order management and monitoring module 1126. In at least one embodiment, order management and monitoring module 1126 can be configured to collect usage statistics about customer usage of subscribed services. In at least one embodiment, statistics can be collected for amounts of storage used, amounts of data transferred, number of users, and amounts and / or changes in system up times and system down times.

[0166] In at least one embodiment, the third-party network infrastructure system 1100 may include an identity management module 1128 configured to provide identity services, such as access management and authorization services within the third-party network infrastructure system 1100. In at least one embodiment, the identity management module 1128 may control information about customers who wish to utilize services provided by the third-party network infrastructure system 1102. In at least one embodiment, such information may include information authenticating the identities of such customers and information describing which actions those customers are authorized to perform with respect to various system resources (e.g., files, directories, applications, communication ports, memory segments, etc.). In at least one embodiment, the identity management module 1128 may also include managing descriptive information about each customer, as well as information about how and by whom the descriptive information may be accessed and modified.

[0167] Figure 12 A cloud computing environment 1202 is shown in accordance with at least one embodiment. In at least one embodiment, the cloud computing environment 1202 includes one or more computer systems / servers 1204 with which computing devices such as personal digital assistants (PDAs) or cell phones 1206A, desktop computers 1206B, laptop computers 1206C, and / or automobile computer systems 1206N communicate. In at least one embodiment, this allows infrastructure, platforms, and / or software to be provided as a service from the cloud computing environment 1202, so that each client does not need to maintain such resources individually. It should be understood that Figure 12 The types of computing devices 1206A-N shown are intended to be illustrative only, and the cloud computing environment 1202 may communicate with any type of computerized device over any type of network and / or network / addressable connection (eg, using a web browser).

[0168] In at least one embodiment, computer system / server 1204, which may be represented as a cloud computing node, is operable with numerous other general-purpose or special-purpose computing system environments or configurations. In at least one embodiment, computing systems, environments, and / or configurations that may be suitable for use with computer system / server 1204 include, but are not limited to, personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputer systems, mainframe computer systems, and distributed cloud computing environments including any of the foregoing, and / or variations thereof.

[0169] In at least one embodiment, computer system / server 1204 can be described in the general context of computer system-executable instructions (such as program modules) executed by a computer system. In at least one embodiment, program modules include routines, programs, objects, components, logic, data structures, etc. that perform specific tasks or implement specific abstract data types. In at least one embodiment, exemplary computer system / server 1204 can be practiced in a distributed cloud computing environment, where tasks are performed by remote processing devices linked through a communication network. In at least one embodiment, in a distributed cloud computing environment, program modules can be located in both local and remote computer system storage media, including memory storage devices.

[0170] Figure 13 The cloud computing environment 1202 ( Figure 12 ) provides a set of functional abstraction layers. It should be understood in advance that Figure 13 The components, layers, and functions shown in are intended to be illustrative only, and the components, layers, and functions may vary.

[0171] In at least one embodiment, the hardware and software layer 1302 includes hardware and software components. In at least one embodiment, the hardware components include mainframes, servers based on various RISC (Reduced Instruction Set Computer) architectures, various computing systems, supercomputing systems, storage devices, networks, networking components, and / or variations thereof. In at least one embodiment, the software components include network application server software, various application server software, various database software, and / or variations thereof.

[0172] In at least one embodiment, the virtualization layer 1304 provides an abstraction layer from which the following exemplary virtual entities can be provided: virtual servers, virtual storage, virtual networks (including virtual private networks), virtual applications, virtual clients, and / or variations thereof.

[0173] In at least one embodiment, the management layer 1306 provides various functions. In at least one embodiment, resource provisioning provides dynamic acquisition of computing resources and other resources for performing tasks within the cloud computing environment. In at least one embodiment, metering provides usage tracking when resources are utilized within the cloud computing environment, as well as billing or invoicing for the consumption of those resources. In at least one embodiment, resources may include application software licenses. In at least one embodiment, security provides authentication for users and tasks, as well as protection of data and other resources. In at least one embodiment, a user interface provides access to the cloud computing environment for both users and system administrators. In at least one embodiment, service level management provides allocation and management of cloud computing resources so that required service levels are met. In at least one embodiment, service level agreement (SLA) management provides pre-placement and acquisition of cloud computing resources in anticipation of future demand for the cloud computing resources according to the SLA.

[0174] In at least one embodiment, workload layer 1308 provides functionality that leverages a cloud computing environment. In at least one embodiment, workloads and functionality that can be provided from this layer include: mapping and navigation, software development and management, educational services, data analysis and processing, transaction processing, and service delivery.

[0175] Supercomputing

[0176] The following figures illustrate, but are not limited to, exemplary supercomputer-based systems that may be used to implement at least one embodiment.

[0177] In at least one embodiment, a supercomputer may refer to a hardware system that exhibits significant parallelism and includes at least one chip, wherein the chips in the system are interconnected by a network and placed in a hierarchically organized housing. In at least one embodiment, a large hardware system that fills a computer room with several racks, each rack containing several boards / rack modules, each board / rack module containing several chips all interconnected by a scalable network is at least one embodiment of a supercomputer. In at least one embodiment, a single rack of such a large hardware system is at least one other embodiment of a supercomputer. In at least one embodiment, a single chip that exhibits significant parallelism and includes several hardware components may also be considered a supercomputer because as feature sizes may decrease, the amount of hardware that can be combined in a single chip may also increase.

[0178] Figure 14A chip-level supercomputer according to at least one embodiment is shown. In at least one embodiment, the main computation is performed within a finite state machine (1404) called a thread unit within an FPGA or ASIC chip. In at least one embodiment, a task and synchronization network (1402) connects the finite state machine and is used to dispatch threads and execute operations in the correct order. In at least one embodiment, a memory network (1406, 1410) is used to access multi-level partitioned on-chip cache levels (1408, 1412). In at least one embodiment, a memory controller (1416) and an off-chip memory network (1414) are used to access off-chip memory. In at least one embodiment, an I / O controller (1418) is used for cross-chip communication when the design is not suitable for a single logic chip.

[0179] Figure 15 A supercomputer at the rack module level is shown according to at least one embodiment. In at least one embodiment, within the rack module, there are multiple FPGA or ASIC chips (1502) connected to one or more DRAM units (1504) that constitute the main accelerator memory. In at least one embodiment, each FPGA / ASIC chip is connected to its neighboring FPGA / ASIC chips using differential high-speed signaling (1506) using a wide bus on the board. In at least one embodiment, each FPGA / ASIC chip is also connected to at least one high-speed serial communication cable.

[0180] Figure 16 A rack-scale supercomputer is shown in accordance with at least one embodiment. Figure 17 An overall system-level supercomputer according to at least one embodiment is shown. In at least one embodiment, see Figure 16 and Figure 17, between rack modules in a rack and across racks throughout the system, high-speed serial optical or copper cables (1602, 1702) are used to implement a scalable, potentially incomplete, hypercube network. In at least one embodiment, one of the accelerator's FPGA / ASIC chips is connected to a host system (1704) via a PCI-Express connection. In at least one embodiment, the host system includes a host microprocessor (1708) on which the software portion of the application runs, and memory consisting of one or more host memory DRAM cells (1706) that are coherent with the memory on the accelerator. In at least one embodiment, the host system can be a separate module on one of the racks, or can be integrated with one of the modules of the supercomputer. In at least one embodiment, a circular topology of cube connections provides communication links to create a hypercube network for a large supercomputer. In at least one embodiment, small groups of FPGA / ASIC chips on a rack module can act as a single hypercube node, increasing the total number of external links per group compared to a single chip. In at least one embodiment, a group includes chips A, B, C, and D on a rack module with an internal wide differential bus connecting A, B, C, and D in a ring organization. In at least one embodiment, there are 12 serial communication cables connecting the rack modules to the outside world. In at least one embodiment, chip A on the rack module connects to serial communication cables 0, 1, and 2. In at least one embodiment, chip B connects to cables 3, 4, and 5. In at least one embodiment, chip C connects to cables 6, 7, and 8. In at least one embodiment, chip D connects to cables 9, 10, and 11. In at least one embodiment, the entire group {A, B, C, D} that makes up the rack modules can form a hypercube node within a supercomputer system, with up to 212 = 4096 rack modules (16384 FPGA / ASIC chips). In at least one embodiment, in order for chip A to send a message out on link 4 of the group {A, B, C, D}, the message must first be routed to chip B using the on-board differential wide bus connection. In at least one embodiment, a message arriving on link 4 of the group {A, B, C, D} destined for chip A (i.e., to B) must also first be routed to the correct destination chip (A) within the group {A, B, C, D}. In at least one embodiment, parallel supercomputer systems of other sizes may also be implemented.

[0181] AI

[0182] The following figures illustrate, but are not limited to, exemplary artificial intelligence-based systems that can be used to implement at least one embodiment.

[0183] Figure 18AInference and / or training logic 1815 is shown for performing inference and / or training operations associated with one or more embodiments. Figure 18A and / or Figure 18B Provide details about the inference and / or training logic 1815.

[0184] In at least one embodiment, inference and / or training logic 1815 may include, but is not limited to, code and / or data storage 1801 for storing forward and / or output weights and / or input / output data, and / or other parameters used to configure neurons or layers of a neural network being trained and / or used for inference in aspects of one or more embodiments. In at least one embodiment, training logic 1815 may include or be coupled to code and / or data storage 1801 for storing graph code or other software to control the timing and / or sequence in which weights and / or other parameter information are loaded to configure logic, including integer and / or floating point units (collectively referred to as arithmetic logic units (ALUs)). In at least one embodiment, code (such as graph code) loads weights or other parameter information into a processor ALU based on the architecture of the neural network to which such code corresponds. In at least one embodiment, code and / or data storage 1801 stores weight parameters and / or input / output data for each layer of a neural network trained or used in conjunction with one or more embodiments during forward propagation of input / output data and / or weight parameters during training and / or inference using aspects of one or more embodiments. In at least one embodiment, any portion of code and / or data storage 1801 may be included with other on-chip or off-chip data storage devices, including the processor's L1, L2, or L3 cache memory or system memory.

[0185] In at least one embodiment, any portion of code and / or data storage 1801 may be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, code and / or code and / or data storage 1801 may be cache memory, dynamic random addressable memory ("DRAM"), static random addressable memory ("SRAM"), non-volatile memory (e.g., flash memory), or other storage. In at least one embodiment, the choice of whether code and / or code and / or data storage 1801 is internal or external to a processor, or includes DRAM, SRAM, flash memory, or some other type of storage, may depend on the available storage on-chip versus off-chip, the latency requirements of the training and / or inference functions being performed, the batch size of data used in inference and / or training of the neural network, or some combination of these factors.

[0186] In at least one embodiment, inference and / or training logic 1815 can include, without limitation, code and / or data storage 1805 for storing backward and / or output weights and / or input / output data corresponding to neurons or layers of a neural network being trained and / or used for inferencing in aspects of one or more embodiments. In at least one embodiment, code and / or data storage 1805 stores weight parameters and / or input / output data for each layer of a neural network that is trained in conjunction with one or more embodiments during backpropagation of input / output data and / or weight parameters during training and / or inferencing using aspects of one or more embodiments. In at least one embodiment, training logic 1815 can include or be coupled to code and / or data storage 1805 to store graph code or other software to control timing and / or sequence in which weight and / or other parameter information will be loaded to configure logic, including integer and / or floating point units (collectively, arithmetic logic units (ALUs)).

[0187] In at least one embodiment, code such as graph code causes weight or other parameter information to be loaded into processor ALUs based on an architecture of a neural network to which such code corresponds. In at least one embodiment, any portion of code and / or data storage 1805 can be included with other on-chip or off-chip data storage, including a processor’s LI, L2, or L3 cache or system memory. In at least one embodiment, any portion of code and / or data storage 1805 can be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, code and / or data storage 1805 can be cache memory, DRAM, SRAM, non-volatile memory (e.g., Flash memory), or other storage. In at least one embodiment, whether code and / or data storage 1805 is internal or external to a processor, or includes a choice of DRAM, SRAM, Flash, or some other storage type, can depend on available storage on-chip versus off-chip, latency requirements of training and / or inferencing functions being performed, batch size of data used in inferencing and / or training of a neural network, or some combination of these factors.

[0188] In at least one embodiment, code and / or data storage 1801 and code and / or data storage 1805 may be separate storage structures. In at least one embodiment, code and / or data storage 1801 and code and / or data storage 1805 may be a combined storage structure. In at least one embodiment, code and / or data storage 1801 and code and / or data storage 1805 may be partially combined and partially separate. In at least one embodiment, any portion of code and / or data storage 1801 and code and / or data storage 1805 may be included with other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory.

[0189] In at least one embodiment, inference and / or training logic 1815 may include, but is not limited to, one or more arithmetic logic units (“ALUs”) 1810 , including integer and / or floating point units, for performing logical and / or mathematical operations based at least in part on or directed by training and / or inference code (e.g., graph code), the results of which may produce activations (e.g., output values ​​from a layer or neuron within a neural network) stored in activation storage 1820 , which is a function of input / output and / or weight parameter data stored in code and / or data storage 1801 and / or code and / or data storage 1805 . In at least one embodiment, the activations stored in activation storage 1820 are generated based on linear algebra and / or matrix-based math performed by ALU 1810 in response to executing instructions or other code, where weight values ​​stored in code and / or data storage 1805 and / or data storage 1801 are used as operands along with other values ​​such as bias values, gradient information, momentum values, or other parameters or hyperparameters, any or all of which may be stored in code and / or data storage 1805 or code and / or data storage 1801 or in another storage on or off-chip.

[0190] In at least one embodiment, one or more ALUs 1810 are included within one or more processors or other hardware logic devices or circuits, while in another embodiment, one or more ALUs 1810 may be external to the processor or other hardware logic devices or circuits (e.g., coprocessors) that use them. In at least one embodiment, ALUs 1810 may be included within an execution unit of a processor or otherwise within an ALU bank accessible by an execution unit of a processor, either within the same processor or distributed across different processors of different types (e.g., central processing units, graphics processing units, fixed function units, etc.). In at least one embodiment, code and / or data storage 1801, code and / or data storage 1805, and activation storage 1820 may share a processor or other hardware logic device or circuit, while in another embodiment, they may be in different processors or other hardware logic devices or circuits, or in some combination of the same and different processors or other hardware logic devices or circuits. In at least one embodiment, any portion of activation storage 1820 may be included with other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory. Additionally, inference and / or training code may be stored with other code accessible to the processor or other hardware logic or circuitry and fetched and / or processed using the processor's fetch, decode, schedule, execute, retire, and / or other logic circuitry.

[0191] In at least one embodiment, activation storage 1820 can be cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other storage devices. In at least one embodiment, activation storage 1820 can be completely or partially internal or external to one or more processors or other logic circuits. In at least one embodiment, the choice of whether activation storage 1820 is internal or external to the processor, or includes DRAM, SRAM, flash memory, or some other storage type, can depend on the available on-chip versus off-chip storage, the latency requirements of the training and / or inference functions being performed, the batch size of the data used in inference and / or training of the neural network, or some combination of these factors.

[0192] In at least one embodiment, Figure 18A The inference and / or training logic 1815 shown in FIG can be used in conjunction with an application specific integrated circuit (“ASIC”), such as the one from Google. Processing unit from Graphcore TM Inference Processing Unit (IPU), or from Intel (e.g., "Lake Crest") processor. In at least one embodiment, Figure 18A The inference and / or training logic 1815 shown in FIG may be used in conjunction with central processing unit (“CPU”) hardware, graphics processing unit (“GPU”) hardware, or other hardware such as a field programmable gate array (“FPGA”).

[0193] Figure 18B Inference and / or training logic 1815 is shown in accordance with at least one embodiment. In at least one embodiment, inference and / or training logic 1815 may include, but is not limited to, hardware logic in which computing resources are dedicated or otherwise used exclusively in conjunction with weight values ​​or other information corresponding to one or more neuron layers within a neural network. In at least one embodiment, Figure 18B The inference and / or training logic 1815 shown in FIG can be combined with an application specific integrated circuit (ASIC) (such as the one from Google Processing unit from Graphcore TM Inference Processing Unit (IPU), or from Intel (e.g., "Lake Crest") processor. In at least one embodiment, Figure 18B The inference and / or training logic 1815 shown in FIG can be used in conjunction with central processing unit (CPU) hardware, graphics processing unit (GPU) hardware, or other hardware such as a field programmable gate array (FPGA). In at least one embodiment, the inference and / or training logic 1815 includes, but is not limited to, code and / or data storage 1801 and code and / or data storage 1805, which can be used to store code (e.g., graph code), weight values, and / or other information, including bias values, gradient information, momentum values, and / or other parameter or hyperparameter information. Figure 18B In at least one embodiment described in

[0065] , each of code and / or data storage 1801 and code and / or data storage 1805 is associated with dedicated computing resources, such as computing hardware 1802 and computing hardware 1806, respectively. In at least one embodiment, each of computing hardware 1802 and computing hardware 1806 includes one or more ALUs that perform mathematical functions (such as linear algebraic functions) solely on the information stored in code and / or data storage 1801 and code and / or data storage 1805, respectively, with the results being stored in activation storage 1820.

[0194] In at least one embodiment, each code and / or data storage 1801 and 1805, and corresponding computational hardware 1802 and 1806, respectively, corresponds to a different layer of a neural network, such that the resulting activations from one storage / computation pair 1801 / 1802 in code and / or data storage 1801 and computational hardware 1802 are provided as input to the next storage / computation pair 1805 / 1806 in code and / or data storage 1805 and computational hardware 1806, mirroring the conceptual organization of the neural network. In at least one embodiment, each of storage / computation pairs 1801 / 1802 and 1805 / 1806 may correspond to more than one neural network layer. In at least one embodiment, additional storage / computation pairs (not shown) may be included in the inference and / or training logic 1815, either after or in parallel with storage / computation pairs 1801 / 1802 and 1805 / 1806.

[0195] Figure 19 The training and deployment of a deep neural network according to at least one embodiment is shown. In at least one embodiment, an untrained neural network 1906 is trained using a training dataset 1902. In at least one embodiment, the training framework 1904 is the PyTorch framework, while in other embodiments, the training framework 1904 is TensorFlow, Boost, Caffe, Microsoft Cognitive Toolkit / CNTK, MXNet, Chainer, Keras, Deeplearning4j, or other training frameworks. In at least one embodiment, the training framework 1904 trains the untrained neural network 1906 and enables it to be trained using the processing resources described herein to generate a trained neural network 1908. In at least one embodiment, the weights can be randomly selected or selected by pre-training using a deep belief network. In at least one embodiment, the training can be performed in a supervised, partially supervised, or unsupervised manner.

[0196] In at least one embodiment, untrained neural network 1906 is trained using supervised learning, where training dataset 1902 includes inputs paired with expected outputs for the inputs, or where training dataset 1902 includes inputs with known outputs and the outputs of neural network 1906 are manually graded. In at least one embodiment, untrained neural network 1906 is trained in a supervised manner, processing inputs from training dataset 1902 and comparing the resulting outputs to a set of expected or desired outputs. In at least one embodiment, errors are then backpropagated through untrained neural network 1906. In at least one embodiment, training framework 1904 adjusts the weights that control untrained neural network 1906. In at least one embodiment, training framework 1904 includes tools for monitoring how well untrained neural network 1906 converges toward a model (such as trained neural network 1908) suitable for generating correct answers (such as results 1914) based on input data (such as new dataset 1912). In at least one embodiment, the training framework 1904 repeatedly trains the untrained neural network 1906 while adjusting the weights using a loss function and an adjustment algorithm (such as stochastic gradient descent) to refine the output of the untrained neural network 1906. In at least one embodiment, the training framework 1904 trains the untrained neural network 1906 until the untrained neural network 1906 achieves a desired accuracy. In at least one embodiment, the trained neural network 1908 can then be deployed to implement any number of machine learning operations.

[0197] In at least one embodiment, untrained neural network 1906 is trained using unsupervised learning, wherein untrained neural network 1906 attempts to train itself using unlabeled data. In at least one embodiment, unsupervised learning training dataset 1902 will include input data without any associated output data or "ground truth" data. In at least one embodiment, untrained neural network 1906 can learn groupings within training dataset 1902 and can determine how individual inputs relate to untrained dataset 1902. In at least one embodiment, unsupervised training can be used to generate a self-organizing map in trained neural network 1908 that is capable of performing operations useful in reducing the dimensionality of new dataset 1912. In at least one embodiment, unsupervised training can also be used to perform anomaly detection, which allows identification of data points in new dataset 1912 that deviate from the normal pattern of new dataset 1912.

[0198] In at least one embodiment, semi-supervised learning can be used, which is a technique in which a mixture of labeled and unlabeled data is included in the training dataset 1902. In at least one embodiment, the training framework 1904 can be used to perform incremental learning, such as through transfer learning techniques. In at least one embodiment, incremental learning enables the trained neural network 1908 to adapt to new datasets 1912 without forgetting the knowledge infused into the trained neural network 1408 during initial training.

[0199] 5G network

[0200] The following figures illustrate, but are not limited to, exemplary 5G network-based systems that may be used to implement at least one embodiment.

[0201] Figure 20 The architecture of a system 2000 for a network according to at least one embodiment is shown. In at least one embodiment, the system 2000 is shown as including user equipment (UE) 2002 and UE 2004. In at least one embodiment, the UEs 2002 and 2004 are shown as smartphones (e.g., handheld touchscreen mobile computing devices that can connect to one or more cellular networks), but may also include any mobile or non-mobile computing device, such as a personal digital assistant (PDA), a pager, a laptop computer, a desktop computer, a wireless handheld device, or any computing device that includes a wireless communication interface.

[0202] In at least one embodiment, any of UE 2002 and UE 2004 may comprise an Internet of Things (IoT) UE, which may include a network access layer designed for low-power IoT applications that utilize short-lived UE connections. In at least one embodiment, the IoT UE may utilize technologies such as machine-to-machine (M2M) or machine-type communication (MTC) for exchanging data with an MTC server or device via a public land mobile network (PLMN), proximity-based services (ProSe), or device-to-device (D2D) communication, a sensor network, or an IoT network. In at least one embodiment, the M2M or MTC data exchange may be machine-initiated data exchange. In at least one embodiment, the IoT network describes interconnected IoT UEs, which may include uniquely identifiable embedded computing devices (within the Internet infrastructure) with short-lived connections. In at least one embodiment, the IoT UE may execute background applications (e.g., keep-alive messages, status updates, etc.) to facilitate connectivity to the IoT network.

[0203] In at least one embodiment, UE 2002 and UE 2004 can be configured to connect (e.g., be communicatively coupled) to a radio access network (RAN) 2016. In at least one embodiment, RAN 2016 can be an evolved universal mobile telecommunications system (UMTS) terrestrial radio access network (E-UTRAN), a NextGen RAN (NG RAN), or some other type of RAN. In at least one embodiment, UE 2002 and UE 2004 utilize connection 2012 and connection 2014, respectively, each of which includes a physical communication interface or layer. In at least one embodiment, connections 2012 and 2014 are shown as air interfaces for achieving communicative coupling and can be consistent with a cellular communication protocol, such as a Global System for Mobile Communications (GSM) protocol, a Code Division Multiple Access (CDMA) network protocol, a Push-to-Talk (PTT) protocol, a PTT over Cellular (POC) protocol, a Universal Mobile Telecommunications System (UMTS) protocol, a 3GPP Long Term Evolution (LTE) protocol, a fifth generation (5G) protocol, a New Radio (NR) protocol, and variations thereof.

[0204] In at least one embodiment, the UEs 2002 and 2004 may also directly exchange communication data via a ProSe interface 2006. In at least one embodiment, the ProSe interface 2006 may alternatively be referred to as a side link interface, which includes one or more logical channels, including but not limited to a physical side link control channel (PSCCH), a physical side link shared channel (PSSCH), a physical side link discovery channel (PSDCH), and a physical side link broadcast channel (PSBCH).

[0205] In at least one embodiment, UE 2004 is shown as being configured to access an access point (AP) 2010 via a connection 2008. In at least one embodiment, connection 2008 may comprise a local wireless connection, such as a connection consistent with any IEEE 802.11 protocol, wherein AP 2010 would include Wireless Fidelity (WFI). ) router. In at least one embodiment, AP 2010 is shown as being connected to the Internet and not to the core network of the wireless system.

[0206] In at least one embodiment, RAN 2016 can include one or more access nodes enabling access to network 2014 for D2D enabled UEs 2002 and 2004. In at least one embodiment, these access nodes (AN) can be referred to as base stations (BS), NodeBs, evolved NodeBs (eNB), giga-NodeBs (gNB), RAN nodes, and so on, and can comprise ground stations (e.g., terrestrial access points) or satellite stations providing coverage over a geographic area (e.g., a cell).

[0207] In at least one embodiment, any of RAN nodes 2018 and 2020 can terminate the air interface protocol and can be the first point of contact for UEs 2002 and 2004. In at least one embodiment, any of RAN nodes 2018 and 2020 can fulfill various logical functions for the RAN 2016 including, but not limited to, RNC functions such as radio bearer management, uplink and downlink dynamic radio resource management and data packet scheduling, and mobility management.

[0208] In at least one embodiment, UEs 2002 and 2004 can be configured to communicate using orthogonal frequency division multiplexing (OFDM) communication signals with each other or with any of RAN nodes 2018 and 2020 over a multicarrier communication channel in accordance with various communication techniques, such as, but not limited to, an orthogonal frequency division multiple access (OFDMA) communication technique (e.g., for downlink communications) or single carrier frequency division multiple access (SC-FDMA) communication technique (e.g., for uplink and ProSe or sidelink communications), and / or variants thereof. In at least one embodiment, OFDM signals can comprise multiple orthogonal subcarriers.

[0209] In at least one embodiment, a downlink resource grid can be used for downlink transmissions from any of RAN node 2018 and 2020 to UEs 2002 and 2004, while uplink transmissions can utilize a similar approach. In at least one embodiment, a grid can be a time-frequency grid, called a resource grid or time-frequency resource grid, which is the physical resource in the downlink in each slot. In at least one embodiment, such a time-frequency plane representation is a common practice for OFDM systems, which makes it intuitive for radio resource allocation. In at least one embodiment, each column and each row of the resource grid corresponds to one OFDM symbol and one OFDM subcarrier, respectively. In at least one embodiment, the duration of the resource grid in the time domain corresponds to one slot, which depends on the downlink slot duration. In at least one embodiment, the minimum time-frequency unit in the resource grid is denoted as a resource element. In at least one embodiment, each resource element in the resource grid can be assigned to a particular physical channel and / or used for transmission of data and control information for UEs 2002 and 2004. In at least one embodiment, a resource grid can be further partitioned into resource blocks and subcarriers. In at least one embodiment, each resource block comprises a collection of resource elements and can be used to transmit one transport block to one UE 2002 and 2004. In at least one embodiment, the use of resource blocks can be arranged vertically in the time domain and natively in the frequency domain.

[0210] In at least one embodiment, a physical downlink shared channel (PDSCH) can carry user data and higher layer signaling to UEs 2002 and 2004. In at least one embodiment, a physical downlink control channel (PDCCH) can carry information about the transport format and resource allocations for the PDSCH channels, among other information. In at least one embodiment, it can also inform UEs 2002 and 2004 about the transport format, resource allocation, and HARQ information for uplink shared channel. Generally, in at least one embodiment, downlink scheduling (which allocates control and shared channel resource blocks to UEs 2002 within a cell) can be performed by any of RAN nodes 2018 and 2020 based on channel quality information feedback from any of UEs 2002 and 2004. In at least one embodiment, downlink resource allocation information can be sent on the PDCCH used for each of UEs 2002 and 2004.

[0211] In at least one embodiment, the PDCCH may use control channel elements (CCEs) to transmit control information. In at least one embodiment, before being mapped to resource elements, the PDCCH complex symbols may first be organized into quadruplets, which may then be permuted using a sub-block interleaver for rate matching. In at least one embodiment, each PDCCH may be transmitted using one or more of these CCEs, where each CCE may correspond to nine sets of four physical resource elements referred to as resource element groups (REGs). In at least one embodiment, four quadrature phase shift keying (QPSK) symbols may be mapped to each REG. In at least one embodiment, one or more CCEs may be used to transmit the PDCCH, depending on the size of the downlink control information (DCI) and the channel conditions. In at least one embodiment, there may be four or more different PDCCH formats (e.g., aggregation levels, L=1, 2, 4, or 8) defined in LTE with different numbers of CCEs.

[0212] In at least one embodiment, an enhanced physical downlink control channel (EPDCCH) using PDSCH resources may be used for control information transmission. In at least one embodiment, EPDCCH may be transmitted using one or more enhanced control channel elements (ECCEs). In at least one embodiment, each ECCE may correspond to nine sets of four physical resource elements referred to as enhanced resource element groups (EREGs). In at least one embodiment, ECCEs may have other numbers of EREGs in some cases.

[0213] In at least one embodiment, the RAN 2016 is shown as being communicatively coupled to a core network (CN) 2038 via an S1 interface 2022. In at least one embodiment, the CN 2038 can be an evolved packet core (EPC) network, a NextGen packet core (NPC) network, or some other type of CN. In at least one embodiment, the S1 interface 2022 is divided into two parts: an S1-U interface 2026, which carries traffic data between the RAN nodes 2018 and 2020 and the serving gateway (S-GW) 2030; and an S1-Mobility Management Entity (MME) interface 2024, which is a signaling interface between the RAN nodes 2018 and 2020 and the MME 2028.

[0214] In at least one embodiment, CN 2038 includes MME 2028, S-GW 2030, Packet Data Network (PDN) Gateway (P-GW) 2034, and Home Subscriber Server (HSS) 2032. In at least one embodiment, MME 2028 can be functionally similar to the control plane of a traditional Serving General Packet Radio Service (GPRS) Support Node (SGSN). In at least one embodiment, MME 2028 can manage mobility aspects of access, such as gateway selection and tracking area list management. In at least one embodiment, HSS 2032 can include a database for network users, including subscription-related information used to support network entities handling communication sessions. In at least one embodiment, CN 2038 can include one or more HSSs 2032, depending on the number of mobile users, device capacity, network organization, etc. In at least one embodiment, HSS 2032 can provide support for routing / roaming, authentication, authorization, naming / addressing resolution, location dependencies, etc.

[0215] In at least one embodiment, the S-GW 2030 may terminate the S1 interface 2022 towards the RAN 2016 and route data packets between the RAN 2016 and the CN 2038. In at least one embodiment, the S-GW 2030 may be the local mobility anchor for inter-RAN node handovers and may also provide an anchor for inter-3GPP mobility. In at least one embodiment, other responsibilities may include lawful interception, charging, and some policy enforcement.

[0216] In at least one embodiment, the P-GW 2034 can terminate the SGi interface toward the PDN. In at least one embodiment, the P-GW 2034 can route data packets between the EPC network 2038 and an external network, such as a network including an application server 2040 (or application function (AF)), via an Internet Protocol (IP) interface 2042. In at least one embodiment, the application server 2040 can be an element that provides applications using IP bearer resources using a core network (e.g., a UMTS packet service (PS) domain, an LTE PS data service, etc.). In at least one embodiment, the P-GW 2034 is shown as being communicatively coupled to the application server 2040 via an IP communication interface 2042. In at least one embodiment, the application server 2040 can also be configured to support one or more communication services (e.g., voice over Internet Protocol (VoIP) sessions, PTT sessions, group communication sessions, social networking services, etc.) for UEs 2002 and 2004 via the CN 2038.

[0217] In at least one embodiment, P-GW 2034 can also be a node for policy enforcement and charging data collection. In at least one embodiment, Policy and Charging Enforcement Function (PCRF) 2036 is the policy and charging control element of CN 2038. In at least one embodiment, in a non-roaming scenario, a single PCRF can exist in the Home Public Land Mobile Network (HPLMN) associated with the UE's Internet Protocol Connectivity Access Network (IP-CAN) session. In at least one embodiment, in a roaming scenario with local traffic breakout, two PCRFs can exist associated with the UE's IP-CAN session: a Home PCRF (H-PCRF) in the HPLMN and a Visited PCRF (V-PCRF) in the Visited Public Land Mobile Network (VPLMN). In at least one embodiment, PCRF 2036 can be communicatively coupled to Application Server 2040 via P-GW 2034. In at least one embodiment, Application Server 2040 can signal PCRF 2036 to indicate a new service flow and select appropriate Quality of Service (QoS) and charging parameters. In at least one embodiment, PCRF 2036 can supply this rule to a Policy and Charging Enforcement Function (PCEF) (not shown) with the appropriate Traffic Flow Template (TFT) and QoS Class (QCI) identifier, which initiates the QoS and charging specified by the Application Server 2040.

[0218] Figure 21 The architecture of a system 2100 of a network according to some embodiments is shown. In at least one embodiment, the system 2100 is shown to include a UE 2102, a 5G access node or RAN node (shown as (R)AN node 2108), a user plane function (shown as UPF 2104), a data network (DN 2106), which in at least one embodiment can be an operator service, internet access, or a third-party service, and a 5G core network (5GC) (shown as CN 2110).

[0219] In at least one embodiment, CN 2110 includes an authentication server function (AUSF 2114); a core access and mobility management function (AMF 2112); a session management function (SMF 2118); a network exposure function (NEF 2116); a policy control function (PCF 2122); a network function (NF) repository function (NRF 2120); a unified data management (UDM 2124); and an application function (AF 2126). In at least one embodiment, CN 2110 may also include other elements not shown, such as a structured data storage network function (SDSF), an unstructured data storage network function (UDSF), and variations thereof.

[0220] In at least one embodiment, the UPF 2104 can serve as an anchor point for intra-RAT and inter-RAT mobility, an external PDU session point interconnected to the DN 2106, and a branching point supporting multi-homed PDU sessions. In at least one embodiment, the UPF 2104 can also perform packet routing and forwarding, packet inspection, user plane enforcement of policy rules, lawful interception of packets (UP collection), service usage reporting, QoS processing for the user plane (e.g., packet filtering, gating, UL / DL rate enforcement), uplink service validation (e.g., SDF to QoS flow mapping), transport-level packet marking in the uplink and downlink, downlink packet buffering, and downlink data notification triggering. In at least one embodiment, the UPF 2104 can include an uplink classifier to support routing of service flows to the data network. In at least one embodiment, the DN 2106 can represent various network operator services, internet access, or third-party services.

[0221] In at least one embodiment, the AUSF 2114 may store data used for authentication of the UE 2102 and handle authentication-related functions. In at least one embodiment, the AUSF 2114 may facilitate a common authentication framework for various access types.

[0222] In at least one embodiment, the AMF 2112 may be responsible for registration management (e.g., for registering UE 2102, etc.), connection management, reachability management, mobility management, and lawful interception of AMF-related events, as well as access authentication and authorization. In at least one embodiment, the AMF 2112 may provide transport of SM messages for the SMF 2118 and act as a transparent proxy for routing SM messages. In at least one embodiment, the AMF 2112 may also provide UE 2102 with an SMS function (SMSF) ( Figure 21 In at least one embodiment, the AMF 2112 may act as a Security Anchor Function (SEA), which may include interaction with the AUSF 2114 and the UE 2102 and receiving intermediate keys established as a result of the UE 2102 authentication process. In at least one embodiment, where USIM-based authentication is used, the AMF 2112 may retrieve security material from the AUSF 2114. In at least one embodiment, the AMF 2112 may also include a Security Context Management (SCM) function that receives keys from the SEA that it uses to derive access network-specific keys. In addition, in at least one embodiment, the AMF 2112 may be the termination point for the RAN CP interface (N2 reference point), the termination point for NAS (NI) signaling, and perform NAS encryption and integrity protection.

[0223] In at least one embodiment, the AMF 2112 may also support NAS signaling with the UE 2102 over the N3 Interworking Function (IWF) interface. In at least one embodiment, the N3 IWF may be used to provide access to untrusted entities. In at least one embodiment, the N3 IWF may be the termination point for the N2 and N3 interfaces for the control plane and user plane, respectively. Thus, it may handle N2 signaling from the SMF and AMF for PDU sessions and QoS, encapsulate / decapsulate packets for IPSec and N3 tunnels, mark N3 user plane packets in the uplink, and enforce QoS corresponding to the marking of N3 packets, taking into account the QoS requirements associated with such markings received over N2. In at least one embodiment, the N3 IWF may also relay uplink and downlink control plane NAS (NI) signaling between the UE 2102 and the AMF 2112, and relay uplink and downlink user plane packets between the UE 2102 and the UPF 2104. In at least one embodiment, the N3IWF also provides a mechanism for IPsec tunnel establishment with the UE 2102.

[0224] In at least one embodiment, the SMF 2118 may be responsible for session management (e.g., session establishment, modification, and release, including tunnel maintenance between the UPF and AN nodes); UE IP address allocation and management (including optional authorization); selection and control of UP functions; configuring traffic steering at the UPF to route traffic to the appropriate destination; interface termination towards the policy control function; policy enforcement and control portion of QoS; lawful interception (for SM events and interface to the LI system); termination of the SM portion of NAS messages; downlink data notification; originator of AN-specific SM information, which is sent to the AN via the AMF on N2; determining the SSC mode for the session. In at least one embodiment, the SMF 2118 may include the following roaming functions: handling local implementation to apply QoS SLAB (VPLMN); charging data collection and charging interface (VPLMN); lawful interception (for SM events in the VPLMN and interface to the LI system); supporting interaction with external DNs to transport signaling for PDU session authorization / authentication by the external DN.

[0225] In at least one embodiment, the NEF 2116 can provide a means for securely exposing services and capabilities provided by 3GPP network functions to third parties, internal exposure / re-exposure, application functions (e.g., AF 2126), edge computing or fog computing systems, and the like. In at least one embodiment, the NEF 2116 can authenticate, authorize, and / or throttle the AF. In at least one embodiment, the NEF 2116 can also convert information exchanged with the AF 2126 and information exchanged with internal network functions. In at least one embodiment, the NEF 2116 can convert between AF service identifiers and internal 5GC information. In at least one embodiment, the NEF 2116 can also receive information from other network functions (NFs) based on their exposed capabilities. In at least one embodiment, this information can be stored in the NEF 2116 as structured data or in a data storage NF using standardized interfaces. In at least one embodiment, the stored information can then be re-exposed by the NEF 2116 to other NFs and AFs and / or used for other purposes, such as analysis.

[0226] In at least one embodiment, the NRF 2120 may support service discovery functionality, receive NF discovery requests from NF instances, and provide information about the discovered NF instances to the NF instances. In at least one embodiment, the NRF 2120 also maintains information about available NF instances and the services they support.

[0227] In at least one embodiment, the PCF 2122 can provide policy rules to the control plane functions to implement them and can also support a unified policy framework to manage network behavior. In at least one embodiment, the PCF 2122 can also implement a front end (FE) for accessing subscription information related to policy decisions in the UDR of the UDM 2124.

[0228] In at least one embodiment, the UDM 2124 can process subscription-related information to support network entities handling communication sessions and can store subscription data for the UE 2102. In at least one embodiment, the UDM 2124 can include two components: an application FE and a user data repository (UDR). In at least one embodiment, the UDM can include a UDM FE, which is responsible for handling credentials, location management, subscription management, and the like. In at least one embodiment, several different front ends can serve the same user in different transactions. In at least one embodiment, the UDM-FE accesses the sub-subscription information stored in the UDR and performs authentication credential processing; user identity processing; access authorization; registration / mobility management; and subscription management. In at least one embodiment, the UDR can interact with the PCF 2122. In at least one embodiment, the UDM 2124 can also support SMS management, where the SMS-FE implements similar application logic as described above.

[0229] In at least one embodiment, the AF 2126 can provide application influence on service routing, access to the Network Capability Exposure (NCE), and interaction with the policy framework for policy control. In at least one embodiment, the NCE can be a mechanism that allows the 5GC and AF 2126 to provide information to each other via the NEF 2116, which can be used for edge computing implementations. In at least one embodiment, network operators and third-party services can be hosted near the UE 2102's attachment access point to achieve efficient service delivery with reduced end-to-end latency and load on the transport network. In at least one embodiment, for edge computing implementations, the 5GC can select a UPF 2104 close to the UE 2102 and perform service steering from the UPF 2104 to the DN 2106 via the N6 interface. In at least one embodiment, this can be based on UE subscription data, UE location, and information provided by the AF 2126. In at least one embodiment, the AF 2126 can influence UPF (re)selection and service routing. In at least one embodiment, based on operator deployment, the network operator may allow the AF 2126 to interact directly with the relevant NFs when the AF 2126 is considered a trusted entity.

[0230] In at least one embodiment, the CN 2110 may include an SMSF, which may be responsible for SMS subscription checking and verification, and relaying SM messages to / from the UE 2102 to / from other entities, such as SMS-GMSC / IWMSC / SMS routers. In at least one embodiment, the SMS may also interact with the AMF 2112 and the UDM 2124 for notification procedures that the UE 2102 is available for SMS delivery (e.g., setting a UE unreachable flag and notifying the UDM 2124 when the UE 2102 is available for SMS).

[0231] In at least one embodiment, the system 2100 may include the following service-based interfaces: Namf: a service-based interface exposed by AMF; Nsmf: a service-based interface exposed by SMF; Nnef: a service-based interface exposed by NEF; Npcf: a service-based interface exposed by PCF; Nudm: a service-based interface exposed by UDM; Naf: a service-based interface exposed by AF; Nnrf: a service-based interface exposed by NRF; and Nausf: a service-based interface exposed by AUSF.

[0232] In at least one embodiment, system 2100 may include the following reference points: N1: a reference point between the UE and the AMF; N2: a reference point between the (R)AN and the AMF; N3: a reference point between the (R)AN and the UPF; N4: a reference point between the SMF and the UPF; and N6: a reference point between the UPF and the data network. In at least one embodiment, there may be more reference points and / or service-based interfaces between NF services within the NF; however, these interfaces and reference points have been omitted for clarity. In at least one embodiment, the NS reference point may be between the PCF and the AF; the N7 reference point may be between the PCF and the SMF; the N11 reference point may be between the AMF and the SMF, and so on. In at least one embodiment, CN 2110 may include an Nx interface, which is an inter-CN interface between the MME and the AMF 2112 to enable interworking between CN 2110 and CN 7221.

[0233] In at least one embodiment, the system 2100 may include multiple RAN nodes (such as (R)AN nodes 2108), wherein an Xn interface is defined between two or more (R)AN nodes 2108 (e.g., gNBs) connected to the 5GC 410, between an (R)AN node 2108 (e.g., gNBs) and an eNB (e.g., macro RAN node) connected to the CN 2110, and / or between two eNBs connected to the CN 2110.

[0234] In at least one embodiment, the Xn interface may include an Xn user plane (Xn-U) interface and an Xn control plane (Xn-C) interface. In at least one embodiment, the Xn-U may provide non-guaranteed delivery of user plane PDUs and support / provide data forwarding and flow control functions. In at least one embodiment, the Xn-C may provide management and error handling functions, functions for managing the Xn-C interface, and mobility support for UE 2102 in connected mode (e.g., CM-CONNECTED), including functions for managing UE mobility in connected mode between one or more (R)AN nodes 2108. In at least one embodiment, mobility support may include context transfer from an old (source) serving (R)AN node 2108 to a new (target) serving (R)AN node 2108, and control of a user plane tunnel between the old (source) serving (R)AN node 2108 and the new (target) serving (R)AN node 2108.

[0235] In at least one embodiment, the Xn-U protocol stack can include a transport network layer built on top of an Internet Protocol (IP) transport layer and a GTP-U layer for carrying user plane PDUs on top of a UDP and / or one or more IP layers. In at least one embodiment, the Xn-C protocol stack can include an application layer signaling protocol, referred to as Xn Application Protocol (Xn-AP), and a transport network layer built on top of an SCTP layer. In at least one embodiment, the SCTP layer can be on top of an IP layer. In at least one embodiment, the SCTP layer provides a guaranteed delivery of application layer messages. In at least one embodiment, in the transport IP layer, point-to-point transmission is used to deliver signaling PDUs. In at least one embodiment, the Xn-U protocol stack and / or the Xn-C protocol stack can be the same as or similar to user plane and / or control plane protocol stacks shown and described herein.

[0236] Figure 22 is an illustration of a control plane protocol stack in accordance with some embodiments. In at least one embodiment, control plane 2200 is illustrated as a communication protocol stack between UE 2002 (or alternatively, UE 2004), RAN 2016, and MME 2028.

[0237] In at least one embodiment, PHY layer 2202 can transmit or receive information used by MAC layer 2204 over one or more air interfaces. In at least one embodiment, PHY layer 2202 can also perform link adaptation or adaptive modulation and coding (AMC), power control, cell search (e.g., for initial synchronization and handover purposes), and other measurements used by higher layers, such as RRC layer 2210. In at least one embodiment, PHY layer 2202 can further perform error detection on the transport channels, forward error correction (FEC) coding / decoding of the transport channels, modulation / demodulation of physical channels, interleaving, rate matching, mapping to physical channels, and Multiple Input Multiple Output (MIMO) antenna processing.

[0238] In at least one embodiment, MAC layer 2204 can perform mapping between logical channels and transport channels, multiplexing of MAC service data units (SDUs) from one or more logical channels into transport blocks (TB) to be delivered to PHY via transport channels, demultiplexing of MAC SDUs to one or more logical channels from TBs delivered via transport channels from PHY, multiplexing of MAC SDUs onto TBs, scheduling information reporting, error correction through hybrid automatic repeat request (HARQ), and logical channel prioritization.

[0239] In at least one embodiment, the RLC layer 2206 can operate in multiple operating modes, including transparent mode (TM), unacknowledged mode (UM), and acknowledged mode (AM). In at least one embodiment, the RLC layer 2206 can perform transmission of upper layer protocol data units (PDUs), error correction through automatic repeat request (ARQ) for AM data transmission, and concatenation, segmentation, and reassembly of RLC SDUs for UM and AM data transmission. In at least one embodiment, the RLC layer 2206 can also perform re-segmentation of RLC data PDUs for AM data transmission, reordering of RLC data PDUs for UM and AM data transmission, detection of duplicate data for UM and AM data transmission, discarding RLC SDUs for UM and AM data transmission, detecting protocol errors for AM data transmission, and performing RLC re-establishment.

[0240] In at least one embodiment, the PDCP layer 2208 can perform header compression and decompression of IP data, maintain PDCP sequence numbers (SNs), perform in-sequence delivery of higher layer PDUs when re-establishing lower layers, eliminate duplication of lower layer SDUs when re-establishing lower layers for radio bearers mapped on RLC AM, encrypt and decrypt control plane data, perform integrity protection and integrity verification on control plane data, control timer-based data discard, and perform security operations (e.g., encryption, decryption, integrity protection, integrity verification, etc.).

[0241] In at least one embodiment, the main services and functions of the RRC layer 2210 may include broadcasting of system information (e.g., included in a master information block (MIB) or system information block (SIB) related to the non-access stratum (NAS)), broadcasting of system information related to the access stratum (AS), paging, establishment, maintenance, and release of an RRC connection between a UE and an E-UTRAN (e.g., RRC connection paging, RRC connection establishment, RRC connection modification, and RRC connection release), establishment, configuration, maintenance, and release of point-to-point radio bearers, security functions including key management, inter-radio access technology (RAT) mobility, and measurement configuration for UE measurement reporting. In at least one embodiment, the MIB and SIB may include one or more information elements (IEs), each of which may include a separate data field or data structure.

[0242] In at least one embodiment, the UE 2002 and the RAN 2016 may utilize a Uu interface (e.g., an LTE-Uu interface) to exchange control plane data via a protocol stack including a PHY layer 2202, a MAC layer 2204, an RLC layer 2206, a PDCP layer 2208, and an RRC layer 2210.

[0243] In at least one embodiment, a non-access stratum (NAS) protocol (NAS protocol 2212) forms a highest stratum of the control plane between UE 2002 and MME 2028. In at least one embodiment, NAS protocol 2212 supports mobility of UE 2002 and session management procedures to establish and maintain IP connectivity between UE 2002 and P-GW 2034.

[0244] In at least one embodiment, an Si application protocol (Si-AP) layer (Si-AP layer 2222) can support functions of the Si interface and include elementary procedures (EPs). In at least one embodiment, an EP is a unit of interaction between RAN 2016 and CN 2028. In at least one embodiment, S1-AP layer services can include two groups: UE-associated services and non-UE-associated services. In at least one embodiment, these services perform functions including, but not limited to: E-UTRAN Radio Access Bearer (E-RAB) management, UE capability indication, mobility, NAS signaling transfer, RAN Information Management (RIM), and configuration transfer.

[0245] In at least one embodiment, a stream control transmission protocol (SCTP) layer (alternatively referred to as a stream control transmission protocol / internet protocol (SCTP / IP) layer) (SCTP layer 2220) can ensure reliable delivery of signaling messages between RAN 2016 and MME 2028 based, in part, on IP protocols supported by IP layer 2218. In at least one embodiment, L2 layer 2216 and L1 layer 2214 can refer to communication links (e.g., wired or wireless) used by RAN nodes and MMEs to exchange information.

[0246] In at least one embodiment, RAN 2016 and one or more MMEs 2028 can utilize an S1-MME interface to exchange control plane data via a protocol stack including L1 layer 2214, L2 layer 2216, IP layer 2218, SCTP layer 2220, and Si-AP layer 2222.

[0247] Figure 23 is a diagram of a user plane protocol stack, in accordance with at least one embodiment. In at least one embodiment, user plane 2300 is shown as a communication protocol stack between UE 2002, RAN 2016, S-GW 2030, and P-GW 2034. In at least one embodiment, user plane 2300 can utilize the same protocol layers as control plane 2200. In at least one embodiment, UE 2002 and RAN 2016 can utilize a Uu interface (e.g., an LTE-Uu interface) to exchange user plane data via a protocol stack including PHY layer 2202, MAC layer 2204, RLC layer 2206, PDCP layer 2208.

[0248] In at least one embodiment, the General Packet Radio Service (GPRS) Tunneling Protocol (GTP-U) layer for the user plane (GTP-U layer 2304) can be used to carry user data within the GPRS core network and between the radio access network and the core network. In at least one embodiment, the transmitted user data can be packets in any format of IPv4, IPv6 or PPP format. In at least one embodiment, the UDP and IP security (UDP / IP) layer (UDP / IP layer 2302) can provide a checksum for data integrity, port numbers for addressing different functions at the source and destination, and encryption and authentication of selected data streams. In at least one embodiment, the RAN 2016 and the S-GW 2030 can utilize the S1-U interface to exchange user plane data via a protocol stack including the L1 layer 2214, the L2 layer 2216, the UDP / IP layer 2302 and the GTP-U layer 2304. In at least one embodiment, the S-GW 2030 and the P-GW 2034 may utilize an S5 / S8a interface to exchange user plane data via a protocol stack including an L1 layer 2214, an L2 layer 2216, a UDP / IP layer 2302, and a GTP-U layer 2304. In at least one embodiment, as described above with respect to Figure 22 As discussed, the NAS protocol supports the mobility of UE 2002 and session management procedures to establish and maintain IP connectivity between UE 2002 and P-GW 2034 .

[0249] Figure 24 Components 2400 of a core network according to at least one embodiment are shown. In at least one embodiment, the components of CN 2038 can be implemented in one physical node or in separate physical nodes, the separate physical nodes including components for reading and executing instructions from a machine-readable medium or computer-readable medium (e.g., a non-transitory machine-readable storage medium). In at least one embodiment, network function virtualization (NFV) is used to virtualize any or all of the above-described network node functions via executable instructions stored in one or more computer-readable storage media (described in further detail below). In at least one embodiment, a logical instantiation of CN 2038 can be referred to as a network slice 2402 (e.g., network slice 2402 is shown as including HSS 2032, MME 2028, and S-GW 2030). In at least one embodiment, a logical instantiation of a portion of CN 2038 can be referred to as a network sub-slice 2404 (e.g., network sub-slice 2404 is shown as including P-GW 2034 and PCRF 2036).

[0250] In at least one embodiment, the NFV architecture and infrastructure can be used to virtualize one or more network functions onto physical resources including a combination of industry-standard server hardware, storage hardware, or switches, which may alternatively be performed by dedicated hardware. In at least one embodiment, the NFV system can be used to perform a virtual or reconfigurable implementation of one or more EPC components / functions.

[0251] Figure 25 2 is a block diagram illustrating components of a system 2500 for supporting network function virtualization (NFV) according to at least one embodiment. In at least one embodiment, system 2500 is shown as including a virtualization infrastructure manager (shown as VIM 2502), a network function virtualization infrastructure (shown as NFVI 2504), a VNF manager (shown as VNFM 2506), virtualized network functions (shown as VNF 2508), an element manager (shown as EM 2510), an NFV orchestrator (shown as NFVO 2512), and a network manager (shown as NM 2514).

[0252] In at least one embodiment, the VIM 2502 manages resources of the NFVI 2504. In at least one embodiment, the NFVI 2504 may include physical or virtual resources and applications (including a hypervisor) used to execute the system 2500. In at least one embodiment, the VIM 2502 may utilize the NFVI 2504 to manage the lifecycle of virtual resources (e.g., the creation, maintenance, and teardown of virtual machines (VMs) associated with one or more physical resources), track VM instances, track performance, faults, and security of VM instances and associated physical resources, and expose VM instances and associated physical resources to other management systems.

[0253] In at least one embodiment, VNFM 2506 can manage VNF 2508. In at least one embodiment, VNF 2508 can be used to perform EPC components / functions. In at least one embodiment, VNFM 2506 can manage the lifecycle of VNF 2508 and track the performance, faults, and security of the virtual aspects of VNF 2508. In at least one embodiment, EM 2510 can track the performance, faults, and security of the functional aspects of VNF 2508. In at least one embodiment, the data tracked from VNFM 2506 and EM 2510 can include, in at least one embodiment, performance measurement (PM) data used by VIM 2502 or NFVI 2504. In at least one embodiment, both VNFM 2506 and EM 2510 can scale up / down the number of VNFs in system 2500.

[0254] In at least one embodiment, the NFVO 2512 can coordinate, authorize, release, and occupy resources of the NFVI 2504 in order to provide the requested service (e.g., to execute an EPC function, component, or slice). In at least one embodiment, the NM 2514 can provide an end-user function package responsible for managing the network, which can include network elements with VNFs, non-virtualized network functions, or both (management of the VNFs can occur via the EM 2510).

[0255] Computer-based systems

[0256] The following figures set forth, but are not limiting of, exemplary computer-based systems that can be used to implement at least one embodiment.

[0257] Figure 26 A processing system 2600 is shown in accordance with at least one embodiment. In at least one embodiment, system 2600 includes one or more processors 2602 and one or more graphics processors 2608, and can be a single-processor desktop system, a multi-processor workstation system, or a server system having a large number of processors 2602 or processor cores 2607. In at least one embodiment, processing system 2600 is a processing platform incorporated within a system-on-chip (SoC) integrated circuit for use in mobile, handheld, or embedded devices.

[0258] In at least one embodiment, the processing system 2600 may include or be incorporated into a server-based gaming platform, including a gaming console, a mobile gaming console, a handheld gaming console, or an online gaming console, including a gaming and media console. In at least one embodiment, the processing system 2600 is a mobile phone, a smart phone, a tablet computing device, or a mobile internet device. In at least one embodiment, the processing system 2600 may also include a device coupled to or integrated into a wearable device, such as a smartwatch wearable device, a smart glasses device, an augmented reality device, or a virtual reality device. In at least one embodiment, the processing system 2600 is a television or set-top box device having one or more processors 2602 and a graphical interface generated by one or more graphics processors 2608.

[0259] In at least one embodiment, one or more processors 2602 each include one or more processor cores 2607 to process instructions that, when executed, perform operations for system and user software. In at least one embodiment, each of the one or more processor cores 2607 is configured to process a specific instruction set 2609. In at least one embodiment, the instruction set 2609 can facilitate complex instruction set computing (CISC), reduced instruction set computing (RISC), or computing via very long instruction words (VLIW). In at least one embodiment, multiple processor cores 2607 can each process a different instruction set 2609, which can include instructions that facilitate emulating other instruction sets. In at least one embodiment, the processor cores 2607 can also include other processing devices, such as a digital signal processor (DSP).

[0260] In at least one embodiment, the processor 2602 includes a cache memory (cache) 2604. In at least one embodiment, the processor 2602 can have a single internal cache or multiple levels of internal cache. In at least one embodiment, the cache memory is shared among various components of the processor 2602. In at least one embodiment, the processor 2602 also uses an external cache (e.g., a level 3 (L3) cache or a last level cache (LLC)) (not shown), which can share this logic among the processor cores 2607 using known cache coherence techniques. In at least one embodiment, the processor 2602 further includes a register file 2606. The processor 2602 may include different types of registers (e.g., integer registers, floating point registers, status registers, and an instruction pointer register) for storing different types of data. In at least one embodiment, the register file 2606 may include general purpose registers or other registers.

[0261] In at least one embodiment, one or more processors 2602 are coupled to one or more interface buses 2610 to transmit communication signals, such as address, data, or control signals, between the processors 2602 and other components in the system 2600. In at least one embodiment, the interface bus 2610 can be a processor bus, such as a version of a Direct Media Interface (DMI) bus. In at least one embodiment, the interface bus 2610 is not limited to a DMI bus and can include one or more peripheral component interconnect buses (e.g., PCI, PCI Express), a memory bus, or other types of interface buses. In at least one embodiment, the processor 2602 includes an integrated memory controller 2616 and a platform controller hub 2630. In at least one embodiment, the memory controller 2616 facilitates communication between storage devices and other components of the processing system 2600, while the platform controller hub (PCH) 2630 provides connections to input / output (I / O) devices via a local I / O bus.

[0262] In at least one embodiment, the memory device 2620 can be a dynamic random access memory (DRAM) device, a static random access memory (SRAM) device, a flash memory device, a phase change memory device, or a device having suitable performance for use as processor memory. In at least one embodiment, the memory device 2620 can be used as system memory for the processing system 2600 to store data 2622 and instructions 2621 for use when one or more processors 2602 execute applications or processes. In at least one embodiment, the memory controller 2616 is also coupled to an optional external graphics processor 2612, which can communicate with one or more graphics processors 2608 in the processor 2602 to perform graphics and media operations. In at least one embodiment, a display device 2611 can be connected to the processor 2602. In at least one embodiment, the display device 2611 can include one or more internal display devices, such as in a mobile electronic device or portable computer device, or an external display device connected via a display interface (such as a DisplayPort). In at least one embodiment, the display device 2611 may include a head-mounted display (HMD), such as a stereoscopic display device used in virtual reality (VR) applications or augmented reality (AR) applications.

[0263] In at least one embodiment, the platform controller hub 2630 enables peripheral devices to connect to the storage device 2620 and the processor 2602 via a high-speed I / O bus. In at least one embodiment, the I / O peripherals include, but are not limited to, an audio controller 2646, a network controller 2634, a firmware interface 2628, a wireless transceiver 2626, a touch sensor 2625, and a data storage device 2624 (e.g., a hard drive, flash memory, etc.). In at least one embodiment, the data storage device 2624 can be connected via a memory interface (e.g., SATA) or via a peripheral bus, such as a peripheral component interconnect bus (e.g., PCI, PCIe). In at least one embodiment, the touch sensor 2625 can include a touch screen sensor, a pressure sensor, or a fingerprint sensor. In at least one embodiment, the wireless transceiver 2626 can be a Wi-Fi transceiver, a Bluetooth transceiver, or a mobile network transceiver, such as a 3G, 4G, or Long Term Evolution (LTE) transceiver. In at least one embodiment, the firmware interface 2628 enables communication with the system firmware and, in at least one embodiment, can be a unified extensible firmware interface (UEFI). In at least one embodiment, a network controller 2634 can enable network connectivity to a wired network. In at least one embodiment, a high-performance network controller (not shown) is coupled to the interface bus 2610. In at least one embodiment, the audio controller 2646 is a multi-channel high-definition audio controller. In at least one embodiment, the processing system 2600 includes an optional legacy I / O controller 2640 for coupling legacy (e.g., Personal System 2 (PS / 2)) devices to the processing system 2600. In at least one embodiment, the platform controller hub 2630 can also be connected to one or more universal serial bus (USB) controllers 2642 that connect input devices such as a keyboard and mouse 2643 combination, a camera 2644, or other USB input devices.

[0264] In at least one embodiment, instances of the memory controller 2616 and the platform controller hub 2630 may be integrated into a discrete external graphics processor, such as the external graphics processor 2612. In at least one embodiment, the platform controller hub 2630 and / or the memory controller 2616 may be external to one or more processors 2602. In at least one embodiment, the processing system 2600 may include the external memory controller 2616 and the platform controller hub 2630, which may be configured as a memory controller hub and a peripheral controller hub in a system chipset that communicates with the processor 2602.

[0265] Figure 27A computer system 2700 is shown in accordance with at least one embodiment. In at least one embodiment, the computer system 2700 can be a system of interconnected devices and components, a SOC, or some combination thereof. In at least one embodiment, the computer system 2700 is formed by a processor 2702, which can include an execution unit for executing instructions. In at least one embodiment, the computer system 2700 can include, but is not limited to, components such as the processor 2702, which employs an execution unit including logic to execute algorithms for processing data. In at least one embodiment, the computer system 2700 can include a processor such as the Intel® processor available from Intel Corporation of Santa Clara, California. Processor family, XeonTM, XScaleTM and / or StrongARMTM, Core TM or Nervana TM microprocessor, although other systems (including PCs with other microprocessors, engineering workstations, set-top boxes, etc.) may also be used. In at least one embodiment, the computer system 2700 may execute a version of the WINDOWS operating system available from Microsoft Corporation of Redmond, Wash., although other operating systems (UNIX and Linux in at least one embodiment), embedded software, and / or graphical user interfaces may also be used.

[0266] In at least one embodiment, the computer system 2700 can be used in other devices, such as handheld devices and embedded applications. Some of at least one embodiment of a handheld device include cellular phones, Internet Protocol (IP) devices, digital cameras, personal digital assistants ("PDAs"), and handheld PCs. In at least one embodiment, the embedded application can include a microcontroller, a digital signal processor ("DSP"), a SoC, a network computer ("NetPC"), a set-top box, a network hub, a wide area network ("WAN") switch, or any other system that can execute one or more instructions according to at least one embodiment.

[0267] In at least one embodiment, computer system 2700 may include, but is not limited to, a processor 2702, which may include, but is not limited to, one or more execution units 2708, which may be configured to execute Compute Unified Device Architecture ("CUDA") ( Developed by NVIDIA Corporation of Santa Clara, California) program. In at least one embodiment, a CUDA program is at least a portion of a software application written in the CUDA programming language. In at least one embodiment, computer system 2700 is a single-processor desktop or server system. In at least one embodiment, computer system 2700 may be a multi-processor system. In at least one embodiment, processor 2702 may include, but is not limited to, a CISC microprocessor, a RISC microprocessor, a VLIW microprocessor, a processor implementing a combination of instruction sets, or any other processor device, such as a digital signal processor in at least one embodiment. In at least one embodiment, processor 2702 may be coupled to a processor bus 2710 that may transmit data signals between processor 2702 and other components in computer system 2700.

[0268] In at least one embodiment, the processor 2702 may include, but is not limited to, a level 1 ("L1") internal cache memory ("cache") 2704. In at least one embodiment, the processor 2702 may have a single internal cache or multiple levels of internal cache. In at least one embodiment, the cache memory may reside external to the processor 2702. In at least one embodiment, the processor 2702 may include a combination of internal and external caches. In at least one embodiment, the register file 2706 may store different types of data in various registers, including, but not limited to, integer registers, floating point registers, status registers, and an instruction pointer register.

[0269] In at least one embodiment, an execution unit 2708, including but not limited to logic for performing integer and floating-point operations, is also located in the processor 2702. The processor 2702 may also include a microcode ("ucode") read-only memory ("ROM") for storing microcode for certain macroinstructions. In at least one embodiment, the execution unit 2708 may include logic for processing a packed instruction set 2709. In at least one embodiment, by including the packed instruction set 2709 in the instruction set of the general-purpose processor 2702, along with associated circuitry to execute the instructions, operations used by many multimedia applications may be performed using packed data in the general-purpose processor 2702. In at least one embodiment, many multimedia applications may be executed faster and more efficiently by using the full width of the processor's data bus to perform operations on the packed data, which may eliminate the need to transfer smaller units of data across the processor's data bus to perform one or more operations on one data element at a time.

[0270] In at least one embodiment, execution unit 2708 can also be used in microcontrollers, embedded processors, graphics devices, DSPs, and other types of logic circuits. In at least one embodiment, computer system 2700 can include, but not limited to, memory 2720. In at least one embodiment, memory 2720 can be implemented as a DRAM device, SRAM device, flash memory device, or other memory device. Memory 2720 can store instructions 2719 and / or data 2721 that can be executed by processor 2702 as signaled by data signals.

[0271] In at least one embodiment, system logic chip can be coupled to processor bus 2710 and memory 2720. In at least one embodiment, system logic chip can include, without limitation, a memory controller hub (“MCH”) 2716, and processor 2702 can communicate with MCH 2716 via processor bus 2710. In at least one embodiment, MCH 2716 can provide a high bandwidth memory path 2718 to memory 2720 for instruction and data storage and for storage of graphics commands, data, and textures. In at least one embodiment, MCH 2716 can direct data signals between processor 2702, memory 2720, and other components in computer system 2700, and can

[0272] In at least one embodiment, the computer system 2700 may use the system I / O 2722 as a proprietary hub interface bus to couple the MCH 2716 to the I / O controller hub ("ICH") 2730. In at least one embodiment, the ICH 2730 may provide direct connectivity to certain I / O devices via a local I / O bus. In at least one embodiment, the local I / O bus may include, but is not limited to, a high-speed I / O bus used to connect peripheral devices to the memory 2720, chipset, and processor 2702. Examples may include, but are not limited to, an audio controller 2729, a firmware hub ("Flash BIOS") 2728, a wireless transceiver 2726, a data store 2724, a traditional I / O controller 2723 including user input 2725 and a keyboard interface, a serial expansion port 2777 (e.g., USB), and a network controller 2734. The data store 2724 may include a hard drive, a floppy disk drive, a CD-ROM device, a flash memory device, or other mass storage device.

[0273] In at least one embodiment, Figure 27 A system comprising interconnected hardware devices or "chips" is shown. In at least one embodiment, Figure 27 An exemplary SoC may be shown. In at least one embodiment, Figure 27 The devices shown in can be interconnected with a proprietary interconnect, a standardized interconnect (e.g., PCIe), or some combination thereof. In at least one embodiment, one or more components of system 2700 are interconnected using a Compute Express Link (CXL) interconnect.

[0274] Figure 28 A system 2800 is shown in accordance with at least one embodiment. In at least one embodiment, the system 2800 is an electronic device that utilizes a processor 2810. In at least one embodiment, the system 2800 can be, in at least one embodiment but not limited to, a notebook computer, a tower server, a rack server, a blade server, a laptop computer, a desktop computer, a tablet computer, a mobile device, a phone, an embedded computer, or any other suitable electronic device.

[0275] In at least one embodiment, system 2800 may include, but is not limited to, a processor 2810 communicatively coupled to any suitable number or kind of components, peripherals, modules, or devices. In at least one embodiment, processor 2810 is coupled using a bus or interface, such as an I 2C bus, System Management Bus ("SMBus"), Low Pin Count (LPC) bus, Serial Peripheral Interface ("SPI"), High Definition Audio ("HDA") bus, Serial Advanced Technology Attachment ("SATA") bus, USB (Revisions 1, 2, 3), or Universal Asynchronous Receiver / Transmitter ("UART") bus. In at least one embodiment, Figure 28 A system is shown that includes interconnected hardware devices or "chips". In at least one embodiment, Figure 28 An exemplary SoC may be shown. In at least one embodiment, Figure 28 The devices shown in can be interconnected with proprietary interconnects, standardized interconnects (e.g., PCIe), or some combination thereof. In at least one embodiment, Figure 28 One or more components of the system are interconnected using Compute Express Link (CXL) interconnect lines.

[0276] In at least one embodiment, Figure 28 The system may include a display 2824, a touch screen 2825, a touchpad 2830, a near field communication unit ("NFC") 2845, a sensor hub 2840, a thermal sensor 2846, an express chipset ("EC") 2835, a trusted platform module ("TPM") 2838, a BIOS / firmware / flash memory ("BIOS, FW Flash") 2822, a DSP 2860, a solid-state disk ("SSD") or a hard disk drive ("HDD") 2820, a wireless local area network unit ("WLAN") 2850, a Bluetooth unit 2852, a wireless wide area network unit ("WWAN") 2856, a global positioning system (GPS) 2855, a camera ("USB 3.0 camera") 2854 (e.g., a USB 3.0 camera), or a low-power double data rate ("LPDDR") memory unit ("LPDDR3") 2815 implemented in at least one embodiment of the LPDDR3 standard. Each of these components may be implemented in any suitable manner.

[0277] In at least one embodiment, other components may be communicatively coupled to processor 2810 via the components discussed above. In at least one embodiment, an accelerometer 2841, an ambient light sensor (“ALS”) 2842, a compass 2843, and a gyroscope 2844 may be communicatively coupled to sensor hub 2840. In at least one embodiment, a thermal sensor 2839, a fan 2837, a keyboard 2846, and a touchpad 2830 may be communicatively coupled to EC 2835. In at least one embodiment, a speaker 2863, an earpiece 2864, and a microphone (“mic”) 2865 may be communicatively coupled to an audio unit (“audio codec and class-D amplifier”) 2864, which in turn may be communicatively coupled to DSP 2860. In at least one embodiment, audio unit 2864 may include, but is not limited to, an audio codec / decoder (“codec”) and a class-D amplifier. In at least one embodiment, a SIM card (“SIM”) 2857 may be communicatively coupled to WWAN unit 2856. In at least one embodiment, components such as the WLAN unit 2850 and the Bluetooth unit 2852 and the WWAN unit 2856 may be implemented as a next generation form factor (NGFF).

[0278] Figure 29 An exemplary integrated circuit 2900 is shown in accordance with at least one embodiment. In at least one embodiment, the exemplary integrated circuit 2900 is a SoC, which can be manufactured using one or more IP cores. In at least one embodiment, the integrated circuit 2900 includes one or more application processors 2905 (e.g., CPUs), at least one graphics processor 2910, and may additionally include an image processor 2915 and / or a video processor 2920, any of which may be modular IP cores. In at least one embodiment, the integrated circuit 2900 includes peripheral or bus logic including a USB controller 2925, a UART controller 2930, an SPI / SDIO controller 2935, and an I / O controller. 2 S / I 2 C controller 2940. In at least one embodiment, the integrated circuit 2900 may include a display device 2945 coupled to one or more of a High Definition Multimedia Interface (HDMI) controller 2950 and a Mobile Industry Processor Interface (MIPI) display interface 2955. In at least one embodiment, storage may be provided by a flash memory subsystem 2960, including flash memory and a flash memory controller. In at least one embodiment, a memory interface may be provided via a memory controller 2965 for accessing SDRAM or SRAM memory devices. In at least one embodiment, some integrated circuits also include an embedded security engine 2970.

[0279] Figure 30A computing system 3000 is shown in accordance with at least one embodiment. In at least one embodiment, computing system 3000 includes a processing subsystem 3001 having one or more processors 3002 and system memory 3004 communicating via an interconnect path that may include a memory hub 3005. In at least one embodiment, memory hub 3005 may be a separate component within a chipset assembly or integrated within one or more processors 3002. In at least one embodiment, memory hub 3005 is coupled to an I / O subsystem 3011 via a communication link 3006. In at least one embodiment, I / O subsystem 3011 includes an I / O hub 3007, which enables computing system 3000 to receive input from one or more input devices 3008. In at least one embodiment, I / O hub 3007 may enable a display controller, included in one or more processors 3002, to provide output to one or more display devices 3010A. In at least one embodiment, the one or more display devices 3010A coupled to the I / O hub 3007 may include local, internal, or embedded display devices.

[0280] In at least one embodiment, the processing subsystem 3001 includes one or more parallel processors 3012 coupled to a memory hub 3005 via a bus or other communication link 3013. In at least one embodiment, the communication link 3013 can be one of many standard-based communication link technologies or protocols, such as, but not limited to, PCIe, or can be a vendor-specific communication interface or communication structure. In at least one embodiment, the one or more parallel processors 3012 form a computationally focused parallel or vector processing system that can include a large number of processing cores and / or processing clusters, such as a multi-integrated core (MIC) processor. In at least one embodiment, the one or more parallel processors 3012 form a graphics processing subsystem that can output pixels to one of one or more display devices 3010A coupled via an I / O hub 3007. In at least one embodiment, the one or more parallel processors 3012 can also include a display controller and display interface (not shown) to enable direct connection to one or more display devices 3010B.

[0281] In at least one embodiment, a system storage unit 3014 can be connected to the I / O hub 3007 to provide a storage mechanism for the computing system 3000. In at least one embodiment, an I / O switch 3016 can be used to provide an interface mechanism to enable connections between the I / O hub 3007 and other components, such as a network adapter 3018 and / or a wireless network adapter 3019 that can be integrated into the platform, as well as various other devices that can be added via one or more add-on devices 3020. In at least one embodiment, the network adapter 3018 can be an Ethernet adapter or another wired network adapter. In at least one embodiment, the wireless network adapter 3019 can include one or more of Wi-Fi, Bluetooth, NFC, or other network devices including one or more radios.

[0282] In at least one embodiment, computing system 3000 may include other components not explicitly shown, including USB or other port connections, optical storage drives, video capture devices, and / or variations thereof, which may also be connected to I / O hub 3007. Figure 30 The communication paths that interconnect the various components in the system can be implemented using any suitable protocol, such as a PCI (Peripheral Component Interconnect)-based protocol (e.g., PCIe), or other bus or point-to-point communication interfaces and / or protocols (e.g., NVLink high-speed interconnect or interconnect protocol).

[0283] In at least one embodiment, one or more parallel processors 3012 include circuitry optimized for graphics and video processing (including video output circuitry in at least one embodiment) and constitute a graphics processing unit (GPU). In at least one embodiment, one or more parallel processors 3012 include circuitry optimized for general-purpose processing. In at least one embodiment, the components of computing system 3000 can be integrated with one or more other system elements on a single integrated circuit. In at least one embodiment, one or more parallel processors 3012, memory hub 3005, processor 3002, and I / O hub 3007 can be integrated into a system-on-chip (SoC) integrated circuit. In at least one embodiment, the components of computing system 3000 can be integrated into a single package to form a system-in-package (SIP) configuration. In at least one embodiment, at least a portion of the components of computing system 3000 can be integrated into a multi-chip module (MCM), which can be interconnected with other multi-chip modules to form a modular computing system. In at least one embodiment, I / O subsystem 3011 and display device 3010B are omitted from computing system 3000.

[0284] Processing system

[0285] The following figures illustrate, but are not limited to, exemplary processing systems that can be used to implement at least one embodiment.

[0286] Figure 31 An accelerated processing unit ("APU") 3100 is shown in accordance with at least one embodiment. In at least one embodiment, the APU 3100 was developed by Advanced Micro Devices, Inc. of Santa Clara, California. In at least one embodiment, the APU 3100 can be configured to execute applications, such as CUDA programs. In at least one embodiment, the APU 3100 includes, but is not limited to, a core complex 3110, a graphics complex 3140, a fabric 3160, an I / O interface 3170, a memory controller 3180, a display controller 3192, and a multimedia engine 3194. In at least one embodiment, the APU 3100 can include, but is not limited to, any combination of any number of core complexes 3110, any number of graphics complexes 3140, any number of display controllers 3192, and any number of multimedia engines 3194. For purposes of illustration, multiple instances of similar objects are referred to herein by reference numerals, where the reference numeral identifies the object and a number in parentheses identifies the desired instance.

[0287] In at least one embodiment, core complex 3110 is a CPU, graphics complex 3140 is a GPU, and APU 3100 is a processing unit that is not limited to integrating 3110 and 3140 onto a single chip. In at least one embodiment, some tasks may be assigned to core complex 3110, while other tasks may be assigned to graphics complex 3140. In at least one embodiment, core complex 3110 is configured to execute primary control software associated with APU 3100, such as an operating system. In at least one embodiment, core complex 3110 is the main processor of APU 3100, controlling and coordinating the operations of the other processors. In at least one embodiment, core complex 3110 issues commands that control the operations of graphics complex 3140. In at least one embodiment, core complex 3110 may be configured to execute host executable code derived from CUDA source code, and graphics complex 3140 may be configured to execute device executable code derived from CUDA source code.

[0288] In at least one embodiment, core complex 3110 includes, but is not limited to, cores 3120(1)-3120(4) and L3 cache 3130. In at least one embodiment, core complex 3110 may include, but is not limited to, any number of cores 3120 and any combination of any number and type of caches. In at least one embodiment, cores 3120 are configured to execute instructions of a particular instruction set architecture ("ISA"). In at least one embodiment, each core 3120 is a CPU core.

[0289] In at least one embodiment, each core 3120 includes, without limitation, a fetch / decode unit 3122, an integer execution engine 3124, a floating point execution engine 3126, and an L2 cache 3128. In at least one embodiment, fetch / decode unit 3122 fetches instructions, decodes such instructions, generates micro-operations, and dispatches individual micro-instructions to integer execution engine 3124 and floating point execution engine 3126. In at least one embodiment, fetch / decode unit 3122 can concurrently dispatch one micro-instruction to integer execution engine 3124 and another micro-instruction to floating point execution engine 3126. In at least one embodiment, integer execution engine 3124 executes, without limitation, integer and memory operations. In at least one embodiment, floating point engine 3126 executes, without limitation, floating point and vector operations. In at least one embodiment, fetch-decode unit 3122 dispatches micro-instructions to a single execution engine in place of both integer execution engine 3124 and floating point execution engine 3126.

[0290] In at least one embodiment, each core 3120(i) has access to an L2 cache 3128(i) included in core 3120(i), where i is an integer representing a particular instance of core 3120. In at least one embodiment, each core 3120 included in core complex 3110(j) is connected to other cores 3120 included in core complex 3110(j) via an L3 cache 3130(j) included in core complex 3110(j), where j is an integer representing a particular instance of core complex 3110. In at least one embodiment, cores 3120 included in core complex 3110(j) have access to all L3 caches 3130(j) included in core complex 3110(j), where j is an integer representing a particular instance of core complex 3110. In at least one embodiment, L3 cache 3130 can include, without limitation, any number of slices.

[0291] In at least one embodiment, graphics complex 3140 can be configured to perform compute operations in a highly parallel manner. In at least one embodiment, graphics complex 3140 is configured to perform graphics pipeline operations such as draw commands, pixel operations, geometric calculations, and other operations associated with rendering images to a display. In at least one embodiment, graphics complex 3140 is configured to perform operations that are not graphics related. In at least one embodiment, graphics complex 3140 is configured to perform graphics related operations and operations that are not graphics related.

[0292] In at least one embodiment, graphics complex 3140 includes, but is not limited to, any number of compute units 3150 and L2 cache 3142. In at least one embodiment, compute units 3150 share L2 cache 3142. In at least one embodiment, L2 cache 3142 is partitioned. In at least one embodiment, graphics complex 3140 includes, but is not limited to, any number of compute units 3150 and any number (including zero) and type of cache. In at least one embodiment, graphics complex 3140 includes, but is not limited to, any amount of dedicated graphics hardware.

[0293] In at least one embodiment, each compute unit 3150 includes, but is not limited to, any number of SIMD units 3152 and shared memory 3154. In at least one embodiment, each SIMD unit 3152 implements a SIMD architecture and is configured to execute operations in parallel. In at least one embodiment, each compute unit 3150 can execute any number of thread blocks, but each thread block executes on a single compute unit 3150. In at least one embodiment, a thread block includes, but is not limited to, any number of execution threads. In at least one embodiment, a workgroup is a thread block. In at least one embodiment, each SIMD unit 3152 executes a different warp. In at least one embodiment, a warp is a group of threads (e.g., 16 threads), where each thread in a warp belongs to a single thread block and is configured to process different data sets based on a single instruction set. In at least one embodiment, predication can be used to disable one or more threads in a warp. In at least one embodiment, a channel is a thread. In at least one embodiment, a work item is a thread. In at least one embodiment, a wavefront is a warp. In at least one embodiment, different wavefronts in a thread block can be synchronized and communicated via shared memory 3154.

[0294] In at least one embodiment, fabric 3160 is a system interconnect that facilitates data and control transmissions across core complex 3110, graphics complex 3140, I / O interface 3170, memory controllers 3180, display controller 3192, and multimedia engine 3194. In at least one embodiment, APU 3100 can include, without limitation, any number and type of system interconnects in addition to or instead of fabric 3160 that facilitate data and control transmissions across any number and type of directly or indirectly linked components that can be internal or external to APU 3100. In at least one embodiment, I / O interface 3170 represents any number and type of I / O interface (e.g., PCI, PCI-Extended (“PCI-X”), PCIe, Gigabit Ethernet (“GBE”), USB, etc.). In at least one embodiment, various types of peripheral devices are coupled to I / O interface 3170. In at least one embodiment, peripheral devices coupled to I / O interface 3170 can include, without limitation, a keyboard, a mouse, a printer, a scanner, a joystick or other type of game controller, a media recording device, an external storage device, a network interface card, etc.

[0295] In at least one embodiment, display controller 3192 displays images on one or more display devices, such as liquid crystal display (“LCD”) devices. In at least one embodiment, multimedia engine 3194 includes, without limitation, any number and type of multimedia-related circuitry, such as a video decoder, a video encoder, an image signal processor, etc. In at least one embodiment, memory controllers 3180 facilitate data transfers between APU 3100 and unified system memory 3190. In at least one embodiment, core complex 3110 and graphics complex 3140 share unified system memory 3190.

[0296] In at least one embodiment, APU 3100 implements a memory subsystem that includes, without limitation, any number and type of memory controllers 3180 and memory devices (e.g., shared memory 3154) that can be dedicated to one component or shared among multiple components. In at least one embodiment, APU 3100 implements a cache subsystem that includes, without limitation, one or more cache memories (e.g., L2 cache 2728, L3 cache 3130, and L2 cache 3142), each of which can be private to a component or shared among any number of components (e.g., core 3120, core complex 3110, SIMD unit 3152, compute unit 3150, and graphics complex 3140).

[0297] Figure 32A CPU 3200 is shown according to at least one embodiment. In at least one embodiment, the CPU 3200 is developed by Advanced Micro Devices, Inc. of Santa Clara, California. In at least one embodiment, the CPU 3200 can be configured to execute application programs. In at least one embodiment, the CPU 3200 is configured to execute host control software, such as an operating system. In at least one embodiment, the CPU 3200 issues commands to control the operation of an external GPU (not shown). In at least one embodiment, the CPU 3200 can be configured to execute host executable code derived from CUDA source code, and the external GPU can be configured to execute device executable code derived from such CUDA source code. In at least one embodiment, the CPU 3200 includes, but is not limited to, any number of core complexes 3210, fabric 3260, I / O interfaces 3270, and memory controller 3280.

[0298] In at least one embodiment, core complex 3210 includes, but is not limited to, cores 3220(1)-3220(4) and L3 cache 3230. In at least one embodiment, core complex 3210 may include, but is not limited to, any number of cores 3220 and any combination of any number and type of caches. In at least one embodiment, cores 3220 are configured to execute instructions of a specific ISA. In at least one embodiment, each core 3220 is a CPU core.

[0299] In at least one embodiment, each core 3220 includes, but is not limited to, a fetch / decode unit 3222, an integer execution engine 3224, a floating-point execution engine 3226, and an L2 cache 3228. In at least one embodiment, the fetch / decode unit 3222 fetches instructions, decodes these instructions, generates micro-ops, and dispatches individual micro-ops to the integer execution engine 3224 and the floating-point execution engine 3226. In at least one embodiment, the fetch / decode unit 3222 can simultaneously dispatch one micro-op to the integer execution engine 3224 and another micro-op to the floating-point execution engine 3226. In at least one embodiment, the integer execution engine 3224 performs, but is not limited to, integer and memory operations. In at least one embodiment, the floating-point engine 3226 performs, but is not limited to, floating-point and vector operations. In at least one embodiment, the fetch-decode unit 3222 dispatches micro-ops to a single execution engine that replaces both the integer execution engine 3224 and the floating-point execution engine 3226.

[0300] In at least one embodiment, each core 3220(i) can access an L2 cache 3228(i) included in the core 3220(i), where i is an integer representing a specific instance of the core 3220. In at least one embodiment, each core 3220 included in a core complex 3210(j) is connected to the other cores 3220 in the core complex 3210(j) via an L3 cache 3230(j) included in the core complex 3210(j), where j is an integer representing a specific instance of the core complex 3210. In at least one embodiment, a core 3220 included in a core complex 3210(j) can access all L3 caches 3230(j) included in the core complex 3210(j), where j is an integer representing a specific instance of the core complex 3210. In at least one embodiment, the L3 cache 3230 can include, but is not limited to, any number of slices.

[0301] In at least one embodiment, fabric 3260 is a system interconnect that facilitates data and control transfers across core complexes 3210(1)-3210(N) (where N is an integer greater than zero), I / O interface 3270, and memory controller 3280. In at least one embodiment, CPU 3200 may include, in addition to or in lieu of fabric 3260, but is not limited to, any number and type of system interconnects that facilitate data and control transfers across any number and type of directly or indirectly linked components that may be internal or external to CPU 3200. In at least one embodiment, I / O interface 3270 represents any number and type of I / O interfaces (e.g., PCI, PCI-X, PCIe, GBE, USB, etc.). In at least one embodiment, various types of peripherals are coupled to I / O interface 3270. In at least one embodiment, peripherals coupled to I / O interface 3270 may include, but are not limited to, a display, a keyboard, a mouse, a printer, a scanner, a joystick or other type of game controller, a media recording device, an external storage device, a network interface card, and the like.

[0302] In at least one embodiment, memory controllers 3280 facilitate data transfers between CPU 3200 and system memory 3290. In at least one embodiment, core complex 3210 and graphics complex 3240 share system memory 3290. In at least one embodiment, CPU 3200 implements a memory subsystem that includes, without limitation, any number and type of memory controllers 3280 and memory devices that can be dedicated to one component or shared among multiple components. In at least one embodiment, CPU 3200 implements a cache subsystem that includes, without limitation, one or more cache memories (e.g., L2 cache 3228 and L3 cache 3230), each of which can be private to a component or shared among any number of components (e.g., core 3220 and core complex 3210).

[0303] Figure 33 An exemplary accelerator integration slice 3390 is shown in accordance with at least one embodiment. As used herein, a “slice” includes a specified portion of processing resources of an accelerator integration circuit. In at least one embodiment, an accelerator integration circuit provides cache management, memory access, environment management, and interrupt management services on behalf of multiple graphics processing engines that are part of graphics acceleration modules. Graphics processing engines can each comprise a separate GPU. Alternatively, graphics processing engines can include different types of graphics processing engines within a GPU, such as graphics execution units, media processing engines (e.g., video encoders / decoders), samplers, and blit engines. In at least one embodiment, a graphics acceleration module can be a GPU with a plurality of graphics processing engines. In at least one embodiment, a graphics processing engine can be a separate GPU integrated on a common package, line card, or chip as the CPU.

[0304] Application effective address space 3382 within system memory 3314 stores process elements 3383. In one embodiment, process elements 3383 are stored in response to GPU invocations 3381 from applications 3380 executing on processor 3307. Process elements 3383 contain processing state for corresponding applications 3380. Work descriptors (WDs) 3384 contained in process elements 3383 can be individual jobs requested by an application or can contain pointers to queues of jobs. In at least one embodiment, WDs 3384 are pointers to job request queues in application effective address space 3382.

[0305] Graphics acceleration module 3346 and / or individual graphics processing engines can be shared by all or a subset of processes in a system. In at least one embodiment, infrastructure for setting up processing state and sending WDs 3384 to graphics acceleration module 3346 to start a job in a virtualized environment can be included.

[0306] In at least one embodiment, a dedicated process programming model is implemented. In this model, a single process owns a graphics acceleration module 3346 or individual graphics processing engines. As graphics acceleration module 3346 is owned by a single process, a hypervisor initializes the accelerator integration circuit for the owning partition and an operating system initializes the accelerator integration circuit for the owning partition when graphics acceleration module 3346 is assigned.

[0307] In operation, a WD fetch unit 3391 in accelerator integration slice 3390 fetches a next WD 3384 including an indication of work to be completed by one or more graphics processing engines of graphics acceleration module 3346. Data from WD 3384 can be stored in registers 3345 used by memory management unit (MMU) 3339, interrupt management circuit 3347, and / or environment management circuit 3348, as shown. At least one embodiment of MMU 3339 includes segment / page walk circuitry to access segment / page tables 3386 within an OS virtual address space 3385. Interrupt management circuit 3347 can handle interrupt events (INTs) 3392 received from graphics acceleration module 3346. Effective addresses 3393 produced by graphics processing engines, when executing graphics operations, are translated to real addresses by MMU 3339.

[0308] In one embodiment, a same set of registers 3345 is replicated for each graphics processing engine and / or graphics acceleration module 3346 and can be initialized by a system hypervisor or operating system. Each of these replicated registers can be included in accelerator integration slice 3390. Exemplary registers that can be initialized by a hypervisor are shown in Table 1.

[0309] Table 1 - Hypervisor-Initialized Registers

[0310] 1 Slice Control Register 2 Real address (RA) plan processing area pointer 3 Authorization Mask Override Register 4 Interrupt vector table input offset 5 Interrupt vector table entry restriction 6 Status Register 7 Logical partition ID 8 Real Address (RA) Hypervisor Accelerator Utilization Record Pointer 9 Storage Description Register

[0311] Exemplary registers that can be initialized by an operating system are shown in Table 2.

[0312] Table 2 - Operating System-Initialized Registers

[0313]

[0314]

[0315] In one embodiment, each WD 3384 is specific to a particular graphics acceleration module 3346 and / or a particular graphics processing engine. It contains all the information the graphics processing engine needs to do its work or work, or it can be a pointer to a memory location where the application has set up a command queue for work to be done.

[0316] Figures 34A-34B An exemplary graphics processor according to at least one embodiment of the present disclosure is shown. In at least one embodiment, any exemplary graphics processor can be manufactured using one or more IP cores. In addition to the illustrated diagram, in at least one embodiment, other logic and circuitry can be included, including additional graphics processors / cores, peripheral interface controllers, or general-purpose processor cores. In at least one embodiment, the exemplary graphics processor is used within a SoC.

[0317] Figure 34A An exemplary graphics processor 3410 of a SoC integrated circuit is shown, which may be manufactured using one or more IP cores, in accordance with at least one embodiment. Figure 34B An additional exemplary graphics processor 3440 of a SoC integrated circuit is shown, which may be manufactured using one or more IP cores, in accordance with at least one embodiment. Figure 34A The graphics processor 3410 is a low power graphics processor core. In at least one embodiment, Figure 34B The graphics processor 3440 is a higher performance graphics processor core. In at least one embodiment, each graphics processor 3410, 3440 can be a variation of the graphics processor 510 of Figure 5.

[0318] In at least one embodiment, the graphics processor 3410 includes a vertex processor 3405 and one or more fragment processors 3415A-3415N (e.g., 3415A, 3415B, 3415C, 3415D through 3415N-1 and 3415N). In at least one embodiment, the graphics processor 3410 can execute different shader programs via separate logic, such that the vertex processor 3405 is optimized to perform operations for the vertex shader program, while one or more fragment processors 3415A-3415N perform fragment (e.g., pixel) shading operations for the fragment or pixel or shader program. In at least one embodiment, the vertex processor 3405 performs the vertex processing stage of the 3D graphics pipeline and generates primitives and vertex data. In at least one embodiment, the fragment processors 3415A-3415N use the primitives and vertex data generated by the vertex processor 3405 to generate a frame buffer for display on a display device. In at least one embodiment, the fragment processors 3415A-3415N are optimized to execute fragment shader programs as provided in the OpenGL API, which can be used to perform similar operations as pixel shader programs provided in the Direct 3D API.

[0319] In at least one embodiment, the graphics processor 3410 additionally includes one or more MMUs 3420A-3420B, caches 3425A-3425B, and circuit interconnects 3430A-3430B. In at least one embodiment, the one or more MMUs 3420A-3420B provide virtual-to-physical address mapping for the graphics processor 3410, including for the vertex processor 3405 and / or the fragment processors 3415A-3415N, which can reference vertex or image / texture data stored in memory, in addition to the vertex or image / texture data stored in the one or more caches 3425A-3425B. In at least one embodiment, the one or more MMUs 3420A-3420B can synchronize with other MMUs within the system, including one or more MMUs associated with one or more application processors 505, image processor 515, and / or video processor 520 of FIG. 5, so that each processor 505-520 can participate in a shared or unified virtual memory system. In at least one embodiment, one or more circuit interconnects 3430A-3430B enable graphics processor 3410 to connect to other IP cores within the SoC via an internal bus of the SoC or via direct connections.

[0320] In at least one embodiment, graphics processor 3440 includes Figure 34A3420A-3420B, caches 3425A-3425B, and circuit interconnects 3430A-3430B of the graphics processor 3410. In at least one embodiment, the graphics processor 3440 includes one or more shader cores 3455A-3455N (e.g., 3455A, 3455B, 3455C, 3455D, 3455E, 3455F, through 3455N-1 and 3455N) that provide a unified shader core architecture in which a single core or type or core can execute all types of programmable shader code, including shader program code for implementing vertex shaders, fragment shaders, and / or compute shaders. In at least one embodiment, the number of shader cores can vary. In at least one embodiment, the graphics processor 3440 includes an inter-core task manager 3445 that acts as a thread dispatcher to dispatch execution threads to one or more shader cores 3455A-3455N and a tiling unit 3458 to accelerate tile-based rendering operations in which rendering operations of a scene are subdivided in image space, for example, to exploit local spatial coherence within a scene or to optimize use of internal caches.

[0321] Figure 35A FIG35 shows a graphics core 3500 according to at least one embodiment. In at least one embodiment, the graphics core 3500 may include Figure 24 In at least one embodiment, the graphics core 3500 may be Figure 34B 3455N. In at least one embodiment, graphics core 3500 includes a shared instruction cache 3502, texture units 3518, and cache / shared memory 3520, which are common to execution resources within graphics core 3500. In at least one embodiment, graphics core 3500 may include multiple slices 3501A-3501N or partitions of each core, and a graphics processor may include multiple instances of graphics core 3500. Slices 3501A-3501N may include support logic including local instruction caches 3504A-3504N, thread schedulers 3506A-3506N, thread dispatchers 3508A-3508N, and a set of registers 3510A-3510N. In at least one embodiment, the slices 3501A-3501N may include a set of additional function units (AFUs) 3512A-3512N, floating point units (FPUs) 3514A-3514N, integer arithmetic logic units (ALUs) 3516A-3516N, address calculation units (ACUs) 3513A-3513N, double precision floating point units (DPFPUs) 3515A-3515N, and matrix processing units (MPUs) 3517A-3517N.

[0322] In one embodiment, FPUs 3514A-3514N can perform single-precision (32-bit) and half-precision (16-bit) floating point operations, while DPFPUs 3515A-3515N can perform double-precision (64-bit) floating point operations. In at least one embodiment, ALUs 3516A-3516N can perform variable precision integer operations at 8-bit, 16-bit, and 32-bit precision, and can be configured for mixed precision operations. In at least one embodiment, MPUs 3517A-3517N can also be configured for mixed precision matrix operations, including half-precision floating point operations and 8-bit integer operations. In at least one embodiment, MPUs 3517A-3517N can perform various matrix operations to accelerate CUDA programs, including enabling support for accelerated General Matrix to Matrix multiplication (GEMM). In at least one embodiment, AFUs 3512A-3512N can perform additional logical operations not supported by floating point or integer units, including trigonometric operations (e.g., Sine, Cosine, etc.).

[0323] Figure 35B A general purpose graphics processing unit (GPGPU) 3530 in at least one embodiment is shown. In at least one embodiment, GPGPU 3530 is highly parallel and suitable for deployment on a multi-chip module. In at least one embodiment, GPGPU 3530 can be configured to enable highly parallel compute operations to be performed by a GPU array. In at least one embodiment, GPGPU 3530 can be directly linked to other instances of GPGPU 3530 to create a multi-GPU cluster to improve execution time for CUDA programs. In at least one embodiment, GPGPU 3530 includes a host interface 3532 to enable connection to a host processor. In at least one embodiment, host interface 3532 is a PCIe interface. In at least one embodiment, host interface 3532 can be a vendor-specific communications interface or communication structure. In at least one embodiment, GPGPU 3530 receives commands from a host processor and uses a global scheduler 3534 to dispatch execution threads associated with those commands to a group of compute clusters 3536A-3536H. In at least one embodiment, compute clusters 3536A-3536H share a cache memory 3538. In at least one embodiment, cache memory 3538 can be used as an upper level cache for cache memory within compute clusters 3536A-3536H.

[0324] In at least one embodiment, GPGPU 3530 includes memory 3544A-3544B coupled to a compute cluster 3536A-3536H via a set of memory controllers 3542A-3542B. In at least one embodiment, memory 3544A-3544B may include various types of memory devices, including dynamic random access memory (DRAM) or graphics random access memory, such as synchronous graphics random access memory (SGRAM), including graphics double data rate (GDDR) memory.

[0325] In at least one embodiment, computing clusters 3536A-3536H each include a set of graphics cores, such as Figure 35A The graphics core 3500, which may include multiple types of integer and floating-point logic units, can perform computational operations at various precisions, including computations suitable for use with CUDA programs. In at least one embodiment, at least a subset of the floating-point units in each compute cluster 3536A-3536H can be configured to perform 16-bit or 32-bit floating-point operations, while a different subset of the floating-point units can be configured to perform 64-bit floating-point operations.

[0326] In at least one embodiment, multiple instances of GPGPU 3530 can be configured to operate as a compute cluster. In at least one embodiment, compute clusters 3536A-3536H can implement any technically feasible communication technology for synchronization and data exchange. In at least one embodiment, multiple instances of GPGPU 3530 communicate via host interface 3532. In at least one embodiment, GPGPU 3530 includes an I / O hub 3539 that couples GPGPU 3530 to a GPU link 3540, enabling direct connection to other instances of GPGPU 3530. In at least one embodiment, GPU link 3540 is coupled to a dedicated GPU-to-GPU bridge that enables communication and synchronization between multiple instances of GPGPU 3530. In at least one embodiment, GPU link 3540 is coupled to a high-speed interconnect to send and receive data to other GPGPUs or parallel processors. In at least one embodiment, multiple instances of GPGPU 3530 are located in separate data processing systems and communicate via a network device accessible via host interface 3532. In at least one embodiment, GPU link 3540 may be configured to connect to a host processor, in addition to or in place of host interface 3532. In at least one embodiment, GPGPU 3530 may be configured to execute CUDA programs.

[0327] Figure 36AA parallel processor 3600 in accordance with at least one embodiment is shown. In at least one embodiment, the various components of the parallel processor 3600 may be implemented using one or more integrated circuit devices, such as a programmable processor, an application specific integrated circuit (ASIC), or an FPGA.

[0328] In at least one embodiment, parallel processor 3600 includes parallel processing unit 3602. In at least one embodiment, parallel processing unit 3602 includes an I / O unit 3604 that enables communication with other devices, including other instances of parallel processing unit 3602. In at least one embodiment, I / O unit 3604 can be directly connected to other devices. In at least one embodiment, I / O unit 3604 connects to other devices using a hub or switch interface (e.g., memory hub 605). In at least one embodiment, the connection between memory hub 605 and I / O unit 3604 forms a communication link. In at least one embodiment, I / O unit 3604 is connected to a host interface 3606 and a memory crossbar switch 3616, where host interface 3606 receives commands for performing processing operations and memory crossbar switch 3616 receives commands for performing memory operations.

[0329] In at least one embodiment, when host interface 3606 receives command buffers via I / O unit 3604, host interface 3606 can direct work operations to execute those commands to front end 3608. In at least one embodiment, front end 3608 is coupled to scheduler 3610, which is configured to dispatch commands or other work items to processing array 3612. In at least one embodiment, scheduler 3610 ensures that processing array 3612 is properly configured and in a valid state before dispatching tasks to processing array 3612. In at least one embodiment, scheduler 3610 is implemented by firmware logic executing on a microcontroller. In at least one embodiment, the microcontroller-implemented scheduler 3610 can be configured to perform complex scheduling and work dispatch operations at both coarse and fine granularity, thereby enabling fast preemption and context switching of threads executing on processing array 3612. In at least one embodiment, host software can authenticate workloads for scheduling on processing array 3612 through one of multiple graphics processing doorbells. In at least one embodiment, the workload can then be automatically distributed across the processing array 3612 by scheduler 3610 logic within a microcontroller that includes scheduler 3610.

[0330] In at least one embodiment, processing array 3612 can include up to "N" processing clusters (e.g., cluster 3614A, cluster 3614B, through cluster 3614N). In at least one embodiment, each cluster 3614A-3614N of processing array 3612 can execute a large number of concurrent threads. In at least one embodiment, scheduler 3610 can allocate work to clusters 3614A-3614N of processing array 3612 using various scheduling and / or work distribution algorithms, which can vary depending on the workload generated by each program or computation type. In at least one embodiment, scheduling can be handled dynamically by scheduler 3610 or can be partially assisted by compiler logic during the compilation of program logic configured to be executed by processing array 3612. In at least one embodiment, different clusters 3614A-3614N of processing array 3612 can be assigned to process different types of programs or to perform different types of computations.

[0331] In at least one embodiment, processing array 3612 can be configured to perform various types of parallel processing operations. In at least one embodiment, processing array 3612 can be configured to perform general-purpose parallel computing operations. In at least one embodiment, processing array 3612 can include logic to perform processing tasks including filtering video and / or audio data, performing modeling operations including physics operations, and performing data transformations.

[0332] In at least one embodiment, processing array 3612 is configured to perform parallel graphics processing operations. In at least one embodiment, processing array 3612 may include additional logic to support the execution of such graphics processing operations, including but not limited to texture sampling logic to perform texture operations, as well as tessellation logic and other vertex processing logic. In at least one embodiment, processing array 3612 may be configured to execute shader programs related to graphics processing, such as, but not limited to, vertex shaders, tessellation shaders, geometry shaders, and pixel shaders. In at least one embodiment, parallel processing unit 3602 may transfer data from system memory via I / O unit 3604 for processing. In at least one embodiment, during processing, the transferred data may be stored in on-chip memory (e.g., parallel processor memory 3622) during processing and then written back to system memory.

[0333] In at least one embodiment, when parallel processing unit 3602 is used to perform graphics processing, scheduler 3610 can be configured to divide incoming workloads into tasks of approximately equal size to better enable distribution of graphics processing operations across multiple clusters 3614A-3614N of processing array 3612. In at least one embodiment, portions of processing array 3612 can be configured to perform different types of processing. In at least one embodiment, a first portion can be configured to perform vertex shading and topology generation, a second portion can be configured to perform tessellation and geometry shading, and a third portion can be configured to perform pixel shading or other screen space operations to produce a rendered image for display on a display device. In at least one embodiment, intermediate data produced by one or more of clusters 3614A-3614N can be stored in buffers to allow transmission of intermediate data between clusters 3614A-3614N for further processing.

[0334] In at least one embodiment, processing array 3612 can receive processing tasks to be executed from scheduler 3610, which receives commands defining the processing tasks from front end 3608. In at least one embodiment, a processing task can include an index into data to be processed, such as can include surface (patch) data, raw data, vertex data, and / or pixel data, as well as state parameters and commands defining how the data is to be processed (e.g., what programs are to be executed). In at least one embodiment, scheduler 3610 can be configured to fetch the index corresponding to a task, or can receive the index from front end 3608. In at least one embodiment, front end 3608 can be configured to ensure that processing array 3612 is configured in an effective state before launching a workload specified by an incoming command buffer (e.g., a batch-buffer, a push buffer, etc.).

[0335] In at least one embodiment, each of one or more instances of parallel processing unit 3602 can be coupled to parallel processor memory 3622. In at least one embodiment, parallel processor memory 3622 can be accessed via memory crossbar 3616, which can receive memory requests from processing array 3612 and I / O unit 3604. In at least one embodiment, memory crossbar 3616 can access parallel processor memory 3622 via memory interface 3618. In at least one embodiment, memory interface 3618 can include multiple partition units (e.g., partition unit 3620A, partition unit 3620B, through partition unit 3620N), which can each be coupled to a portion of parallel processor memory 3622 (e.g., a memory unit). In at least one embodiment, the plurality of partition units 3620A-3620N are configured to be equal to the number of memory cells, such that the first partition unit 3620A has a corresponding first memory cell 3624A, the second partition unit 3620B has a corresponding memory cell 3624B, and the Nth partition unit 3620N has a corresponding Nth memory cell 3624N. In at least one embodiment, the number of partition units 3620A-3620N may not be equal to the number of memory devices.

[0336] In at least one embodiment, memory units 3624A-3624N may include various types of memory devices, including dynamic random access memory (DRAM) or graphics random access memory, such as synchronous graphics random access memory (SGRAM), including graphics double data rate (GDDR) memory. In at least one embodiment, memory units 3624A-3624N may also include 3D stacked memory, including but not limited to high bandwidth memory (HBM). In at least one embodiment, render targets such as frame buffers or texture maps may be stored across memory units 3624A-3624N, allowing partition units 3620A-3620N to write portions of each render target in parallel to efficiently use the available bandwidth of parallel processor memory 3622. In at least one embodiment, local instances of parallel processor memory 3622 may be eliminated in favor of a unified memory design utilizing system memory in combination with local cache memory.

[0337] In at least one embodiment, any of the clusters 3614A-3614N of the processing array 3612 can process data to be written to any memory unit 3624A-3624N within the parallel processor memory 3622. In at least one embodiment, the memory crossbar 3616 can be configured to transmit the output of each cluster 3614A-3614N to any partition unit 3620A-3620N or another cluster 3614A-3614N, which can perform other processing operations on the output. In at least one embodiment, each cluster 3614A-3614N can communicate with a memory interface 3618 via the memory crossbar 3616 to read from or write to various external storage devices. In at least one embodiment, memory crossbar switch 3616 has connections to memory interface 3618 for communicating with I / O unit 3604, as well as connections to local instances of parallel processor memory 3622, thereby enabling processing units within different processing clusters 3614A-3614N to communicate with system memory or other memory that is not local to parallel processing unit 3602. In at least one embodiment, memory crossbar switch 3616 may use virtual channels to separate traffic flows between clusters 3614A-3614N and partition units 3620A-3620N.

[0338] In at least one embodiment, multiple instances of parallel processing unit 3602 can be provided on a single plug-in card, or multiple plug-in cards can be interconnected. In at least one embodiment, different instances of parallel processing unit 3602 can be configured to interoperate with each other, even if the different instances have different numbers of processing cores, different amounts of local parallel processor memory, and / or other configuration differences. In at least one embodiment, some instances of parallel processing unit 3602 can include higher precision floating point units relative to other instances. In at least one embodiment, a system incorporating one or more instances of parallel processing unit 3602 or parallel processor 3600 can be implemented in a variety of configurations and form factors, including but not limited to desktop, laptop, or handheld personal computers, servers, workstations, game consoles, and / or embedded systems.

[0339] Figure 36B FIG3 shows a processing cluster 3694 according to at least one embodiment. In at least one embodiment, the processing cluster 3694 is included within a parallel processing unit. In at least one embodiment, the processing cluster 3694 is Figure 36Aone of the processing clusters 3614A-3614N. In at least one embodiment, processing cluster 3694 can be configured to execute many threads in parallel, where the term “thread” refers to an instance of a particular program executed by a particular group of one or more processing clusters. In at least one embodiment, Single Instruction Multiple Data (SIMD) instruction issue techniques are used to support parallel execution of a large number of threads with no or negligible program overhead. In at least one embodiment, Single Instruction Multiple Thread (SIMT) techniques are used to support parallel execution of a large number of generally synchronous threads using a common instruction unit configured to issue instructions to a group of processing engines within each processing cluster 3694.

[0340] In at least one embodiment, operation of processing cluster 3694 can be controlled via a pipeline manager 3632 that allocates processing tasks to SIMT parallel processors. In at least one embodiment, pipeline manager 3632 receives instructions from scheduler 3610 and manages execution of those instructions by graphics multiprocessor 3634 and / or texture unit 3636, in at least one embodiment, graphics multiprocessor 3634 is an exemplary instance of a SIMT parallel processor. However, in at least one embodiment, various types of SIMT parallel processors of differing architectures can be included within processing cluster 3694. In at least one embodiment, one or more instances of graphics multiprocessor 3634 can be included within processing cluster 3694. In at least one embodiment, graphics multiprocessor 3634 can process data and a data crossbar 3640 can be used to distribute processed data to one of a number of possible destinations, including other shader units. In at least one embodiment, pipeline manager 3632 can facilitate distribution by specifying destinations for processed data as a function of the destination’s source in either a fixed function or switched fabric manner. Figure 36A

[0341] In at least one embodiment, each graphics multiprocessor 3634 within processing cluster 3694 can include an identical set of functional execution logic (e.g., arithmetic logic units, load store units (LSUs), etc.). In at least one embodiment, functional execution logic can be configured in a pipelined manner in which instructions are issued at a first stage, passed through stages of the pipeline with

[0342] ​In at least one embodiment, instructions delivered to processing cluster 3694 constitute a thread. In at least one embodiment, a set of threads executing across a set of parallel processing engines constitutes a warp. In at least one embodiment, a thread group is a group of threads executing the same program, although each thread within a thread group can be at different instruction points within the program. In at least one embodiment, a thread group is associated with a same set of instruction boundaries. In at least one embodiment, a thread group includes fewer threads than are available processing engines within graphics multiprocessor 3634. In at least one embodiment, when a thread group includes fewer threads than the number of processing engines within graphics multiprocessor 3634 available to be invoked, one or more of the processing engines can be idle while a cycle of the thread group is being processed. In at least one embodiment, a thread group can include more threads than are available processing engines within graphics multiprocessor 3634. In at least one embodiment, when a thread group includes more threads than the number of processing engines within graphics multiprocessor 3634, a processing engine can be invoked on a single thread multiple times within one cycle. In at least one embodiment, multiple thread groups can execute concurrently on a graphics multiprocessor 3634, where each thread group is associated with a separate and distinct set of instruction boundaries.

[0343] In at least one embodiment, graphics multiprocessor 3634 includes internal cache memory, to perform load and store operations. In at least one embodiment, graphics multiprocessor 3634 can bypass internal cache and use register file memory, such as Ll cache 3648, within processing cluster 3694. In at least one embodiment, each graphics multiprocessor 3634 can also have access to L2 cache within a partition unit (e.g., partition units 3620A-3620N) that is shared among multiple processing clusters 3694 and can be used to transfer data between threads. Figure 36A In at least one embodiment, graphics multiprocessor 3634 can also have access to off-chip global memory, which can include one or more of local parallel processor memory and / or system memory. In at least one embodiment, any memory outside of parallel processor 3602 can be considered global memory. In at least one embodiment, processing cluster 3694 includes multiple instances of graphics multiprocessor 3634, which can share common instructions and data stored in Ll cache 3648.

[0344] In at least one embodiment, each processing cluster 3694 can include an MMU 3645 that is configured to translate virtual addresses into physical addresses, as is known to those skilled in the art. In at least one embodiment, one or more instances of MMU 3645 can reside within Figure 36A3618. In at least one embodiment, the MMU 3645 includes a set of page table entries (PTEs) that are used to map virtual addresses to physical addresses of tiles (more on tiles below) and optionally to cache line indices. In at least one embodiment, the MMU 3645 may include a translation lookaside buffer (TLB) or cache that may reside within the graphics multiprocessor 3634 or L1 cache 3648 or processing cluster 3694. In at least one embodiment, the physical addresses are processed to assign surface data access locality for efficient request interleaving between partition units. In at least one embodiment, the cache line index may be used to determine whether a request for a cache line is a hit or a miss.

[0345] In at least one embodiment, the processing clusters 3694 can be configured such that each graphics multiprocessor 3634 is coupled to a texture unit 3636 to perform texture mapping operations, which may involve, for example, determining texture sample locations, reading texture data, and filtering the texture data. In at least one embodiment, texture data is read from an internal texture L1 cache (not shown) or from an L1 cache within the graphics multiprocessor 3634, and texture data is retrieved from an L2 cache, local parallel processor memory, or system memory, as needed. In at least one embodiment, each graphics multiprocessor 3634 outputs processed tasks to a data crossbar 3640 to provide the processed tasks to another processing cluster 3694 for further processing or to store the processed tasks in an L2 cache, local parallel processor memory, or system memory via a memory crossbar 3616. In at least one embodiment, a pre-raster operations unit (preROP) 3642 is configured to receive data from the graphics multiprocessor 3634 and direct the data to a ROP unit, which can communicate with a partitioning unit (e.g., a partitioning unit) as described herein. Figure 36A In at least one embodiment, the PreROP 3642 unit can perform optimizations for color blending, organize pixel color data, and perform address translation.

[0346] Figure 36C A graphics multiprocessor 3696 is shown in accordance with at least one embodiment. In at least one embodiment, the graphics multiprocessor 3696 is Figure 36B36. In at least one embodiment, the graphics multiprocessor 3696 is coupled to the pipeline manager 3632 of the processing cluster 3694. In at least one embodiment, the graphics multiprocessor 3696 has an execution pipeline that includes, but is not limited to, an instruction cache 3652, an instruction unit 3654, an address mapping unit 3656, a register file 3658, one or more GPGPU cores 3662, and one or more LSUs 3666. The GPGPU cores 3662 and LSUs 3666 are coupled to cache memory 3672 and shared memory 3670 via a memory and cache interconnect 3668.

[0347] In at least one embodiment, the instruction cache 3652 receives a stream of instructions to be executed from the pipeline manager 3632. In at least one embodiment, the instructions are cached in the instruction cache 3652 and dispatched for execution by the instruction unit 3654. In one embodiment, the instruction unit 3654 can dispatch instructions as thread groups (e.g., warps), assigning each thread of the thread group to a different execution unit within the GPGPU core 3662. In at least one embodiment, the instructions can access any local, shared, or global address space by specifying an address within the unified address space. In at least one embodiment, the address mapping unit 3656 can be used to convert addresses in the unified address space into different memory addresses that can be accessed by the LSU 3666.

[0348] In at least one embodiment, register file 3658 provides a set of registers for the functional units of graphics multiprocessor 3696. In at least one embodiment, register file 3658 provides temporary storage for operands for the data paths of the functional units (e.g., GPGPU core 3662, LSU 3666) connected to graphics multiprocessor 3696. In at least one embodiment, register file 3658 is divided between each functional unit such that a dedicated portion of register file 3658 is allocated to each functional unit. In at least one embodiment, register file 3658 is divided between the different thread groups being executed by graphics multiprocessor 3696.

[0349] In at least one embodiment, the GPGPU cores 3662 may each include an FPU and / or ALU for executing instructions of the graphics multiprocessor 3696. The GPGPU cores 3662 may be architecturally similar or the architectures may differ. In at least one embodiment, a first portion of the GPGPU core 3662 includes a single-precision FPU and integer ALU, while a second portion of the GPGPU core includes a double-precision FPU. In at least one embodiment, the FPU may implement the IEEE 754-3608 standard for floating-point arithmetic or enable variable-precision floating-point arithmetic. In at least one embodiment, the graphics multiprocessor 3696 may additionally include one or more fixed-function or special-function units to perform specific functions, such as copying rectangles or pixel blending operations. In at least one embodiment, one or more of the GPGPU cores 3662 may also include fixed-function or special-function logic.

[0350] In at least one embodiment, the GPGPU core 3662 includes SIMD logic capable of executing a single instruction on multiple sets of data. In at least one embodiment, the GPGPU core 3662 can physically execute SIMD4, SIMD8, and SIMD16 instructions, and logically execute SIMD1, SIMD2, and SIMD32 instructions. In at least one embodiment, the SIMD instructions for the GPGPU core can be generated by a shader compiler at compile time, or automatically generated when executing a program written and compiled for a single program multiple data (SPMD) or SIMT architecture. In at least one embodiment, multiple threads of a program configured for a SIMT execution model can be executed by a single SIMD instruction. In at least one embodiment, eight SIMT threads performing the same or similar operations can be executed in parallel by a single SIMD8 logic unit.

[0351] In at least one embodiment, memory and cache interconnect 3668 is an interconnect network that connects each functional unit of graphics multiprocessor 3696 to register file 3658 and shared memory 3670. In at least one embodiment, memory and cache interconnect 3668 is a crossbar interconnect that allows LSUs 3666 to implement load and store operations between shared memory 3670 and register file 3658. In at least one embodiment, register file 3658 can operate at same frequency as GPGPU cores 3662, resulting in very low latency for data transfers between GPGPU cores 3662 and register file 3658. In at least one embodiment, shared memory 3670 can be used to enable communication between threads executing on functional units within graphics multiprocessor 3696. In at least one embodiment, cache memory 3672 can be used to store data for threads executing on functional units and texture data for textures accessed by these threads. In at least one embodiment, shared memory 3670 can also be used as a program managed cache.

[0352] In at least one embodiment, parallel processor or GPGPU as described herein is communicatively coupled to host / processor cores to accelerate graphics operations, machine learning operations, pattern analysis operations, and various general purpose GPU (GPGPU) functions. In at least one embodiment, GPU can be communicatively coupled to host processor / cores by a bus or other interconnect (e.g., a high speed

[0353] General-Purpose Computing

[0354] The following figures illustrate, without limitation, exemplary software configurations used in general-purpose computing to implement at least one embodiment.

[0355] Figure 37A software stack of a programming platform is shown, in accordance with at least one embodiment. In at least one embodiment, a programming platform is a platform for utilizing hardware on a computing system to accelerate compute tasks. In at least one embodiment, a software developer can access a programming platform through libraries, compiler directives, and / or extensions to a programming language. In at least one embodiment, a programming platform can be, but is not limited to, CUDA, Radeon Open Compute Platform (“ROCm”), OpenCL (OpenCL TM ), SYCL, or Intel One API.

[0356] In at least one embodiment, software stack 3700 of a programming platform provides an execution environment for application 3701. In at least one embodiment, application 3701 can include any computer software capable of launching on software stack 3700. In at least one embodiment, application 3701 can include, but is not limited to, artificial intelligence (“AI”) / machine learning (“ML”) applications, high performance computing (“HPC”) applications, virtual desktop infrastructure (“VDI”), or data center workloads.

[0357] In at least one embodiment, application 3701 and software stack 3700 run on hardware 3707. In at least one embodiment, hardware 3707 can include one or more GPUs, CPUs, FPGAs, AI engines, and / or other types of computing devices that support a programming platform. In at least one embodiment, software stack 3700 can be vendor-specific and compatible only with devices from a particular vendor, e.g., with CUDA. In at least one embodiment, software stack 3700 can be used with devices from different vendors, e.g., with OpenCL. In at least one embodiment, hardware 3707 includes a host connected to one or more devices that can be accessed via application programming interface (API) calls to perform compute tasks. In at least one embodiment, in contrast to a host within hardware 3707, which can include, but is not limited to, a CPU (but can also include a computing device) and its memory, a device within hardware 3707 can include, but is not limited to, a GPU, FPGA, AI engine, or other computing device (but can also include a CPU) and its memory.

[0358] In at least one embodiment, the programming platform's software stack 3700 includes, but is not limited to, a plurality of libraries 3703, a runtime 3705, and device kernel drivers 3706. In at least one embodiment, each of the libraries 3703 may include data and programming code that can be used by a computer program and utilized during software development. In at least one embodiment, the libraries 3703 may include, but are not limited to, pre-written code and subroutines, classes, values, type specifications, configuration data, documentation, help data, and / or message templates. In at least one embodiment, the libraries 3703 include functions optimized for execution on one or more types of devices. In at least one embodiment, the libraries 3703 may include, but are not limited to, functions for performing mathematical, deep learning, and / or other types of operations on the devices. In at least one embodiment, the libraries 3803 are associated with corresponding APIs 3802, which may include one or more APIs that expose the functions implemented in the libraries 3803.

[0359] In at least one embodiment, application 3701 is written as source code that is compiled into executable code as follows in conjunction with Figure 42 3701. In at least one embodiment, the executable code of application 3701 can run at least in part on an execution environment provided by software stack 3700. In at least one embodiment, during the execution of application 3701, code that needs to run on the device (as opposed to the host) can be obtained. In this case, in at least one embodiment, runtime 3705 can be called to load and start the necessary code on the device. In at least one embodiment, runtime 3705 can include any technically feasible runtime system capable of supporting the execution of application 3701.

[0360] In at least one embodiment, runtime 3705 is implemented as one or more runtime libraries associated with a corresponding API (shown as API 3704). In at least one embodiment, one or more such runtime libraries may include, but are not limited to, functions for memory management, execution control, device management, error handling, and / or synchronization, among others. In at least one embodiment, memory management functions may include, but are not limited to, functions for allocating, deallocating, and copying device memory, and transferring data between host memory and device memory. In at least one embodiment, execution control functions may include, but are not limited to, functions for launching a function on the device (sometimes referred to as a "kernel" when the function is a global function callable from the host), and functions for setting property values ​​in buffers maintained by the runtime library for a given function to be executed on the device.

[0361] In at least one embodiment, the runtime library and corresponding API 3704 can be implemented in any technically feasible manner. In at least one embodiment, one (or any number of) APIs can expose a low-level set of functions for fine-grained control of a device, while another (or any number of) APIs can expose such a higher-level set of functions. In at least one embodiment, a high-level runtime API can be built on top of the low-level APIs. In at least one embodiment, one or more runtime APIs can be language-specific APIs layered on top of a language-independent runtime API.

[0362] In at least one embodiment, the device kernel driver 3706 is configured to facilitate communication with the underlying device. In at least one embodiment, the device kernel driver 3706 can provide APIs such as API 3704 and / or low-level functions that other software relies on. In at least one embodiment, the device kernel driver 3706 can be configured to compile intermediate representation ("IR") code into binary code at runtime. In at least one embodiment, for CUDA, the device kernel driver 3706 can compile non-hardware-specific parallel thread execution ("PTX") IR code into binary code for a specific target device at runtime (caching the compiled binary code), which is sometimes also referred to as "final" code. In at least one embodiment, doing so can allow the final code to run on a target device that may not have existed when the source code was originally compiled into PTX code. Alternatively, in at least one embodiment, the device source code can be compiled into binary code offline without the device kernel driver 3706 compiling the IR code at runtime.

[0363] Figure 38 According to at least one embodiment, Figure 37 3800. In at least one embodiment, the CUDA software stack 3800, on which the application 3801 can be launched, includes a CUDA library 3803, a CUDA runtime 3805, a CUDA driver 3807, and a device kernel driver 3808. In at least one embodiment, the CUDA software stack 3800 executes on hardware 3809, which may include a CUDA-enabled GPU developed by NVIDIA Corporation of Santa Clara, California.

[0364] In at least one embodiment, the application 3801, the CUDA runtime 3805, and the device kernel driver 3808 can perform similar functions as the application 3701, the runtime 3705, and the device kernel driver 3706, respectively. Figure 373806 . In at least one embodiment, the CUDA driver 3807 includes a library (libcuda.so) that implements the CUDA driver API 3806. In at least one embodiment, similar to the CUDA runtime API 3804 implemented by the CUDA runtime library (cudart), the CUDA driver API 3806 may expose, but is not limited to, functions for memory management, execution control, device management, error handling, synchronization, and / or graphics interoperability. In at least one embodiment, the CUDA driver API 3806 differs from the CUDA runtime API 3804 in that the CUDA runtime API 3804 simplifies device code management by providing implicit initialization, context (similar to process) management, and module (similar to dynamically loaded libraries) management. In contrast to the high-level CUDA runtime API 3804, in at least one embodiment, the CUDA driver API 3806 is a low-level API that provides finer-grained control over the device, particularly with respect to context and module loading. In at least one embodiment, the CUDA driver API 3806 may expose functions for context management that are not exposed by the CUDA runtime API 3804. In at least one embodiment, the CUDA driver API 3806 is also language-independent and supports, for example, OpenCL in addition to the CUDA runtime API 3804. Furthermore, in at least one embodiment, the development libraries, including the CUDA runtime 3805, can be considered separate from the driver components, including the user-mode CUDA driver 3807 and the kernel-mode device driver 3808 (sometimes also referred to as a "display" driver).

[0365] In at least one embodiment, the CUDA libraries 3803 may include, but are not limited to, mathematical libraries, deep learning libraries, parallel algorithm libraries, and / or signal / image / video processing libraries, which can be utilized by parallel computing applications (e.g., application 3801). In at least one embodiment, the CUDA libraries 3803 may include mathematical libraries, such as the cuBLAS library, which is an implementation of the Basic Linear Algebra Subroutines ("BLAS") for performing linear algebra operations; the cuFFT library for computing fast Fourier transforms ("FFTs"), and the cuRAND library for generating random numbers, among others. In at least one embodiment, the CUDA libraries 3803 may include deep learning libraries, such as the cuDNN library for primitives for deep neural networks and the TensorRT platform for high-performance deep learning inference, among others.

[0366] Figure 39 According to at least one embodiment, Figure 37ROCm implementation of the software stack 3700. In at least one embodiment, the ROCm software stack 3900 on which the application 3901 can launch includes a language runtime 3903, a system runtime 3905, a thunk 3907, a ROCm kernel driver 3908, and a device kernel driver 3909. In at least one embodiment, the ROCm software stack 3900 executes on hardware 3909, which can include a GPU that supports ROCm, which is developed by AMD Corporation of Santa Clara, California.

[0367] In at least one embodiment, the application 3901 can perform similar functions as the application 3701 discussed above in conjunction with Figure 37 In at least one embodiment, the language runtime 3903 and the system runtime 3905 can perform similar functions as the runtime 3705 discussed above in conjunction with Figure 37 In at least one embodiment, the language runtime 3903 and the system runtime 3905 differ in that the system runtime 3905 is a language-agnostic runtime that implements the ROCr system runtime API 3904 and utilizes a Heterogeneous System Architecture (“HAS”) runtime API. In at least one embodiment, the HAS runtime API is a thin user-mode API that exposes interfaces for accessing and interacting with AMD GPUs, including functions for memory management, execution control dispatching of kernels through the architecture, error handling, system and agent information, and runtime initialization and shutdown, among others. In at least one embodiment, the language runtime 3903 is an implementation of a language-specific runtime API 3902 layered on top of the ROCr system runtime API 3904 as compared to the system runtime 3905. In at least one embodiment, a language runtime API can include, without limitation, a Portable Compute Interface (“HIP”) language runtime API, a Heterogeneous Compute Compiler (“HCC”) language runtime API, or an OpenCL API, among others. In particular, the HIP language is an extension of the C++ programming language with functionally similar versions of CUDA mechanisms, and in at least one embodiment, the HIP language runtime API includes functions similar to the CUDA runtime API 3804 discussed above in conjunction with Figure 38

[0368] ​In at least one embodiment, thunk (ROCt) 3907 is an interface that can be used to interact with underlying ROCm drivers 3908. In at least one embodiment, ROCm drivers 3908 are ROCk drivers, which are a combination of AMDGPU drivers and HAS kernel drivers (amdkfd). In at least one embodiment, AMDGPU drivers are device kernel drivers for GPUs developed by AMD that perform similar functions to those discussed above in connection with Figure 37 In at least one embodiment, HAS kernel drivers are drivers that allow different types of processors to more efficiently share system resources via hardware features.

[0369] In at least one embodiment, various libraries (not shown) can be included in ROCm software stack 3900 above language runtime 3903 and provide similar functionality to CUDA libraries 3803 discussed above in connection with Figure 38 In at least one embodiment, various libraries can include, but are not limited to, math, deep learning, and / or other libraries such as a hipBLAS library that implements similar functions to CUDA cuBLAS, a rocFFT library similar to CUDA cuFFT for computing FFTs, etc.

[0370] Figure 40 An OpenCL implementation of software stack 3700 is shown in accordance with at least one embodiment. Figure 37 In at least one embodiment, OpenCL software stack 4000 on which application 4001 can be launched includes an OpenCL framework 4005, an OpenCL runtime 4006, and drivers 4007. In at least one embodiment, OpenCL software stack 4000 executes on hardware 4008 that is not vendor-specific. In at least one embodiment, because device is supported by different vendors, specific OpenCL drivers can be required to interoperate with hardware from such vendors.

[0371] In at least one embodiment, application 4001, OpenCL runtime 4006, device kernel drivers 4007, and hardware 4008 can perform similar functions to those discussed above in connection with Figure 37 In at least one embodiment, application 4001 also includes OpenCL kernels 4002 that have code to be executed on a device. In at least one embodiment, application 4001, OpenCL runtime 4006, device kernel drivers 4007, and hardware 4008 can perform similar functions to those discussed above in connection with

[0372] In at least one embodiment, OpenCL defines a "platform" that allows a host to control devices connected to the host. In at least one embodiment, the OpenCL framework provides a platform layer API and a runtime API, shown as platform API 4003 and runtime API 4005. In at least one embodiment, runtime API 4005 uses contexts to manage the execution of kernels on devices. In at least one embodiment, each identified device can be associated with a respective context, which runtime API 4005 can use to manage the device's command queue, program objects and kernel objects, shared memory objects, etc. In at least one embodiment, platform API 4003 exposes functions that allow device contexts to be used to select and initialize devices, submit work to devices via command queues, and enable data transfer to and from devices. In addition, in at least one embodiment, the OpenCL framework provides various built-in functions (not shown), including mathematical functions, relational functions, image processing functions, etc.

[0373] In at least one embodiment, a compiler 4004 is also included in the OpenCL framework 4005. In at least one embodiment, source code can be compiled offline before executing the application or compiled online during execution of the application. In contrast to CUDA and ROCm, OpenCL applications in at least one embodiment can be compiled online by compiler 4004, which is included to represent any number of compilers that can be used to compile source code and / or IR code (e.g., Standard Portable Intermediate Representation ("SPIR-V") code) into binary code. Alternatively, in at least one embodiment, OpenCL applications can be compiled offline before executing such applications.

[0374] Figure 41 Software supported by a programming platform according to at least one embodiment is shown. In at least one embodiment, programming platform 4104 is configured to support various programming models 4103, middleware and / or libraries 4102, and frameworks 4101 that applications 4100 can rely on. In at least one embodiment, application 4100 can be an AI / ML application implemented using, for example, a deep learning framework (in at least one embodiment, MXNet, PyTorch, or TensorFlow), which can rely on libraries such as cuDNN, NVIDIA Collective Communications Library ("NCCL"), and / or NVIDIA Developer Data Loading Library ("DALI") CUDA libraries to provide accelerated computation on the underlying hardware.

[0375] In at least one embodiment, the programming platform 4104 can be a combination of the above Figure 38 , Figure 39 and Figure 40 One of the CUDA, ROCm, or OpenCL platforms described. In at least one embodiment, programming platform 4104 supports multiple programming models 4103, which are abstractions of the underlying computing system that allow for the expression of algorithms and data structures. In at least one embodiment, programming models 4103 can expose features of the underlying hardware in order to improve performance. In at least one embodiment, programming models 4103 can include, but are not limited to, CUDA, HIP, OpenCL, C++ Accelerated Massive Parallelism (“C++ AMP”), Open Multi-Processing (“OpenMP”), Open Accelerators (“OpenACC”), and / or Vulcan Compute.

[0376] In at least one embodiment, libraries and / or middleware 4102 provide implementations of abstractions of programming models 4104. In at least one embodiment, such libraries include data and programming code that can be used by computer programs and utilized during software development. In at least one embodiment, such middleware includes software that provides services to applications in addition to those that can be obtained from programming platform 4104. In at least one embodiment, libraries and / or middleware 4102 can include, but are not limited to, cuBLAS, cuFFT, cuRAND, and other CUDA libraries, or rocBLAS, rocFFT, rocRAND, and other ROCm libraries. Additionally, in at least one embodiment, libraries and / or middleware 4102 can include NCCL and ROCm Communication Collectives Library (“RCCL”) libraries, which provide communication routines for GPUs, MIOpen libraries for deep learning acceleration, and / or Eigen libraries for linear algebra, matrix and vector operations, geometric transformations, numerical solvers, and related algorithms.

[0377] In at least one embodiment, application frameworks 4101 rely on libraries and / or middleware 4102. In at least one embodiment, each application framework 4101 is a software framework used to implement a standard structure for application software. In at least one embodiment, AI / ML applications can be implemented using frameworks such as Caffe, Caffe2, TensorFlow, Keras, PyTorch, or MxNet deep learning frameworks.

[0378] Figure 42 shows compiled code to run on Figures 37-40In at least one embodiment, compiler 4201 receives source code 4200, which includes both host code and device code. In at least one embodiment, compiler 4201 is configured to convert source code 4200 into host executable code 4202 for execution on the host and device executable code 4203 for execution on the device. In at least one embodiment, source code 4200 can be compiled offline before executing the application, or compiled online during execution of the application.

[0379] In at least one embodiment, source code 4200 may include code in any programming language supported by compiler 4201, such as C++, C, Fortran, etc. In at least one embodiment, source code 4200 may be included in a single-source file having a mixture of host code and device code, with the location of the device code indicated therein. In at least one embodiment, the single-source file may be a .cu file including CUDA code or a .hip.cpp file including HIP code. Alternatively, in at least one embodiment, source code 4200 may include multiple source code files, rather than a single source file, in which host code and device code are separated.

[0380] In at least one embodiment, compiler 4201 is configured to compile source code 4200 into host executable code 4202 for execution on a host and device executable code 4203 for execution on a device. In at least one embodiment, compiler 4201 performs operations including parsing source code 4200 into an abstract system tree (AST), performing optimizations, and generating executable code. In at least one embodiment where source code 4200 comprises a single source file, compiler 4201 may separate device code from host code in such a single source file, compile the device code and the host code into device executable code 4203 and host executable code 4202, respectively, and link device executable code 4203 and host executable code 4202 together in a single file, as described below with respect to Figure 26 discussed in more detail.

[0381] In at least one embodiment, host executable code 4202 and device executable code 4203 may be in any suitable format, such as binary code and / or IR code. In the case of CUDA, in at least one embodiment, host executable code 4202 may include native object code, while device executable code 4203 may include code in a PTX intermediate representation. In at least one embodiment, in the case of ROCm, both host executable code 4202 and device executable code 4203 may include target binary code.

[0382] Other variations are within the spirit of the present disclosure. Thus, while the disclosed technology is susceptible to various modifications and alternative configurations, certain illustrated embodiments thereof are shown in the drawings and have been described in detail above. However, it should be understood that there is no intention to limit the disclosure to one or more specific forms disclosed, but on the contrary, the intention is to cover all modifications, alternative configurations, and equivalents falling within the spirit and scope of the present disclosure as defined by the appended claims.

[0383] Unless otherwise noted or clearly contradicted by the context, the use of the terms "a" and "an" and "the" and similar references in the context of describing the disclosed embodiments (particularly in the context of the appended claims) should be interpreted as covering the singular and plural, rather than as definitions of terms. Unless otherwise noted, the terms "include," "have," "include," and "contain" should be interpreted as open-ended terms (meaning "including but not limited to"), unless otherwise noted. The term "connected" (when unmodified, refers to a physical connection) should be interpreted as partially or completely contained within, attached to, or connected together, even if there is some intervention. Unless otherwise noted herein, references to numerical ranges herein are intended only to be used as a shorthand method of referring to each individual value falling within the range, and each individual value is incorporated into the specification as if it were separately recited herein. In at least one embodiment, unless otherwise noted or contradicted by the context, the use of the term "set" (e.g., "item set") or "subset" should be interpreted as a non-empty set comprising one or more members. Furthermore, unless otherwise indicated or contradicted by context, the term "subset" of a corresponding set does not necessarily mean a proper subset of the corresponding set, but rather a subset and a corresponding set may be equivalent.

[0384] Unless expressly indicated otherwise or clearly contradicted by context, conjunctions such as phrases of the form "at least one of A, B, and C" or "at least one of A, B, and C" are understood in context to generally refer to an item, clause, or the like, which may be A or B or C, or any non-empty subset of the set A, B, and C. In at least one embodiment of a set having three members, the conjunctions "at least one of A, B, and C" and "at least one of A, B, and C" refer to any of the following sets: {A}, {B}, {C}, {A, B}, {A, C}, {B, C}, {A, B, C}. Thus, such conjunctions are not generally intended to imply that certain embodiments require the presence of at least one of A, at least one of B, and at least one of C. Additionally, unless expressly indicated otherwise or contradicted by context, the term "plurality" refers to plurality (e.g., "a plurality of items" refers to a plurality of items). In at least one embodiment, the number of items in the plurality of items is at least two, but may be more if expressly indicated or indicated by context. Further, unless stated otherwise or clear from context, the phrase "based on" means "based at least in part on" rather than "based solely on."

[0385] Unless otherwise indicated herein or clearly contradicted by context, the operations of the processes described herein may be performed in any suitable order. In at least one embodiment, processes such as those described herein (or variations and / or combinations thereof) are performed under the control of one or more computer systems configured with executable instructions and are implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that are collectively executed on one or more processors by hardware or a combination thereof. In at least one embodiment, the code is stored on a computer-readable storage medium in the form of a computer program that, in at least one embodiment, includes a plurality of instructions that can be executed by one or more processors. In at least one embodiment, the computer-readable storage medium is a non-transitory computer-readable storage medium that excludes transient signals (e.g., propagated transient electrical or electromagnetic transmissions) but includes non-transitory data storage circuits (e.g., buffers, caches, and queues). In at least one embodiment, code (e.g., executable code or source code) is stored on a set of one or more non-transitory computer-readable storage media (or other memory for storing executable instructions) having executable instructions stored thereon, which, when executed by one or more processors of a computer system (i.e., as a result of being executed), causes the computer system to perform the operations described herein. In at least one embodiment, the set of non-transitory computer-readable storage media includes a plurality of non-transitory computer-readable storage media, and one or more of the individual non-transitory storage media in the plurality of non-transitory computer-readable storage media lacks all of the code, but rather the plurality of non-transitory computer-readable storage media collectively store all of the code. In at least one embodiment, the executable instructions are executed so that different instructions are executed by different processors, and in at least one embodiment, the non-transitory computer-readable storage medium stores instructions, and the main central processing unit ("CPU") executes some instructions, while the graphics processing unit ("GPU") executes other instructions. In at least one embodiment, different components of the computer system have separate processors, and different processors execute different subsets of instructions.

[0386] Thus, in at least one embodiment, a computer system is configured to implement one or more services that individually or collectively perform the operations of the processes described herein, and such a computer system is configured with applicable hardware and / or software that enables the implementation of the operations. Furthermore, the computer system implementing at least one embodiment of the present disclosure is a single device, and in another embodiment is a distributed computer system comprising multiple devices operating in different ways such that the distributed computer system performs the operations described herein and such that no single device performs all of the operations.

[0387] The use of any and all at least one example or exemplary language (e.g., "such as") provided herein is intended merely to better illustrate embodiments of the present disclosure and does not limit the scope of the disclosure unless otherwise required. No language in the specification should be construed as indicating any non-claimed element as essential to the practice of the disclosure.

[0388] All references, including publications, patent applications, and patents, cited herein are hereby incorporated by reference to the same extent as if each reference were individually and specifically indicated to be incorporated by reference and were set forth in its entirety herein.

[0389] In the description and claims, the terms "coupled" and "connected," as well as their derivatives, may be used. It should be understood that these terms may not be intended as synonyms for each other. Rather, in at least one embodiment, "connected" or "coupled" may be used to indicate that two or more elements are in direct or indirect physical or electrical contact with each other. "Coupled" may also mean that two or more elements are not in direct contact with each other, but still cooperate or interact with each other.

[0390] Unless expressly stated otherwise, it is understood that throughout this specification, terms such as “process,” “calculate,” “compute,” “determine,” etc., refer to the actions and / or processes of a computer or computing system or similar electronic computing device that processes and / or converts data represented as physical quantities (e.g., electronic) in the registers and / or memories of the computing system into other data similarly represented as physical quantities in the memories, registers, or other such information storage, transmission, or display devices of the computing system.

[0391] In a similar manner, the term "processor" may refer to any device or portion of memory that processes electronic data from registers and / or memory and converts that electronic data into other electronic data that can be stored in registers and / or memory. As one of the non-limiting embodiments, a "processor" may be a CPU or a GPU. A "computing platform" may include one or more processors. As used herein, in at least one embodiment, a "software" process may include software and / or hardware entities that perform work over time, such as tasks, threads, and intelligent agents. Likewise, each process may refer to multiple processes to execute instructions continuously or intermittently, sequentially, or in parallel. The terms "system" and "method" may be used interchangeably herein, as long as a system may embody one or more methods, and a method may be considered a system.

[0392] In this document, reference may be made to obtaining, acquiring, receiving, or inputting analog or digital data into a subsystem, a computer system, or a computer-implemented machine. In at least one embodiment, the process of obtaining, acquiring, receiving, or inputting analog and digital data can be accomplished in a variety of ways, such as by receiving data as parameters of a function call or a call to an application programming interface. In some implementations, the process of obtaining, acquiring, receiving, or inputting analog or digital data can be accomplished by transmitting data via a serial or parallel interface. In another implementation, the process of obtaining, acquiring, receiving, or inputting analog or digital data can be accomplished by transmitting data from a providing entity to an acquiring entity via a computer network. Reference may also be made to providing, outputting, transmitting, sending, or presenting analog or digital data. In various examples, the process of providing, outputting, transmitting, sending, or presenting analog or digital data can be accomplished by transmitting data as input or output parameters of a function call, an application programming interface, or an interprocess communication mechanism.

[0393] Although the above discussion illustrates one implementation of at least one embodiment of the described technology, other architectures may be used to implement the described functionality and are intended to fall within the scope of this disclosure. Furthermore, although specific allocations of responsibilities are defined above for discussion purposes, the various functions and responsibilities may be allocated and divided in different ways depending on the circumstances.

[0394] Furthermore, although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter claimed in the appended claims is not necessarily limited to the specific features or acts described. Rather, the specific features and acts are disclosed as example forms of implementing the claims.

Claims

1. A cooling system comprising: One or more distribution manifolds formed within at least one panel of the server unit housing, the one or more distribution manifolds positioned to receive liquid and direct the liquid to one or more outlets, the one or more outlets coupled to one or more cold plates, the cold plates having an inlet port and an outlet port, wherein the inlet port is positioned to couple to a first outlet of the one or more outlets and the outlet port is positioned to couple to a second outlet of the one or more outlets. 2 . The system of claim 1 , wherein the one or more outlets include quick connect fittings for engaging mating quick connect fittings of the one or more cold plates. 3 . The system of claim 1 , wherein the one or more distribution manifolds comprise flow circuits formed within a void space formed in the at least one panel, the void being positioned between an inner panel portion and an outer panel portion. 4 . The system of claim 3 , wherein the one or more flow circuits comprise flexible tubing disposed within the void space.

5. The system of claim 1, wherein the one or more distribution manifolds comprise flow circuits integrally formed within the thickness of the at least one panel. 6 . The system of claim 1 , wherein the first of the one or more outlets is a supply to the one or more cold plates and the second of the one or more outlets is a return from the one or more cold plates.

7. The system according to claim 6, further comprising: a supply line extending through the one or more panels to the supply; as well as A return line extends through the one or more panels to the return, wherein the supply line diameter is different than the return line diameter, or the supply line length is different than the return line length.

8. A cooling system comprising: A housing having an integral distribution manifold formed in at least one panel of the housing, the integral distribution manifold comprising at least one distribution inlet, at least one distribution outlet, and at least one cooling outlet, wherein the at least one panel is movable relative to the other panels forming the housing.

9. The system of claim 8, wherein at least one of the at least one distribution inlet or the at least one distribution outlet is coupled to a server-level manifold via a quick-connect fitting.

10. The system of claim 9, wherein the at least one panel is at least one of axially movable or pivotally movable relative to other panels forming the housing, the axial or pivotal movement exposing an interior portion of the housing.

11. The system of claim 8, wherein the at least one cooling outlet is directly coupled to an associated cold plate inlet.

12. The system of claim 11, wherein each of the at least one cooling outlet and the associated cold plate inlet comprises a blind-mate quick-connect fitting that prevents flow without a connection to an associated fitting.

13. The system of claim 8, further comprising: One or more sensors for determining at least one of flow rate, pressure drop, temperature, or leak status.

14. The system of claim 8, further comprising: One or more flow controllers are positioned along the flow passage of the integrated distribution manifold, the one or more flow controllers regulating fluid flow based at least in part on one or more cooling parameters.

15. A cooling system comprising: a first panel for forming at least a first portion of the housing, the first panel including one or more integral first flow channels; a second panel for forming at least a second portion of the housing, the second panel including one or more integral second flow channels; as well as A cold plate is positioned within the housing, the cold plate having an inlet port and an outlet port, wherein the inlet port is positioned to couple to the one or more integral first flow channels and the outlet port is positioned to couple to the one or more integral second flow channels.

16. The system of claim 15 , wherein each of the first panel and the second panel is movable relative to the cold plate, movement of the first panel from a first position to a second position forming a direct coupling between supply ports associated with the one or more integral first flow channels, and movement of the second panel from a third position to a fourth position forming a direct coupling between return ports associated with the one or more integral second flow channels.

17. The system of claim 15, wherein the first panel further comprises: inner panel; as well as An outer plate is separated from the inner plate by a gap, the one or more integral first flow channels being positioned within the gap.

18. The system of claim 15, wherein the one or more integral first flow channels include a plurality of outlets, each of the plurality of outlets having a blind-mate quick-connect fitting.

19. The system of claim 15, further comprising: One or more sensors for determining at least one of flow rate, pressure drop, temperature, or leak status.

Citation Information

Patent Citations

  • Leak detection and response system for liquid cooling of electronic racks of a data center

    CN110557925A

  • Cooling module design for servers

    US20200404812A1